Vision's LLM Moment: an evening with Robert Geirhos (Google DeepMind)

Robert Geirhos, Staff Research Scientist at Google DeepMind

upcoming: August 18th 2026 @ 5pm - 6:30pm
Robert Geirhos is coming to Heidelberg to talk about one of the more interesting questions in AI right now: are video models quietly turning into general-purpose vision models, the same way LLMs became general-purpose language models?

Think about how much language models changed things. A few years back you needed a separate model for every job: one for translation, a different one for summarizing, another for answering questions. Then LLMs showed up and you could suddenly do all of it by just asking. The recipe behind that shift was pretty plain once you saw it: take a big generative model and train it on a huge pile of web data.

Robert's argument is that the same thing might be starting to happen in vision, and video is where to look.

His team has been investigating Veo 3, and it keeps doing things nobody trained it to do. It can pick objects out of a scene, find edges, edit images, reason about how physical objects behave, work out what you can actually do with an object, and even solve mazes and symmetry puzzles. None of that was the training goal. It just emerged, the same way many LLM abilities did.

If that pattern holds up, it is a big deal. It hints at a future where one vision model handles most perception tasks instead of requiring a new specialized model every single time.

This evening is worth joining whether you build with these systems, research them, or are simply curious about where they are headed. You will hear the case from one of the people actually making it.

About Robert

Robert is a Staff Research Scientist at Google DeepMind in Zürich, working on reaching visual intelligence through video models like Veo 3. He is one of the senior authors of the paper this talk is based on, "Video models are zero-shot learners and reasoners."

If you have been around computer vision for a while, you probably already know some of his earlier work: the 2019 paper showing that image classifiers lean on texture far more than shape, and "Shortcut learning in deep neural networks" from 2020. Both are cited constantly. Along the way, he has received the ELLIS PhD Award, a NeurIPS Outstanding Paper Award, and multiple oral presentations across ICLR, NeurIPS, ICML, and VSS.

How the evening runs

Robert will speak for about 60 minutes, followed by an open Q&A. Afterwards, stick around to meet people and chat over drinks and snacks.

Event Info

Please help us plan ahead by registering for the event on our Meetup event page.

What? Vision's LLM Moment: an evening with Robert Geirhos (Google DeepMind)
Who? Robert Geirhos, Staff Research Scientist at Google DeepMind
When? Tuesday, August 18th 2026, 5:00 PM to 6:30 PM CEST
Where? BioQuant, Room SR041, Im Neuenheimer Feld 267, Heidelberg
Registration Meetup event page