Why Frontier AI Still Sees Like a Toddler, w/ Andrew Dai

17 Jun 2026 · 43 min · 17 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Frontier AI’s visual reasoning is still weak—often no better than a preschool/elementary child—especially when tasks require counting and spatial/combinatorial reasoning beyond ~5 objects. The episode argues current “frontier” multimodal models rely heavily on pattern matching and lack a “visual chain of thought,” so they struggle with complex spatial tasks, physical/engineering-style reasoning, and long-horizon counting.

Guests

Andrew Dai (Andrew Dye), co-founder and CEO of Elorian; previously at Google Brain and DeepMind, including work connected to Gemini and sparse mixture-of-experts systems. Hosts: Corey Knowles and Grant Harvey (Neuron AI Explained).

Key claims

Visual reasoning scales less effectively than language because vision isn’t naturally next-token prediction in 1D. Better results require integrating language + visual reasoning in one model and building a visual chain-of-thought. Elorian trains via synthetic data plus vendor-collected manual labels of how people solve these tasks.

Notable examples

counting place settings for 20 people; metro maps (London) described only in text; tangled USB cords; folding/rotating/spatial matching benchmarks; counting marbles/jar volume with varying sizes; UI layout and engineering design workflows (hundreds of hours in mechanical CAD for one component).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Visual Limitations of Frontier Models

0:00 to 1:00

Explore the reasoning abilities of AI models compared to children.

“An elementary school kid can beat all the Frontier models.”

Understanding Visual Reasoning Challenges

2:05 to 5:05

Discussion on the complexity of visual reasoning and AI's limitations.

“Well, I guess to get started, the thing that when I first went and looked, you guys up that really jumped out at me was your core claim that frontier models still can't reason images much better than a three-year-old.”

The Impact of AI in Design and Visual Tasks

5:05 to 7:06

Examination of AI's role and challenges in design and visual tasks.

“and say I'm only allowing you to use language, no diagrams, no visuals whatsoever.”

Challenges with Multimodal AI Models

8:04 to 12:14

Exploration of where multimodal AI models struggle and their reasoning abilities.

“The spatial combinations, obviously, there's a huge number of spatial combinations as soon as you get more than five objects or so.”

Visual Reasoning Strategies and Future Opportunities

12:14 to 14:01

Discussion on visual reasoning strategies and the future of AI development.

“Which is this sequence of words, which is very, very good for communication, but very different to the visual world, right?”

Counting and Visual Reasoning in AI

14:01 to 16:46

Explore how children count and the implications for AI visual reasoning.

“But if you think about it, this data isn't really available on the web.”

Visual Reasoning vs. Lingual Reasoning

16:46 to 18:16

Discuss the differences between visual and language reasoning capabilities in AI.

“Visit outshift.com to learn more about the Internet of Cognition.”

Challenges of Spatial Reasoning for AI

18:16 to 20:46

Understand the difficulties AI faces with spatial reasoning tasks and their implications.

“Yeah, the paper specifically was like, you know, there's a spatial ceiling for folding, reflection, rotation, and spatial matching.”

Integrating Visual and Linguistic Reasoning

20:46 to 23:42

Learn about the integration of visual and linguistic reasoning in AI models.

“There's a 2D for a bunch of problems like recognizing characters and a bunch of things.”

The Future of AI Models and Device Operation

23:42 to 25:56

Discuss the potential for AI models to operate on edge devices versus the cloud.

“So could these models be small enough to operate on device, like on a satellite or on a robot?”
Show all 17 chapters

Evaluating Progress in Visual Reasoning AI

25:56 to 28:00

Examine how progress in AI visual reasoning is evaluated and the challenges involved.

“But you keep pushing bigger, too, at the same time to further the advancement itself.”

Evaluating Progress in AI Models

28:00 to 29:19

Explore the challenges and importance of evaluation methods for AI models.

“You've probably seen like models are being released today and they're still using evaluation data sets that are three, four years old, or maybe even longer.”

Challenges of Benchmarking AI Models

29:20 to 31:09

Discuss the issues surrounding benchmark saturation and the need for modern evaluations.

“And we also test our model across a wide range of evals.”

Visual Reasoning and Its Potential

31:10 to 36:15

Understand the implications of visual reasoning in AI for future technologies and product development.

“Basically, I don't know if you saw the post from Dan Shipper about after automation where he talks about how if you just change the frame a little bit of an eval, you revert everything back to zero.”

Future of AI and Product Development

36:16 to 37:55

Envision the impact of visual reasoning on product timelines and engineering labor.

“It's the longer you wait, the more you'll pay, it seems like.”

The Role of Agents in Engineering

37:56 to 41:40

Investigate the potential for AI agents to enhance engineering workflows and labor.

“reasoning, the life cycle, the timeline for physical products and just physical things gets much more compressed.”

Closing Remarks and Call to Action

42:21 to 42:44

Encouragement to engage with the podcast and subscribe for more content.

“Well, if you haven't yet, please hit the like and subscribe buttons below to help us continue to bring you the most fascinating people in the AI MySpace.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00An elementary school kid can beat all the Frontier models. If you have an elementary school kid that can beat any model, that seems to indicate we're not quite there yet, at least in the visual world. So let's say there's a table and there's some place settings around it and I ask you how many people is the place settings set up for, right? And then I give you a certain amount of time to look at that image. And this is also at the level of which current Frontier models can operate at. So they use mostly pattern matching, very well known behavior. Same with kids. We talked to an engineering firm who's spending, that team is spending hundreds of hours in the design software.

0:45And this one component is taking them like 200 to 300 hours of mechanical engineering time. It's very odd to think about as a software engineer, because, you know, easy to automate, but these like more traditional things still take a long time.

1:02Corey:Welcome humans to the Neuron AI Explained. I'm Corey Knowles, and I'm joined today by the man, the myth, and the follow-up question, Grant Harvey. How are you today? I do ask a lot of follow-up questions, so it's an accurate description of me. Well, today we're going to talk about how AI can write code, pass exams, summarize the internet, But you show it a real-world image that requires some amount of spatial judgment, maybe physical reasoning or engineering context, and suddenly that magic starts to wobble. Well, Elorian is a new AI research and product lab focused on that problem, visual reasoning.

1:36Corey:And we're joined today by its co-founder and CEO, Andrew Dye. Andrew spent years at Google Brain and DeepMind, including work connected to Gemini and sparse mixture of expert systems. Put simply, he's one of the researchers whose career runs straight through the modern AI stack, from pre-training and instruction following to sparse scaling and Gemini-era multimodal systems. Basically, we're incredibly excited to talk to him. Quick note that today's episode is sponsored by Dell Technologies and NVIDIA. You're going to hear more about them here in just a little bit. Andrew, welcome to the Neuron.

2:07Corey:Glad to have you here. Thanks, Kari. It's an honor to be on here. Awesome. Well, I guess to get started, the thing that when I first went and looked, you guys up that really jumped out at me was your core claim that frontier models still can't reason images much better than a three-year-old. And I was wondering, what do you mean by reason about images? And what is it that models are failing to do today? Yeah, I would say that reasoning is kind of like a complex subject, as you can imagine. I think you see some of these video generation models and they say these models do reasoning and therefore they create great video you can also some people might think of oh it can identify exactly what species this flower is as a kind of reasoning and yeah reasoning is involving all these problems but really what we're talking about is more complex reasoning so one example can be if I ask you a question about a picture.

3:11For example, let's say there's a table and there's some place settings around it and ask you how many people is the place settings set up for? And then I give you a certain amount of time to look at that image. So I would guess that if it's only, say, a table for two or three people, it would take you less than a second to say, well, this is set up for two people or three people or four people. It's a signal that will tell you the answer. And this is also at the level of which current Frontier models can operate at. So they use mostly pattern matching so that you can instantly pattern match and get an accurate number if it's less than five or so.

3:48It's also called, yeah, it's a very well-known behavior. Same with kids. But then if I show you the image and it's laid out for like 20 people, say it's a really long table, like Harry Potter style environment or banquet, then I'm going to guess it's going to take you longer than one second to count those place settings. And it is exactly in these types of scenarios that these models fall flat. So anything we say that takes longer than a second, these models just can't handle.

4:17Corey:What's the difference between, say, like an image description when you give a chatbot, like Chatubichi or Claude, an image, and you ask it to describe it versus actual visual reasoning? Like what is the difference? Yeah, so visual reasoning is kind of more than just saying what is in an image, right? So we can take this to a more complex level. So in the place setting scenario for a banquet, you can say that it's set up for 20 people. Everyone has a fork, a knife, and a plate. But there are lots of scenarios where description just doesn't work. So, for example, say there's a metro map, right? 2D metro map.

4:56Say it's London. London has one of the most complex metro maps in the world. for you to fully describe that metro map and say I'm only allowing you to use language, no diagrams, no visuals whatsoever. I'm expecting it's going to be pretty hard for you to have every detail of that metro map laid out. So there are just things that are simply much easier to describe with visual imagery. The language mazes are another one. or like say you have a tangle of USB cords and you want to figure out which one to unplug to get your power adapter. It's going to be very hard to describe that situation in text.

5:40Corey:I don't even know how to do that without physically pulling them apart. So here's the thing about enterprise AI right now. Every leadership team has the pilot. Every company has the proof of concept, but actually scaling AI from the developer's desktop to the data center all the way out to the edge, that's where things start to break. In fact, 95 % of organizations say they can't even get their data ready for AI workloads. Not the model, the data. That's why we teamed up with Dell AI Factory with NVIDIA to build a resource hub we're actually excited about. It includes articles, videos, case studies that cover the full AI journey from strategy and data foundations to infrastructure decisions to real-world use cases and ROI.

6:22Corey:You'll find stories on what an AI-native factory looks like when physical AI hits the production floor to where hybrid AI workloads should actually run and what sovereign AI means when you're operating under real compliance constraints. And the whole thing is backed by the industry's first end-to-end enterprise AI portfolio. AI-ready workstations, servers, storage, networking, all jointly engineered with NVIDIA. So you can start small on a ProMax workstation, scale out to the data center, and get to production up to 86 % faster than going it alone. So if you're tired of hype and just want to understand what it actually takes to make Gen AI work, head over to the Enterprise Guide to Scalable AI Hub on techrepublic.com or click the link in the description of this video.

7:06Corey:I guess maybe a slightly more technical direction here. Where do current models, specifically let's say multimodal models, where do they break first? Is it around spatial relationships, object permanence, physical causality? Are there different areas like that where it's like they do well at this, but this is just awful? Yeah, spatial relationships they do very well at. One way I like to think about this is the difference between pattern matching, pattern recognition versus more complex reasoning. So a lot of these models, anything that is pattern matching, they can do very well. And that includes identifying what species this flower is or what breed of dog this is.

7:49All you need to do there is pattern matching. But pattern matching, of course, breaks down when you have things that are more combinatoric, when you have many possible combinations of some things, many possible ways things can be arranged. You can describe this as a spatial reasoning problem, or you can just describe it as just a reasoning problem. The spatial combinations, obviously, there's a huge number of spatial combinations as soon as you get more than five objects or so. So that's where it shows up quite obviously, but it also shows up in things that you don't need to strictly describe as spatial, but has some spatial element like counting, following chords, navigating, navigation, these kinds of problems.

8:29Corey:That makes sense. I feel like that's a really, you know, in a practical setting, that's really important for, let's say, like even just a simple use case of, hey, I'm trying to design a user interface. And, you know, can the agent that I'm working with actually look at the user interface and know, you know, what everything is, like keep track of all of it, reason about it? Because anecdotally, I've seen some of the top models lie to my face about an image, especially in visual design work, when you needed to actually see what it's doing. So what's going on there and how can this be improved based on what you're building?

9:03Yeah, design, I think, is one of the most interesting areas for us in design. It is definitely much more about pattern matching, right? To create a great UI, you need to understand space, 2D space, how to arrange things so things are intuitive and clear. And you don't end up with an interface from like the 90s or something like that.

9:27Corey:unless that's your goal some people do like that oh yeah they're creating like a retro game or retro experience yeah yeah so yeah design props up everywhere so we talk currently about like these models replacing white collar work right but white collar work some of it is coding definitely not the majority and then there's some math and then there's a lot of like document analysis summarization writing but the rest of white collar work is all visual like a part of that is design going from web design through to marketing design but also engineering mechanical engineering things like iPhones EV batteries Starship rockets and I'm pretty sure that no one is designing the next iPhone using code designers just they don't code they don't you can't design with code these things like they have a presence in the physical world and they're like every angle every centimeter right for an iphone every millimeter matters and code just isn't precise enough to do that in a nice way for the designer and then there's other things like understanding photographs understanding charts bar charts line graphs say like a stock market charts right these things are all inherently visual and the way a lot of these current models are working with these this that the chain of reasoning or chain of thought is entirely in the tech space, which leads to a lot of these issues that we see with hallucination and other things happening.

10:56Corey:So is what you're building essentially creating some sort of visual chain of thought? Is that part of it? Yeah, exactly. That's a core part of what we're doing. And we believe that the visual chain of thought is a necessary component. So say when you're buying a sofa, you're probably going to visualize how it looks like in a living And when an architect is designing a new floor plan for a new office building, they might need to figure out how a wheelchair can navigate through the space is everything. And I think it's very hard to do that just through text. Wow. Well, you know, we've scaled text models.

11:32Corey:We've scaled multimodal data. We've scaled compute. Why hasn't scaling solved this already? What makes the visual world more stubborn, I guess? Yeah, great question. So all frontier models model language fundamentally different to vision. So language is done with this pre-training next token prediction setup, which is the setup we proposed 11 years ago. We had the first paper on pre-training and fine tuning back in 2015. And that approach has turned out to be very scalable once it's combined with transformers from 2017. And obviously, we've just been able to scale and scale that approach. But vision is very different.

12:11So language naturally lends itself to next token prediction because we communicate in 1D, right? Which is this sequence of words, which is very, very good for communication, but very different to the visual world, right? Visual world is inherently, we're perceiving it through 2D or 3D through our eyes. So squashing the visual world into a 1D sequence to predict the next token in visual space doesn't really make sense. And so I don't think we've really hit upon the right setup there to have vision models that scale. So you do see that scaling is not as effective for vision as it is for language.

12:53But that also creates opportunity because for us, it means we don't need to build trillion parameter models to have a state-of-the-art visual model. That's great. There's a lot of opportunity to find something that does scale.

13:08Corey:Nice. Does this look something like, you know, instead of predicting the next token, you're predicting the next image, predicting the next frame? How do you actually do the visual reasoning part of it? Right. So the visual reasoning part of it, we're taking a more, like I would say, guided approach. so for kids say kids that are like preschool age if you ask them how many how many balls are in a box if there's three or two say there's three two one or four four balls usually they can get the answer right immediately instantly without any kind of reasoning or anything and they have a pretty decent accuracy and that's because that's from pattern matching because with an arrangement of three balls is very you just intuit what the number is yeah you don't even think about it it just exactly exactly but if you have more than five balls if you start having 10 or 20 what you see these kids doing is they they start pointing to the balls to track what they've counted already and what they have left to count for adults obviously we sometimes do this in our mind but for kids you can still see this happening and you see this happening for various other tasks like following a line or following a maze, they will use their finger to point or they will use a pen to draw.

14:30But if you think about it, this data isn't really available on the web. No one has like counting data on the web because it's just so obvious to adults how to do it. And also it's more work to edit an image, right? And so this kind of data isn't available. So what we're doing is we're both constructing the data through synthetic data And we are also working with vendors to collect this data where people manually label exactly how they are solving that problem.

15:01Corey:Wow. That just really made me think of, like, you used to see these contests in a store where there would be, like, a great big jar full of marbles. Yes. And you have to, you know, it'd be like, whoever gets the closest gets the jar or gets, you know,$50 or something. And this makes me think about the difference in counting what's laying flat in a box versus needing to be able to intuit the volume possible inside this jar, for example, and be able to intuit how many M &Ms or marbles or whatever are likely in the jar. Is that kind of the sort of difference we're talking about? Yeah, that's exactly right.

15:38That's a great example because I've also seen problems where there are varying sizes of marbles, balls in a jar. Some of them are like jawbreakers, you know, size of a small fist and some are smaller. And then you get a peek at, you know, how many there are, but then you have to do a little bit of basic statistics and volumes and everything. But yeah, this is exactly the kind of problem that needs this combination of visual reasoning and linguistic reasoning.

16:09Corey:So that's awesome. That's really what I was wondering. I remember those from like movies and seeing them as a kid back in the 80s and 90s in places. It was a thing. Are you hitting the limits of siloed AI? Just as humans once transformed society by sharing intent, knowledge and innovation, AI faces a similar inflection point. To achieve distributed superintelligence, we must move beyond simply scaling up. We need to scale out, too. OutShift by Cisco is building the Internet of Cognition, an open infrastructure enabling agents and humans to collaborate in real time. Visit outshift.com to learn more about the Internet of Cognition.

16:50Corey:So you worked on Gemini, which was designed as a multimodal model pretty much from the beginning, if I recall correctly. And what did that experience teach you about the limits of today's multimodal training recipes? Yes, I was a data area lead for Gemini. I've been leading Gemini since day one. And what we've seen is that language and coding capabilities have really advanced rapidly, both in Gemini and outside. But we do see that visual capabilities, multimodal capabilities on the reasoning side have been slower to advance. on the generation side obviously as being in leaps and bounds there's a lot of great generation models like VO and Genie and others that can generate very high quality video but this isn't necessarily reflected in the reasoning part and there was a benchmark that came out a few months ago that exemplified that that benchmark showed that these models say the capabilities are still at a level of a preschooler yeah that was a visual reasoning benchmark right from yeah Yeah, we have a question about that, actually.

17:59Corey:Yeah. Yeah, and so, yeah, even an elementary school kid can beat all the frontier models. So to me, when we're talking about AGI and already being an AGI in the visual world, if you have an elementary school kid that can beat any model, that seems to indicate we're not quite there yet, at least in the visual world. Yeah. Yeah, yeah. Yeah, the paper specifically was like, you know, there's a spatial ceiling for folding, reflection, rotation, and spatial matching. Why are those problems so difficult or why are they AI's weakness? Yeah, so I think it comes from a combination of things. One is, like I said before, about these models only having the text chain of thought and no visual chain of thought.

18:46But another reason is that if you look at the development of computer vision, it's been very focused on just a few tasks. So these are object recognition. That's why Google Lens is so great. And we can identify any plant or flower. And another one is like localization, seeing where things are in image, segmentation, detection. These are things that are very useful for specific use cases, like surveillance, monitoring, like when your doorbell goes off and when your doorbell app goes off and says, oh, there's a mouse running across your yard or a scroll or something. That's all the traditional setups and objectives.

19:33But these like more complex visual reasoning tasks, like spatial reasoning tasks, has, I'd say there's been less attention paid to those. I think partly because they are so hard And another part I'd say is probably they do require some degree of like text reasoning. And that text reasoning just didn't exist. One example is one and a half years ago, all one didn't exist. We didn't even have text reasoning. Sounds like a lifetime ago, but it was only 18 or 20 months ago. That's crazy. Yeah. And then we had that and now and then we have great math and coding models. and yeah that line of progression it just was clear to me while I was while I was in Gemini that the next step would be incorporating visual reasoning now that we have a very good foundation

20:23Corey:that makes a lot of sense when I think about how I think of you know if I take a piece of paper and I start to fold it there's some part of it that I intuit there's some part of it that I I don't know if I say in my head if I fold it this way it's gonna look like this but or if it's just memory like there are a lot of parts of your brain that you're putting to work there when you think about how something will fold or or rotate it makes me wonder if there needs to be like a 3d element of this or like you have to train models in 3d for them to understand the space i don't know the technical behind that but curious what your thoughts are there yeah that's uh great yeah that's very perceptive i'd say yeah i i agree that there is an element of like 3d modeling in our brains.

21:08There's a 2D for a bunch of problems like recognizing characters and a bunch of things. But yeah, I agree there's some kind of 3D modeling going on. So that needs to be represented in the chain of thought as well.

21:21Corey:Yeah, because like I see that light just above your head, that strip light. And I don't think that's a light that's crooked on a ceiling. I know that that light is at that angle because it's going away, that it's getting farther away from you. Wow. really interesting so it sounds like high thoughts but you have to think of it that way like you have to think about all the things you take for granted to try and think how how they think i always think could you could you watch a video while you folded that paper i think might be a good test like could you truly do something else while that's going on or or is it a thing that you are actually consciously considering as you go i would be curious to know and honestly i don't know it's I'm just spitballing here ideas.

22:04Corey:Speaking of, something I'm curious about is like, what's the goal for a Lorient? Are you seeing this as like, we're going to build a foundation model? Is this going to be a reasoning layer that would exist within other models? Or maybe some type of something I'm not even considering that maybe doesn't fit neatly into the buckets we know already? Yeah, so we are building a visual reasoning frontier model. And we believe that it's very important to integrate a language and visual reasoning into the same model. So our model can be called independently and it will do all these kinds of reasoning by itself.

22:40And yeah, as part of that, like I gave with the example with the floor plan or the sofa, it will need to generate images. It will need to be able to edit images. Eventually, it might need to like rotate things in 3D space, but that'll probably be later. very very but yeah these are all things that we believe needs to be done at the same time as a text-based reasoning and the text-based reasoning is important to connect with humans right because that's how we communicate right that's how you tell a model to do something or that's how you you get the model output if it's a question answering setup so that's why we're not just doing a visual train of thought that's why it needs to be quite integrated into the text part as well.

23:25Corey:I suppose it makes sense then why the ability to communicate with it via text is the place to start. Like as in terms of AI development as a whole, like to do the other things, you will need to be able to communicate with it. So you, I guess, start there. Makes sense. Yeah, exactly. Yeah. One saying, wow, was it Gemini that I thought nailed this quite well is that Gemini is spread into several areas as a data area there's a eval area and really data and eval is about communicating with the model right it's the only way we have to communicate with the model in a sense totally yeah yeah well it's kind of interesting that we're starting with language it makes sense from for where we're at now as a society but it seems almost like every other you know intelligent being animal human what have you started with maybe perhaps a visual reasoning first so it's kind of funny that we're doing it in reverse if you think of it that way yeah because they don't all have language you're right yeah like well i was just thinking about flies you know in my in my house today it's like well the fly you know it can't communicate with me but it can see me it can navigate this 3d space and and you know its brain is i don't know the order of magnitude but much much smaller than than ours and it still has that kind of encoded in it so it is interesting to build this like to build the visual reasoning second you would think maybe you know everything else built it first yeah yeah that's a great observation but yeah ultimately i think the reason this happened is as the compute costs get bigger there's definitely a need for this to be economically useful and it's going to be very hard to interact with a fly brain even if we could well related to this so some of your stated use cases include engineering design satellite analysis and robotic navigation.

25:14Corey:So could these models be small enough to operate on device, like on a satellite or on a robot? Or will it be, you know, still like cloud for the frontier and, you know, you stream it to the device? How are you envisioning that? I think initially we are building something for the cloud, for streaming. But yeah, I think it's definitely like edge devices, like there's a huge range of applications there. So we will be looking to build smaller distilled models as well that are compact. That seems to be the path for AI is that, you know, you're always pushing the frontier, which means it's going to be large.

25:52Corey:But then like as it gets large, you start working to make it smaller and more condensed. But you keep pushing bigger, too, at the same time to further the advancement itself. So I think what you're working on is really, really unique. And it's not a conversation we've had before. So we're both really excited to talk about this. We're starting basic and working our way up. Exactly. Related to this, so you've also worked around sparse mixture of experts or MOE style scaling. Does visual reasoning require a different kind of specialization than language reasoning? Almost like different experts for the different elements like geometry, physics?

26:29Yeah. Yeah, the mixture of experts architecture that's currently popular around all the labs is quite language-tuned and language-specific. So as soon as you have multimodal inputs, there are some odd things that happen to the experts. So the allocation becomes uneven. And this is because fundamentally the underlying representations, call them activations or embeddings or whatever, But fundamentally, they are different between visual inputs versus linguistic inputs. Because obviously for language inputs, they fit very well into this neatly organized embedding space. But for multimodal inputs, it's not the case.

27:12These days, people don't use these kind of discrete representations. They use continuous representations. So fundamentally, the behavior is quite different, especially with MOE setup. So there are more things to think about in terms of optimization, balancing the experts, all those kind of things. So we're doing some adjustments to the architecture there. And we can do it because we are specializing our model to these multimodal use cases. Whereas if you needed to trade off and balance multimodal with coding, with math, with language, then you can't make the kind of changes that we want to make.

27:47Corey:That makes sense. Yeah, because it's almost like a jack of all trades, master of none. kind of situation if you're trying to do general purpose it's like maybe you can get away with a lot more when you're changing it yeah in a niche setting yeah exactly something i've wondered is how do you evaluate progress with what you're working on like benchmarks are tricky they can be gamed models learn them a little bit over time and real world imagery is messy it could be anything so like how do you know when your model is improving yeah yeah that's right evaluation is a hot topic and also very hard with these models.

Read the full transcript

28:27You've probably seen like models are being released today and they're still using evaluation data sets that are three, four years old, or maybe even longer. So yeah, we believe that this should be kind of an expiration date to these evaluations. And there needs to be much more work on evaluations. So one example is ArcGGI, which people have been using to say, all these models are really good at multimodal reasoning. But RKGI 1 is only 32 by 32 pixels, which is quite far from any sensible real world scenario. Yeah, yeah, exactly. Other than like Atari games, maybe.

29:03Corey:Yeah. So I think there definitely needs to be more work there and people building more evals. For us, we are also building evals as well. Internally, we hope to release eval externally to contribute to the community later on. And we also test our model across a wide range of evals. But I think, yeah, I would hope that the industry invests more on evals, just like they're investing so much on data, right? Well, like you said, data and evaluation is the way we communicate with the models. It's like the only way we communicate with them. So if you're only communicating in one direction, that's not...

29:46Exactly, exactly. So there needs to be much more attention there and much more protection against like bench maxing and things like that. Yeah. Current models, whether intentionally or not, they are training on these evals. If they've been out on the internet for a while, it's hard to even detect them, you know, training data.

30:05Corey:Well, are those two things in conflict though? Like can, can you at the same time publish and release more evals and then trust them and not bench max on them? Because I thought one of the reasons they don't release benchmarks publicly is like so that, you know, the models aren't benchmarking essentially. Yeah. But the thing is, people are still releasing the benchmark results. So even if things are not being released publicly, the model providers will still release benchmark results. And it's better for them at least to release some newer evals and older evals, right? older evals, they would have been forked, copied, modified like hundreds of times over the past few years.

30:49That's just inevitable. So I think we should still be releasing more modern evals, but we should also be keeping some of the evals held out, you know, secret as a test set to prevent some limited amount of benchmarking. But ultimately, I think the best solution to benchmarking is to just regularly release evals.

31:08Corey:Yeah, and newer ones and fresher, Basically, I don't know if you saw the post from Dan Shipper about after automation where he talks about how if you just change the frame a little bit of an eval, you revert everything back to zero. And so we keep changing the frame of what we're evaluating these models on, and it keeps resetting the scores. So the more that you do that, that can actually be a good thing because then you're resetting expectations and you're pushing it further and further. I suppose there's a cost element there too, though, because then if you actually want to see how your model does, you could wind up theoretically retesting it repeatedly because every time you change the model, the old numbers become useless, right?

31:51Corey:Like you're not evaluating Apple's devils. I mean, unless you're able to somehow like have new questions that are similar but not, I don't know. Yeah, it's not an easy topic. It's not? Yeah, ideally you should be re-evaluating old models when you have new evals, just to know we're really making progress. But really, in machine learning, what we say is that the evals drive progress. So if you don't have the right evals, then you're not really driving progress in that direction. That's what we talk about when we have benchmark saturation. That eval is not so useful for progress. Tinging topics slightly before we wrap up here, would you say that visual reasoning would you say your visual reasoning model is another path that's completely separate from language models it's something that would be built into models via like a multimodal situation or is it part of a system and or of models where the agent could actually call it when it's ready to review imagery or when it needs it or all of the above because it's flexible enough that it could fit in all those scenarios what do you picture there right And now what we're building is tuned towards visual problems, any kind of visual problem.

33:09So currently, if you have a pure text problem or a problem without any visual element, then we'd expect you to use one of the other models on the market. But if there is a visual element to it, we'd expect you to use our model. And that's just because we are building a specialist model and we believe that's the way things are heading in. And like you said, it's very hard for a generalist model, which has to make these trade-offs compared to a specialist one. And yeah, you can, in theory, use this as a layer on another model. But we think that'll be like a suboptimal experience, both from the user side and from the actual modeling side.

33:49Because then you can imagine the internal representation for what this problem is, what this image is. Like a maze doesn't line up with the text-based reasoning that happens later on and you get like hallucinations. And hallucinations are a bit more common when these representations start going to some bad places.

34:12Corey:Well, you all launched with a pretty healthy seed round and strategic participation from NVIDIA and other major AI researchers. I'm curious, now that you have that level of backing, what does that make possible for you that wasn't before? Yes, these models, even though we don't need to scale to training some parameters, we still need to train some sizable models to be competitive with the state of the art there. Visual reasoning, like I said, is kind of a new area. So it does need more research, more experiments to understand what's the right way to do it, what kind of data you need. So we need to spend like that on a sizable amount of compute.

35:01So we have some compute contracts now that have us guaranteed compute for the next three years. But we're also starting to build our own cluster because ultimately like having our own cluster gives us a lot of control over the networking, things like that. And those elements are really important to building a great model.

35:23Corey:Makes sense. that's great though because that's that it's all about just helping you get to that next step yeah i know chasing compute right now is is got to be wild yeah and uh people say now compute is an advantage it's like if you start now it's a lot harder to get it's just a lot harder to secure compute prices have been going up like really crazy amounts but yeah and locking it in for three years i mean yeah yeah yeah you know that was uh we've grant and i have talked a little bit about that debate a couple of years ago between open ai and anthropic where like open ai started locking down these like long-term compute deals two years ago and and i think that's kind of something that really kept them afloat this year because it's getting increasingly more expensive It's harder to get.

36:15Corey:It's more scarce. It's the longer you wait, the more you'll pay, it seems like. Yeah, exactly. And so, yeah, I'm glad we started this company. When we did, if we waited half a year later or a year later, then it would be much harder and we'd need much more capital. But, yeah, given that we started in December, we've been able to get up and running. And we're already training models, already getting state-of-the-art results on some benchmarks compared to the frontiers. So, yeah, we can make progress. Do you have an ETA when you might be ready to bring something to market or even a pipe dream goal?

36:55Yeah, we're targeting before the end of this year. Yeah, so it's fairly fast compared to some of the other Neo Labs. But AI competition is very intense and we want to stay at the frontier. It's very important too. progress quickly.

37:15Corey:You certainly have the background for it. You've been doing this for a good amount of time. You've worked on a ton of amazing papers. I mean, I don't even need to list them all, Palm and Palm 2 and all this stuff. I guess my question, one of my last questions here is five years from now, what changes if visual reasoning works the way that you want it to work? What new class of products become obvious? And is this more like a first step towards a true, let's call it like an interaction model or true world model? I know those terms have been thrown out a lot recently and have gotten confused a little bit, but what do you see it looking like in the future, five years from now?

37:52Yeah, five years from now, I would say as a result of advances in visual reasoning, the life cycle, the timeline for physical products and just physical things gets much more compressed. So instead of waiting, you know, several years for a new model of car to come out, maybe there's a brand new model every year and each one is lighter and more performant. Same with robots and batteries that the timescale for these just gets vastly compressed. And yeah, I think that's how we measure progress these days, right? How often does your battery capacity go up? How many lenses are in your iPhone? These are things that we can measure, but these take very, right now, they take a very long time to develop.

38:56Corey:Now it's about better, smaller, faster, cheaper, I would say, is the long-term goal is like serving the maximum amount of quality you can, as light as you can, as affordable as you can, with as high quality as you can, I guess. Yeah, but I think a lot of people don't realize that a lot of this is still bottlenecked by human labor. There just aren't enough like mechanical engineers or designers in the world. And so there are people we talked to an engineering firm who's spending, that team is spending hundreds of hours in the design software, modifying this component for like a platform that they're building.

39:35And this one component is taking them like 200 to 300 hours of mechanical engineering time. And yeah, it's very odd to think about as a software engineer because, you know, easy to automate, but these like more traditional things that still take a long time.

39:50Corey:Wow. Well, do you envision this working like an agent that then like they can work with and is directly in the software that they're using? Is it maybe even a new interaction paradigm? Because I know like speaking of Apple, they've been working on trying to do AR glasses. They have like a cool demo with visual. I forget what it's called, but the vision, you know, AR, VR kind of mixed reality thing that they're doing. Like, is it something where it makes it possible to do this much quicker? Like you give the spec to the agent and then the agent can get you like that 90 % or 100 % prototype. Are they like subbing in for human labor where now all of a sudden I have 10 engineers on my team when before I had just me or 100 or how do you see that?

40:34Corey:I mean, I know it's kind of hard to predict, but. Yeah, I see it being part of an agentic workflow for sure. And yeah, I don't think it's going to replace these mechanical engineers or these technical designers because from what I understand, there's a big shortage for this kind of skill in the market. You just can't convince people to do mechanical engineering compared to software engineering. The barrier to entry is much lower. Anyone can learn coding and pick that up really quickly. but to learn things like how to use, you know, some of these packages like sorted works. You have to learn like pretty advanced operations like extrusion, rotation of a 3D object.

41:16It's just like a completely different set of skills. And a lot of people just aren't interested in learning that. So there's a huge shortage.

41:22Corey:Well, maybe that'll change because, you know, maybe the coding, the labor market starts to like get really competitive and wages go down. And then people say, you know, we need to go into mechanical engineering for all the reasons you just said. exactly exactly go back yeah well Andrew thank you so much for joining us today it's been a fascinating conversation thanks so much for having me on the show it's a pleasure absolutely where can people go to keep up with you and the work that Elorian is doing yeah so we have our website elorian.ai and we also have accounts on ex Elorian.ai and LinkedIn and there we're posting some of the work that we're doing and our plans That's awesome.

42:04Corey:Thanks again to Dell Technologies and NVIDIA for sponsoring today's episode. For our full content hub on AI-native factories, hybrid AI workloads, data readiness, and sovereign AI, head on over to the Enterprise Guide to Scalable AI Hub on techrepublic.com, or click the link in the description of this video. Well, if you haven't yet, please hit the like and subscribe buttons below to help us continue to bring you the most fascinating people in the AI MySpace. Also, pop by the Neuron.ai to sign up for our daily newsletter and check out the blog content that we publish throughout the week every week.

42:37Corey:And as always, thank you for joining us. That's all we have for today. So farewell for now, humans.

From the publisher

AI can write code, pass exams, and summarize the web, but ask it to reason through a real-world image, and the magic often breaks. Andrew Dai, co-founder and CEO of Elorian, joins The Neuron to explain why visual reasoning may be one of the biggest unsolved problems in AI.


Andrew spent years at Google Brain and DeepMind, including work connected to Gemini and sparse mixture-of-experts systems. Now, he’s building Elorian around a simple but powerful idea: if AI is going to understand the physical world, it needs more than text-based reasoning layered on top of images.


In this episode, Corey and Grant talk with Andrew about why frontier models struggle with counting, navigation, design, engineering, charts, and physical reasoning; why scaling language models hasn’t solved vision; what a “visual chain of thought” might look like; and how better visual reasoning could accelerate robotics, satellite analysis, product design, and mechanical engineering.


Sponsored by Dell Technologies and NVIDIA. Learn more at techrepublic.com/hubs/the-enterprise-guide-to-scalable-ai/.


Sponsored by Outshift: Visit https://outshift.cisco.com/?utm_campaign=fy26q3_outshift_ww_paid-media_ioc-neuronai-outshift_podcast&utm_channel=podcast&utm_source=podcast to learn more about the Internet of Cognition.


Subscribe to The Neuron for more conversations with the people building the future of AI.

More from The Neuron: AI Explained

All 106 episodes
Why Frontier AI Still Sees Like a Toddler, w/ Andrew DaiThe Neuron: AI Explained · 43 min
Listen in VO