In short
Google AI: Release Notes - Episode Summary
Episode Title
Building Gemini's Coding Capabilities
Hosts and Guests
- Host: Logan Kilpatrick
- Guests:
- Connie Fan (Product Lead for Gemini's coding capabilities)
- Danny Tarlow (Research Lead for Gemini's coding capabilities)
Episode Overview In this episode, the hosts discuss the development of Gemini, one of the leading AI coding models. The conversation explores early goals, the concept of "vibe coding," strategies for handling large codebases, and the future of programming languages in the context of AI advancements.
Key Topics Discussed
- Early Coding Goals
- The development journey of Gemini and the aspirations that shaped its coding capabilities.
- Importance of moving beyond traditional competitive programming metrics to focus on real-world developer needs.
- Ingredients of a Great Coding Model
- Data Methodology: Emphasis on multi-file edits and making significant changes beyond simple code completions.
- Contextual Understanding: The model’s ability to navigate large repositories and understand the broader context within which code exists.
- The Rise of Vibe Coding
- Introduction to "vibe coding" where developers express their needs more abstractly, enabling the model to generate code based on general concepts rather than specific instructions.
- Potential for the model to empower both professional developers and those with less coding experience to build projects.
- Code as a Reasoning Tool
- Discussion on how code can serve as a universal problem-solving tool, going beyond traditional coding tasks.
- Exploration of the model's ability to solve more complex coding challenges and provide solutions in various contexts.
- Evaluation of Coding Models
- Importance of gathering feedback from internal and external developers to improve the model iteratively.
- Challenges in measuring performance across different programming languages and frameworks.
- Strategies for Large Codebases
- Development of methods to handle extensive codebases, allowing the model to effectively understand and manipulate large amounts of code.
- The agentic approach, where the model acts autonomously to handle coding tasks.
- Performance Across Programming Languages
- Assessment of how the model performs in different languages, highlighting areas of strength and weakness.
- Consideration of the implications of relying on popular languages like Python and JavaScript versus supporting a wider variety of programming languages.
- Future of Programming Languages
- Speculation on the emergence of new programming languages specifically designed for AI integration.
- Discussions on the potential reinforcement of existing languages due to their favorable positioning within the AI ecosystem.
Short-Term Improvements and Future Directions
- The conversation also touches on:
- Plans to enhance the tool-calling functionality within coding contexts.
- Continuous feedback integration to refine model capabilities and user experience.
- The need to balance between model performance and user interaction styles.
Conclusion The episode wraps up with a reflection on the collaborative effort within the Gemini team and beyond, emphasizing the importance of collective input from various stakeholders in achieving significant advancements in AI coding capabilities.
Additional Resources
- Watch the episode on YouTube: [Google AI: Release Notes - Building Gemini's Coding Capabilities](https://www.youtube.com/watch?v=jwbG_m-X-gE)
---
These notes provide a comprehensive overview of the discussions and insights shared in the episode, focusing on the development and future of AI coding models like Gemini.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Let's talk about all the stuff that got us to the point of having really great coding models. Everyone on the team is always so curious and genuinely just wants to make the model better. You can also just, you know, see so many possibilities as the next frontier of what's going to become possible. The external reaction has been, we have one of the best coding models, if not the best coding model in the world right now. People should have access to this incredible intelligence and be able to do more. Team Gemini, great note to end on.
0:53Hey everyone, welcome back to Release Notes. Today we're talking with Connie Fan, who's the product lead for Gemini's coding capabilities, and Danny Tarlow, who's the research lead for Gemini's coding capabilities. Thanks for being here and hanging out, guys. Thanks for having us. Yeah, I love that talk. I'm excited. So the sort of context for this conversation is we have great coding models. And sort of the thread of the conversation is let's talk about all the stuff that got us to the point of having really great coding models. We released the sort of Gemini 2.5 Pro variant late March, which folks were really excited about.
1:26Maybe you can both sort of just walk us through the sort of year leading up to the code sort of breakthrough moment for Gemini, if you will, and what it took to get there and sort of, yeah, what was happening behind the scenes to get us to the place where I think the external reaction has been, we have one of the best coding models, if not the best coding model in the world right now. What was the bit that changed so that we could make all the progress that we've made? Yeah, I think if you were to say there were three goals, it would have been competitive programming number one and folks like open ai did an amazing job of setting early evals like human eval around um trying to measure these models coding capabilities at all but but um someone who's really good at leak code isn't necessarily someone who's like the strongest teammate right so so that was a bit divergent from what um folks ultimately needed not to anthropomorphize the models too much, but competitive programming was one.
2:30LMSys really set the tone a lot as well. And that's not really what you're actually doing day to day. And then if there was a third that was more productive, it was still sort of limited in a code completion space, not everything that we know the models can now do and everything we still hope the models can do in the near future. So the first two being, to your point, not really reflective of what developers are actually doing. And the third one being not quite ambitious enough for them would be how to characterize the goals. Danny? Yeah, I think that's a good answer. I think maybe you could expand that out and just say, I think one of the things we looked at when we first started a year or so ago on this was just like, do we have the fundamentals right in what we're doing?
3:22From a code perspective or just generically? Just from a model building perspective. Okay. From a, there's many, many parts, both within people who are looking at code and people who are sort of tangentially working on code or people who are just building general model capabilities across Gemini. And really, any individual capability is about the coming together of all of these things. And so when I say fundamentals, I mean what Connie is saying in terms of, do we have the right focus and when we say we want to make the model better at coding does that mean the same thing to everybody working across all of these different things I think that's really important that you know in order to make sure we all are pushing in the same direction and pushing things in the right way but I also just mean in terms of you know the model itself and in terms of you know if we have an idea about hey you know we observe that the model has a shortcoming in this space and then we think we have good ideas about how to address that or which part of the process needs to be fixed in order to do that i think i would say if the fundamentals are not right then you go and you try to make this change but then you don't get the outcome that you want in terms of the model's behaviors and the reasons for that there's many many reasons for that and i think those are kind of like you know the very first things to be looking at uh understanding you know Why aren't we getting the performance that we want in this space and sort of tracing that back as a first step?
4:52Can we opine on the point for a second of like why the competitive programming bit doesn't generalize? Because my intuition and to use like the human analogy, like I feel like the folks who are like, you know, world class competitive programmers like tend to be like sought after as software engineers and folks in that position from like working at great companies and like going and doing engineering work there. I've never done competitive programming. I actually don't really have a good sense of what the tasks look like, but why does that not generalize well from a thing to hill climb from a model perspective?
5:23I don't want to be too negative about it. I think it is a very interesting, challenging problem that really stresses the model's coding capabilities. It's something important. It's something we do need our models to be able to do. I think it's just when you think about what it means for a model to be good at coding, one example is working inside of the context of a large repository. That's something that doesn't show up at all in a competitive programming environment. You're just working in a very self-contained thing. Here's your problem description, start from scratch, build me a relatively short solution.
6:00Compare that to what a software developer is doing on a a day-to-day basis is like, there's this bug report coming in about some crash that I don't know. It could be spread across, you know, a hundred different locations in the code base, and I need to go and figure out. So it's just the sort of the set of capabilities that the model needs is much larger than just what's encapsulated in competitive programming. All right. So we talked a little bit about sort of the history of the Gemini coding effort, but would love to know, like from today's perspective, like what are the things we're focus on to like, you know, the, the main ingredients of what makes a great coding model.
6:37And we don't have to go into all the technical detail, but like from a North star, is it like we're focused on like completion tasks, but then also like entire code-based tasks or like how does all of that break down? Yeah. I mean, I think it breaks down always to data methodology. Um, And so they're shifting to meet each other accordingly. But on the data side, I think to Danny's point that repo context is so important. That's probably the biggest shift that you'll see. Is it just completion anymore? No. I think we specifically care about making multi-file edits, bigger changes than just here are umpteen lines for you.
7:31Here is what you really wanted to do if you had an hour at a time to sit down and do it yourself in the context of the code base that you're working out of. Yeah, that makes sense. Yeah, I like to take inspiration from just where people are using our models and where they're really getting joy and value out of that. And then try to think carefully about it and say, you know, where is this going to be in three months? Where is this going to be in six months? And sort of like try to be ahead of the curve in that. And so, you know, I think there's a lot of really interesting stuff happening in the vibe coding space of, you know, people are doing really cool stuff already with the current generation of models, but you can also just, you know, see so many possibilities as the next frontier of what's going to become possible there and thinking about that.
8:23And then also thinking about the, you know, the more professional developer and like Connie is saying about the repo context and all of the considerations that come up, not just in the code editing part of the developer lifecycle, but, you know, where do people really spend their time and struggle and where are the difficult challenges that come up in the entire software development process? I think that's the hardest part of our job. Maybe it's not the hardest part of Danny's job, because he does all the smart research stuff too. But like getting the timing right and like making the right bets is constantly a question for us and um you know we were to your point that 2.5 pro is being really well received we were looking back at the like hopes and dreams we'd written down a year ago six months ago and like I was patting ourselves on the back anyway and you're like you know what yeah yeah that that was good but I think we're going through a cycle right now where we're like, oh man, we really hope we're making the right bets and we'll find out in six months.
9:27I feel like the intuition of building things to resemble the developer workflows, just to hit on this point again, it is really interesting. I have two questions related to this. One, why did we not do that to begin with? And was it just like, if you look at the previous era of coding models, very completion heavy, Was it just that the models were like not good enough and like the next token prediction paradigm just like stopped working as you tried to do like larger edits? Or was it like an architectural thing that the models were like, all the previous code models were like very diff slash completion focused?
10:05I think my take is there just wasn't enough mindshare on it. I think it's something that, you know, I think there had been a lot of work in various pockets of the research world and whatnot on these kinds of problems. But I think they weren't getting the same level of mainstream attention that they are now, maybe because the model capabilities weren't strong enough. And so, you know, crossing that threshold of I really trust this to be more involved in more parts of the software development process might not have been there. But I think if we had a time machine and we'd go back in time for years, then yeah, I think we could accelerate the attention that was placed on these.
10:49yeah i i i mean dan even knows in which part of an eyelid but it feels like the models were capable if you put the effort in but you need to suspend disbelief that someone would even like want to just come with a couple natural language sentences and walk away with a web app like yeah that wasn't clear i think to folks that that was going to be such a dominant use case and i think Like for us now, it's like this idea of vibe coding, like it doesn't even have to stop at web apps. Like there's so many things as long as it's possible for the person to validate the final artifact and feel confidence in the output that can be vibed.
11:34So, yeah, I think we just want to keep keep pushing along that dimension. Yeah. One other random comment, which is it's also just interesting to see the disconnect between like how like people who are learning to code and this like this thread of the like whole repo task and like being in a large code base. Like if you think about like what you're taught when you learn to code often, it's like not like you're not in university and they're not dropping you into like a mono repo and being like, okay, now go solve this bugs. It's like all these like very contrived discrete tasks, which actually, interestingly, I feel like map to some of the early model capabilities.
12:12And then as soon as you graduate school, if you go through a traditional computer science degree program, your most first jobs, you're going to be in some large code base that you're not building something zero to one. And it's very interesting that the model development process actually kind of tracks that human learning process of early code stuff. But I'm also to go back to the point that you were both making about vibe coding stuff um how has that become something that we're thinking about from a code perspective is it like materially different than you know the north star from six months ago pre-andre carpathy tweeting about vibe coding um was that something that we were already thinking about yeah i mean i think at least in In my mind, I've always had this division of the kinds of users and audiences that we care about and that we want to make sure that we're building the right capabilities for.
13:11One of them is the more sophisticated professional developer working in the enterprise kind of environment like you're talking about. And then the other is this idea of what we're seeing with VIVE coding, but it's been around for longer than that, of the tools that we're building are going to help people who are not professional programmers or who have a little bit of programming experience, you know, expanding out what they're going to be able to do with programming. I think we've seen that as an important slice of what we're trying to do for a while. I think, you know, the recent expansion and excitement is, you know, maybe sort of teaching us some new things or giving us some new inspiration of what you might be able to do in that space, though.
13:58nice totally yeah i mean like maybe we didn't have a pithy thing to call it so thank you andre for that um but i think like the greatest hopes and dreams are always to uh take this skill set that is like very concentrated in this geography we're currently sitting in um bay area and like just empower people who don't have it to do basic things for themselves like we'll always need professional developers and that's not by any means a substitute but um people should have access to this incredible intelligence and be able to do more and and i think we've always cared about that yeah as if uh andre needed yet another hit uh but yeah he he got the vibe coding it.
14:49I'm interested to talk about sort of the connection between how we think about the coding pillar within Gemini's role in bringing new capabilities to other areas inside of the Gemini ecosystem, if you will. And maybe, I don't have a good sense of this. I'm curious from both of you, instruction following as an example, does the, I don't even know off the top of my head if we have like an IF pillar or something like that. But like, are they looking to like, you all from a code perspective to like help upstream and like hill climb on the general capability? Or like, from our perspective, is it like, you know, the thing we care about is code and these like very discrete specific tasks.
15:32And like, that will generally hopefully translate to these like other capabilities as well. But like, it's not, and you know, maybe insert multimodal as another example, like, is it how connected are all these different capabilities from like a model perspective? Yeah, I mean, I think it's extremely interconnected is the short answer in terms of, you know, I think like, you know, we're talking about the model capabilities getting better at code. This is undeniable of what's happened. But, you know, like how did that happen and which pieces of work across all of Gemini helped to make that happen is definitely a sort of a team effort spread across the entire set of sort of both capability focused and more kind of like horizontal focused efforts within the project.
16:28And so to the extent though of like how much does improvement in coding contribute back to other parts of capabilities, I think there's also like many ways and many degrees to which that happens. I think in some ways, you know, there are some capabilities where, you know, maybe you have a problem and you weren't thinking about representing it as a coding problem, or like, you know, you want to solve a word problem to help somebody with homework. And maybe it's useful to convert that first into code that says, you know, here I understand what, you know, who has how many apples and what gets multiplied together.
17:07And then you can actually, you know, maybe execute some code in order to do that, or reasoning in code space is a natural way to solve problems that are outside of the coding sphere. I think that's a very interesting way. How much does that happen? Because just to opine on this point for a second, it feels like there's a, and I don't know what, I don't know if we have analysis of like what percentage of queries, like could you, user questions or prompts, could you actually just like write a bunch of code under the hood to like solve whatever the underlying problem they have is? But have we thought about like I feel like today it's like somewhat reserved where like if you're like oh do this thing the model will it's probably not going to generate code unless you like really ask it to but like do we from like a north star perspective is there a world where we just like are generating code for every query to like build you some bespoke software to like solve whatever the thing is that you're interested in it's hard for me to imagine that for you know write me a poem about something something Like a polar generator app?
18:05I don't know. That's interesting. Yeah. Yeah. I mean, there's obvious exceptions, but I do feel like code is everything in the context of LLMs. And a couple examples that I like to hark on with Danny, so sorry to hark on them again. Like, I think if you look at our query stream today, people recently in the US would have asked a lot of like give me tips on how to do taxes like help me with this general concept like make sense of this like purely natural language based concept for me but at the core of it if you were to take their holistic problem um and you were to write them a bare bones tax calculation um it'd be hard to even call it software but just solve their problem for them Like that is, I think, totally doable via generating a lot of code to your point of maybe the prompt didn't specifically ask for like, write me a mini QuickBooks or something like that, you know?
19:09But that's the spirit of what they're really asking for when you look across all their prompts. I think that's a future we can probably orient around. and we should build that i feel like this would be a fun experiment to like what happens if you just like force enable the model to like always build you software to answer your query and like what is how do users feel about that i feel it could be like really really like and i feel like that seems somewhat outside the state space of like what developers would be doing with this like i feel like it's not a developer workflow it's like an end user workflow but i feel like that'd still be super interesting totally um totally and and like i think even beyond that of moving beyond like an individual person's needs and an individual developer's needs.
19:55My favorite pseudo eval recently that I liked to hark on with Danny was, I saw this post online of like someone who's not really into AI being like, if AI is so smart, why can't it bring the cost of eggs down? And I was like, why can't it bring the cost of eggs down? You know, like, theoretically, this is a lot of publicly accessible, but hard to reach data that you can write, just write some code under the hood for, to your point, and solve a problem that could meaningfully help a lot of people. That's not a world we're in yet. And that's not the capabilities we have. But it sounds like a really tough email.
20:36It's like, it's like, how much new Gemini model, how much did you lower their price of eggs how much did you help people writ large how much did you just help society writ large that's the eval that is interesting to think about um yeah i feel like that's that's like a great we could have a whole conversation about how hard the eval story is and i also think like i'm curious if there's any threads like maybe this is the thread on this around just like how um like future north stars for coding from an eval perspective like does today's set of evals like capture how we think models are going to develop or like i've seen some of the other ones in the ecosystem where i don't fully grok what they're doing but it's like how much economic value do you create as a developer doing these tasks like how much money could you make in these contexts or like an agentic system like how do we think about that from a eval perspective yeah i mean the way i like to think about it is we want to understand and predict and make bets about where the real world value is going to be.
21:38We want to make sure we're not kind of like over optimizing for that. I think one of the things that we want to do with how we steer the model capabilities is really make sure we're tackling you know the hardest most crucial bits of the capabilities around this thing so we don't want to just like you know spin off some quick wins because if we got really good at x then we could make a quick buck i think it's we want to stay focused on you know what are the core fundamental challenges in this space that are useful for sort of the real world value and then let's you know try to to push the model's capabilities as directly as possible in those directions and so then like what is the eval that you need to go alongside that um i think you know there's always a range of of put something out there into an A-B test, see what happens in the real world.
22:33That's always going to be the most reflective representative thing. But then we always try to make different trade-offs in our evaluations of things that are easier to measure. You don't have to ship your model to a new startup and have them run for a year building on top of it and see how successful they are. That's not really practical. And so, you know, we try to be pragmatic about what proxies we can find. Yeah. How difficult of a problem is that? And I ask this through the lens of I feel like every coding surface now as sort of like the agentic harnesses and like what they're doing specifically changes or like maybe they don't have any agentic harnesses or whatever it is, is sort of using code models in a different way.
23:17And I think even if you look at the spectrum of like Google products of like how folks use some of our internal coding stuff, how folks use like Colab as an example or Project IDX or Project Firebase Studio, like there's a huge spectrum of different of those different coding use cases. How like I feel like there's a lot of capabilities where like maybe they're not being used that like chat, chat is chat. And like, you know, there's not all these like different interesting form factors of how people are doing that, at least right now. How much is that like a challenge from making the models better at coding so that it generalizes across all the different ways that people are using code models when I feel like there's actually seems like a pretty reasonable distribution?
23:59I think that's the main challenge of what we do is, you know, we are trying to build capabilities that work for all of the use cases simultaneously. And so we need to figure out, you know, what are the things that we can do that will generally help the model's capabilities across all of these different sort of like real world use cases. And that I would say is really the difficult part in all of this. So tying to like this broader thread about like all the breadth of use cases and the breadth of feedback that we're getting from a coding perspective, I think one of the coolest things about Google is we have like literally 100 ,000 plus engineers across like every, you know, language and skill set and different stack, et cetera, et cetera.
24:50How does that help from a code quality standpoint? Yeah. Historically, we haven't leaned on this enough, but we are starting to, and it's amazing. um a hundred thousand of the smartest most brilliant sometimes most opinionated engineers in the world and and to danny's point that we are still lacking evals and we don't want to pigeonhole ourselves into uh over rotating around any one thing to have this broad exposure for sort of like a known gap in coding evals right now of nuanced tastes of professional developers um getting that live feedback is incredible and there's gradients of it with the baby tests but even just going back to the idea of vibe evals like um a vibe eval that i want to set i we have a couple of vibe evals but like um jeff dean truly one of the best engineers to have ever engineered um and he's so nice about everything but what's the day we get a gemini output in front of him And he's not just like, oh, good.
25:58But like, wow, that was incredible. And that made me more effective today. Or you had Emma on the podcast a little while back. And when he's like opinionated, Emma's opinionated. But like, let's like, like, let's pretend he's on call and something's on fire and it's going terribly. And he like needs to like turn it around. When is he going to like actually trust that Gemini isn't going to mess it up? And like in this critical moment where he's feeling the stress, like he would actually prefer to have Gemini in the loop than not to have it in the loop. Like these individuals who really represent like, OK, this is a new tier of capability for the model.
26:38It's incredible to be able to interact with them directly. And of course, so many others as well. And then the whole broad base of Google engineers is it's just so lucky that we can we can talk to them. And we're all one company. Yeah. And Denny, how do we balance that with external developer reception or external developer feedback? Do we see a big divide between what sort of internal Googlers tell us about coding models versus the external developer ecosystem? Or is it pretty similar? I think it's pretty similar. I think there's obviously different segments of users. So you know, like the developers that Connie is talking about here, we probably don't have high bandwidth communication channels to external specific use cases there.
27:26And so you can get sort of different kinds of richer bandwidth signals just directly by talking to somebody and saying hey try this or try that yeah um but in terms of what are the things that matter in terms of the model's capabilities for these use cases i don't think there's anything pulling in different directions here i think it's um you know kind of clear the the set of capabilities that you need you know and there's different for different segments of the population but for the you know analogous external versus internal developer, I think they're pretty similar in terms of what they're looking for.
28:04Yeah. To pull on this thread around the people who are using coding models, how do you think about the AI skeptics of people who aren't that interested in coding models or love writing code by hand? And I feel like, again, through this lens that developers are opinionated for you know in the most positive sense um is it like a you know to win the next you know and number of developers is it like we need to make the models just like better in general or is it like there's a different product experience you think we need to build in order to like showcase to developers that this stuff really works well and um is going to be valuable for them i mean we're very model oriented people here right so at least for me it's um I think it's really motivating to sort of like look at this, to take exactly what Connie said before of, you know, what is it that they're looking for and what is it if they are a skeptic that's holding them back?
29:01And then I think we've seen a few cases of converts recently as well. And then, you know, to understand, okay, this is what changed about the model and sort of, you know, it got to this level of nuance and depth in terms of its understanding at that point it won over the trust for you know this particular slice of use cases I think it's a great sort of like hill climbing metric in some sense on its own is just how do we take the more and more challenging the more and more skeptical people and what is it that we need the model to be able to do in order to win them over I think we enjoy the existence of skeptics they give us something to orient around and um yeah we need to make the model better but not just broadly better like they're likely skeptics because the model is amazing at the things that they genuinely need and care about and so we it's it's our job to figure out what that is and how to do it and and for most of those folks does it tend to be or just like generally for classes of areas where the model doesn't perform well does it like is it um pocketed into like different use cases or is it like different languages like the models really there's tons of python code in the world the model's really good at python c plus plus is harder because it's there's less code or something like that like how does like do from a programming language and like framework perspective is there yeah what what how do you all think about that i think there will be languages like cobalt that's it's just really hard for us to get data we need cobalt data because i I feel like that's the whole, the like running bit of like the cobalt engineer, because there's, you know, a hundred thousand banks or whatever that are all still running cobalt or something.
30:44Yeah, the white whale. But for the most part, like the languages that you mentioned, I don't think we have a dearth of data, but I think we have not always historically optimized our mixture properly. and back to the idea of figure out what's actually needed right now, what people actually care about, what will they actually use the model for, work backwards from there. That's sort of the name of the game. Yeah, I think my experience with some of these people is everybody has a different kind of challenge in their mind of like, okay, if the model can do this, then I will be impressed. But it seems like everybody has their own.
31:24some are you know very basic things of like I want you to do this very routine thing in my workflow but I want you to do it perfectly other people are like I think Connie's example of everything is on fire am I gonna turn to you in desperation like you know that's another another might be you know just the complexity and sort of how accurate and coherent the sort of the reasoning and the solution is in a fairly complex problem. And so I think, I don't know, my experience has been with when you really talk to these people and you say, okay, what is it that like test this out? We think the model has gotten much better.
32:03Do you believe it? It's interesting to see the different ways that different people go. And it's also hard to predict. And that's another reason for us to sort of keep pushing in generality, right? I don't think we can optimize for these seven people and then Apple also convince, you know, the next seven people as well. I think we really just need to be pushing the sort of the overall model capabilities in a way such that we just get these like bigger and bigger jumps. Yeah. One thing that like I just want to give a shout out to our entire code team about is this like shared mentality that I just felt as Danny was speaking of like if there's negative feedback, I think no one on the team is ever defensive.
32:41Everyone on the team is always so curious and genuinely just wants to make the model better. and sometimes I feel defensive on the team's behalf because I see how hard everyone works and how brilliant these people are. I'm like, what do you mean you're still not happy? But then the researchers themselves get the feedback and it's never anything but like, okay, how do I make this better? And it's just like the most amazing atmosphere to be in and the most incredible team. I love that. I feel like the focus on helping make developers happy is I need to come hang out with you all more and spend time with the team.
Read the full transcript
33:16Just a meta comment about this thread of just like different ecosystems and different languages that the model is good at. Something that I've thought about is, you know, as the model writes more and more code, and I think like Cursor put out something about like a billion lines of code or something every day that's being written just through Cursor. And as the model is like, I feel like generally pretty good at like concentrated in certain languages and frameworks. Like what's your both meta guess as far as like, does Python and JavaScript or TypeScript become like the last programming languages ever?
33:52Because like they just run away with it because the models become so good at writing Python and TypeScript. And like, you know, it's just this like reinforcement flywheel effect that all the code in the world is Python and JavaScript. And then it makes it really hard to like bootstrap future languages, perhaps. Have you thought about that or? Yeah, I mean, I've heard this argument before and especially the sort of, I guess, the angst of like a new programming language developer that you're starting from behind because you don't get the AI assistance sort of working as well. and as people continue to as people rely on this more I wouldn't discount though the sort of like a how good models might be able to be by just giving it the specification of a programming language in context and then as the models get stronger you would think that they would be able to get more and more skilled with the sort of taking in more complex longer descriptions of how a new language works.
34:55And then B, you know, we're in such a sort of like a flux period of how people are using programming languages and what they're doing with it. I wouldn't be surprised to see a new programming language emerge and say, you know, for the AI age, this is actually the new thing. And it's somewhere that comes out of left field. I don't think that's impossible. That's so interesting. As someone who's done, who's spent a bunch of time on programming language stuff i feel like it is uh it is kind of daunting to think about like bootstrapping a new community and making the mod make now making the models better at that thing but i do think this would actually be a really interesting eval is if you like we make our own internal programming language derived from you know choose whatever you know stack you want um and you give the model that context like with each new coding model we release like just dropping all the information about that model in context, like, is it good?
35:50Can it like follow the paradigm of that language that's like totally not represented in the training data? I feel like that'd actually be an interesting, like how well is the coding capability generalizing versus, yeah, versus it just getting good at Python and TypeScript or something like that. I don't know. You all have to make it now. So this is a taking feature request session. This is feature request session. Awesome. So one of the threads that we've talked about so far is just like how capabilities from the rest of the Gemini world sort of translate towards code capabilities. And it feels like long context has been like a really, really interesting thread with the most recent models.
36:33And Danny, I'm curious, like your take on sort of the combination of coding capabilities and long context and like the whole, you know, drop a million line repo into the context window of the model and have it go and do a bunch of stuff versus, yeah, some like smaller, more focused code edits and stuff like that. Do you have a perspective on the use case and what works best? Yeah, I mean, I think it's quite interesting. all that's happening there. I think in my mind, I organize it in terms of kind of like problem versus solution strategy. And so the problem that we're trying to solve here is we want people to be able to work with more and more complex code bases.
37:17For example, if you take this version of long context and you might have a case where you can't fit a code repo into 32 ,000 tokens or you can't fit it in a million tokens or you can't fit it in 10 million tokens. but still, you know, we want it to feel from the perspective of interacting with the model, like it has this capability. When you think about solution strategies for that, then indeed, there's the one which is, let's make our long context capabilities really, really good. Let's throw everything we can in the code base in context and then solve things in kind of like a single step you know here's my code base I need a new feature implemented please I'll put me the edits to make that happen the other thing that's happening though sort of in a very promising and very interesting way is all of this agentic coding work and the kind of like agents interacting with users in different ways or being more autonomous and you know if you think about sort of like, you know, you can take the anthropomorphic view of like, think of the model as a person and how does a person solve this?
38:29Then, you know, if I dump you into a large code base, then you're going to use code search. You're going to look at the file hierarchy. You're going to sort of jump around. You're going to read some code, then you're going to search for other things. This is kind of the solution strategy that the agentic approaches are taking. So I think both of them are interesting and you know maybe they're not so different you can sort of mix and match the two of them together so i think it'll be interesting to see how these two things play out i think the the interesting thing about the dump it all in context and then you know the model is just magical and its capabilities i think that's something that also when it works it also feels uh i think very surprising as a user you know there's more of that kind of like magic of wow i just overloaded you with anything that i can imagine and it really snaps you out of this like thinking of the model as a a person or that and i like that aspect of it as well of it's you know it lets us think about sort of the capabilities of model independent of how we might think about a person uh going about solving that problem yeah it's very interesting because it's um it has this magic feel but then it also has this just like very impractical feel like i feel like As you keep scaling up, you're never going to be able to take Google's monorepo and put it in.
39:47It would just be so expensive to do one inference pass of however many, many, many millions of lines of code are in the monorepo to solve some bug. I feel like the agentic approach has to be what ends up scaling. But I would also be curious to see how much that diverges. like does the model find ways of solving the problems agentically that like the sort of human developers aren't doing like it like weird you know you see this in like rl examples where the model finds like weird arbitrage approaches in chess or any of these like other game strategies like do you think we'll see that in code where the model's just like doing all types of like wacky non-human developer things but it actually like solves the problem well yeah i don't see why not I think one important thing about this space though is at least in these in the current state of the world in the professional developer context the communication back to the user is important as well so it's you know yeah let's let it do crazy sort of like new strategies in terms of how it goes and figures out where the bug is and what the problem is but when it comes time to like you know make a change and and fix it it needs to do it in a way that is interpretable and understandable.
41:05I'm thinking about all the bad edge cases of this, which is the model's like, oh, good strategy. Delete all the code and then you can't have bugs, which is like the quintessential example of like the wrong approach to solving this problem. But this has been awesome. I'm curious about sort of what is next. Like we have a great coding model. We've made it better. People love it. Like what else can we do? Like, is it just keep hill climbing on a bunch of benchmarks or like new capabilities from a coding standpoint or like how are we thinking about the sort of North Star in the future direction for Gemini's coding capabilities?
41:41It feels really nice to be sitting here right now in a moment where like the model is surprisingly good. Like I see the posts online of like is Gemini making a comeback? Is Google making a comeback? And it's so heartwarming to see and also feel like we know we have good stuff in the pipeline and we know that it's there's still good stuff coming and I hope this ages well um and but we also know that like the competition is so fierce and no one that anyone can surprise us with something amazing but but we are confident in in what we have going on um yeah it's setting up the right benchmarks and hill climbing them like i that that greatly diminishes i think when i say it that way what our team is doing but um i i think creating the right benchmarks for ourselves internally and what we really believe in um making sure the team is set up to to hill climb those that's that's just the only way they get there so we spent a bunch of time talking about sort of the long-term future and also the history of everything we've done.
42:53What is sort of the short-term improvement, you know, line of sight to like making the model better at X, Y, and Z things from y 'all's perspective? And specifically, like I know we've gotten, there's always tons of feedback coming in about Gemini models, the 03251, lots of feedback, the 0506 preview model, lots of feedback like what does the short-term uh improvement horizon look like yeah i mean i think going back to the sort of code capabilities come from all over gemini in terms of making this happen i think one thing we heard very clearly from the first 2.5 pro release was there were some issues around the tool calling functionality particularly like in a code-based context and and you want the sort of agentic model to be making code edits for you, these kinds of things, we weren't getting the reliability and the behaviors that we want.
43:50So there's been a bunch of work from people outside code areas as well across Gemini to try to improve this. And I think we do see quite a nice improvement in the most recent release. But we see a number of continued improvements in that direction. I think, you know, once you start really looking carefully at the model's behaviors, you start seeing things of like, that's not quite the smoothest user interaction in the terms of the way that that worked. Or, you know, I wish the model would have been better at this category of use cases here, of how people are asking about and trying to use it. And so I think really fine tuning that capability and making sure that we sort of continue the trajectory that we had in the last two releases would be great here.
44:42How do we think about the models like style? Because I know one of the big the May release of the model, lots of like sort of visual improvements and how it does web UI and stuff like that. Is there a sense of like, you know, this is almost like model personality and we need to like maintain sort of the taste that the model has? Or like, is that is that a sort of capability you think we can stay sort of a consistent taste profile? Or like, will that change from like a model to model basis as far as like what, you know, a good website looks like as one example? Yeah. I think catering to high taste people is something that we really care about.
45:24and it differs a little bit for us. So for the web development example, I think even today, oftentimes you get a model generated UI and you're just so impressed that an alum did this. And that's amazing. You still feel a little bit of a magic moment, but you're still kind of deep down. You're like, an alum did this. Like this wasn't like a beautiful, well-designed website by professional standards. And so around that specifically, there are very targeted methodologies for improving the visual layouts and and huge huge shout out to um jeremiah will on our team others who've been really pushing this capability um but i think the style means so much to so many people and even just in interacting with the model while you're coding.
46:20Someone raised a fun example to me where the model had already gotten it wrong twice and that's obviously not the ideal behavior. You just want the model to get it right. But because Gemini started being a little cheeky, being like third time's the charm or like oh this is really tricky. Let's try this again. Like I saw some of these folks. Yeah, like it makes somehow makes you like forgive Gemini for like not being the best code model right out the gate and just working with you and so I think that style as well like tone personality it's not necessarily something that you think of as critical to coding but it is in the way people embrace these models and I think we start to see different preferences between professional developers and someone who's just starting to learn like no one wants to be made to feel like they're stupid ever but especially not when you're just getting started and you're just trying your best so So yeah, and style, long way to answer your question.
47:12We certainly care about style. It manifests for us in methodology in different ways, if it's visual, if it's text, but we care a lot about it. I love that. I want my requests, and I don't think this can happen at the model level. This is more of a product thing, but I feel like I always want like five options on the visual side. It's just like, I feel like it's the taste bit is like, I want the model to exercise taste, but then I'd like my taste to sort of choose which of the distribution, the taste distribution actually gets built. That's on you. Put that in AI Studio. We can put it in AI Studio.
47:45I feel like that'd be fun. Very token expensive. Any sort of aha moments from a code perspective? I'm actually curious when... I remember seeing a bunch of the chats when the earlier 2.5 Pro model was in the works and seeing some of the results. and it did seem super exciting, but like, what was that, what was that moment like? Was that the expectation, like models done training, you start playing around with it and like everyone is sitting there being blown away that this thing is really great at coding? Or like, what did, what was that experience or what did that look like? I think for me, one of the first examples where I thought it was really, really cool is one of the things I like to do is just in the vibe coding world, I'd say, you know, I'm going to sit down, I'm going to make a game from scratch and I'm going to iterate for, I don't know, 30 minutes or an hour and just kind of see how it feels to try to put more and more burden on the model and see what it's capable of.
48:44One of the examples that I was looking at is kind of like a platformer game, but it had done something where it had put a platform that wasn't reachable by the character. So it was like impossible to finish a level. and so I asked it you know please go through all five different levels figure out which platforms are reachable and not and then make an adjustment to you know move things around so that it is playable and the character can do it and so you know this was one of the earlier experiences with the pro the 2.5 pro thinking and sort of seeing it go through and say okay I'm going to identify all of the platforms across all of the levels.
49:24I'm going to do the little calculations in terms of, you know, this is how high the character jumps and this is the sort of calculation of if it's reachable. And then just to see it kind of, you know, do the thinking for about a minute, go through all of these cases and then successfully make the edit and then, you know, go back and run it again. And it was sort of, it had made the change that I asked of it. I think that was the case where it was like, okay, like this is the, you know, the capability, the sort of like the thinking and the vibe coding parts of things coming together, the coherence with the model doing exactly what you wanted it to do, it doing something tedious and sort of like a little bit of math mixed in.
50:03I thought for me, that was like a really nice moment of like, wow, you know, this is cool. This is when you paint Connie and you're like, we're shipping. A low-key actor. He was like, here, this amazing thing happened. I was like, whoa. But yeah, to the point of shipping, I think like, did everyone know 2.5 Pro was going to be an amazing model? I think after the pre-trained base came out, everyone was so excited. But after the first post-train checkpoint, I don't think it was obvious to me that, like, we were going to hit all the code goals that we'd wanted. And I think it took a couple go-arounds before it was like, oh, my gosh, this is going to happen.
50:42And I think when people internally were like, consensus, oh, my gosh, this is happening. Truly, like, that's like, all right, let's ship it. But it was a bit of a surprise, I think. I don't think it was so obvious that like sequentially, okay, we completed this training step. Now we have confidence. And then we completed this one. It's like, yes, we know this is going to be good. I think it might have crept up on everyone a little bit just how good it was. I love that. What like initial AI coding experiences and like how both of you sort of experienced that to begin with? Like, was it co-pilot?
51:15Was it something like another tool? Like, what did you have a moment where you were like, this actually feels possible and like we should dedicate, you know, research direction and like time of life into making that possible for both of you? Yeah, I mean, I guess I started this early in the sense of when I was doing my PhD, I was really focused on things like structured data, more complex data than just like image classification or just like sequences of text. and I was interested in the complexity and the outputs that our machine learning models are generating. And as a grad student, you would write your introduction paragraph in a paper, and you would say, one day we will try to tackle modeling computer programs as an output that exemplifies the complexity in what we're trying to generate and model with our models.
52:14And so I I think it was 2013 is when I went and I did a postdoc and they gave me an intern. And so this was my first kind of free attempt to define a project and fully scope it into it with somebody else. And that's when I started working on building generative models of source code. Did it work early on? Like, was it, you know, for folks who see today's coding tools and like, you know, you're getting entire fully fledged things that work actually most of the time out of the box now with our, with our most recent models. Like, what did it actually look like for those early ones? Was it just like very small text completion or was it still doing like whole program generation?
52:58Yeah. I mean, it, maybe it was a bit more uneven. And it was certainly way worse. And the kind of, you know, problems that people were trying to tackle is, like, one of the first things we were interested in is this notion of you define variables and then you use the same variable throughout the course of a program. That's just like a basic rule of programming languages. If you would have given this to a, you know, recurrent neural network of the time, the sort of the precursor to the transformers, it wouldn't really know that concept of just like the basic rules. You could build more structured models where you're sort of bringing more human knowledge into the design of the architecture and whatnot.
53:37And then you could build some models which you generate samples which look more realistic and we capture some of these things. But it was definitely light years away from where we are now in terms of sort of how meaningful and how realistic things you could generate. Yeah, I love that. I haven't got more questions on this one, but Connie, I'm curious for you. what was the like moment where you're like I want to do code stuff have you always been working on this from uh inside of DeepMind no um yeah in my time in DeepMind I've always been working on code um before that I didn't work on code but I'd always wanted to work at DeepMind that was always my dream job as and and I think when opportunities finally came up um I also just wanted to do something that was useful to people and I honestly wasn't sure at the time like in a decade from now would all the LLM stuff matter um or would it kind of just wash out and be like that was weird um but but code was the one vertical that felt like undeniably useful and like it was going to matter and it was already mattering why what was your connection was it just like that some early like product stuff was working in that sound show or what did you feel that way yeah I think it was like the the first vertical that people were both getting just like sheer enjoyment from of like those like personal aha magic moments as hobbyists but also actually adopting it at work already and breaking out of like the standard chat modality and really integrating it into their workflows like that just felt so promising and like to your original question of when did I personally feel it And I was honestly running on blind belief and fumes when I joined this team of like, this is going to be so amazing someday.
55:21But like, I didn't feel my personal magic, like, oh my gosh, I can absolutely code now with all these tools. And like, I can just express whatever I wanted English moment until honestly using cursor and not our models. but like uh yeah maybe almost a year ago now i finally had that moment but i'd already i'd already taken this job and was just hoping that um well not just hoping like genuinely believing that this is all going to work out someday but took it before i felt like an aha moment personally yeah i love that i feel like so many people um code was the aha moment for them seeing that like llms like would be a useful product creation in the world and would be a useful you know product vertical to create value.
56:03And it is, I think there's so many use cases for people who like use AI, but it's like not actually that useful. And I feel like code has always been the one case in my mind that's like clearly beneficial for the world and useful. And there's like lots of value being created. So it's, it's awesome to hear. So we've been talking about this whole thread of like the Gemini capabilities and how code benefits across, you know, we make other things better code gets better vice versa um why not historic historically prior to the you know large foundation model paradigm there's like lots of really domain specific coding models and in a lot of ways like code feels different than a bunch of other paradigms why not have a code specific model and i feel like it's it's become somewhat obvious as we've talked about these use cases like it there's a lot of ways it doesn't make sense but like are there places where it does make sense maybe really narrow product specific tasks like i do think you see eg cursor is very open about training their own completion model and it's amazing that they've done that and it's very purpose-built for one narrow slice of things but to danny's platform or example um where you need to have some concept of physics and some concept of math or my favorite vibe check is build me a Taylor Swift ranker app.
57:24Like you just need a sense of world knowledge that isn't just code and it's behind everything that people want to code. So no, for the broad swath of tasks, I don't think they benefit from just a code model. Yeah. I mean, I think it's, you know, it's hard to rule out anything ever or whatnot, but But this question of what does that even mean, like, I think Connie's kind of saying this as well, but like, you know, if we think about where things are going is, you know, code means more and more parts of the software development process. It means connecting to all sorts of different pieces of information as a part of the coding process, some of which are code specific things, some of which are not.
58:12um if if what you're kind of implying with a code specialist is we're going to sort of like up code capabilities here and then we're going to down capabilities somewhere else it's a little bit unclear where you want to down weight those capabilities and so I think the just thinking about it as look all of this is interconnected we're all working together on sort of like a general set of capabilities of the model and we're gonna you know try to find ways to make everything play together in the best possible way and sort of have that that like really nice generalist model I think feels to me like a great way to be going right now yeah and I feel like that proof's in the pudding like we made a great coding model it's also good at a bunch of other things and like we continue to be able to hill climb so I'm happy that's the approach uh that we're taking this has been a sort of a very, very fun conversation.
59:10I need to hang out more with the coding team. Thank you both for all the hard work. And thanks to the rest of the code team for all the hard work to make great models that developers love. Thanks for having us. And thank you for also just being always on the forefront of collecting all the feedback out there, interacting with folks. You are a part of the code team. You are part. That's the easy part. You all do the hard work. I just get to send you people complaining or bringing the models. I got one more penguin. I'm kidding. Thank you. And then, yeah, just a final shout out to sort of everybody outside of code and the Gemini team as well.
59:45Like it really is, you know, a lot of different things coming together that produce the capabilities that we have. So it's cool that we're sort of all kind of moving kind of in lockstep in a way across the entire development in a way that it's like, yeah, we get all the great capabilities across the board and code gets a lot better. is definitely a sort of a collective force. Yeah, this was a ton of fun. Thank you. Thank you both for having this conversation. I think it's always fun to chat coding stuff and hopefully we'll have a ton more progress to come back and talk about soon.
From the publisher
Connie Fan, Product Lead for Gemini's coding capabilities, and Danny Tarlow, Research Lead for Gemini's coding capabilities, join host Logan Kilpatrick for an in-depth discussion on how the team built one of the world's leading AI coding models. Learn more about the early goals that shaped Gemini's approach to code, the rise of 'vibe coding' and its impact on development, strategies for tackling large codebases with long context and agents, and the future of programming languages in the age of AI.
Watch on YouTube: https://www.youtube.com/watch?v=jwbG_m-X-gE
Chapters:
0:00 - Intro
1:10 - Defining Early Coding Goals
6:23 - Ingredients of a Great Coding Model
9:28 - Adapting to Developer Workflows
11:40 - The Rise of Vibe Coding
14:43 - Code as a Reasoning Tool
17:20 - Code as a Universal Solver
20:47 - Evaluating Coding Models
24:30 - Leveraging Internal Googler Feedback
26:52 - Winning Over AI Skeptics
28:04 - Performance Across Programming Languages
33:05 - The Future of Programming Languages
36:16 - Strategies for Large Codebases
41:06 - Hill Climbing New Benchmarks
42:46 - Short-Term Improvements
44:42 - Model Style and Taste
47:43 - 2.5 Pro’s Breakthrough
51:06 - Early AI Coding Experiences
56:19 - Specialist vs. Generalist Models
