In short
Eye On A.I. Podcast Episode Summary
Episode Title
#289 Eiso Kant: How Reinforcement Learning and Coding Could Unlock Human-Level AI
Podcast Description "Eye on A.I." is a biweekly podcast hosted by Craig S. Smith, featuring discussions with leaders in artificial intelligence who are shaping the future of technology. This episode focuses on the journey towards achieving human-level AI through reinforcement learning and coding.
Episode Overview In this episode, Craig interviews Eiso Kant, Co-Founder of Poolside, to explore the intersection of reinforcement learning and software development as a pathway to human-level AI. The conversation covers a range of topics from the evolution of AI in coding to the framework of Poolside’s AI systems.
---
Key Themes and Discussions
- The Missing Ingredient for Human-Level AI
- Eiso argues that while current AI capabilities (like code autocompletion) are impressive, true human-level intelligence requires more than just language modeling.
- The integration of reinforcement learning into AI systems is crucial for progressing toward advanced intelligence.
- Eiso Kant’s Journey
- Eiso’s background as a computer geek and programmer influenced his pursuit of AI in software development.
- He cites his early explorations into language modeling and reinforcement learning as pivotal moments that shaped his vision for AI.
- Software Development as a Training Ground for AI
- Software development is presented as an ideal environment for training AI due to its need for complex reasoning, long-term planning, and understanding of the digital world.
- Poolside focuses its models specifically on coding to enhance AI capabilities.
- Reinforcement Learning from Code Execution (RLCF)
- RLCF is a method where the AI learns from executing code, receiving feedback about its performance and making iterative improvements.
- This process allows the AI to tackle complex coding tasks and enhances its reasoning abilities over time.
- The Rise of Agentic AI
- Agentic AI systems are evolving to handle more complex tasks beyond simple coding, such as implementing features and debugging.
- These systems can operate in an environment, taking actions based on previous feedback.
- Challenges in Current AI Models
- Eiso highlights that while models are good at generating code, they often struggle with reasoning in coding tasks, particularly when handling multi-step processes.
- The need for continuous refinement and enhancement based on real-world feedback is emphasized.
- Making Software Creation Accessible
- Poolside aims to democratize software development, making it accessible to non-developers through user-friendly AI tools.
- The company envisions a future where anyone can create software with the help of advanced AI systems.
- Overcoming Model Limitations
- Addressing the limitations of current models is crucial for achieving the desired levels of performance and reliability in software development.
- Poolside’s approach is to develop comprehensive models that adapt and learn from real-world coding environments.
- Poolside’s Full-Stack Approach to AI Deployment
- Eiso discusses Poolside’s focus on delivering AI solutions for enterprise environments, emphasizing scalability and security.
- The goal is to enable organizations to utilize AI effectively while maintaining oversight and compliance.
- Future Outlook for AI and Software Development
- The conversation ends with a reflection on the future of AI in knowledge work and software development.
- Eiso stresses the importance of adapting to changes as AI capabilities evolve and the need for continuous learning and improvement.
---
Key Takeaways
- Reinforcement Learning is pivotal for moving from current AI capabilities to achieving human-level intelligence.
- Software Development serves as an effective training ground for AI because of its complex demands.
- Agentic AI systems are evolving to perform more advanced tasks, moving beyond simple coding.
- Continuous Improvement through feedback mechanisms is essential for enhancing AI reasoning and performance.
- Poolside's Mission is to make coding accessible to everyone while providing tailored solutions for enterprise clients.
---
Additional Resources
- For more information about Poolside and its offerings, visit [poolside.ai](https://poolside.ai).
- Follow Craig Smith on X: [craigss](https://x.com/craigss)
- Eye on A.I. on X: [EyeOn_AI](https://x.com/EyeOn_AI)
Conclusion This episode of the "Eye on A.I." podcast provides valuable insights into the state of AI in software development and the potential of reinforcement learning to unlock new capabilities. Eiso Kant’s expertise and vision shed light on the future of AI and its implications for developers and enterprises alike.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00If you want to get highly capable foundation models in software development, you can focus only on code. Software development is a representation of understanding the world. So our models, they're very capable general purpose models, by the way. They're quite good at writing poems and helping you with all other things that are on software development. But what we do is we focus all of our efforts earlier in the training towards software development. And we put huge amounts of emphasis in our training data and our approaches to training on reasoning. Where it is now moving towards is increasingly more agentic.
0:32And what this means is that the tasks are becoming more complex, higher abstracted, not just write a function, but try to implement a whole feature, try to figure out a complex bug. And here it also means that it's a multi-step process. Well, first of all, Craig, thank you for having me here today. And it is definitely worth going a little bit back in my background. I'm at my core computer geek. I've been programming for most of my life. And in 2015, I read an article by Andrei Karpathy. and it was titled the unreasonable effectiveness of recurrent neural nets and i have to say that article pretty much launched me down a rabbit hole uh of language modeling and now it's important to note that this is 2015 because in 2016 the next thing that really caught my attention was like for a lot of the world alpha go and when alpha go came out it thankfully was the unreasonable effectiveness of reinforcement learning and i think those two things that happen in pretty short succession really framed my thinking till this date so in 2016 i pivoted the company i was building to focus entirely on building ai that was capable of writing code this was using language models did you say 20 2015 to 2016 yeah before transformers yeah before transformers so this is at the time we were training models with lstms so the precursor to transformers and we were starting to do some of the world's first work around reinforcement learning via code execution back then we used to call it rl via compilers and and you can kind of see why those two things from those two origins kind of came together at that moment i was kind of lucky to get the right information in front of me at the right time and over the course of the following four years of working on this with a whole team i built up an incredibly strong conviction that language modeling this done went to transformers post lstms was going to be able to go all the way to human level capabilities but it was not going to be able to do so just by predicting the next token.
2:48It was going to have to go hand in hand with reinforcement learning. And along that way, by the end of 2017, I met now my co-founder, Jason. Jason was the CTO at GitHub. It was about two years before they got acquired by Microsoft, and he made an acquisition offer to buy that company. We had the world's first code completion models that were working, and he saw kind of a similar future that I believed in, was that the AI was going to approximate human level capabilities and software development, given enough time and resources. Turned down the offer, but became good friends and kind of never stopped talking.
3:24And at some point had a podcast together for years, which is just a great way for two grown men to meet with each other on a regular schedule. And, you know, fast forward, that company didn't succeed. We spent about four years on it and we were too early. we actually had quite a bit of technology working but it was actually challenging to get developers willing to connect their editor to the internet sounds crazy today but not that long ago that was the that was a big part of the challenge um and come you know life happened and then come around november 22 hits you know 2022 and chat gpt comes out and that point everything that would always believe them was going to happen was starting to happen and that kind of was to some extent the thing that looked like the biggest failure in my career i think i spent years working on with a whole research team like this strong conviction that didn't end up succeeding as a company and then now all of a sudden you know it's taking off in the world and at that point to both me and jason as we're very you know still speaking on a weekly basis it was very clear what the next 10 years were going to look like it was now just going to be a massive acceleration of closing the gap between models and human intelligence and even beyond.
4:45And so we started asking ourselves the question of like, what does it take to go into that race? And we looked around and we realized that the sentiment in the world was all we need to do to reach AGI is to scale up parameter size and more data and just do more language modeling. And that wasn't our view. Our view was, yes, skill massively matters. It has a direct relationship with intelligence and models and kind of empirically shown till date and continues to show. Our view was that it had to go hand and hand with reinforcement learning. And that's what we built Poolside around. That's why we started this company.
5:22And you built it as a code generation platform, not as a step towards AGI. So both. When we wrote something on our website two years ago on day one, and it's still there, poolside.ai slash vision, It's like, what's our view for the next years? And we laid out three steps. Step one, make AI capable to assist everyone in building software. Step two, allow anyone to build software with AI. Step three, generalize to all domains. And very clearly state that we're in the rate, this company exists to be in the race to AGI. And the reason we focused on software development is that we felt it was going to do two things for us.
6:01One, it was going to unlock the first area where AI was going to have massive economic impact. and a reason from the fact that we saw AI getting increasingly more capable in the domain, but also that software developers and people who build software have traditionally always been on the frontier of adopting technology early. So it was obvious to us that it was going to be the first place that was going to have big impact in terms of productivity and changing the way we work. The second part of it was that we wanted to kind of put blinders on. We knew that if we were going to say, hey, we're general purpose for everything, from writing poems to helping with medical knowledge to, you know, going down software development, we were going to spread ourselves too thin.
6:44And software development is a great proxy for intelligence. It requires understanding the world. It requires complex reasoning. It requires planning over long-term objectives. It requires interacting with the digital world in front of us. It requires visual understanding. Like it requires a lot of what makes up, not all, but a lot of what makes up, you know, valuable human intelligence. And so our view was by pushing the frontier there, we were going to naturally converge on the same point as everyone at SIAID, I think, in the future will, which is often referred to as AGI. Right. And the things like GitHub co-pilot that's based on open AI's models or some of the other.
7:30Most of the other code generation platforms are based on the general foundation models, I think. You guys built proprietary models from scratch, is that right? Trained solely on code? No, that second part is not correct. The first part is, you know, entering the race to AGI, we build foundation models from the ground up. But if you want to get highly capable foundation models in software development, you can't focus only on code. software development is a representation of understanding the world so our models they're very capable general purpose models by the way they're quite good at writing poems and helping you with all other things that are on software development but what we do is we focus all of our efforts earlier in the training towards software development and we put huge amounts of emphasis in our training data and our approaches to training on reasoning on the ability for models to take longer term objectives and successfully work through those both by reasoning through the information that's presented to them, but also by tool use, by being able to actually interact with their environment.
8:44And so, but because all of us live in a world with a fixed parameter budget, I can't serve a 10 trillion parameter model cost efficiently to users. This is not going to happen on today's hardware. So if you think about every model company being able to serve a maximum size model doesn't matter if it's a sparse MLE or dense but effectively a certain amount of sense that you can spend per inference request right because we all operate on effectively the same you know three different types of hardware that are there so we all have the same cost profile don't get me wrong it can be 20 or 30 percent more efficient here or there and deep seek did a great job and like but we're all on the same on the same budget and our view has always been is let's use that parameter budget far more towards software development by pushing it towards software development related capabilities earlier in its training by applying more training compute to software development related knowledge tasks etc and a lot of that has to do with our work in reinforcement learning yeah and you you use uh reinforcement learning from uh code execution right which is um i mean there's been a lot of work on rl aif um you know feedback from other models but this is specifically uh at uh running code that's been written to see whether or not it executes is that right yeah absolutely so if we our view has always has been is that in the pre-training stage of models, and I think this is generally held so, we're pushing these models on the most general task you could possibly have, right?
10:26Predicting the next token. But when we talk about skills like software development, they operate inside an environment, right? The code gets written, it gets executed, it gets tested, it needs to run in a full system. We as developers are modifying it. We're finding, we're getting feedback, from errors and bugs that we're introducing ourselves. And all of that feedback makes an iterative process to build software. Our view that experiential iterative process should be very kin to how models learn. So step one for us two years ago was building a code execution environment. That's grown a lot. By now, we are starting to approximate a million repositories, the real world code bases that are fully containerized with their test suite that you can make any changes in and have tens of millions of revisions as a total set and you can have you can define tasks synthetic but also humanly written tasks in there have the models explore the solution space and with rl it's kind of always the same thing right you have a task you have a number of samples that you're rolling out in terms of the possible trajectories to solving that task, and then you've got the rewards from when it successfully solves it or fills.
11:40And so by having this extremely large environment, our job becomes increasingly improving the quality of the tasks that we generate, most of which synthetic, increasingly improving the reward signal that we bring to the model. The obvious one is the unit tests are passing, but there's a lot more reward signal that can join on top of that. And then always making sure that the diversity of that environment is constantly growing so that you're both having more diverse tasks and more challenging tasks. And now as models are getting more capable, you know, with our latest generation of model, that is no longer just a single or multi-turn, you know, rollout.
12:20Now it's an agent that is entering into that environment and getting access to more tools and doing more complex things to learn. yeah and the initial training you have two models right two two main foundation models what's the difference between them and then the question i was going to ask is on training are you training them in the traditional traditional it's also new but uh you know feeding it uh huge amounts of data feeding the transformer algorithm huge amounts of data for it to encode knowledge in its weights or are you like designer stout and deep seek training directly with reinforcement learning so if you look at the different stages of training the foundation model kind of the it's the usual term the traditional pipeline right is we're pre-training on effectively the web right after you've done a lot of work in terms of optimizing the web rewriting it with models making it more coherent tagging it waiting at tons of experiments and then we have kind of the next token prediction stage uh and that happens throughout the majority of the of the pre-training budget along the way you'll you'll shift some of the distributions of your data sets differently towards the end of training and then you've kind of have the post-training step right and and when in post-training you find yourself with techniques that are a mix between supervised fine tuning.
13:52So like you said, providing a data sets with examples, uh, often there's a more conversational. So the model gets used to this conversational style of back and forth, you know, user assistant. And then you have the reinforcement learning component, which is, you know, uh, giving the model, these sets of tasks in these environments and rewarding it, you know, when it successfully completes them or successfully we have steps in the process that are correct. And then you might do some later SFT again. And that's kind of a standard pipeline, I would say, of 2025 model building for everybody. Now, there are places where all of us in the industry are innovating to push further than that.
14:34I would say where we have done and really put our focus on is how much further can we push to reinforcement learning in that. what's the limit of doing RL on verified rewards or code execution, deep exploration in what is the limits of reinforcement learning in non-verified rewards. This is the one part where we don't go into too much detail yet because we hold this a little more proprietary. And then it's what we internally often refer to as bread and butter model building. Effectively, what you're doing in model building is you're improving the effectiveness of your data and you're improving the effectiveness of your compute.
15:15And there is a ton of work that you are constantly doing. If that's your latest sweep over an improved data set, over architecture, over attention mechanisms. But it was important for us that we moved out of the world of artisanal model building, where every single one of these became its own project and effort by researchers, to having really a model factory. How will we be able to actually go very quickly from these ideas to results? How can you run a sweep with, you know, a thousand different variations of hyperparameters or different data setups? How can you make this deterministic so that your results from idea to experiment are like perfectly traceable?
15:57And by now, two years in, building models is more about working on the factory that builds models than it is about the latest artisanal idea that you're pulling components from to try to make work. Yeah. On the RLCEF, code execution feedback, is that a continual loop that the model writes code, executes, gets a result, feeds a result back to the model? to its training, writes code again, executes so it's getting better and better at writing code. How does that work? So that's close to where it started. So where it started was here's a repository. Here it is containerized. It means the code can be executed and tested.
17:01And it started with defining some very concrete tasks, like let's remove a function, hide the function from the model, give it the instruction, back translate that function into an instruction to try to write it, have it think of its thoughts and then its solution and do maybe 10 or 15 or 20 samples and then score the ones that are correct and negatively score the ones that fail. And there's different RL algorithms. The latest popular one is the GRPO where you group these rollouts so you take advantage of both the positive and the negative samples. And that's where for us RL started. where it is now moving towards is increasingly more agentic.
17:39So RLCF with agents in the loop. And what this means is that the tasks are becoming more complex, higher abstracted, not just write a function, but try to implement a whole feature, try to figure out a complex bug. And here it also means that it's a multi-step process. So the agent is in the container. It no longer just only edits code. It can run commands. it can search things it can open files it can read them it can store things it can execute different you know binaries so you can see it trying to pull a dependency and and create it from source again and try to install it so you you see that the agents are effectively becoming closer to having access to the same tools we have as developers right we don't just write code we make a lot more changes in the system and the rewards are increasingly more on successfully completing a longer range task.
18:35RL is a very finicky beast. And getting reinforcement learning stable means that you're always on the border of looking for tasks that are complex enough for the model to be able to learn from, but not so complex that it can never get any solution right. So there's nothing to learn. Not unlike us as humans, by the way. Yeah, yeah. That's interesting. Who else is using this kind of feedback from code execution? because it's the first time i've seen it but that doesn't mean that it isn't the standard among so when i started in the space in 2016 i think we were the first to ever even look at it i'm sure good ideas pop up in many places i'm sure there are others but it wasn't really a thing two years ago when we started poolside i don't think anyone was focused on rl almost at all not just on code execution feedback today it is clear that almost every reasoning model that we see out there, if that's the DeepSeq one you refer to, or from OpenAI or Anthropic, et cetera, we'll have a version of this as part of our training loop.
19:40What I would say is unique about us is the skill at the extremely large scale of our environment for code execution and the vast amount of diversity in terms of tasks and what the model can learn. and and now i would say increasingly more our work is is also going beyond that like what what does rl look like for places that are unverifiable so not just code and then software development because one of the points of view that we really hold is that knowledge is important to learn you need knowledge to understand the way you need representations of knowledge but there is something slightly more universal about successful reasoning and thought over a domain.
20:29And if you can push those capabilities in a domain like software development, you see that it's kind of like an onion. It kind of, or like a better analogy is a stone dropped in the water that has a ripple effect on everything else that sits nearby. So improving, you know, the capabilities in code, improve the capabilities in math, improve with the capabilities and reasoning in other domains, all as far as legal and others going further. And I don't think this is unique anymore. I think others have seen this as well. And so I think right now at the frontier, it's a race of scaling up reinforcement learning compute for everybody.
21:06You said at the beginning that you're building really for developers, but that the long-term goal is to build a system where anybody could write software, presumably with natural language. The current iteration of the product with this RLCEF, Reinforcement Learning with Code Execution Feedback, does that mean that it won't stop working or won't produce a result to the user until the code that it's written executes successfully? So if you think about this from a model and product working together perspective, there is something really interesting in how RLCF has evolved to right now. In the beginning, RLCF was something we did in our training part of code bases.
22:07And we had specific kind of training code for that. But now the agent that is going into the RL loop is the exact same agent that we're going to start shipping as product to users. And so, and by the way, we've already seen examples of this in the world. If you've seen these deep research products that are out there, they're effectively an agent trained for deep research in the RL loop. And then that agent gets, you know, put a UI on top of it so that it becomes available to users. And this is why it's increasingly becoming clear to us that for general purpose software development agents, the most capable agents are going to come out of the foundation of all companies.
22:47That doesn't mean that there won't be very capable specific agents built on top of models by many other people. But because the agent itself, the tools it has access to, the prompts that surround it, that is what's actually going in the training loop, that agent becomes increasingly more capable. For the end user, how you interface with AI for software development really has to do with the fact that AI is not yet at human level capabilities. So everything that we're building around these interfaces today are usually to deal with the fact that the models are still highly fallible. They still make quite a lot of mistakes.
23:23And that means that we are often, you know, building things around them so that you can review the code, that you can come back, that you can have a multi-term conversation, that you can tell it when it was wrong, that you give it. you know feedback from the environment that you're in since your test field how do you pass that along but we are increasingly on a trajectory where that level of you know lots of code and features built around the model are becoming like thinner and thinner yeah yeah because the i'm not a coder and i've been talking to people about this long before the initial gen ai applications hit the market.
24:05There was a guy in machine programming at Intel who used to talk about this. And at that time, it seemed like a distant dream. And then, you know, just a few years later, it's happening. But again, not being a coder, my problem in using any of the code generation tools is I can't spot a mistake. I can't read the code myself. And so it'll hit, it'll give me an output. I run it, get an error, send the error back to the model. The model says, oh, yes, of course. Well, this is what's wrong. Fixes it. I run the code, hit another error, and it's just this endless loop. It never ends. It seems to me, though, that that with RLCF, that could be overcome if you let the model work in a container, as you say, not, you know, giving the user the output until it's executed or not.
25:18Maybe it says, you know, I'm sorry, I can't get this run. You're heading in the right direction, Craig. Like an agent is essentially a model loop environment, after it's MLE, right? And the environment is where it's operating in and the model is what mainly matters and the rest is just the code around it for loop and the environments where you provide the tools to the model. And so what we're starting to see already is that there's increasingly larger tasks of longer duration with higher complexity that models are able to iterate on to try to successfully get to or give up at some point and say, hey, I can figure this out.
25:59And that's kind of the world that we're moving towards, right? We're moving towards where from a multi-turn chat experience back and forth, the world where, hey, I have this task. Can you go and do this? Step one off of being the model asking the right clarification questions, right? Ask the human is a tool in itself, right? For the model, like I can go and ask someone and then go off and try. And if the environment is well set up, like the container, as you described, that's where you're going to see a lot of work getting done. But it's important to note that we're not yet living in a world where models reach the same level of capabilities that we have.
26:36So for certain types of tasks, this absolutely can go successfully to completion 10 out of 10 times. But a lot of tasks is not fully there yet. And this is where this notion of kind of test time compute often gets referred to, giving the model more time to essentially run more inference to try to get to a solution. But it's kind of like tomorrow, if I give you a quantum physics problem, I'm assuming you're not a quantum physicist, Craig. You know, an hour or 500 hours of test time compute is probably not going to get you to solve it. And that's the same with models, right? They have real limitations and gaps still.
27:12They're just not always as obvious, obviously, boundaried as ours are. They often fail in surprising ways. Yeah. And you're using transformers. You mentioned LSDM. I had Sepp Hochreiter, I can never pronounce his name, on the program. And he's continuing his research on LSDMs. And for listeners, he was the guy that came up with long, short-term memory with that algorithm that was the standard for a long time. and he believes that you can widen the memory window and that there's still a lot of applications. And then there's this Mamba, which is layers of...
28:14Yeah, state-space style approach. Thank you, yeah, yeah. And then you have a transform that kind of sums up what's gone on and then another block of these state space layers. Are you using either of those or are you sticking with transformers? So it's a really good question. So we have coming back to the factory approach, right? So and specifically what you're talking about, like attention mechanisms and model architecture. they're an important detail and it's one that you spend time experimenting on but the way to think about them is that in many cases those details unlock either inference speed or training speed and and potentially you know that goes off in hand in hand with longer context windows right so there's this longer working memory of a model we've done quite a lot of work around RNN style attention.
29:17You mentioned Mamba, you have also RWKV, one of the co-authors works at Pulside. We've had quite a bit of success with RNN style attention. In the end, what you kind of see is that the different attention mechanisms in combination as a hybrid often kind of really good results. That's kind of what you were talking about with Mamba when there's, you know, the hybrids where you add transformer blocks and kind of global attention layers. These are interesting details, and I can geek out about them for hours, but I think it's also really important to realize that they are optimizations. They're not fundamental breakthroughs.
29:59We definitely, in our horizon, could still have fundamental breakthroughs on architecture that can have massive impact on either the memory of models, because right now we don't truly have. we either have this you know pre-trained like training embedded updating the weights and activations or just using the working memory if that's an rnn state or mamba style state or transformer it's it's all effectively you know this passes once the inference call is done and so we're those are things that continue to push but also at this point it does feel like We are at a moment where, we've been saying this for a little while, that we don't need a fundamental breakthrough in architecture to be able to get all the way to human level capabilities in coding or software development.
30:50It'll be massively helpful. And there are definitely architecture breakthroughs that if they happen could upend our entire industry. That all of a sudden doesn't mean it requires these billions of dollars anymore to train like models at the frontier. but right now like the unreasonable effectiveness of neural nets is really there you scale them up you scale up more compute if that's on either you know traditional next token prediction because you have the data and synthetic data generation or if that's on rl which i think is increasingly going to be a bigger part of compute budget of training a model so much so that i'm willing to make the prediction that it becomes the largest part of the compute budget of training a model in the following years that you can go all the way.
31:33But yes, those things matter because at the end of the day, the effectiveness of our compute often relates to the architecture. Yeah. You were saying that as smart as the models are, they're not at human level yet. They're certainly at human level with natural language. What is it about code that prevents them from getting to that level? So I think what we see is by having trained on most of the web, and now all of us have been rewriting the web into better and cleaner and synthetic forms of it, is that you definitely get an incredible amount of knowledge encoded. You get an incredible language understanding.
32:29But the web is an output product. It's the final article written. It's not the thought process that went to it. It's the final research paper. It's not every single step of thinking of an experiment that gets to that, right? It's Einstein's relativity theory, but not the hundreds or thousands of hours that he spent thinking it through and the things that didn't work and how. And I think this is really important because the process of the creation of work turns out to be a really important training data set that doesn't exist in enough data that NextTokenPrediction alone will get there. You've got the limit.
33:07If you had infinite data and infinite compute, you could get to AGI. NextTokenPrediction is an incredible strong optimization pressure in learning of the model. But since we don't have, you know, trillions of tokens of thought process of correct reasoning and math and encoding and medicine and on all of these areas, we see that, in my view, the reasoning component, thinking, I see reasoning as a subset of thinking, right? Reasoning is goal-oriented. It requires an outcome. While thinking can be much more broader. that is something that isn't well represented and because we don't have that we see that models can spectacularly fill on things that we consider very simple but my best my favorite example of this is if you look at how they do math so if i take something simple like a large number multiplied by another large number and i throw it into l any llm today the number that it will output will be wrong but it will only be off by maybe you know five percent left or right 3 385 times 9 80 or 802 it will actually be direction correct and the truth is if you gave me that number and required an instant response from me probably be worse but let's say you know hopefully i get directionally correct when we get models to reason about it in the way that we do which is we apply the little algorithm that we learned in school, you know, on how to do that math sum, it gets a correct output.
34:42That's the part that is underdeveloped in models. And reinforcement learning now offers the promise and is trying to offer the results that we can develop that complex reasoning. A software development in coding is just a really great task that requires a lot of complex reasoning and multi-step reasoning to do something correctly. And often because it operates in environments that are so much larger than our working memory, a huge code base, an entire system, that you want the model to be an agent. You want it to be able to interact with the feedback it gets when it makes a mistake because you can't reasonably expect it today to be perfect.
35:19And neither are we, right, by the way. So that's the part that is missing this. And this doesn't just hold true in coding, it holds true in almost every advanced knowledge work domain where you apply today. Yeah. And I mean, I was thinking that in the code is deterministic. It would be easier for a model to follow the steps and understand the reasoning behind the steps. But maybe that's overly naive. I wish it was that case. Now, I do think that because it is deterministic, it makes the training with reinforcement learning to improve that level of reasoning a lot easier than in non-deterministic domains.
Read the full transcript
36:14right you have you can do a lot to be able to say did this code compile was it correct did it pass unit tests there's another notion that i often refer to internally as as large language model arbitrage which is that models are better to reason about code what it is supposed to do what the successful inputs to outputs are than it is about writing it the other way around and by the way i guess so are we i can you know it's it's easier to actually observe something it and reason about what it's supposed to do than actually create that thing. And the combination of this model arbitrage with the fact that it has deterministic outcomes and it can be executed makes it the world's greatest target for RL.
36:57That's what got me excited about it in 2016. It's what got me to start this company to get over Chasen. And it's now, I think, what's leading a lot of the improvements that we're seeing in reasoning capabilities of models, not just in code across the board, because it turns out that reasoning is a representation that we're learning in models that is kind of, you know, touches every other representation of knowledge. Yeah. So where does the product stand now? And I have to ask also, I use Manus, you know, the Chinese multi-agent, autonomous multi-agent software that operates in a virtual website in the cloud so that it's not using your computer while it's reasoning.
37:52And, you know, with mixed results. But where does that stand in the universe of code generation? Is that just an engineering exercise that's interesting, or do you think that they're doing anything? And when you talk about agents in Poolside, is it a similar kind of architecture where it's operating in a virtual environment and returns a response without you having to take your tab up and everything? I haven't used Manus, but I have seen some of the demo videos. And I think it's a great showcase of where model capabilities are combined with a great environment. To the point earlier, model loop environment.
38:52And that environment is where you provide the tools, is where you allow the model to run in that loop and execute. If we look at where we are today, so we started before models were at agentic level capabilities, We started with this back and forth multi-turn kind of chat conversation that a developer has in the editor or on a web assistant with poolside. And then we moved towards that model presenting kind of a plan and code changes, so native edits that it was making in the files. Now we are increasingly moving to an agentic world where the agent itself is effectively a runtime that can be interfaced with from the editor, can be interfaced with from a CLI tool, can be interfaced from an API call.
39:39Right now, it's still very much about having it run in your local environment. Our work is definitely towards remote execution environments. How do we get these agents to run in remote execution environments? Because a lot of code that we develop and work on doesn't run on our laptop. It runs somewhere on a Kubernetes cluster, or runs somewhere in CI. And so I do think across the board, capable agents in software development will follow this pattern everywhere. This is not just an us thing. I think this will happen. Our job really at CoolSight is to do two things really well. One is to really focus on building the most capable models for software development and keep scaling up to push that frontier of what's possible.
40:25And then second is how do we bring all of that product experience around the model for two things. One, to allow others in the future to build on top of us. And second, how do we have our own view on user experience? Because the user experience is constantly evolving in the world as models are getting more capable. And bring that to end users. And here we've made a decision three years ago that we are still very, very glad that we did. It's to really focus on the enterprise. We want to bring poolside out over time to every developer in the world. But we looked at one of the most complex environments that have the largest number of developers working in it.
40:59And it's enterprises. right a u.s bank in new york and a 50 000 software developers and in those environments if you're looking at a view where in 12 months people are running thousands or tens of thousands maybe hundreds of thousands of agents how do you manage that chaos how do you orchestrate those agents how do you audit log them how do you monitor and observe them how do you allow people in your organization to elastically scale them up and down because our view is is that we're moving to a world where agents are effectively an elastic AI workforce.
41:34And how does that relate to PoolSys product in that you're moving towards AI orchestration or writing? We've been building the whole stack, right? So that's what we've been doing for two years. So think about it. The model, everything that serves the model, the API middle layer that exposes it out, the user interfaces. Today, we focus on VS Code, IntelliJ, Visual Studio, and then the CLI and the web, and then the API layer, which is critical how people interact. And so now what is adding to that is an admin and orchestration side of it. And I think this is, frankly, where you'll see everyone in the industry is already moving towards this.
42:19It's clear now that the gap between agents and models, they're effectively the same thing, just agents to wrap around the model is closing with our level of capability so we can entrust them for longer duration tasks, which means we want to be able to manage them, you know, more centrally and with more oversight and all of the, you know, relevant enterprise features that are needed for that. Yeah, but at the bottom layer, you're still talking about code generation. Is that right? And we talk exactly full software development because software development goes beyond code generation. It's about helping you build a product, you know, PRD document.
42:59It's about helping you monitor the logs in a system. Right. Our view is agents are going to be running everywhere. They're going to be working synchronously with developers, asynchronously for tasks to be sent off and also autonomously. They'll be running in your CI. They'll be running in your containers. They'll be observing your logs. I think we are still underestimating the surface area in the world of where agents are going to be running. Yeah. But again, on code generation, is the intent then that developers in a large enterprise would be using poolside to write code as a partner or assistant in the way that people are using GitHub Copilot?
43:44but absolutely absolutely and they already are yeah we're already deployed in enterprise today where where developers are working side by side with poolside uh moving now from a multi-turn chat experience to an increasingly more agenda experience right and and how do you evaluate your effectiveness i mean this is a big topic these days evaluation uh everyone trains to a benchmark and then comes out with the product says look ours beats theirs uh but but that doesn't necessarily uh transfer to the to the user experience how how are you evaluating poolside how do you sell poolside uh to replace copilot or it's a great question so from an we can do a whole episode on evals uh evals you know um are essentially the way that we break them down and i think most in their industry is benchmarks that you have that you hill climb you think they're really representative of the broad range of capabilities you're trying to improve in your model evaluations and benchmarks that you run that you don't want to hill climb but you want to understand if something is going wrong throughout your model training.
45:05These things shouldn't drop down all of a sudden. And then you have your famous vibe checks and red teaming. This is where it's really users, internal, external, paid, free, that are spending time with the versions of your model behind your product, and they're giving their feedback on, you know, what is it better at, what is it worse at. And a lot of things can't be caught initially in evals. effectively what you're learning from your vibe checks you're going to go build evaluation sets for it and then capturing your next generation and your next generation if that's you learn oh your model is stubborn okay how do you eval that so eval are a living set of things it's important that they range very broadly so this is everything from general reasoning to specific coding capabilities to you know personality and and sense of the consistency of the model and so you're constantly improving on that you have a dedicated team for it but you also have internally an evaluation framework that everyone is constantly contributing to they're never perfect in our space and if you would only hill climb a benchmark and you would make every decision around that you can end up with a model that seems great on paper but isn't great to use and we've seen some examples in our industry of that so it's an art that we're all trying to increase and you make more of a science.
46:27In terms of how do we sell poolside, a couple of things. So we focused on enterprises, places with more than 5 ,000 developers. We do this across finance, public sector and defense and kind of core strategic industries, so big industrials, big tech. We work very closely with Amazon Web Services. We're one of very few companies who have a first-party partnership with them. I mean, if an enterprise is looking to purchase Polside. It fully retires their commit spend that they have with Amazon. And that's something that I believe only four companies historically have had with them. Could be wrong, four or five or so.
47:06That means that we have a very strong joint go-to-market motion. But outside of AWS, we also are willing to bring Polside on-prem. This is something we do quite a bit in defense and in public sector, where we're increasingly more bringing it on a server in a data center or what we do in defense and government, a workstation that goes into a lab or a SCIF. But the notion is at the end of the day that because we are able to bring our models behind the firewall of customers, we also see an increasing future where the model weights change for the customer, where we can start bringing reinforcement learning to the customer side.
47:49There's still a lot of work to do there. We're early in that. But when you are going to be running tens of thousands of agents in the enterprise environment, you want that shared knowledge of all the trajectories that they take, every thought, every decision, every interaction with someone in your company, every action, code change, etc. You want that to be able to, over time, change the weights of the model so that the model becomes your company model. And that's something we feel very strongly about. We think that over the next couple of years, we're not going to just see a central model with a big context window.
48:21we'll see models that are increasingly going to adapt to the environment that they're working in. Yeah. I mean, when you say model or product, these are proprietary. You're not open source, are you? We're not open source, no. We think that it's really valuable that the world is building open source models, but from the capital investments that we're making right now, we just haven't seen that to be a strategy for us. Yeah. But when you install something on-prem, how do you protect your IP? I mean, this is... It's a really good question. We've never been too worried about the weights. At current model capabilities, you are constantly working on the next generation of your model.
49:08And so if I take the worst case scenario and risk manage it, if somebody leaks the weights of our model on a Torrent website, A lot of things would have to happen for that to happen. Inside an enterprise, which is quite regulated and pretty solid in terms of how people operate, someone would have to take that, take it out the building, upload it. But even if that happens, three, four, five months later, we're on to the next generation of model. In a year from now, that model is obsolete. And I think there's very fair concerns in the future when models reach certain levels of capabilities that you want to think differently about this.
49:42And there are technical solutions that you can take towards this. But right now, we've actually not taken that same level of paranoia about weights as maybe some others have in the industry. Yeah. Another question, and I'm coming up to an hour, so I don't want to keep you too long. But at the coding level, is there some metric about how often poolside returns a function that's executable and how often it fails? and are you working at the i know you were saying that you're working in a larger uh context so you're working on features and the whole software stack but ultimately it comes down to writing uh executable functions yeah so yeah oh across the stack so So in our own training, everything of this gets measured.
50:51And if we think at the customer side, we expose every metric possible that we can gather. We expose both in APIs, save so that they can import it in BI tools. And this is not just in places where code can't be executed because it's not an agent, because it's being suggested. We also actually save everything from do people accept the changes, do they reject them? I think it's really important for our customers to be able to have really a granular visibility into how AI is helping their teams, where it's failing and where it's successful, and be able to break that down across any dimension they like, if that's a programming language or a team or specific project.
51:29we're still in the realm where where ai you know makes tons of mistakes and you want to be able to actually very clearly observe where that sits because we're definitely living in a world and and will for some time where we're still the primary actors right we're instructing ai uh but yes it's critical and and this is frankly one of the things i think i've really enjoyed seeing our customers be very happy about because we're by providing that level of transparency they can make much more like informed decisions in terms of, you know, where they need to invest in adoption of AI, because it's already taking a part of the team and it's making the massive productive, but not yet everyone.
52:08And places where like, hey, this doesn't work well enough yet. Let's wait till the next generation of models. Is there anything I haven't asked about that you think listeners should know? I think you have touched upon this, but I think it's important to take a step back from all the noise in our market because there's a lot there's a new tool every month and then at every three months i mean it's risen to popularity uh but if you really take a step back and look at the next five years if you hold the same assumption that we hold that model capabilities will converge with human level capabilities and software development and frankly the vast majority of knowledge work that we do behind a laptop then that means a lot of things are going to fundamentally change in terms of how we work and how we structure.
52:55And whatever the latest noise in the market is, bring it back to, is this still relevant in that world? And then you'll find that certain things are highly relevant and are going to become increasingly so. And other things you probably won't hear about in two or three months or two or three years. And I think that's from the best advice that hopefully I can give, which is measure against, will this be relevant when AI reaches this level of capabilities. And it will be a good filter to kind of sift through everything that's happening right now. Yeah. Yeah, well, that's a good point to end on. Yeah, fascinating.
53:34And if people want to give you a spin, they go to poolside.ai, is it? Yeah, so right now, we're only available in enterprises that we work with. So we're not generally available. we want to make sure we get there but if you're working in a large enterprise do definitely reach out to our team from there and we're always happy to then to find a way to engage and it's definitely our goal to make sure Poleside becomes available to everybody outside of the enterprise as well
From the publisher
How do we get from today’s AI copilots to true human-level intelligence? In this episode of Eye on AI, Craig Smith sits down with Eiso Kant, Co-Founder of Poolside, to explore why reinforcement learning + software development might be the fastest path to human-level AI.
Eiso shares Poolside’s mission to build AI that doesn’t just autocomplete code — but learns like a real developer. You’ll hear how Poolside uses reinforcement learning from code execution (RLCF), why software development is the perfect training ground for intelligence, and how agentic AI systems are about to transform the way we build and ship software.
If you want to understand the future of AI, software engineering, and AGI, this conversation is packed with insights you won’t want to miss.
Stay Updated:
Craig Smith on X:https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI
(00:00) The Missing Ingredient for Human-Level AI
(01:02) Eiso Kant’s Journey
(05:30) Using Software Development to Reach AGI
(07:48) Why Coding Is the Perfect Training Ground for Intelligence
(10:11) Reinforcement Learning from Code Execution (RLCF) Explained
(13:14) How Poolside Builds and Trains Its Foundation Models
(17:35) The Rise of Agentic AI
(21:08) Making Software Creation Accessible to Everyone
(26:03) Overcoming Model Limitations
(32:08) Training Models to Think
(37:24) Building the Future of AI Agents
(42:11) Poolside’s Full-Stack Approach to AI Deployment
(46:28) Enterprise Partnerships, Security & Customization Behind the Firewall
(50:48) Giving Enterprises Transparency to Drive Adoption




