#193 Itamar Friedman: How CodiumAI is Making Bug-Free Code a Reality

13 Jun 2024 · 54 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye On A.I. Podcast Episode Notes

Episode Title

#193 Itamar Friedman: How CodiumAI is Making Bug-Free Code a Reality

Podcast Overview Host: Craig S. Smith Description: Eye on A.I. is a biweekly podcast that discusses significant advancements in artificial intelligence and considers their global implications.

Episode Summary In this episode, Craig Smith interviews Itamar Friedman, co-founder and CEO of CodiumAI, a company dedicated to developing intelligent coding systems. The focus is on achieving bug-free code through innovative tools and methodologies.

Key Concepts & Tools Introduced

  • CodiumAI's Vision: Aiming for zero bugs and issues in software development.
  • Tools Developed:
  • Codiumate: A tool for individual developers that helps in planning, writing, and analyzing code.
  • PR Agent: A tool designed for team collaboration that assists in code review and issue detection just before merging code.
  • AlphaCodium: A research project that outperforms DeepMind’s AlphaCode using fewer large language model (LLM) calls without fine-tuning.

Discussion Highlights

  • Flow Engineering: A novel concept introduced by CodiumAI which facilitates more effective software development by structuring the coding process.
  • Future of Coding:
  • Discussion on how AI can automate coding tasks and ensure code quality.
  • The potential for developers to eventually code using natural language, particularly for simple tasks.
  • Emphasis on the importance of a verification platform to ensure code generated by AI meets specifications.

Insights from Itamar Friedman

  • Development Background:
  • Experience in chip and system verification, which influences his approach to software verification.
  • Previous roles include founding two startups, with one acquired by Alibaba Group.
  • On AI and Coding:
  • AI's role in enhancing the coding process is more complex than just generating code. It requires verification and integration into existing workflows.
  • The vision is to create a holistic platform that supports various aspects of software development.

Challenges with AI in Coding

  • Traditional LLMs struggle with hallucinations and errors that can lead to larger coding issues.
  • Itamar suggests that a system that incorporates human-like decision-making processes can mitigate these challenges.

Product Integration and Future Enhancements

  • CodiumAI is developing a multi-agent platform with various tools tailored for different stages of the coding lifecycle.
  • Plans for additional tools and enhancements to improve code integrity and testing processes.

Final Thoughts

  • For Developers: Embrace AI tools to improve productivity, but remain critical of overly simplistic representations of AI capabilities.
  • For Enterprises: Identify specific pain points in software development and seek tailored solutions rather than adopting technology for its own sake.

Call to Action Listeners are encouraged to explore the tools developed by CodiumAI to enhance their coding practices and stay ahead in the evolving landscape of AI in software development.

Additional Information

  • Sponsor: NetSuite by Oracle offers a flexible financing program for enterprises looking to streamline their financial management processes. Visit [netsuite.com/EYEONAI](https://netsuite.com/EYEONAI) for more details.

---

For further information and updates, follow Craig Smith and Eye on A.I. on their respective Twitter accounts:

  • Craig Smith: [@craigss](https://twitter.com/craigss)
  • Eye on A.I.: [@EyeOn_AI](https://twitter.com/EyeOn_AI)

---

This concludes the detailed notes from episode #193 of Eye On A.I. on CodiumAI and the future of coding with AI.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00CodingMate is meant to help planning a task, writing code that is high quality, according to high integrity, fitting to the task. And then most importantly, like analyzing the code that was written or right now or a week ago or years ago and provide additional tests for it. While the peer agent is meant for last, you know, review just before you merge your code, helps you to analyze, describe, review, find issues at the point of the merge. So these are two tools. But like I mentioned, we're building a full platform and more tools are coming. There's going to be an exciting announcement a week from now and another one a week after.

0:42So we're going to be a multi-agentic system. That's Codium AI. Hi, I'm Craig Smith, and this is Eye on AI. In this episode, I speak with Itmar Friedman, co-founder and CEO of Codium AI. Codium AI is building a multi-agent platform to help developers achieve the ambitious goal of zero bugs and issues in their code. Itmar shares his vision for the future of coding, where AI assists developers by automating tasks, catching errors, and enhancing the software development process. We discuss Codium AI's product, Codium Mate, and how it leverages generative AI to improve code quality. ITMAR also provides insights into the company's research, such as their Alpha Codium project and the concept of flow engineering.

1:38I hope you find the conversation as useful as I did. By popular demand, NetSuite has extended its one-of-a-kind flexible financing program for a few more weeks. Head to netsuite.com slash ionai. That's ionai, E-Y-E-O-N-A-I, all run together. Again, head to netsuite.com slash ionai, netsuite.com slash ionai, where it's one of a kind flexible financing program. So first of all, great to be here, Greg. It's a pleasure to have the opportunity to talk about an intelligent coding system of the future. Specifically, my name is Itamar Friedman, the CEO of CodeUmi. And at CodeUmi, we focus on helping developers reach zero bugs.

2:32Zero bugs, it's a bit ambitious target, but that's what we aim to do. I think it's one of the hardest problems that's being tackled these days with regard to software development. Actually, if I have to add two more words to our vision of zero bugs, I wouldn't add them to the left, but to the right. What we're trying to do is zero bugs and zero issues, zero bugs and issues. And for that, that's what we do. And in order to do that, we don't think it's going to be like one tool or one interface that's going to solve it. We think it's a complete code integrity platform. We're building a multi-agent platform for code integrity.

3:14We call it integrity because we want to encapsulate testing and issues and best practices. And that's what we do already today. First, we introduced two tools. The first one was introduced a year ago with already 500 ,000 installation and ID and JetBrains. Actually, I saw that we crossed the 600 ,000 just recently. And the second tool we released like half a year later, somewhere July, August 23, is what we call the peer agent, which is a Git app for GitHub, Bitbucket, you name it, GitLab, etc. The first tool is meant for developers, single developers, to download from the ID JetBrains VS Code marketplaces.

4:02And the second tool is more for teams that they can download via Docker or just download from our open source and connect or directly from the GitHub marketplace, by the way, and connect it to their Git. And that's more for the teams. They have different functionalities. One is meant to help with planning a task. The first one in the ID that we call CodingMate is meant to help planning a task, writing code that is in high with high quality, according to high integrity, fitting to the task. And then most importantly, like analyzing the code that was written or right now or a week ago or years ago and provide additional tests for it.

4:49while the peer agent is meant for last review just before you merge your code. It helps you to analyze, describe, review, find issues at the point of the merge. So these are two tools, but like I mentioned, we're building a full platform and more tools are coming. There's going to be an exciting announcement a week from now and another one a week after. So we're going to be a multi-agentic system. That's Coding AI. Oh, that's amazing, actually. This coding generation space has really fascinated me because I'm not a coder, first of all. And I'm hoping the day will come when I'll be able to speak to type into a platform and it'll write code accordingly.

5:39And it feels like that's coming. What was your background before Codium? mate or Kodium? Yeah, sure. I'll have to share about myself, but I have to say that I think that Kodium is my life mission, of course, in addition to the family. And so I usually identify with that. So my background, I actually started, not my first job, but one of my first was actually in chip and system verification. And when you're like in Mellanox, one of the most successful Israeli companies eventually sold to NVIDIA, a big portion of NVIDIA today. So what I did there is basically verifying that when you take a piece of hardware, it bids in and bids out according to the spec.

6:28And interestingly, once because the specification is so well defined, then you can even do apply techniques such as formal verification. You can almost 100 % mathematically prove that a certain module, a hardware module, is working as expected. That was my early first steps in my career. I also did a bachelor and master degree in machine learning and optimization, etc. And that kind of led me to think, how can we apply some techniques? You know, the hardware people have a lot to learn from the software agility etc but that's what people say but also the software people have a lot to learn from the hardware in my opinion for example testing verification etc so so i was thinking can we apply like into software what we did the hardware and the answer is like it's complicated because as opposed to bits and bits out now you have human in human out like the the entrance point is like user stories you know like this is how people develop software and then there's interfaces.

7:40So that's like my career started by doing that and then moving to machine learning and doing software and thinking about this process. Eventually, I founded two companies. I was two CTOs and both of them VC-backed startups. And then one of them, the last one, was acquired by Alibaba Group and I joined Alibaba Cloud mostly and we built amazing stuff, including, by the way, significant failures, but we also had a couple of huge successes. One of them was building an O2ML solution, fine-tuning machine learning models, LLMs, vision models, etc. That was during 2017 to 2021. We realized that AI, and before that, by the way, I was embarrassed to call AI, I always call it machine learning, but finally AI is arriving.

8:36And now finally we can start trying to apply testing more holistic approaches for testing verification due to software. And that's like after a few years Alibaba, I left and started Codium AI with my partner and also like a meaningful team from my past. And that was like BC before ChatGPT, but just before that, like it was August 2022. Yeah, things happen so fast in this space. So first of all, before we talk about the products, on code generation generally, I mean, so much has happened with generative AI. and I presume that you're using generative AI. And you mentioned the user story that it starts with, software starts with telling a user story.

9:35Do you think it's possible that the day will come when users, obviously highly trained users, or developers who understand software structure and the precision of coding language can code in natural language. And one of the reasons I just got a pitch from a company that's developed, and it sounds like they're not the only one, a new language that's a further extraction, closer to natural language, that nonetheless is precise enough that it can then code or write code. I mean, where do you think we are on that journey, and do you think it's solvable? Short answer, it's going to happen, but not now.

10:35and now not next year and or at least not you know it it wasn't solved yet despite all the videos that we see on x or tiktok or instagram whatever okay so now the longer answer so i think like first of all we can we need to define we need to like separate between the upworks like stuff, upward-like tasks that are separate, independent tasks that you, for example, build me a small website or change my page from here to there to the convoluted, complicated enterprise software. And it's totally different scenarios. It's different in how the tech stack looks, what's the software development life cycle etc and even what tools are will be used and what kind of like uh ai is going to be used so if i go first to the like like professional consumer uh approach like uh for example freelancers uh etc then i think like we saw a few videos that then also being debunked, but despite being debunked, I think in three to five years, I think those companies, and there are more than one, maybe even more than ten, are working on this angle of enabling you to do upwork style of work.

12:15Then I think in three to five years, you would be able to monitor as a professional developer, you would be able to monitor like AI developer that is doing a lot of work for you and all you need to do is babysit and here and there fix it. For some tasks, it could be working well enough that even a technical product manager or technical person that is not necessarily a developer might get a good result and communicate with that agent. It's coming. So that's, I want to say, if you saw the videos, it's not working right now, but I do see a path that it will work in three to five years. Now, for enterprise, it's going to be a totally different solution, a totally different approach.

13:10And software enterprise, it's not just about providing you a piece of code. It's like going and approving the entire process that software needs to go in order to be moved to production and enterprise. Let me elaborate on that. But first, let's imagine the world. Choose a Fortune 100 company. You name it. Think about one. Imagine one. Now, imagine a technical product manager sitting in that company and trying to tell an AI about a new feature that he or she wants to add. and that AI generates that code. And then do you imagine click a button and it's moved to production? Like just like that? I cannot.

13:52I think we're missing one very important part of the platform in order to enable that. It's a platform that all it's meant to do is to verify the correctness of that AI, that the code that was generated by the AI that matches the specification that was required by the technical product vendor as well as saying that nothing else was broken. So the world I imagine is three to five years in enterprise. Like, yes, a technical product manager writing a specification with one or even five tools depends on each one of them might be specializing. And then a platform that's like a dashboard, imagine, check, check, check, check, check, all of this specification or check all the other critical mission, check, all the things, okay, we're ready to go.

14:39and that's how I do imagine it happens three to five years where in some cases either the code generation part or the code integrity part will say it doesn't work, send it to a developer. Okay, and then a developer comes into, so developers are not going to disappear, it's just going to have more challenging tasks than before because the simple ones or 50 % of them could be automated with technical product managers doing their work. So that's how I see the future evolving. And that's, of course, aligned with our vision and mission at Coding. We are the solution, the code integrity solution. We are actually the leader there.

15:28And we're planning to enable any company. We want to be in a ubiquitous tool that any company can utilize. with any software stack. We can integrate to any CI, any Git platform, any ID, any code spec like Jira or whatever. And then we want to enable this future we talked about. If we would be able to be ubiquitous for any enterprise, then they can use whatever code generator they want and get that confidence that they can release code that's generated by AI. I just wrote a piece for Forbes, I think, in which I mentioned Codium, Codium 8. In that piece, I talked to a company that automates unit testing for Java.

16:23And that solution sounded the best among the co-testing generation platforms that I was investigating. And it uses reinforcement learning. as does AlphaCode, the code generation research project by DeepMind. I don't think they've productized that yet. First of all, on the one that I sort of focused on in that piece, it's called DIFFBlue. What's your view of that? Is that a solution for unit test generation? They only do Java. that could be applied to other languages? And is there anything in their approach that's similar to what you're doing for test generation? Yeah, so first of all, I didn't try their tool rigorously, but I do like what they're doing from what I see.

17:34And basically, I think there are a couple or a few companies in the unit testing area where we also provide solutions. and I like it because I think it's one of the testing in general is one of the most painful parts let me open the brackets there, I think the world's divided into two either you don't do enough testing in code review for code review for example you have 20-40 minutes quickly review, you don't do enough of that because you hate it you like writing code, you just want to write your code, or that's one part of the world, or you're in the other part of the world that you do spend a lot of time in testing and coding, and you still hate it.

18:16So the fact that people hate this part, although they know it's critical, whether it's unit testing, integration testing, etc., that show how important it is. So I actually like it. There are a few companies there. In one sense, it makes sense to me to do user reinforcement learning. I think they use it for trying to understand how they can increase code coverage. At the same time, I think these companies that started prior to 2020. Some of these technologies are kind of outdated or, to say more nicely, not exploiting the latest breakthroughs in AI. And for example, let me give you a different solution called TestLM by Facebook.

19:01It's, again, a research paper that was not public, like open source, was not publicized, was not made public, similar to AlphaCode by DeepMind, was not made public. And what they did there is kind of a similar solution. They went over pieces of code and took tests that already exist and enhanced them automatically. And I think that their solution uses the most up-to-date AI, much better, much more friendly, faster, et cetera. And it happens to be that one of the tools that we're going to release is actually very similar to that approach of Facebook. And that leads me to two things I want to say, I think, about the field.

19:54First, I think it's a big problem until today that you have so many tools and companies doing a small part of what we call code integrity or code testing, but it doesn't work. You need a holistic solution. It's not a perfect equivalent, but before Datadog, you had mini tools like CloudWatch by AWS that give you a portion of the cloud observability. But you needed a holistic tool like Datadog that has sensors all around the place that gives you the full image. And I think testing or code integrity also requires such holistic. You can't just have just for one language or one piece of the testing, which is unit test.

20:40You need a holistic solution. Usually, most companies have few languages just in one company, and they have more than unit test requirement in order to actually solve bugs, etc. So that's what we're aiming at. That's why I told you we're building a multi-agent system. Two agents are out. Two more are coming soon. It's going to be quite a few. Each one of them has different flows. We already have more than 25 flows, 100 in a year. Now, the second thing I want to mention is that DeepMind and Meta are definitely like a huge AI powerhouse. So I would definitely follow them with their alpha code and test LLM.

21:20Having said that, we kind of want to build real products. so while DeepMind released a research work that you cannot reproduce we released AlphaCodium which is our take on AlphaCode to reach better results and open sourced and got a lot of buzz we got Kripati talking about it but most importantly it's open source that you can reproduce and try it and similar to TestLLN by Meta I think by the time that this podcast will be released, it will be out, and we will release an open source of it. So if you like to be in the meta-level engineering where they have such a system, you would be able to use our open source for that.

22:06So for us, we take this advanced research and make it into tools that actually can be used. So back to your point, I definitely agree that these kind of companies like Gitlu, they have very advanced, mature solutions, but we need a more holistic that has many flows in order to give us the confidence that one day we can generate AI and really believe that we can publish it to production because another tool did a rigorous work through different processes and not just unit testing. And that, I think, could be an extremely important property of an intelligent coding system of the future. Can you talk a little bit about AlphaCodium?

22:51We spoke previously, and as I recall, you were saying that it beat AlphaCode or at least had better results in, I think, public. Okay, sure. So first, let's define the context. Yeah, I learned too much from the LLMs. Let's lay the ground on what's the context. So similar, there is a website you can play chess against humans or computers, etc. So also a website platform enables you to participate in code contests. It's called Code Forces. And DeepMind have defined a benchmark that is mostly being extracted from that platform. It's called Code Contest Benchmark. and then they tried to automate a machine that solved these problems of the code contest.

23:50This is very deep-minded to do. They're amazing. We obviously admire them. I don't know if it started, but one of their very early works that were published was AlphaZero and AlphaGo So, you know, copying and doing better than the world leader, the world champion. And then later, the Alpha Fold and Alpha Zero, etc. I probably missed a few or confused about it. They did quite a few, and then they did Alpha Code. And I think that's a game changer because AlphaGo is like a bit of a closed system. Once you have Alpha Code, it means that AI can generate code. In the most perfect situation, it can generate additional code that improves itself, etc.

24:43So I'm just giving why I think it's a different ballgame. So I totally see why it's a deep mind-like project. So it's not surprising they did Alpha Code 1 and 2. so basically AlphaCode was doing quite well, AlphaCode 1 was doing quite well, but AlphaCode 2 published on December 23, even did better than the majority of developers, but they have a problem there in AlphaCode 1 and also AlphaCode 2. The problem is that first they extensively run LLMs in order to solve do LLM inference in order to solve one problem in alpha code one it could be like a million times uh for one problem uh inferences and the second thing is they they fine-tune their models for for a specific contest and i think this is a problem because it's show how not generic enough in my opinion their solution is uh so this is the background we uh built on top of their some of their ideas that's why we wanted to give them credit we call it alpha codium and we did uh uh we did better than them if you put a point on on a on a graph in a sense that with less lm calls uh and with no fine tuning no fine tuning we do uh better than them and that's i think a huge achievement because it means that probably it's it's generalizable you can you can take it to different problems with now fine tuning probably could work.

26:18Probably, I mean, because we didn't benchmark on many things, but we are integrating a concept into our tools and we see it work. So how did we do that? To simplify two main things, one, something we called flow engineering, like from prompt engineering, moving towards flow engineering, and I can elaborate. And the second thing, the fact that we integrated into these flows, like code testing and code reviewing methods and aspects that we have a meaningful know-how there. With that flow engineering concept and testing and code reviewing, our system AlphaCodium did better than a majority of developers with a click of a button.

27:05You know where it's going towards. That's our AlphaCodium number one. Let me know if to double-click on flow engineering or other stuff. So Alpha Codium is a research project or is it integrated into your product? Yeah, great question. So it is a research tool that's first of its kind in the sense that you could use it to reproduce our results and actually compete in code forces while others that competed there, for example, DeepMind OpenAI did not. like you can use your tool so easily um so it is a research tool having said that we published like a paper and a blog post and and more like uh material about principles of what we applied in alpha corneum and that we we are integrating this concept we're integrating into our variety of of agents.

28:07And let me give you one example. A week ago or so, we released, and it actually was accepted really well on social media and by our customers and users, we released a new concept of coding with AI. What you can do is first you can plan with our agent. This is, by the way, not so new. I'm admittedly saying like, I can't recall exactly, but I think, like, for example, GPT Engineer, amazing team by Lovable, shout out here. They, for example, introduced the concept of planning before you start a task, coding task, for example. So our tool, the CodeUmates, the IDE, first and ask you to plan a task before you start coding.

29:03Now when we when Codiumate helps the developer to plan a task it does it in a way that's similar to Alpha Codium. Okay it tells the developer give me your task but then now now let's iterate to make sure the specs that you're trying to achieve are clear that you're writing, that you're planning, like a new piece of code that is testable, etc. It helps you build a plan, by the way, that is connected to your existing codebase. Our agent, Codium, it actually helps you enhance existing codebase, which is relatively rare. Most auto-agent, we talked about it, helps you build new stuff. Here it helps you to enhance existing code.

29:52And after it helps you planning step by step, now you can use CodeUmate code completion, which is similar to Copilot, etc., but the code completion is task-aware. So when you come to a file, it almost automatically completes what you wanted before you even needed to prompt it because it takes a look into the plan and works accordingly. So that's a similar concept and behavior to what's happening in AlphaCodium. In AlphaCodium, what we did, very different from others where most solution are doing prompt engineering. Give me the problem. Let's find the best prompt. Iterate over it to provide a solution.

30:37In AlphaCodium, what we did, no. Let's work like a developer. What a developer would do is like 100 steps or so to solve the problem. Okay, let's read the problem. Let's see I understand it. Let's think about all the edge cases of it. Yeah, okay, now that I know about it, let's think about which different solution I can apply. Let's think about the pros and cons of the solution. Okay, let's imagine a pseudo code. Let's think if this pseudo code is actually fitting the problem. Okay, let's try to build a solution. Let's think of, et cetera, there's a larger step there in our flow that was engineered.

31:13We don't let AI decide about this step That's different than what was the common case before, where AI decided about the step with chain of thoughts, tree of thoughts, other techniques. And this way of thinking that helps us to build AlphaCodium also is being applied to our agents with various flows. And one of them is the one I described in CodiumAid that helps you plan a task and then also execute it with a code completion that is task aware. obviously like similar to AlphaCodium when after you complete writing you can also test your code etc that's also something that Codium might help so bottom line recap AlphaCodium is a research tool you can actually find full code full open source reproduce it's not meant for enterprise that's playing around and then the concept is being integrated into our various agents In that explanation, you explain flow engineering, which is the planning, right?

32:20Do you use LLM for that? So, basically, giving again a bit more context, what we're basically doing is shifting from system one syncing to system two syncing. like uh unfortunately daniel hallman just passed away uh i think a couple of months ago and that was really unfortunate and he like framed us uh like in his book uh friend this term coined this term like system one system two right uh thinking fast and slow that that's the name of the book and the idea here is that uh system one is like intuition right it's like it's like lm inference. It's a bit of what we're doing here in the conversation.

33:13We're kind of like an LLM now. The context of what I said two sentences ago flows and prompts me to continue what I'm doing. That's the system one. And that's what we do with LLMs. That's what they're so good in conversation. And that simple question answering because that's basically what we're doing. I admit, I'm not sure I'm smarter when applying system one than the best models that are out there. And the system two applies that you have a process, you have a flow where it helps you to make more intelligent decisions. And I think as people, we applied it many times when you're thinking about career paths.

34:00We don't do that when we drive, but when we're at work and want to think about an architecture. And there's a lot of research about boards, like a board of executives, etc., where if they come into a meeting and they just need to make a decision, yes, no, somebody presents and they vote, they will get very bad results comparing to a flow where it was decided how decisions are going to be made. The fact that decision is not going to be made over one option, whether three, just as an example, and like it really boosts the result so all that background coming to say that what we're saying is that when the board of directors are coming to make a decision they actually do come with a flow with a process that maybe they learn from experience or from stanford or whatever and it's quite rigid like this is the flow and within the flow everything changed what is the topic we want to this?

35:02What is the decision? What is the option? There's a lot of flexibility there. And that's what we did with the flow engineering Alpha-Lacodium. We designed a flow. It's rigid. Each step following the next. But within it, we use LLM to generate options that fits for each step, to generate the analysis for each step. But we do not let it decide about what is the flow and what is the type of the decision that needs to be made. And I claim that we human and good processes also do that. And I actually talked with a couple of professors around the U.S. and the States, and they were like, they smile, et cetera.

Read the full transcript

35:44That's what they teach also, like how to make decisions. Copilot and these LLM-based code completion or code generation systems are trained on, And I mean, the underlying of the foundation models trained on, you know, massive kind of undifferentiated text data that includes a lot of code. And then they're fine-tuned on code. But the reason they are code completion and not code generation, purely code generation systems, is they tend to hallucinate even in small ways or generate errors and encode that those kinds of errors can then cascade into larger errors.

36:46Why don't, and you can talk about the LLMs that you're doing in this, why doesn't someone train an LLM on a massive corpus of verified, clean code so that it doesn't have any bad code in its training set? And then wouldn't it then be much better in reading new code? Yeah, great question and idea. So first of all, I think I hallucinate when I talk quite often if you don't let me use my system too. For example, when I think about something, I want to verify it. Go Google it, search, ask people. That's what I call system too because I'm actually using my tool to make a decision. Otherwise, I think I hallucinate quite a lot, unfortunately.

37:45I try to, by the way, to put probability on what I'm saying. I make mistakes often. Actually, always. Especially when applying System 1. I wouldn't blame the LLM. It's whether how we use them. You can use LLM today as a System 2 and ground it. like a check, put a flow to check the results. And if it's wrong, similar to I could be wrong if you let me use only my system one, like let it fix the answer for a different tool or additional information found it wrong. Okay, I just wanted to say that it doesn't contradict, in my opinion, what you said. I think there are companies and solutions that try to clean data and then training the model, and then it will bring better results.

38:41You can try to do that for code generation, and it might produce better code. For example, I heard that GPT-4 Turbo, the latest update does better on code than the previous one, etc. But you won't. You just won't solve the problem with inference once. It's like you'll get better and better results, but you won't solve it. Because eventually you have to have a system. you want to call it agent or not doesn't matter basically we use agent because we do use tool etc but the more important concept is the flow idea and the LLMs the more they become better it enables a system to reach its capacity, reach its potential but you will need a system to eventually get to a point where you generate code that works okay uh that and by the way when you say good code that's different between different companies different people would consider different codes so so again you will need a system that knows to extract what good code means for different uh like uh people different companies and again that's like sorry that's sorry for pushing like a like a shameless plug but that's what we're focused on, extracting and understanding the specific enterprise requirement and best practices in order to influence an LLM.

40:09Now, last thing I want to say is that would you get the idea of cleaning, cleansing code, extracting, collecting, cleansing code for better code completion, but there are other opportunities here to clean, extracting, clean code for better code review, for better code testing, etc. And actually, there's a huge potential because current models are actually quite stuck in that. Sorry. Because think about the way they're trained. They're trained of, here's a piece of code, here's a missing token, try to complete. This is almost perfect for code completion. The task that is being applied, code completion, actually fits really well the task of the way that these models are trained.

40:59But when you think about code testing or code reviewing, you don't just complete, you take a piece of code and now you need to look at it from the side and say, does this work or not? So the reason that the models are doing something, even without fine tuning, something reasonable, is because they do understand something intrinsically, implicitly about the world. And they're also reading a lot of blog posts, etc. that does talk about it, but still, they're not good enough. And then if you do collect data that is relevant specifically for code review, for code testing, and train it accordingly, not necessarily the same way that we're talking, maybe a bit differently, some of the secret sauce, then you can get powerful models or fit for that.

41:51I can tell you that we have Fortune 100, even a fortune 10 company running our model uh on-prem like air gaps because that's what they care about uh and they get a model that is i would say according to our benchmark at least as good for gpd4 turbo uh uh last edition for our task and arguably maybe even better on-prem like uh or a gap. So that's like unimagine, like unhear or you can't do that with CodeLama or something like that. Let's talk about your products. So is it two distinct products or is it sort of two problems that you focus on? We see it as what's publicly available. We have more that our clients are using and we're going to publish some of it soon.

42:44So what you see is like two agents that are part of one platform. In our Enterprise Edition, by the way, they are connected and they share information, but they're free and teams, what's publicly available, you don't see the connections, they're separate. But they are covering different pieces of your stock development lifecycle, like the point of where you touch your code in different places. One is Kojimate's ID plugin for JetBrains and VS Code, for example. and the other is we call peer agent which you connect to your github github bitbucket even azure devops which is a contribution we got from very one of the biggest companies in the world because it's open core and even though they are the creator of one of the top assistant out there they're still using our our tool because it's so unique and different so we We got even a contribution from there to connect to Azure DevOps.

43:46So basically, our open source or free are independent subproducts, but the full platform and enterprise get, they're actually a system that connects between them and share data. Yeah, so to put it simple, developers are writing code, and they love writing it. and now what they hate doing is analyzing it testing it that's what it does for them it creates analysis and of what we call behavior coverage like natural language description of branch coverage that like actually what it's doing like and then from there you can generate tests and give you analysis i just talked with a few developers for example in mercadolibre which It's one of the companies we have dozens of users from there.

44:38And that's what he told me. We just love the fact that code that we didn't see for a week or two or someone else's code, we immediately come and understand it and get analysis for it and then generate tests that will take me 10 days. I can get it done in two hours, et cetera. Because you can create a huge amount of tests in a short time, sometimes 10 days, sometimes maybe it's one day. but within two hours you'll get like a much better coverage so that's what the id extension id plugin which is called codemate uh the github what it gives you it's more than a dozen of commands by the way the id code might also have more than a dozen uh but uh i just i focused on the testing for for a second before that i described the plan tasking and the code completion there are a dozen more than a dozen of commands now focusing on the pr agent there's more than a dozen commands that you can run on top of the Git platforms.

45:38And that's including describing a pull request or analyzing your pull request and then building a walkthrough and a dictionary or book. And that's extremely important for the persona, the code reviewer, not necessarily the person who developed the code. And here, let me share a meme that is very famous. If you give me, as a reviewer, if you give me 20 lines of code change, I'll give you five comments. 50 lines, I'll give you two comments. 500 lines looks good to me. This is a problem. Okay? Right? So we want to empower the reviewer. So you'll see a dozen of commands, a half of them are for the reviewer to analyze and understand better the PR.

46:30And then what we help there is at the point of merge, we help the reviewer or the AI catch things before they are merged. So bugs and issues, et cetera, and best practices. And this is both of the tools, especially the PR agent, it's highly customizable to a company's best practice, which is also like an enterprise edition, even a more dominant property. Yeah, so that's what we do in CodeMEI. Yeah. And can you talk about the, are both sides using generative AI and what models are you using? Are you using your own LLMs or are you accessing or can the user choose the foundation model that they want to use?

47:24And then are you using reinforcement learning in some way with either side of this? So first of all, we enable a few modes of using models. we enable using open ai or soon other models that are reached finally to the level of or close to the level of open ai models we're using our own model but this this enablement is done only for the enterprise that they can choose otherwise for the free users we enable only open ai by the way they don't need to enter their keys they use our open ai uh so it's it's like mostly for free and um and and basically uh we use the best of breed of open ai models you might think that gpd4 true with anyone is the best in everything but there's other trade-off to to consider uh and and that that's about that but but again this is uh free for for enterprise we do enable from a selected set of models.

48:33We have two optimization processes. One is for the model itself and the other is more for the system where we generate flows. And I wouldn't say that we per say, like use that traditional reinforcement learning PPO or whatever techniques that is used to train models to play games. For example, our training process is much more similar to LoRa style, like fine tuning and training LLMs. It might be slightly different because of things I mentioned before. So not per se like the traditional. It's weird to say traditional because reinforcement learning is relatively like the latest epic or epic is like three, five years.

49:32But that's what I mean, like traditional, not the Q learning or something like that. It's not being used by us. What's sort of the message that you would like listeners to take away and feel to be promotional? I mean, that's fine. Yeah, actually, let me say two things. One is for if you are a developer, engineer, etc. And then you're seeing all these either videos or me talking, etc. I'm not saying that it's the same thing. I think of it more or more pragmatic. I would say that I wouldn't be actually so afraid. By the way, some of the people that are creating buzz is interesting. It's GitHub, right?

50:15They're showing features where you just write what you want to do and do-do-do-do-do. Everything is being done for you. So I would say, actually, I wrote Thomas, the CEO, that I'm quite surprised. With Copalette, they were so pragmatic and suddenly they're trying to, I think, push the border with reality a bit. So I think that I would be less, I would be like very thoughtful and skeptic about some of the videos people see. And I wouldn't like so quickly hurry that, you know, being afraid of your job, being taken at the same time. I think you must like start exploring and exploiting AI because how you're going to develop in five years is going to be very different.

51:07Smart be like a cockpit. By the way, similar to how you do cloud today. Think about how you did cloud 30 years ago. Today, cloud is a lot about orchestration. And code development is going to be a lot of orchestration. And you need to go into the details only once in a while, similar to how it's being done on cloud. And so that's my suggestion for a developer and for the enterprise. What I suggest is think what's your biggest problem. Is it developer happiness, which we all care about? or is it like bugs in production that might cause like outages or a huge amount of, like you can lose reputation or lose a lot of loss of productivity because half of your sprints are fixing stuff.

51:56So I would say like choose what's your pain and then like find the relevant AI tools, not the other way around. Don't go around, oh, I want to use Gen AI. I hear that a lot. no go with what's the pain i want to solve and then find the right hammer and not okay there's a cool hammer here let's let's see if it fits for me so that's that's my suggestion uh only slightly promotional because i think like eventually it is on the side of a bit slightly promotional because eventually i think that what we see is that eventually people understand that that there's enough research and if you put thought in it the real way to to boost productivity is via improvement of quality of your code and quality of your processes.

52:37And that's what we do, coding AI, C-O-D-I-U-M-A-I. That's it for this episode. I want to thank Idmar for his time. If you want to read a transcript of today's conversation, you can find one on our website, IonAI. That's E-Y-E hyphen O-N dot A-I. In the meantime, remember, the singularity may not be near, but AI is changing your world. So pay attention. By popular demand, NetSuite has extended its one-of-a-kind flexible financing program for a few more weeks. Head to netsuite.com slash ionai. That's ionai, E-Y-E-O-N-A-I, all run together. Again, head to netsuite.com slash ionai. NetSuite.com slash IonAI, where it's one of a kind flexible financing program.

From the publisher

This episode is sponsored by Netsuite by Oracle, the number one cloud financial system, streamlining accounting, financial management, inventory, HR, and more.

 

NetSuite is offering a one-of-a-kind flexible financing program. Head to  https://netsuite.com/EYEONAI to know more. 



In this episode of the Eye on AI podcast, join host Craig Smith as he sits down with Itamar Friedman, co-founder and CEO of CodiumAI, a pioneering company in the realm of intelligent coding systems.

 

Itamar shares the ambitious vision of CodiumAI, where developers can aim for zero bugs and issues in their code. Dive into the innovative solutions CodiumAI offers, including Codiumate and PR Agent, tools designed to assist developers in planning, writing, analyzing, and reviewing code with high integrity. 

 

Explore the concept of flow engineering and how CodiumAI's research project, AlphaCodium, outperforms DeepMind's AlphaCode by using fewer LLM calls and no fine-tuning. 

 

Itamar delves into the future of coding, where AI not only generates code but ensures its quality through robust verification systems.Itamar offers valuable advice for both developers and enterprises on navigating the evolving landscape of AI in software development.

 

Tune in to gain insights into how CodiumAI is leading the way in creating a future where coding is more efficient, reliable, and free from bugs.

 

Don't forget to like, subscribe, and hit the notification bell for more on groundbreaking AI technologies.



Stay Updated:

Craig Smith Twitter: https://twitter.com/craigss

Eye on A.I. Twitter: https://twitter.com/EyeOn_AI

 

More from Eye On A.I.

All 266 episodes
#193 Itamar Friedman: How CodiumAI is Making Bug-Free Code a RealityEye On A.I. · 54 min
Listen in VO