A developer's toolkit for SOTA AI

12 Jul 2023 · 42 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Practical AI - A Developer's Toolkit for SOTA AI

Episode Overview In this episode of the Practical AI podcast, hosts Chris Benson, Varun Mohan, and Anshul Ramachandran discuss the development of Codeium, a free AI-powered toolkit for developers. They explore how this tool aims to streamline the modern development process in generative AI and large language models (LLMs), and how it is built on insights gained from GPU software development.

Guests

  • Varun Mohan: CEO and Co-Founder of Codeium
  • Anshul Ramachandran: Lead of Enterprise and Partnership at Codeium
  • Chris Benson: Co-host of Practical AI

Key Topics Covered

Background of the Guests

  • Varun Mohan:
  • Previous work at Neuro focused on autonomous vehicle technology and large-scale offline deep learning workloads.
  • Founded ExaFunction, which developed GPU virtualization software, ultimately leading to the creation of Codeium.
  • Anshul Ramachandran:
  • Also worked at Neuro and transitioned to ExaFunction.
  • Joined Varun in recognizing the need for better tools for engineers using AI in software development.

Codeium's Mission

  • To provide developers with a free toolkit that leverages in-house models and infrastructure rather than simply being another API wrapper.
  • Focus on creating tools for the entire software development lifecycle (e.g., autocomplete, in-IDE chat, and documentation support).

Challenges in GPU Software Development

  • GPUs are hard to virtualize compared to CPUs, leading to inefficiencies in managing resources.
  • The cost and scarcity of GPUs (e.g., NVIDIA models) pose significant challenges for companies scaling their AI workloads.

Trends in Generative AI

  • The rise of tools like GitHub Copilot has shown the potential of AI in software development, but many users find limitations in their applicability at work.
  • There’s a trend toward developers wanting more robust tools that can handle the complexities of their codebases.

Unique Features of Codeium

  • Contextual Understanding:
  • Offers double the context capacity for autocomplete compared to competitors.
  • Capable of contextual awareness across the entire codebase.
  • Personalization:
  • Allows for fine-tuning of models based on the specific codebases of enterprises, improving relevance and reducing errors.
  • User Experience:
  • Built-in functionality to streamline workflows, such as applying diffs directly in the IDE, rather than requiring manual copying and pasting.

Key Takeaways

Competitive Edge

  • Codeium aims to be more than just an autocomplete tool by addressing the entire development workflow.
  • Unlike competitors, Codeium emphasizes safety and privacy with self-hosted solutions, ensuring that enterprise data remains secure.

Future Outlook

  • Varun and Anshul believe that AI tools like Codeium will enhance productivity while still requiring human oversight. They emphasize the importance of building tools that genuinely assist developers without taking away their control.
  • The industry is evolving rapidly, and Codeium is committed to adapting and improving its offerings based on user feedback and advancements in technology.

Generative AI Use Cases

  • Developers can use Codeium for:
  • Autocompleting code snippets.
  • Generating documentation and unit tests.
  • Refactoring code and facilitating peer review processes through integrated chat features.

Conclusion This episode of Practical AI provides valuable insights into the development of AI tools designed specifically for software engineers. Codeium positions itself uniquely in the market by focusing on developer needs while ensuring data privacy and optimizing workflows.

For further discussion and resources:

  • [Codeium](https://codeium.com)
  • Related blog posts on GitHub Copilot's limitations and the future of AI in software development.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:06Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related technologies are changing the world, this is the show for you. Thank you to our partners at Fastly for shipping all of our pods super fast to wherever you listen. Check them out at Fastly.com. And to our friends at Fly, deploy your app servers and database close to your users. No ops required. Learn more at fly.io.

0:43Welcome to another edition of the Practical AI Podcast. My name is Chris Benson. I'm your co-host today. Normally, we would have Daniel Whitenack joining us, but Daniel has just gotten off a plane. He flew halfway around the world, and we decided to give him a break from today. I would be, he was more lucid than I would be under the same situation. Today, I wanted to dive right in. we have a super cool topic. It is not dissimilar from some of the other general things we've been talking about, but I have two guests today. I'd like to introduce Varun, who is the CEO and co-founder of Codium, and Anjul, who is the lead of their enterprise and partnership.

1:24Welcome to the show, guys. Thanks for having us. Thanks for having us, Chris. Hey, you're welcome. I'm really interested in learning more about Codium. When Daniel lined you guys up. He's like, Chris, you got it. He sent me this thing saying, you got to look at this. This is really cool and everything. And I'm like, get him on the show. He's like, I'm already doing that. So, so really glad to have you guys on. And he's going to be bumming that he missed the conversation because he was pretty excited about it. And so I guess I wanted to, before we even dive into Codium and the problems it's trying to solve and such, if you guys could each just tell me a little bit about how you found yourself arriving at this moment, kind of a little bit about your background, how you got into AI, and how this became the thing.

2:09Varun, if you want to kick off and then Anshul afterwards. So maybe I can get started. It actually starts in 2017. I started working at this company called Neuro that does autonomous goods delivery. So it's an AV company. There I sort of worked on large-scale offline deep learning workloads. So as you can imagine, an autonomous vehicle company needs to run large-scale simulation. They need to basically be able to test their ML models at scale before they can actually deploy them on a car. And sort of in 2021, I left Neuro and started ExaFunction, which is the company that is building out this product, Codium.

2:46And ExaFunction started out building GPU virtualization software. So you can imagine for these large scale deep learning applications, one big problem is GPUs are scarce, they're expensive, and also hard to program. And sort of what ExaFunction started building was solutions and software to make it so that applications that ran on GPUs were more effectively using the GPU hardware. And we realized that our software with ExaFunction was best applicable to generative AI tech and started building out Codium around a year ago. Very cool. And before I dive in, because I have several questions for you, but I want to give Anshul a chance to introduce himself here.

3:26Go ahead, Anshul. Surprisingly, my story is actually quite similar. I was also working at Neuro. So Varun and I used to work together back in the day. I was not actually working on, you know, the ML infrastructure side of things. That was something that, you know, Varun had hands-on on. But, you know, I decided to kind of also join the team at ExoFunction. And I think, yeah, as Varun mentioned about, you know, a year ago, I think we noticed there was like, I think three things kind of happened at the same time that we noticed that led us to Codium, right? I think the first one is that, you know, we're engineers.

3:57All of us here are engineers. and we had all tried the GitHub co-pilots and all these cool AI tools for code in their beta. And we're like, wow, this is absolutely going to be the future of software development. But at the same time, it's still scratching the surface of potentially everything that we do as engineers. So that's, I think, number one, I think that we realized. Then number two was talking to a lot of our friends at these bigger companies or anything like that. A lot of them were just saying like, oh yeah, it's cool. I've tried it for my personal project, but I can't use it at work, right?

4:29My work's not allowing me to use that. So it was like the second thing we heard. And the third thing was exactly what Varun alluded to. We were building like ML infrastructure at scale for really large workloads. Like when this entire generative AI wave started coming, we're like, wow, we're actually kind of sitting on the perfect infrastructure for this. So I think all of those three things kind of combined together for us to be like, do you know what? Let's build out an application ourselves and build an application that we as engineers are customers ourselves, right? And that ended up becoming Kodium.

4:59As you were getting into doing GPU software, what was in general some of the challenges that you were seeing? You know, with NVIDIA has their various software supporting things like that. Clearly, you saw that there was a need for something beyond that. Can you talk a little bit about just the layout that you saw in the environment before you got to all the generative stuff and the fact that you had infrastructure? What positioned you for that? And what was the thing that you decided that you needed to address? Maybe I can take it a step back of why these GPU workloads are just a little bit annoying compared to CPU workloads.

5:32One of the really sort of unique things about GPUs is that unlike CPUs, it's kind of tricky to virtualize. Like one common thing that we have with CPUs is you can put a bunch of containers on a single VM, and then you can kind of make use of the CPU compute like effectively. Right. You can basically dump 10 applications onto a CPU and it's perfectly fine. For GPU, it's a little bit more messy because the GPU doesn't have a ton of memory. So you can't just load up infinitely many models on there. Like let's imagine you have a GPU with 16 gigs of memory and each of these models takes like 10 gigs.

6:06You can't really even put two applications on there. So then that already becomes a big issue. And that's sort of what a lot of these large deep learning workloads were struggling with. So when I was at Neuro, one big problem we had was we had around like tens of models, but we had these workloads that needed hundreds of GPUs, some of them even thousands of GPUs. And we struggled to basically make it so that we were even able to use the hardware properly. And then, you know, you could imagine the complexity then stacks with now we're in a state where companies have trouble even getting access to 10 GPUs because of NVIDIA sort of scarcity issues.

6:42And then also the cost of a GPU is like not like a CPU. It's like significantly more expensive. Like the cost of a single H100 chip is well over 30 grand. So these aren't like very cheap chips. So there's like a big need at the time to figure out how do we leverage the hardware properly. And sort of that's what we had to build software for. And just to clarify for me, was that why you were still at Neuro? Or was that after you started ExaFunction? Yeah. So while I was at Neuro, we sort of worked through, or I sort of led a team that sort of built software that kind of fixed these problems. But ExaFunction was focused on generically, how do we make sure deep learning-based applications could best leverage GPUs?

7:21That's sort of what we started out building, actually. And then Codium came out from that, actually. Gotcha. Tell me a little bit about, as you have been right in the middle of this progression, just to frame it for a second, if you look at the last couple of years in particular, and the pace of change has been so much. And so you were right there starting at Neuro and then creating ExaFunction, seeing some of the challenges, could you talk a little bit about how the industry was evolving and changing as you were seeing it so that we can get a sense of kind of how you moved toward Codium, you know, to give a little bit of the history instead of just starting from where that is.

8:01Can you talk a little bit about, you know, the itches that you were scratching and why it led that direction? What did this AI industry look like to you? Yeah. So when we started, like, you can just imagine everything was a lot more smaller scale, right the hyperscalers or the cloud providers just didn't have nearly as much gpus like if you ask them like what fraction of cloud spend is gpu spend it's probably like very small single digit percentage points maybe even less than that at the time so this is like a very small workload for them when we sort of started both me and anshul started in a row in like 2018 but then over time this grew a ton like we could see it from the training workloads these were no longer like even single node training workloads like back in the day a single gpu node that had maybe like eight V100s or something was like considered a lot of compute.

8:46And suddenly now we were able to witness the fact that this was slowly becoming eight A100 nodes. And then more than eight of these nodes were necessary then even to train these models. And similarly, to prove out that these models were capable, like in an actual production setting, you needed to run offline testing at massive scales, like on the order of like 5 ,000 to 10 ,000 T4s scales, which is like kind of incredible in terms of raw flops. So we were able to see this hockey stick happen in front of us. And then that's sort of what made us want to start ExaFunction in the first place. We realized that there were going to be large deep learning workloads.

9:22One interesting fact is for us, like for just the ExaFunction GPU virtualization software that we ended up selling to enterprises, we ended up managing over 10 ,000 GPUs on GCP in a single GCP region. So we ended up managing more than 20%. And we realized that, that, hey, this was only going to keep growing. Like when we talked to the cloud providers, they were only going to keep growing the number of GPUs. And we realized, I guess the interesting thing was in the future, generative AI was going to be potentially the largest GPU workload though. That was the big thing we realized was GPT-3 came out, which was, I guess, in 2021 now.

9:57Gotcha. So, but you had already, at that point, were you already in Dexa function? Had it already started at that point? Yeah, it had already started and we were sort of selling GPU virtualization software to large autonomous vehicle and robotics companies. Gotcha. And so basically, if I'm understanding correctly, the whole generative tsunami just kind of landed on you when you were already sitting in that space doing GPU virtualization already. So you just managed to land right in front of the wave, it sounds like. Yeah. So we started working on Codium like maybe four or five months ago before ChadGPT.

10:31It was interesting just because we realized that an application like GitHub Copilot was going to be one of the largest GPU workloads, period. Like, I don't know if you've probably tried the product out. It's like every time you do a key press, you're going out to the cloud and doing trillions of computations, right? So it's like a massive workload. And we had like, as Anshul said, the perfect infrastructure to basically run this at enormous scale. Not to mention we were in love with the product from day one. Like we were early users of the product the moment it came out in 2021. Very cool. And so as generative is starting to take off kind of with chat GPT hitting the world and really changing things quite rapidly, you know, I think people are still shocked at how fast things have moved.

11:12You would start Codium already. What kind of synergy were you starting to see there in terms of knowing that you have one of presumably many, many GPTs coming and other similar generative models. You had just gotten into Codium. Can you talk a little bit about what that was and what were you putting together in your minds to recognize the opportunity that it was? Yeah. So I think one of the great things about the entire chat GPT wave is that everyone was using it. This is a thing where literally every individual is using AI. And so it helped us in general, right? You know, like a big wave raises all ships kind of thing.

11:51It really helped us, we weren't really going out and now telling people like, hey, a tool like Codium can help productivity because that was kind of just now assumed by everybody. Like, oh, yeah, if I do any kind of, you know, knowledge work, then there's potential for AI to help. Right. And I think so from that sense, when the star like, you know, chat GPT wave really came about, that overall kind of just like helped us in terms of convincing people to even try the product. The other thing that we we recognize is that we were positioning ourselves very specifically from the beginning, when it comes to code.

12:22Code is actually a very interesting modality. It's not like your standard chat GPT where you have a long context that a user puts in and then it produces context coming out. Code is interesting in the sense that, as we mentioned, it's an autocomplete that's like a passive AI rather than like an AI that you're actually instructing the model to do something. It's happening every keystroke. So it has to be a relatively smaller model. You have these hundreds of billions of parameter models being used. It has to be relatively low latency. And then code itself is interesting. If you have a cursor in the middle of a code block, the context both before and after your cursor really matters.

13:01It's not just what comes before. So there's all these interesting situational constraints about code that you put all these things together and we realize that, OK, all these chat GPT waves and conversational AIs are happening. That's great. but we're still not going to be like, you know, rolled over by that because we're kind of focusing on a very specific application and modality of LLMs that was pretty unique in many ways.

13:44Could you take a moment as we're diving into Codium Ingenerative AI and its unique capabilities there and just differentiate a little bit about for those, you know, so many people have tried Copilot. And so it's kind of inevitable that you're going to get that comparison to some degree. Can you talk a little bit about what Copilot's not doing for generative AI or how you're approaching it that allows you to show people this as a better way forward from your perspective? I mean, we have tons of respect for the Copilot team. I'm just going to start with that. As Bruin said, we were all early users of it.

14:20Definitely not putting you into conflict with them. That just is a starting point for people. Absolutely. Yeah, I think the way we kind of view this, and I kind of alluded to this earlier, is that writing brand new code with autocomplete is really just one small task that we do as engineers. We refactor code. We ask for help. We write documentation. We do PR reviews. And so kind of our general approach has always been, let's try to build an AI toolkit rather than an AI autocomplete tool. Got it. So we can get more into this, into the weeds here, but like autocomplete is just one of our functionalities that we provide, right?

14:55We provide like an in IDE chat. So I think like chat GPT, except integrated with the IDE, natural language search over your code base using like embeddings and vector stores in the background. So like we're really trying to expand, like how can we address like the entire software development lifecycle? So I think that's probably the most obvious difference with a tool like Copilot from like an individual developer point of view. But then the other thing which really kind of builds off of all the infrastructure that Vern was mentioning earlier is that we were already deploying, you know, ML infrastructure in our previous customers' private clouds.

15:27Like we already had all this expertise of how can we take actual ML infra, deploy it for a customer in a way that, you know, they can fully trust the solution because, you know, we're not getting any of their data. And so another really big differentiator for us was like, okay, I think this might actually be a tool that enterprises can use confidently and safely because we have the infrastructure to do the deployment in a manner that they would be open to using. So I think that was like the other differentiator when it came specifically to enterprises, but we can dive more into that later. No, that sounds good.

15:57I want you to connect one more thing for me. Going from being able to deploy the infrastructure and helping your customers in that way to Codium as a tool, what's the leap there that got you from one to the other? How did you get from infra-focused to Codium-focused? Oh, yeah. I think we had to do a full 180 when we started. We went from a full infra-service company to let's create a product for consumers. It was a full 180 in terms of product. A pivot. Yeah, full, in some degree, a pivot, because we knew that, you know, eventually, okay, we'll deploy to customers VPCs. That sounds great. But, like, if we're going to ship something to a customer, we had to be, like, super confident that it was a product that would work well, right?

16:37Because we're getting no feedback from their developers. And so we actually first focused for the first, like, six or so months of Codium just building out, like, an individual tier, right? Any developer can go try it. We can see how they like it, right? like try our new capabilities, get feedback from an actual community, do all these like community building things that we hadn't really done as like, you know, infra as a service company. But that was like a really huge focus for us. And, you know, we've grown our actual Codium individual plan to like over 100 ,000, you know, active developers using us for like, you know, many hours a day because you code for that long if you're a developer.

17:11You know, that's like plenty of feedback to us, right? Plenty of people actually using the tool telling us like, yeah, this is good. This isn't good. Like, oh, you tried pushing a new model? That's worse. Like all those things we actually learned so that we can get a product that's good. So that was like the, I guess, the intermediate period, right? Really learning from actual developers what is a good product and what is not. And I think that's like that's always going to be a key kind of part of our development cycle. You're coming into this with this rich knowledge and infrastructure for customers.

17:40That's a huge area of expertise. It's an area of expertise that even though you're moving forward into the kind of the Codium era, if you will, in my words, that is a skill set and level of expertise that very few organizations have deeply that you would have had there. How did that inform you in terms of Codium and differentiation against whether it be Copilot or other tools that are out there or just, you know, developers, you know, throwing things into chat GPT? What did that background give you that gave you that differentiation in the marketplace? Yeah. So I think when we started, the thing we started with is like, no one cares if we have better infrastructure once you're a product.

18:20Like if we have better infrastructure, that's great. But if that makes a product that's the same, no one should care. They just assume that you should. Yeah. So what we started with is we set a very high bar for ourselves. Codium is an entirely free product. So like for the individual user, it's something that they can install and use immediately for free. There are unlimited. There's like no limits at all. So like when it comes to autocomplete, you can use it as much as you want. And this is, by the way, They forced us to do things where infrastructure is as efficient as possible. Just to give you a sense of the numbers we're talking about here, we process over 10 billion tokens of code a day.

18:55That might sound like a large number. That's like over a billion lines of code a day that we process for our own developers. We're forced to do this entirely for free. And then on top of that, we probably have one of the world's largest chat applications also because it's in IDE as well. And all of this put together has allowed us to build a very, very scalable piece of infrastructure such that we're the largest users of our own product. We're the largest user of our own product. We learn the most from our users. And we can then take those learnings and deploy in a very cost-effective, very efficient and optimized way to our own enterprise users.

19:28It's one of those things where we force ourselves to learn a lot from an individual plan and then take all those learnings and actually bring them over to the enterprise. And a lot of the learnings we were only able to make because we place like very, I would say like annoying infrastructure constraints in ourselves by saying, hey, you guys got to do this entirely for free, basically. And we're committed to building, Codium is going to be a free product forever. Actually, the individual plan will always be free. And it's one of those things where our users are just always like, how are these guys even doing it?

19:56Like, what are they even doing to make this happen? And most of our users, by the way, are users that have churned off of Copilot. We have spent very little, if not anything on marketing. So it's just one of those things where our users are like, how do we make this free? We take the approach of, we think some of the best products in the world are free, like products at Google, right? They're entirely free. Google doesn't tell you all the time that they have the best infrastructure, but they do have the best infrastructure. It just so happens to be the case that that shows itself off in the best product.

20:21And we could talk a little bit more about how we take our sort of focus on infrastructure and make a much better enterprise product as well. But like, that's the way we sort of look at it. It's like, how do we deliver materially better experiences with our infrastructure? And our users shouldn't care that we actually did that. You've brought it up. You got to go there now, man. Go ahead and dive right into it. I guess like one of the interesting things, like just going to how we run one of the world's largest LLM applications, what that sort of focus forced us to do is give it a single piece of compute, like let's say a single node or a single box of GPUs, we can host the most number of users on there.

20:53So like, let's say a large company comes to us, they can be confident that whether they're on-prem or they're in VPC, we can give them a solution where the cost of the hardware is not going to dominate the cost of the software itself. Because right now there's kind of this misunderstanding that the GPUs are really expensive, which is true, they are. But the trade-off is they have a lot of compute. Modern GPUs like A100s can do 300 teraflops of compute, which is some ungodly number. That's a crazy number compared to what a modern CPU can do. And we can leverage that the best. And we've sort of been forced to do that.

21:28If we didn't do that properly, we'd have outages with our service all the time. Because of that, enterprises trust us to be like the best solution to run in their own tenant in an air-gapped way, which is fantastic because that's like the way that we can build the most trust and deploy these pieces of technology to them the most effectively because they don't want to ship their code outside of the company. Anshul can talk a little bit more about how we leverage things like fine-tuning as well. It's like a purely infrastructure problem that's very unique to us versus like any other company as well.

21:56Anshul, do you want to sort of take that? I mean, yes. I think, you know, as Viren said, there's a lot of things that we do from like the individual infrastructure point of view so that we can do crazy things like make it all free for all of our individual users. But once we actually self-host, there's actually a lot of things that you can do that just any other tool can't do without being self-hosted. And one of the ones that everyone just mentioned is personalization. If you're fully hosted in a company's tenant, you can use all of their knowledge bases to create a substantially better product.

22:27I think the way we generally think about it is that you have a generic model that's good. It's learned from trillions of tokens of code in the public corpus. But if you think about any individual company, they have themselves hundreds of millions of tokens of code that has never seen the light of day. And that's actually the code that's the most relevant for them if they want to write any new code. Think of all the internal syntax, semantics, utility functions, libraries, DSLs, whatever it might be. And a model like a Copilot or a Codium, by the nature of it having to be low latency, can only take about 150 or so lines of code as context.

23:03So this is not like one of those chat GPTs or GPT-4s where you're putting in files and files of context. It's really small where you can put in. And so there's really no way for a single inference to have full context of your code base without actually fine-tuning the base model that we ship to them on all of their local code. And so we've actually done a bunch of studies on how this actually massively reduces hallucinations and all these other things that you always hear coming up with LLMs. But things like this, things like providing more in-depth analytics, all these things actually come up by being self-hosted.

Read the full transcript

23:39And as Rune mentioned, these are all at the core, to some degree, an infra problem. How do you actually do fine-tuning locally in a company's tenant? That's actually an infra problem that we're happy to talk more about. But maybe I'll pass it back to you, Chris. Actually, I'm about to ask a follow-up about that because you've got me really thinking about some of the use cases in my own life on that. And so with the self-hosting model and you're able to now kind of like, you know, OpenAI, I said, you know, with chat GBT4, there's only so far we're going to go because we've kind of we've used the public corpus of knowledge out there on the Internet.

24:14You know, so there's only so much more vertical scaling you can do on the model learning. And so, you know, you're touching on the fact that there's so much hidden IP and code, hidden information and code that is of huge value, particularly to the company that it's in, because it's representing their business model and the way their business has evolved over time. And so if I'm understanding you correctly, you're basically saying that your solution can take advantage of that on their behalf and really, really hone against it. what are some of the limits on privacy? Are they able to do that? Because that's a big topic.

24:51We've actually talked about it on the show before about, you know, in this generative AI age with IP concerns and privacy concerns and, you know, getting the lawyers involved. Are you able to do the training on their site and keep it to the customer entirely? Or do they have to let their IP out and stuff? How do you approach that problem? Yeah, I mean, so one of just the answer to any question of like, does any IP leave coding for enterprises? The answer is always no. So in pretty much every part of the system, our guarantee is to actually be able to deploy this whole thing fully air-gapped. We've even deployed in places like AWS CovCloud, which is entirely, doesn't even have a connection with the internet kind of scenario.

25:32So nothing ever leaves there. To address some of the points you brought up there, Chris, yeah, I mean, we're not the only ones who are saying, oh no, the data that a company has privately is super important. and is potentially even more important than the size of the model. I think a good example of this is actually meta. Instead of using a GitHub Copilot or any generic system, they decided in, I guess, classic meta fashion to train their own autocomplete model internally using all of their code. And they actually published a paper, I think, a few weeks back. And their model was, in terms of size, I think 1.3 billion parameters, like, small in respect to the LLM world.

26:11and it just massively outperformed GitHub Copilot on pretty much every task. There's certainly corroborating evidence to what we're saying about fine tuning that doing this actually does lead to materially better performances for the user in question. Now, does that meta model going to be good for everyone else to code? Probably not, but that's also not the whole point. And in terms of being able to fine tune locally, yeah, we're able to do this completely locally. And again, it comes down to scale of data. our base model has been trained on trillions of tokens of code, right? That's a lot. That's why we need this multi-node GPU setup to do all this training.

26:49But an actual company, if they have, say, even 10 million lines of code, that's about 100 million or so tokens. There's a huge order of magnitude difference still between this pre-training and the fine-tuning, which is why we can do this kind of locally on, actually, surprisingly, whichever hardware they choose to provision for serving their developers. So again, this comes to some of our ML inference background and all the stuff that we know how to do. We actually can do fine-tuning and inferences on that same piece of hardware. So we don't actually ask companies to provision more hardware. And even more critically, we are able to do fine-tuning during any idle time of that GPU.

27:28So whenever that GPU is not being used to perform an inference, it's actually doing back-prop steps to continuously improve the model. You know, fine-tune is just one aspect of like a larger kind of personalization system. But, you know, we've instrumented all this on hardware using our infra-roots to actually create a system that is relatively easy to manage. It's not like a crazy amount of overhead for any company to manage or use Codium, but still get like, you know, the maximum possible wins from these AI tools. Okay, so that is super cool. And you mentioned things like GovCloud, which I have actually worked in because of in my day job quite a bit.

28:04And I can think of a whole bunch of other use cases for me personally, which begs the question about kind of going back for a moment because we are practical AI and we like to always give some practical routes for people into that. So if we're going to go back toward the beginning of the conversation for a moment, and we have some folks that are listening to this right now, and they've been using Copilot for a while. They're probably putting code into ChatGBT and trying to accelerate there with varying degrees of success. They've been experimenting with BARD, and BARD's gotten better on code lately, obviously.

28:39And so, so many people that I talk to are still very frustrated with kind of the workflow of the whole thing. And recognizing that there are these, you've outlined these differentiators, you know, from Copilot and other competition out there in a friendly competition kind of way. Talk a little bit about some of the specific generative AI use cases that would be good if someone was in that position where they're like, yeah, I'm using the stuff, but I'm not, I'm a little bit frustrated with it. I don't have it down. and if they were to give Codium that chance and dive in on it, can you give me several kind of layout the use cases on what are they going to get when they move in from a very practical, like for me now as the coder perspective, what will that look like?

29:24What are they bonusing? And maybe give me a couple of different ones because I'm really curious and selfishly, I'm probably going to go try each of these that you're telling me. So I'm scratching my own itch by asking the question. I think you pointed out like, yeah, workflows and the user experience for a lot of AI tools, like everyone's still kind of trying to figure it out, right? We're still in very early days of these AI applications. And this is our learnings of trying to become a product company. We're actually taking like the UX quite seriously, right? And this is actually what the individual plan is great to get feedback on.

29:53I think very, you know, concretely, I think a lot of people have that frustration of like having to copy a code block over to chat GPT, write out a full prompt and like, you know, remember the exact prompt that they typed in before that gave them a good result and then copying the answers back in and then making modifications like that workflow is clearly kind of broken so when we actually built our chat functionality into the ide we're like okay what are all the parts here that can get totally streamlined right and so we actually did things like you know on top of every function block there's little like code lenses that are just these small buttons that someone can like click like explain this function it'll automatically pull in all that relevant context open it up in the window you're not copying anything over and it's like writing you know it out in human text or if you say like refactor a function or add doc strings right or write a unit test these are all just like small little buttons or you know preset prompts that you can just then click it'll do this generation on the side and then we even have a way of clicking like apply diff and because we know where we pull the context in we can apply diff right back into the context right and so you're not copying things back and trying to like resolve merge conflicts like all these things are done kind of automatically.

31:01So there's a lot of really cool things you can actually do when you start bringing these things into the IT where developers are. And we spent a lot of time really thinking, as you said, from a workflow point of view, how do you make this like super smooth? Varun, could you talk a little bit about maybe some specific tasks that you're seeing people doing? When we talk about generative and it's expanded and, you know, from LLMs and we're, you know, we're doing things in video, we're doing things in natural language, all of the different modalities are gradually being addressed with these different models and different tools that are being built around it.

31:34Could you talk a little bit about, you know, what are people trying to code right now? What specifically is Codium helping them? Like what, not just about Codium, but the actual use cases themselves so that they go, ah, I can see a path forward. I can go do that. I know how to generate this or that or the other with generative AI in Codium. Can you talk a little bit about those and something of a specific level? So interestingly, just a little bit about multi-modality, I think we're maybe a little bit far from leveraging, I guess, other modes beyond text for code. I think maybe that'll happen, but I think there's not enough evidence right now yet.

32:11For autocomplete, just to be open about sort of the functionality we have, we have autocomplete, we have search, and we have code-based aware chat, right? So for, we recognize right now that of the usage, autocomplete accounts for more than 90 to 95 % of the usage of the product. It's because chatting is not something people do like even every day, potentially. They might open it up once every couple of days, but autocomplete is something that's like always on very passively helpful and people get the most value out of it, which is kind of counterintuitive. I think people don't recognize that immediately, but when people are doing autocomplete, we've recognized there's like two modalities, right?

32:46Of the way people type code. There's a modality of accelerating the developer, which is like, Hey, I kind of know what I'm going to type and I just want to tab complete the result. And then there is also an exploration phase, which is like, I don't even know what I'm trying to do based on that. I write a comment. This is like a classic thing where like my behavior writing code is materially changed because of tools like Codium, where I'll write a comment and I kind of just hope and pray that it pulls in the right context so that it gives me the best generation possible. So in my mind, for the acceleration case, Codium is like very helpful, right?

33:16It can like auto-complete a bunch of code, but to make the exploration case, that's where of the true magical moment comes in where I had like no clue at like how I was going to use a bunch of these APIs. And that's sort of what we're focused on trying to make really better, whether that be in chat, as well as with autocomplete, how do we make it so that we can build the most knowledgeable AI that is maximally helpful and also minimally just like annoying. The interesting thing about Codium as a product or these autocomplete products is they get a little bit of getting used to, but even despite the fact that they write wrong things, it's not very annoying because you can very easily just say, I don't want this completion or it didn't like write an entire file out and you need to go and correct a bunch of functions.

33:57It was like a couple of lines or maybe like 10 lines of code. You can very easily validate that it's correct, right? That comes back to then what Anshul was saying, which is how do we make sure we can provide always the maximally helpful sort of AI agent? The answer is have the best context possible. And a couple of nitty gritty details we do is currently our context and we'll write a blog post about this is double what Copilot's is. We allow double the amount of context for autocomplete than what they do. The second thing is we're able to pull context throughout the code base. And this is actually that same piece of technology that is pulling context throughout the code base through search and all these other functionalities is getting used as part of chat for code base aware chat, which is something that Copilot doesn't even have today yet.

34:40The third piece is finally for a large enterprises. How do we make it so that these models actually semantically understand your code, which is where fine tuning comes in. It's like for us, context gets us a lot of the way, but it doesn't get us all the way. Because you can just imagine, even with double the context, let's say we can pass in a thousand lines of code. For a company with 10 million lines of code, we're scratching four orders of magnitude less code than the company actually has. So this is where our vision is like, we want to continually ramp up the amount of knowledge these models have and the ways in which they can be helpful.

35:12I don't know if that answered the question there. It did actually your acceleration versus exploration analogy. That was for me personally, different people get different things that really clarified for me where I might be using Copilot or where I would go and use Codium on that because I do struggle on the exploration side myself. It's a lot easier on the acceleration yet end of the line and the line, you know, and crank through that fast, which I've been able to do with these other tools. But I have struggled on the exploration side because I kind of want to do a thing and I'm kind of trying to figure it out and I'm just going to kind of see where my fingers lead on that.

35:48And having that ability to support that in the way you described, that gave me a very clear understanding from my standpoint. So I'd like to ask each of you where this is going, both in the large and in your specific concern with Codium. You know, things have never moved faster than they're moving right now in terms of how fast these technologies are progressing. And Daniel and I have a habit, we were commenting on our last episode about this, we have a habit of saying, yeah, we recently mentioned this thing and that we'd get to it, but then we turn around and we end up talking about that we just got there way faster than we ever anticipated.

36:25With the speed of generative AI, and you're already creating these amazing tools and stuff like that, and you're having to stay out front, where's your brain taking you at night? You know, when you, when you stop and you chill out and have a glass of wine or whatever you do, and you're kind of just pondering, what does the future look like? And I'd like to know both from your own specific personal standpoints in terms of your product and that, but just the generative AI world in general, how do you see it going forward? I'd love your insights. Yeah. I think, um, the classic question and then the grand scheme of things are like, Oh my God, is like generative AI just going to like totally get rid of my job or completely like invalidate and i think for us we will be the first people to say that you know we we do think like ai would just be like the next step in a series of at least in code or a series of tools that have had made like developers more productive right that have led them to be able to focus on more kind of interesting parts of software development and you know be an assistant right with all these tools are called ai assistant tools i think for a reason you know we're definitely not at a place yet, I don't think for a while, where there isn't going to be like a human in the loop, like in control, you know, guiding the AI and what to do.

37:36So from that kind of respect, like the doomsday scenario, I don't want to speak for a minute, but I think we're like pretty far from that mentality. But we do think, I think, you know, we wouldn't have gotten into Codium if we didn't genuinely think that there was just so many things that we do as a day-to-day as engineers that are just a little frustrating, boring, kind of take us out of the flow state, you know, slow us down. Those all seem like very prime, ripe things to try to address with AI. And I think that's kind of our general goal. I think there's a lot more capabilities to build. I don't think search, chat, these are going to be the last, I guess, building blocks that we build.

38:12We have more capabilities coming up that we're super excited about. But yeah, it's also going to be a thing where, as you said, this is moving super quickly. We have research, open source, There's like applications all developing at the same time at breakneck speed. And so I think part of what we're also looking forward to is like, how can we also just like educate like, you know, at least software developers on the best way to use AI tools, how to like best make the most use of it so that they are part of the way for it and that they also can get a lot of value. Well said. Varun? Yeah. Maybe if I was to just say like, you were asking me what the big worry is.

38:47For me, the big worry is there's going to be a lot of like exciting new demos that people end up building. And obviously for us as a company, we need to make strategic bets on like, hey, this is a worthwhile thing for us to invest in. For instance, I think a couple months ago, there was an entire craze on agents being able to write like entire pieces of code for you and all these other things. For us, though, we had lots of enterprise companies that were sort of using the product at the time and recognize that the technology just wasn't there yet. Right. Like take a code base that's like 100 million lines of code or 10 million lines of code.

39:18it's going to be hard for you to write C++ that's like five files that compiles perfectly and then also like uses all the other libraries when you have context that's like, you know, five files. It's not going to be the easiest problem. And I think that's maybe an example, but for us, we've currently, I would say just a pat on the back over the last eight months, iterated like significantly faster than every other company in this space, just in terms of the functionality. But we need to make strategic bets on what the next thing to sort of work on is at any given point. And we need to be very careful about like, hey, this is like a very exciting area, but is it like actually useful to our users, right?

39:53Like, is it actually useful in that, hey, like maybe we could do something where a great example is given a PR, we generate a summary. And I think Copilot has tried building something like this. And we tried using the product that Copilot had, and it was just wrong a lot of the times. And I think that would have been an interesting idea for us to pursue and keep trying to make work. But then there is a diminishing returns. And I think Anshul and I have seen this very clearly in autonomous vehicles, where we had a piece of technology that was kind of just not there yet. Like it needs a couple more breakthroughs of machine learning to kind of get there.

40:26And the idea of building it five years in advance, right, you shouldn't be doing that. You just 100 % shouldn't be building a tool when the technology just isn't there yet. And that is something that keeps me up at night is like, what are the next things we need to build while keeping in mind of this is what the technological capability set is like today, if that makes sense. It does. And it's a very practical AI perspective, if you will. So very fitting final words for the show today. Well, Varun and Anshul, thank you very, very much for coming on the show. It's fascinating. I got a lot of insight, a lot of new things to go explore from what you just taught me.

41:02And I appreciate your time. Thank you for coming on. Thanks for having us. Thanks a lot, Chris.

41:14Thank you for listening to Practical AI. Your next step is to subscribe now, if you haven't already. And if you're a longtime listener of the show, help us reach more people by sharing Practical AI with your friends and colleagues. Thanks once again to Fastly and Fly for partnering with us to bring you all ChangeDog podcasts. check out what they're up to at fastly.com and fly.io and to our beat freaking residents breakmaster cylinder for continuously cranking out the best beats in the biz that's all for now we'll talk to you again next time

From the publisher

Chris sat down with Varun Mohan and Anshul Ramachandran, CEO / Cofounder and Lead of Enterprise and Partnership at Codeium, respectively. They discussed how to streamline and enable modern development in generative AI and large language models (LLMs). Their new tool, Codeium, was born out of the insights they gleaned from their work in GPU software and solutions development, particularly with respect to generative AI, large language models, and supporting infrastructure. Codeium is a free AI-powered toolkit for developers, with in-house models and infrastructure - not another API wrapper.

Join the discussion

Changelog++ members save 1 minute on this episode because they made the ads disappear. Join today!

Sponsors:

  • Fastly – Our bandwidth partner. Fastly powers fast, secure, and scalable digital experiences. Move beyond your content delivery network to their powerful edge cloud platform. Learn more at fastly.com
  • Fly.io – The home of Changelog.com — Deploy your apps and databases close to your users. In minutes you can run your Ruby, Go, Node, Deno, Python, or Elixir app (and databases!) all over the world. No ops required. Learn more at fly.io/changelog and check out the speedrun in their docs. 
  • Typesense – Lightning fast, globally distributed Search-as-a-Service that runs in memory. You literally can’t get any faster! 

Featuring:

Show Notes:

Something missing or broken? PRs welcome!

More from Practical AI

All 157 episodes
A developer's toolkit for SOTA AIPractical AI · 42 min
Listen in VO