In short
Podcast Summary: Sourcegraph and the Frontier of AI in Software Engineering with Beyang Liu
Episode Overview In this episode of *Software Engineering Daily*, host Sean Falconer interviews Beyang Liu, CTO and co-founder of Sourcegraph. The discussion centers around Sourcegraph's powerful code search and intelligence tool that aids developers in navigating large codebases and the evolving role of AI in software engineering.
Key Topics Discussed
Sourcegraph's Core Functionality
- Code Search and Intelligence:
- Advanced search capabilities to find references, functions, and dependencies across repositories.
- Integration with development workflows to streamline code reviews and team collaboration.
- Historical Context:
- Beyang Liu discusses the founding of Sourcegraph, influenced by experiences in large enterprise environments, particularly at Palantir.
The Evolution of Software Engineering
- AI’s Impact:
- AI has the potential to change software development by generating code but the fundamental problem remains: software engineering does not scale efficiently.
- The "mythical man month" problem persists where adding more engineers to a software project can slow down progress due to increased communication overhead.
Challenges in Software Development
- Understanding vs. Writing Code:
- The bottleneck in software engineering is understanding and reading code, not just writing it.
- Generating code with AI can lead to increased technical debt if the generated code is not well understood by the team.
- Scaling Complexity:
- As software projects grow, managing dependencies and communication becomes increasingly complex.
AI’s Role in Addressing Challenges
- Reading and Understanding Code:
- Sourcegraph aims to enhance understanding through context retrieval and validation layers.
- AI can assist in code generation but must be paired with human oversight for effective implementation.
- Automating Code Reviews:
- Sourcegraph is developing a code review agent to automate and enhance the review process by applying defined rules and invariants.
- Use of declarative coding where rules are specified to maintain architectural vision and code quality without manual reviews for every change.
Future of AI in Software Engineering
- Automation Level:
- Beyang predicts that by 2025, AI will take a more central role in coding, reducing the need for constant human input while maintaining quality through validation processes.
- Impact on Junior Engineers:
- Junior engineers will need to shift their focus from traditional coding techniques to validating and understanding what AI-generated code does.
- New skills will emerge emphasizing testing and understanding outputs rather than just writing code.
Conclusion
- Sourcegraph aims to redefine how software is developed, tackling longstanding issues in engineering efficiency and effectiveness. The convergence of AI and human expertise is seen as pivotal for future innovations in the field.
Key Takeaways
- AI Integration:
- AI can assist in coding but should not replace the understanding and oversight required for maintaining code quality.
- Automation in Reviews:
- Automating code reviews can help enforce standards and reduce the burden on senior engineers, allowing them to focus on strategic development.
- Changing Skillsets:
- The rise of AI in coding may change the skills needed in software engineering, prioritizing validation and high-level thinking over traditional coding skills.
Closing Remarks Beyang Liu encourages listeners interested in tackling the challenges of navigating large codebases to consider Sourcegraph, as the company continues to innovate in the developer tool landscape.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Sourcegraph is a powerful code search and intelligence tool that helps developers navigate and understand large codebases efficiently. It provides advanced search functionality across multiple repositories, making it easier to find references, functions, and dependencies. Additionally, Sourcegraph integrates with various development workflows to streamline code reviews and collaboration across teams. Biang Liu is the CTO and co-founder at Sourcegraph, where he has worked for the past 12 years. In this episode, he joins the show with Sean Falconer to talk about the frontier of leveraging AI and software engineering.
0:37This episode is hosted by Sean Falconer. Check the show notes for more information on Sean's work and where to find him.
0:56Bion, welcome to the show. Great to be here. Thanks for having me, Sean. Yeah, absolutely. So you've been working on Sourcegraph now for over a decade. I guess like how have things evolved and what are some of the big pivotal moments in the company's evolution? Yeah, definitely. So, I mean, as with many things, it's like some things stay the same. Some things change drastically. Obviously, the impact of AI has, I think, introduced in everyone's minds the potential of changing the way that software is built. And there's certainly a lot more bot or computer generated code now, a lot more use of large language models to write all sorts of functionality across many different verticals and industries building software.
1:37But there's still some of the fundamental problems that still exist, which are kind of the same as the ones that we were originally started to tackle. So the fundamental problem of software development, not really having economies of scale. So like in contrast to every other industry, every other industry, you build a bigger factory, get more efficient, the stuff gets cheaper to produce as you scale. Software engineering is basically the opposite, right? Like as you grow, things get less efficient. So you've got mythical man month. Yeah, exactly. The mythical man month still applies today, even in the age of LLMs and AI that we live in.
2:14And, you know, this harkens back to the founding story of Sourcecraft, which is really born out of my co-founder Quinn and my work inside very large enterprise code bases. We were both engineers at Palantir. He was before deployed. I was just a regular engineer, but we're both going into the field a lot and working with a lot of Palantir's like Fortune 500 customers. I shouldn't say a lot. We worked with like the first two that Palantir landed. We were kind of on this like SWAT team of sorts that was trying to open up a new line of business for Palantir. And that was our introduction to very large, messy codebases.
2:52I mean, Palantir's codebase itself at that point was probably like five, six years old. So, you know, it wasn't as bad as the codebases that we were being deployed into, but it was still enough to have, you know, a bit of cruft, a non-trivial amount of tech debt. Was that their sort of like desktop version at the time, like basically running some sort of Java suite? Palantir's code base? Yeah, yeah. Yeah, yeah, yeah. So, well, actually, it was kind of a combination of things. So, you know, originally, the project I started on was like a Java thick client, because in those days, you have to remember, like this is before React, before kind of like the modern web development stack, Backbone JS and CoffeeScript were new at the time.
3:35So like that was the hot new thing. and so to get the richness of interactivity they needed for a lot of the like data analysis workflows that they're trying to enable java swing was was the only option when they got started but then by the time we started we got put on quickly on on this like upstart team and so we were using web technologies on that team so you know front-end javascript and and that sort of thing but we had to integrate in the context of these very large like banking code bases because those were the customers they were trying to go after and that was just a huge pain it was just like you couldn't get anything done.
4:07It took like weeks, sometimes even months just to like, ask around and acquire enough understanding of how the code fit together to even, you know, begin to think about like what you should be building or how that fit into the broader picture. And that was our first introduction that kind of gave us this insight into the scope of the problem, because like we were being deployed, Palantir was brought in to solve like very high priority problems. We were talking about like, you know, things that were costing banks like millions of per day that were on some level existential if they weren't solved within a finite amount of time.
4:41This is all kind of like hangover from the 2008 financial crisis and cleaning that up. And so, yeah, it was kind of crazy to us that these institutions had all the money in the world to throw at solving this problem, could not solve the problem. And I just think that speaks to the fundamental problem that faces any software engineering at scale, which is software engineering really doesn't scale. That's still the problem that we're here to tackle. So we've kind of tackled it for a lot of customers at some level. And I think being able to integrate LLMs and AI into our tech stack, I think, makes it possible to actually sort of solve the mythical man month problem in the next decade or so, which is really exciting.
5:19Can you elaborate a little bit on that? Like, how do you see AI helping solve that problem? Because presumably, I think the problem with scaling engineering is you end up with this large scale dependency graph, essentially. Yeah. With tight coupling, essentially, between people. Communications break down, everything just grinds to halt. Yeah. Could you potentially recreate the digital version of that if you were using AI as well? Because you're going to end up with a lot of these sort of digital dependencies, essentially. Oh, totally. And I think where the attention is today is all about AI-generated code, which is great.
5:52I mean, it's truly mind-blowing. that you can create a whole simple web application basically in one shot with any of the big frontier models today. You can do a lot with them inside existing code bases too, in terms of updating and editing the code with the right context. But I think the bottleneck, people still don't fully grok that the bottleneck in large-scale software engineering is not on the writing code side, It's on the understanding and reading code side. And so to your point, it's like right now, everyone's just like taking code out of the LM and basically pasting into their code. That maybe, you know, increases by an order of magnitude the volume of code you're able to generate in a given amount of time.
6:36But, you know, does that actually move the needle? Like each line of code in, you know, the engineering discipline, we tend to view as more a liability rather than an asset. Like the functionality you get from a line of code is an asset to your users, an asset to the business. But the fact that you had to write a line of code in order to do that is kind of like a liability. It's like a column in the tech debt stack. And so I think there's a very real world where like the amount of code in the world just like balloons. And there's all this like slop, which wasn't even written. There was never a human brain that conceived of it and understood how it worked, at least with like human written code.
7:08Like somebody at some point understood how it worked. With the AI generated code, like that might not ever be the case. And then it might even exacerbate this problem of like, hey, the existing code is a tangled spaghetti mess. to the point where, you know, the AI can't work in it. And I as a human can also not work in it. And so, yeah, I think that remains the challenge ahead of us. And that's one of the challenges that we want to solve. And that is sort of this like fundamental challenge of the mythical man month. So for those who haven't read the book, it's kind of like a classic about the nature of software development.
7:39It's core thesis. It was written in the seventies, basically, you know, someone looked at the progression of software projects inside these like mainframe companies back in the day and noted the fact that like there was this weird phenomenon where when they would add people to a software project in order to speed things up, it would actually end up slowing things down because it introduced just more communication overhead and a loss of kind of like coherence or unifying vision in the code base. You know, a lot of like abstractions that would clash with each other, things, pieces that wouldn't fit nicely together.
8:11And so the conclusion there was, you know, like software development, it's more like birthing a baby rather than like churning out widgets in a factory. And I think, I don't know if it's from this book or somewhere else, but it's like, you know, you can't have nine women make a baby in one month, right? Like there's no way to speed that up because there's all these like serial dependencies. And the task is fundamentally like, it requires like one person to shepherd it along. And the analogy applies to software. It applied to software development back then. It applied to software development, you know, in 2019 before ChatGPT, and it still applies today.
8:43But I think the, the possibility, the potentiality that we see and are excited about is that with the benefit of LMs, you actually have a lever around human intelligence to be able to go to like an architect or like a senior engineer and say like, we will give you this tool that you can wield to enforce the type of coherence and standards and consistency across the code base that you need to keep things clean as you add more contributors and you scale your engineering org. I don't think there's a way around the scaling because the fact of the matter is, if your software becomes successful, you're going to have more users get on it.
9:19You're going to have more customers. Those customers are going to have feature requests that tailor the software to their needs. And so you kind of have to spin up more people to build all those features. The challenge is how do you manage the complexity and maintain that coherence of vision as you add more contributors into the project? In that case, this idea, maybe you can leverage AI to do a lot of the code generation, but then you sort of have, you know, this like architect that's acting like the moderator or the orchestrator. They understand how the pieces go together. They understand what the standards are.
9:49And they're there to do essentially the understanding piece of this where AI is maybe not ready to help us with today. Is that the idea? Yeah, that's exactly right. In an ideal world, you would have like one very smart individual, you know, reviewing all the code that goes in your code base today, ensuring that it fits architecturally into the vision that they had in mind. Now, people don't do that because if you try to do that, you literally could not fit the amount of reviews you'd have to do. That person would basically have to spend all their time just doing that. And even then, they would only cover a fraction of the code that's being committed into the code base.
10:23On the generation side, I think that AI for coding, I think, is one of the areas where AI is finding a lot of success today. And And in a lot of ways, that seems surprising in some sense. But I also think that it also makes a lot of sense in a number of ways as well, because there's a lot of examples, there's a lot of code that's been digitized, that's available for training. And the other thing is that one of the hard problems, I think, when working with large language models is it's hard to kind of benchmark and eval them versus if you're building like a purpose-built model, you know what the inputs and outputs relatively should be.
10:56So I can essentially create a set where I can evaluate the model. I make changes, I can reevaluate the model. It's hard to do that with a general purpose model. But if I'm doing code generation, it constrains the universe and I can actually build like a relatively good benchmark and eval set to test against. I'm curious, you know, from your perspective, like, do you think that is one of the, you know, primary reasons why AI encoding is finding such success? Or is there also other factors that at play? I definitely think that that is one of the primary reasons why it's finding such success. I think like valid code is much more constrained, a much more constrained domain than valid language, because you have this kind of like Oracle that tells you first, you know, does it compile?
11:37If it doesn't compile, you have something that like tells you exactly the error that's preventing it from compiling. And then if it does compile, you also have these things called unit tests that validate the correctness of the function. And so that I mean, you can use that in a variety of ways, like you can use that inference time to verify and validate the output that you get to make sure that it's correct. You can also use it at training time. You have a way to generate a high quality set of synthetic training data by basically taking like the output of an existing model, running it, and then checking the errors that it produces when it attempts to generate code that conforms to the user query or prompt.
12:14And so you put those two and two together and you have a vector for making models much more precise and useful at inference time and also a way to improve the models at the training step itself, which I think a lot of the frontier labs have been doing. And if anything, it's almost like the fact that this exists for code, it's great for the coding domain, but it also probably generalizes like logical reasoning capability outside the coding domain. I remember when JetGPT was first released, there was a lot of speculation that a lot of the emergent capabilities came when they started training the model, not just on natural language, but also on code.
12:52Because there's certain logical referential patterns in code that map to logical reasoning that then can also generalize to logical reasoning patterns in natural language as well, which is really interesting. So that kind of helps us with the generation part of this. So what is Sourcegraph doing to kind of help with this understanding problem? Like if I'm using AI or even, you know, I have engineers, humans that are generating all this code, how do I actually, as that thing grows over time, maintain an understanding of it? Yeah. So there's kind of two ways where we differentiate in the understanding domain and make our product, you know, just like much more powerful and impactful on like big, messy production code bases.
13:37One is at the kind of like context retrieval layer. And then the other one is at the kind of like validation and verification layer. So the way to think about this is if you're trying to use AI to generate code, you know, first you want to give it an example so that it gets within the rough ballpark of what is right or what is reasonable in your code base. Like without the benefit of contextual snippets, it's basically analogous to asking a human to write code for your code base, but not actually letting them look at any of the other code that exists in your code base. So like, what would a human do in that case, they would just go to Stack Overflow and like, write code that, you know, fits your description, but uses just like the common open source frameworks that are well represented on GitHub or Stack Overflow.
14:21And that just doesn't fly within a lot of like messy production code bases, because, you know, the longer code bases existed, the more kind of divergent and special snowflakey it becomes. So the first thing we do is we use our code search engine, which is our first product, our original product, which is kind of like a Google for code and does much more than search. It also does like code navigation and it has access to a couple more like structured databases about code. These like ontologies of code and technical context. We bring those relevant snippets into the context window and that helps steer the LLM to generating something that is much more within this distribution of what is acceptable in your organization.
14:59So what we find is that there's a lot higher acceptance rates for the code that we generate inside the organization. It doesn't have to be like reworked or rewritten as much because it has the benefit of that context. And then on the second side of things is like the validation verification layer. We've been partnering with a lot of our enterprise customers to figure out how to like automate a lot of the review that is done inside their organization. So, you know, code review is this process that's, you know, kind of like half box checking exercise, but half like, you know, actually there's, there's some like useful things that we want to catch.
15:33And I think the way that humans do it, oftentimes it becomes like 98 % box checking and like maybe 2 % actually catching stuff because the box checking part just takes up all the time and everybody wants to get over with. So they can get back to like actually building stuff, which is more fun. And also something you get more credit for, like no one, no one gets a promotion for doing a great, you know, code review. Right. But machines can do it much more thoroughly into much higher quality than we can. So one of the things that we kind of got pulled into is we found all these customers of ours were basically hitting our APIs and building their own internal code review bots against them.
16:08And so we're like, wait a minute, there seems to be some commonality across these things. And so we started building a code review agent that's now in early access. I just gave a talk about this at the AI Engineer Summit in New York with one of our partners, booking.com. And the interesting thing here is I think it actually is, it opens up like a new paradigm for modifying or constraining what code goes into your code base. Our kind of like catchphrase way of describing this is like, it's almost like a declarative style of coding where you declare these rules and invariants that you want to hold in different parts of your code base.
16:42These are things that like your architect or your senior engineer could specify. Like I wrote this class this way. It should only ever invoke this other component in this other way. Or if you ever see this pattern code, this is an anti-pattern, rewrite it to be more like this. Change all your map functions to for loops or maybe vice versa, depending on what your overall philosophy is. Now you can define these rules in one place and just have them automatically enforced. You don't have to go and manually review every piece of code looking for these sets of patterns and anti-patterns. What's the representation of those rules?
17:14Are you using ontologies as the underlying representation? The representation right now is mostly just like natural language, plus a way of specifying like which set of files in the code base this rule should apply to. So there's kind of like a selector for like, hey, this is the subdirectory or this is the logical project where this rule should hold. And then the rules themselves, I think right now they're mostly natural language. I think in the next phase of this, there'll probably be like, some will be natural language and some will be natural language descriptions, but implemented by the invocation of like a precise tool.
17:46Like, you know, this particular pattern, the AST should never exist. I've described it in natural language, but like, you know, translate this to a, you know, tree sitter query or something like that. And that's how you can like test and verify. In the context retrieval or the code search, you mentioned that you're using ontologies there. So is that ontology like a base model representation of the concepts? Or are you starting with the base and then making something that's like customer domain specific as well? We think of it as like a knowledge graph that is built around the structure of the code.
18:19So like this basic skeleton of the knowledge graph is really the reference graph in the code. So you have definitions and you have references. And that's all code is at the end of the day. It's a bunch of definitions and then they're referenced in other places. and this is the knowledge graph that you're kind of like implicitly walking when you as a human go in and try to understand how something works right like you maybe like you do a couple searches to get within the rough ballpark but then after that it's like go to definition go to definition find references you know hover over the symbol to see what the docs are and so like that knowledge graph is very useful for humans it's also very useful for surfacing potential relevant context for lms as well.
18:58And then around that kind of like core code graph, there's other things that also come into play. So there's like a long tail of technical knowledge that is stored in things like issue trackers, corporate chat, production logs. And we've built a way into our system to integrate these other pieces of knowledge. So we had this thing called open context, which is basically like a simple protocol for saying like, hey, there's this knowledge source, you know, maybe it's your issue tracker, and here's how you query it, here's how you select the specific item is very strong parallels to MCP now, but kind of like predates MCP by year.
19:32We built an open context to MCP integration. I think we might like merge the two in our platform moving forward. But this was something that we had very early because we recognized the strong need to pull in relevant context from all these different other sources that were ultimately connected back to the code graph in some way. Like you have an issue tracker that pertains to a certain area of the code base. You have a production log that has a stack trace and all those things in the stack trace map to lines of code at some specific revision. So constructing that knowledge graph just allows you to have this map of how do I pull in all the relevant pieces of context when a user asks a question or wants to generate some code.
20:13Yeah. I think having developed something similar to MCP before MCP became public makes a ton of sense because if you're doing any of this stuff with relatively complicated systems, you're just going to end up with your AI having to pull in data from all kinds of different places. Like you can't have essentially a gravitational model for data where you're just dumping it all into a wake and then surfacing it from there. You have to essentially tap it directly from the source in order for it to be the latest information and not have to deal with people like re-architecting the data. You're not going to be storing your source code in your lake anyway.
20:47It's not going to make sense somewhere else essentially. So you have to liberate the data from all those different places. Yeah, totally. And it helps to have like a structured knowledge source, right? Like these two things are complementary, right? Like LLMs have this kind of like basic form of reasoning and the ability to integrate context. But what they don't excel at is structuring things down to like a very precise, you know, spec or description. And you saw this with like, you know, especially like the Gen 1 LLMs where like they had trouble emitting valid JSON. on. Now that after billions of dollars of training, they can finally do that.
21:23But it's still a very inefficient way to acquire that structure. It's far better to just put it into a symbolic representation and have the LLM walk that and then use that as a context fetching mechanism to complement its core memory and reasoning abilities. In terms of the commit review or PR review agent, can you walk me through what does that workflow? So I submit a commit of PR, this agent monitoring my PRs, then what happens essentially? So there's kind of like two steps to this. There's like the rule definition process, and then there's the rule enforcement stage. So the rule definition process just actually happens in these files, these like source graph rules files.
22:02And you can put these in a particular directory or subdirectory or repository, you know, wherever they're defined, you can define like a pattern of like what files these rules apply to. And so that's something that you define in your code base and you define once. And once you define them, we're actually building multiple hooks into the software development lifecycle where they are enforced. So one of the things we're doing is we're actually bringing these rules into context in the editor. So when you're generating code with AI in your editor, using our editor extension, the rules are part of the background context of the LLM.
22:36And so the LLM can take into account the rules when generating code. Now, that's not 100 % guarantee that they'll actually follow the rules. Oftentimes, they ignore them, but it's better than nothing. It saves you some time. It's kind of like shifting left this process of like getting feedback and constraining the code that you're creating to be valid or acceptable in your code base. Now, the second place where they're enforced is at code review time. So you push up your patch for review. And then after that, you tag in our code review agent. And then it goes and takes a look at the rules that apply to each hunk that's been changed and then post comments to that hunk, noting problems that break the rules or suggestions on how to fix the code so that it conforms to the rules.
23:19Is my understanding there, basically every source file has some set of rules that represent its essentially file-related context. And then you can use that as a way to essentially evaluate the commit against the lines that changed and what the rules are for that specific line, that specific context. It's not every file. It's more, I would say, at the subdirectory level. Okay. Because it might get kind of tedious to have to define. You probably will have a lot of duplicative rules if you had to define it per file. But you can define it, think of it as roughly analogous to an owner's file or a code owner's file where you're not necessarily defining that per file, but you can say like, oh, this subdirectory belongs to this project.
Read the full transcript
23:59And these are the set of rules that should be enforced for this project or only apply this rule to TypeScript files or TSX files in this code base. Is that how you help essentially limit the size of what potentially has to go into the context window? Yes. Like we're not shoving the entire rule set across the entire code base into the context window to review, you know, a single hunk. We sort of fetch the rules that are kind of like relevant to that particular part of the code. And those rules, are those LM generated to start with or are they manual? They're human written. Okay. Yeah. So that's kind of like key step to having something like this set up.
24:35Yeah, exactly. Again, the basic idea here is to give your senior engineers, your architects, the people who want to define the vision for how the code base architecture should be. It basically gives them a lever to enforce that vision. because the pattern that happens across every large scale code base is you have like your senior engineers who have, they have enough scar tissue to know, you know, there's certain like patterns and architectural designs that like you should incorporate to make the code base, you know, workable to make it continue to be the type of place where like people can contribute.
25:11But then you have like junior engineers who have less experience. They're just focused on like building the feature. And they often commit code that breaks those architectural constraints. And it may check the box for what you want to ship in that quarter or in that sprint cycle, but you're taking on tech debt. The code base itself is getting messier. And the incremental next feature gets tougher to build. And so it's this losing battle because the juniors outnumber the seniors. And sometimes the worst case scenario is that when the code base gets so messy, you have to do something like a large scale, you know, code refactor or code migration.
25:46And these can stretch on for like months, even years. Like we have a customer that had a feature flag retirement initiative that was scheduled to take, I think it was like 11 years or something like that. It's kind of like this mind blowing, like you were like, that's like, you know, at least one era of tech, like who knows whether we'll still even be around, you know, then, But like those sorts of projects have to be led by senior engineers because they're the only ones who have the historical context and the technical depth to execute them. But the problem is like once you put your senior engineers on those projects, then like who's building stuff for the user?
26:22Who's actually improving the UX? Like your best and brightest minds are focused on, you know, future flag retirement instead of like building an awesome user experience. And so like, yeah, the vision for the code review agent and just more generally like the declarative system of rules is to give these senior engineers a way to express these constraints and just have them automatically enforce so that they don't have to spend 100 % of their time enforcing these rules. They can actually like build new things as well. So in most of these situations that we're talking about, this is, you know, an existing code base that was built sort of from scratch by humans.
26:58Yeah. So people have this historical context. They know the decisions that were made to get to the place that they're at. They know sort of those like rules that they want to enforce. What about for net new code? How does it work in that situation where I'm starting from scratch and I'm going to leverage AI to do a lot of my generation? And does that sort of impact person's ability to know what's going on and create those types of rules down the road? Yeah. So first of all, I would say like in many organizations, they're like no code is truly an island. it's kind of like even if it's a new project it needs to talk to some existing system needs to integrate with some existing system maybe there's like a common framework that is the way you do things because if you do it that way you get like observability for free you get robustness for free get scale for free inside that organization right because like that's the platform that they built internally so there even if you're starting from scratch it's not truly from scratch and you still want the benefit of, you know, the context of other similar projects so that you can like pattern match against those.
28:01And also the context of what the rules and constraints are, like how to, how to use particular APIs, you know, common pitfalls or foot guns that you want to avoid. And, you know, things like that for truly, truly like from scratch problems. Then I would say that, you know, rules are less relevant, right? Like if you're just creating like a flappy bird clone that you want to share on Twitter and get a lot of likes, then it's like, who cares? Right? Like, So I was talking to someone at one of our customers recently who had this very insightful thing to say, which was technical debt is only debt when you have to touch the code that's messy.
28:36So his idea is like if you're writing a throwaway piece of code, like if all you want is like for that code to do this one thing, you're not going to come back and have to like upgrade it or modify it later. Then, yeah, who cares about how messy it is as long as it does the job, right? That's not a one shot app. Yeah, it's a one-shot app. It's a single-use app. That's not real technical debt because you don't care. You're never going to have to dig through that. It's not going to cost you any time in the future because you're never going to dive into that code again. So in those cases, who cares about the rules?
29:03If I bit into existence and then if it works, it works, then you go on and forget about it. But most of the software we use day to day is not like that, right? Like if you're actually building something that is going to be driving, you know, millions of users or, you know, hundreds of millions of revenue through it, you're going to want to evolve it and change and update it over time. And for that, you do need to think more thoughtfully about like the architectural considerations. You know, it's not like you need to define all the rules up front. You know, these things sort of like evolve organically over time.
29:32But I do think you want to have a system in place such that, you know, when you start to scale when your software becomes successful and you start to hire more people or bring more contributors into it, you want to have a lever, I think, as like the keeper of the vision of that code to be able to like enforce your vision across all the new minds that are going to be like ramping up on it. And presumably a senior person is coming with their own sort of history of projects that are maybe not the same project, but they're going to apply some of those patterns essentially in a net new project, essentially.
30:05Yeah, yeah, exactly. Around the feature flag retirement problem, where you end up putting your best and brightest on it, because it touches so many different systems, you need someone who has sort of like a very comprehensive understanding of the code base and the impact if I remove something. Can AI help us with that? Yeah, it absolutely can. In fact, that was another use case. This is another agent that we're building in partnership with this customer. It's not a review agent. It's actually an agent that can run a large scale code migration. in this case, like the feature flag retirement thing.
30:35There's kind of a spectrum of how difficult these problems are to tackle, right? So like initially we went into this thinking like, oh, it's just dead code removal. Like you could almost do that pre-AI, right? Just like look at if this is referenced and just run one of those like standard like AST level checkers and remove it. You know, why can't you do that? And then we started like whiteboarding this out and it became clear to us that this was like far from trivial because essentially like all these feature flags had these like implicit dependencies. And, you know, when they walked us through like how humans were cleaning things up, it was like, okay, you had to find location of the feature flag, but then you have to like kind of walk up the application step, you maybe cross an API boundary and see if, you know, this thing is triggered.
31:17And sometimes it goes through like, you have to identify like the set of APIs that are like in the middle of the path between like the thing that you can change and the thing that, you know, you want to actually like influence in the end user experience. And then you have to like tag those as well. So it became this like kind of involved task. But what we realized was like, there was sort of like an 80-20 rule that applied here. So like maybe like 80 % of the sites that needed updated, you know, they weren't at the level where you could just like AST remove them, but they were probably at the level where, you know, a simple agent with some kind of like heuristic criteria plus like a decently intelligent LLM could probably figure it out.
31:57You know, like 80 % were quote unquote easy. And then you had maybe like, you know, 80 % of the remaining 20 % that were kind of like medium. And then you had this final like 2 % that were really hard. And so the way what we proposed to them was like, why don't we just take the 80 % that's easy first? And, you know, that will substantially reduce the amount of like tech debt and make it so that you don't have to like, because like tech debt breeds tech debt, right? Like if you're, if you'd hack around like an existing thing, oftentimes you add spaghetti code to work around the existing spaghetti code and you just end up with more spaghettis.
32:28Why don't we tackle the easy 80 %? Then we'll build a more involved agent for the remaining 80 % of the 20%. And then for the last 2%, maybe that still needs to be manual. But by now, we've scoped it to 2 % of the original problem. And so instead of taking 11, 12 years, we can bring this down to a year or maybe less than a year end to end. Yeah. I mean, I think that also, if 80 % of the cases are relatively simple, it's a heavy lift on your resources to have sort of your best engineer, like doing the easy stuff, right? So if you can alleviate that pain through automation, that's hugely valuable.
33:04I think it comes down to intelligent deployment of automation. What can you do reliably essentially with this technology? And then for the things that are not reliable, rely on the human in the loop expertise to go and do that work. In my PhD work, I worked on large scale data integration problems. And this is years before large language models and so forth. But a lot of what I did there was bringing together essentially sort of standard ML to solve the easy part of the integration problems while using sort of leveraging human in the loop to do the hard part of it. So I think these systems are like highly valuable and they'll probably be here for a long time.
33:39Sort of this like human in the loop plus AI system. It's like you want center chess level, essentially. Like how can I use AI plus a person to do something that is not possible by a single person to do? Yes, yes. Yes. And at Palantir, we used to call this human computer symbiosis. I think that was something that Shom, who is now I think the CTO, I think he coined that term. And it's exactly right. Like the stuff that you can do a combination of like human plus computer is always going to far exceed what you can do with computer alone or with human alone. And I think that's the thing that maybe some people don't realize about AI.
34:10It's like AI, it definitely moves the frontier of like what the computer can do. But all that means is that you can just do more with a combination of human plus computer. It's not like a zero sum game. I think it's just like it grows the pie of what we're able to do as like a species. I was having this conversation the other day that if you suddenly you could have essentially billion dollar startups that were a single person, because they were able to leverage all this AI to do all these different things. Well, that just means that you have a lot more companies doing like amazing things. It's not like there's less companies just because you can do it with less people.
34:40Yeah. It's like today, if you could like teleport them back in time, there'd be a billion dollar company powered by like one person. But like in the future, they're going to be in a competitive landscape where like if one person can do that, then so can like 10 other one person companies. And so it just means that like on the consuming side, we're just going to get a lot more like useful stuff. So in terms of AI encoding today, like a lot of it is, and this is true, I think of even outside of coding, but a lot of it's like assistive technology, essentially, like it's a co-pilot, there's still a person involved.
35:11How far away do you think we are from having fully autonomous written code that's a large part of our code bases? Oh, I think that's already happening. So that's what we're working on right now in conjunction with these enterprise partners of ours. So this kind of automated... Actually, I forgot the final step of that rules-based system, which is you have the rules that are enforced in your editor and then at review time. and then it's kind of like what do you do to keep the code base into that state and so our vision is just to have like all these like bots and agents in the background constantly doing things like updating to the latest version because a lot of these rules are not like static a lot of the rules are just like you know use the latest stable version of you know the npm packages right and like that in order to keep that rule updated you have to be like constantly doing stuff in the background.
36:02And right now you don't do that nearly as often as you should, because that takes time. And oftentimes it's like senior engineering time, which is probably like the most valuable, non fungible resource inside your organization. But if you can have a computer do it, you can just have a computer do that, like fix whatever, like breaking changes happen. And over time, I actually think like the vast majority of code written today kind of falls into that bucket, where it's like, it's not really interesting code. It's not creative code, but it's just like glue code or it's a configuration code that needs to be updated.
36:33And in fact, some of our customers with the highest percentage of AI generated code are using it specifically for tasks like that, where it's like updating configuration or updating these things that are like kind of boilerplate-y or like not that interesting, but still like critical, like they're on the critical path for keeping things like clean and secure and stable. So I think that world is kind of already here today in a lot of the organizations that we work with. I think what we'll start to see in this year specifically is much, much more automation within the editor. So I think now like the latest generation of LLMs, you know, like we're recording this on February 28th.
37:12And like within the past month, there's, there's already been like a couple new models that have dropped that I think substantially move the needle as to what they're capable of. And so I think like 2024, we're still very much in like the human in at every stage of the loop, like code assistant mode, I think 2025 is when we kind of like shift into like, oh, now it's more like the LM taking the driver's seat for a lot of these things. And then a human just needs to pop in every now and then when you need to do something like that's a very like special or specific or like, you know, quote unquote, out of distribution of what the LM has seen in their training.
37:47How do you think this changes the nature of like the junior engineer's job? Like how do they, if they're leveraging these tools primarily, right? How do they get to a place where they have sort of that senior level understanding? I think it definitely changes the path to getting there. I think there still is a path, which is at the end of the day, you have to like validate like you as a human pushing the code own the responsibility of ensuring that it's correct. And so how do you validate that? In the pre-AI world, you kind of validate it through the process of writing it. It's like going through the process of like thinking through how to write the code gives you a certain understanding and a certain confidence in the correctness.
38:20and in theory you write unit tests too to validate that but like you know people do that infrequently and unit test coverage is always like spotty right so like the pre-ai way of doing this is like well you wrote it so your brain must have understood it at some point so we have a reasonable amount of confidence that it's like mostly correct yeah and someone reviewed it sort of reviewed it yeah yeah yeah and then i think the post-ai world looks more like well the ai i vibed this code into existence the ai generated it i don't really know how it works but how do i validate that it does work at a sort of like gray box level, right?
38:52Like, and so then now let me generate a unit test with the AI and have the AI explain what the unit test is, is testing for. And let me augment the test cases. I can actually generate a far more comprehensive test suite in a shorter amount of time now with an LLM than I could previously. So I'm actually going to do that now. And if it finds bugs, you know, maybe that's the point which I like kind of like dive into the weeds a little bit and understand what's breaking and how to fix it. So I think there still is, I'm not as pessimistic as some other people who are like, oh, you know, the bottom rung of the ladder has been eliminated.
39:23I think there's still like a way to kind of like, understand what's going on, like, you're still gonna have to do that for some percentage of things, because you have to verify the correctness. I just think that like the bar for test coverage and correctness verification just gets higher now, because higher is now feasible. And we'll just have like a different way of like learning code understanding capabilities. Like, I do think that like the next generation of programmers is going to be much less good at what I call like linesmithing, which is like writing new code from scratch line by line.
39:54The more important skill moving forward will just be like thinking about it at like a higher level, at least like a function level of like, what are the inputs and outputs I want to verify? And how do I test this properly? Yeah. I mean, it probably changes the level of abstraction. So you're working at a level of abstraction earlier in your career than maybe you are now. And that's been going on for a long time, Like even as you move to higher level languages, like people were arguing about assembly versus C, C versus Java, Java versus Python and what you're giving up by, you know, using these sort of higher level languages.
40:27And we've been having that debate for 50 years. So, you know, in a lot of ways, this is, it's a bit of a step function, but it's kind of like the next, you know, natural evolution of that is where now the coding interface becomes natural language. It shifts the job to like, how do I validate that this is correct? how do I stitch these boxes together in a way that makes sense that's going to be scalable and all this sort of stuff. Yeah. Like what I analogize it to is like, you know, when I was in grade school, they still taught us how to, you know, navigate the Dewey decimal system in the library, right?
40:55Because like in those days, that was a critical skill for any knowledge worker to have. Like at some point, you're gonna have to like read a book to acquire the knowledge. And that book is going to be like buried inside some vast, large building that you have to like go to. And then you're gonna have to like, go through all these index cards. And you know, It's going to take you a while to find that one book. And these days, it's just like, who does that anymore? You just go to Google or ChatGPT and type in your question. You get an answer. Yeah. Amazingly, this is the second podcast I've done in the last month where the Dewey Decimal System has come up.
41:26Something I haven't thought about since I was probably eight years old. Yeah. I mean, it's like anyone who went to grade school in the 80s or 90s, that's part of your pre-training data, right? This was also funny enough, someone who had previously worked at Palantir. So Palantir is incepting the idea of Dewey Decimal System, bringing it back. Well, we're getting close on time here. Is there anything else you'd like to share? No, I mean, the only thing I like to say to folks listening is we're Sourcegraph. We build developer tools for massive production code bases. I think we are well positioned to kind of define how software is built in the next era.
42:01We want to solve this problem that has plagued software development since its inception. It's like mythical man month level problem. So if that's of interest to you, we are hiring. And if you are one of those engineers that is toiling away inside a large, messy code base, give us a look. We always love hearing from people who have like challenges or pain points that we might be able to solve. Awesome. Well, beyond, thanks so much for being here. Cheers. Thanks, Sean.
42:37Thank you.
From the publisher
Sourcegraph is a powerful code search and intelligence tool that helps developers navigate and understand large codebases efficiently. It provides advanced search functionality across multiple repositories, making it easier to find references, functions, and dependencies. Additionally, Sourcegraph integrates with various development workflows to streamline code reviews and collaboration across teams. Beyang Liu is the CTO
The post Sourcegraph and the Frontier of AI in Software Engineering with Beyang Liu appeared first on Software Engineering Daily.
