Agent Swarms and Knowledge Graphs for Autonomous Software Development with Siddhant Pardeshi - #763

10 Mar 2026 · 1 h 16 min · 30 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Episode Summary: Agent Swarms and Knowledge Graphs for Autonomous Software Development with Siddhant Pardeshi - #763

Podcast Overview Title: The TWIML AI Podcast Host: Sam Charrington Guest: Siddhant Pardeshi, Co-founder and CTO of Blitzy Episode Focus: Autonomous software development systems capable of producing production-ready code at scale.

Key Themes & Discussions

Introduction to Blitzy

  • Blitzy's Mission: To enhance software development speed through autonomous development systems.
  • Core Technology: Utilizes a hybrid graph-plus-vector approach to enable AI agents to efficiently navigate large codebases.

The Concept of Autonomous Development

  • Code as a Commodity: Siddhant emphasizes that the ability to generate code is not the challenge; ensuring acceptance and adherence to standards is key.
  • End-to-End Autonomy: A distinction between AI-assisted coding and fully autonomous systems that can autonomously generate tested and validated software.

Challenges in Current Development Paradigms

  • Code Acceptance: The primary challenge is not just generating code, but ensuring it meets security, maintainability, and testing standards.
  • Complexity in Existing Codebases: Differences between greenfield (new projects) and legacy systems present unique challenges.
  • Human in the Loop: Despite advances in automation, human oversight remains crucial, particularly for complex tasks that require nuanced understanding.

Agent Engineering and Context Management

  • Dynamic Agent Personas: Agents can adapt their functions based on the task requirements rather than being static.
  • Context Engineering: Importance of providing agents with the right context to minimize information loss over extensive tasks.

Hybrid Graph-plus-Vector Approach

  • Graph Database Utilization: Aids in mapping relationships within codebases, enabling efficient querying and navigation.
  • Dynamic Recruitment of Agents: Multiple agents can be orchestrated flexibly to handle tasks in parallel, improving efficiency.

Evaluating Model Performance

  • Real-World Evaluations: Siddhant stresses the need for evaluations that mimic real-world scenarios rather than relying on traditional benchmarks.
  • Complexity Assessment: Tools are used to evaluate code for maintainability and security, beyond just functional testing.

The Future of Development

  • Continuous Learning and Improvement: Blitzy’s system benefits from user feedback and ongoing interactions, enhancing its performance over time.
  • Monitoring Trends: Observations on the future of autonomous development and the competitive landscape among AI models.

Key Takeaways

  • Autonomous Development is Evolving: Advances in AI can significantly reduce development times while ensuring quality.
  • The Importance of Context and Agent Adaptability: Effective context management and dynamic agent roles are essential for tackling complex software projects.
  • Evaluations Matter: Real-world performance evaluations are crucial in determining the effectiveness of AI in software development.
  • Documentation and Transparency: Ongoing documentation and communication with users help alleviate concerns regarding the use of AI in code generation.

Conclusion This episode highlights the transformative potential of AI in software development, as discussed by Siddhant Pardeshi. The conversation emphasizes the need for companies to adapt to new technologies while maintaining standards of security, maintainability, and efficiency in code generation.

Further Reading: For more insights and case studies from Blitzy, visit their website or follow their updates on social media.

---

For complete show notes and additional resources, visit [TWIML AI Podcast Episode #763](https://twimlai.com/go/763).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Recruiting Swarms of Agents

0:48 to 1:30

Discover how Blitzy uses multiple swarms of agents for efficient coding.

“and use the database as part of the orchestration layer.”

Siddhant's Journey to Blitzy

1:50 to 4:12

Explore Siddhant's background, NVIDIA experience, and the inception of Blitzy.

“I am excited to meet you and I'm really looking forward to digging into your experiences at Blitzy where you're working on autonomous development.”

The Impact of AI on Software Development

4:12 to 6:20

Understand the transformative effects of AI in software development.

“It certainly is true that one of the areas where AI is having the most impact today is in software development.”

The Challenges of Code Acceptance

6:20 to 9:30

Learn about the complexities of achieving code acceptance in AI-generated code.

“Code that is ready for production is a completely different story, right?”

Engineering Context and Agents

9:30 to 14:00

Dive into the concepts of context engineering and agentic engineering for better AI outcomes.

“And that also is an exciting opportunity for the future.”

Introduction to Agentic Engineering and Context Engineering

14:00 to 14:54

Learn about agentic engineering and context engineering in AI systems.

“So that's the part that is agentic engineering where you recruit the right agent with the right set of prompts and tools, with the right level of prompt engineering for the right task, right?”

Understanding Autonomous Development Workflows

14:54 to 17:08

Explore the process of autonomous development and its workflow.

“Like, talk through the process from the perspective of a customer or user and what they're doing and what, you know, what they see, the way in which they're engaged.”

Challenges in Traditional Development Methods

17:08 to 20:32

Examine the limitations and challenges of traditional software development.

“And then if you were to put that plan forward, eventually when it tries to compile, it will make a mistake.”

Innovations in Code Base Navigation and Search

20:32 to 23:26

Learn about new techniques for navigating and searching large code bases.

“overcome, you know, all these many challenges?”

The Future of Multi-Agent Systems in Development

23:26 to 28:00

Discover how multi-agent systems can optimize software development workflows.

“But it's like you're able to reduce your search space using semantic.”
Show all 30 chapters

Dynamically Recruiting Agent Swarms

28:00 to 29:09

Learn how to efficiently recruit and orchestrate multiple AI agents to execute tasks.

“Like I just said earlier in the call, you can't even build a 100K line C compiler with Cloud Code, even with Opus 4.6.”

Concurrency Challenges with Multiple Agents

29:10 to 31:09

Explore techniques to prevent conflicts and ensure smooth operation among numerous agents.

“And we've been able to apply that successfully.”

Graph Databases in Agent Design

31:10 to 33:54

Understand the role of graph databases in managing code dependencies and agent operations.

“And what is the version of that library?”

Dynamic Agent Design and Persona Development

33:55 to 36:10

Discover how agents dynamically adapt their roles and the importance of persona in performance.

“So we write the tools and then the agents look at the spec and even like portions of the spec, right?”

The Role of Prompt Engineering in Agent Performance

36:11 to 38:45

Learn how prompt structure and agent persona impact performance in complex tasks.

“And we saw a lot of that, I think, early on.”

Evaluating Agent Performance and Context Dependence

38:46 to 41:00

Examine the importance of evaluation methods and context in understanding agent capabilities.

“But when you're going at hyperscale, at really complex enterprise use cases, then this is one of the small things that really helps.”

Limitations of Agents and Performance Benchmarks

41:01 to 42:00

Understand the limitations of agent frameworks and the challenges of performance benchmarking.

“with agents is task and context dependent.”

Real-World Model Performance vs Leaderboards

42:00 to 44:09

Learn how real-world performance of AI models differs from leaderboard metrics.

“So they're now testing on Sweebench Pro.”

Importance of Effective Evaluations in AI

44:10 to 46:39

Discover the significance of robust evaluation techniques for AI model development.

“So we're also trying to build our own, like trying to make public our own internal evals and contribute to this space.”

Task Decomposition and Model Selection Challenges

46:40 to 48:56

Understand the complexities in deciding which AI model to use for specific tasks.

“And then you run the model on that task.”

Evaluating Software Autonomy and Quality

48:57 to 52:40

Explore how to assess the quality and autonomy of code produced by AI tools.

“So we've talked quite a bit about how you approach automated or autonomous development.”

Maintaining Code Quality and Security

52:41 to 56:00

Learn about the importance of maintainability and security in AI-generated code.

“But taking a step back to what you talked about, right?”

AI in Code Quality and Documentation

56:00 to 56:52

Learn how AI enhances code quality and addresses documentation gaps.

“But then, you know, we've, again, because you have the graph database that can calculate the relationships between the code, between, you know, code, you understand what's going on in every single line of code, right?”

Preventing Development Failures with AI

56:52 to 59:20

Discover how AI checkpoints improve software development success rates.

“So we've incorporated these things such that when you get code back, it's maintainable, it's well-documented, it checks all the boxes for security, we've checked against everything using web search and all of that.”

Human Elements in AI Development

59:20 to 1:02:42

Understand the role of humans in the AI development process and their concerns.

“Stop development, test everything, fix gaps, then move forward.”

Managing Risks in AI-Powered Development

1:02:42 to 1:07:26

Explore how to address enterprise concerns regarding code deployment risks.

“if you design a system around the humans, and you have a human in the loop, it is extremely difficult to take the human out of the loop, right?”

Adapting to Rapid Technological Changes

1:07:26 to 1:10:00

Learn how to build technology that evolves with fast-paced advancements.

“you can write tests that matter to you, right?”

Exploring Autonomy in Software Development

1:10:00 to 1:11:04

Learn how Blitzy enhances software development using a self-reinforcing knowledge graph.

“Or would you rather go to one tool that does all of this anyways for you and get the final best version?”

Graph Databases and User Feedback

1:11:04 to 1:13:18

Discover how graph databases manage user feedback efficiently in autonomous systems.

“We get signals if you accept the PR, if you make edits, all that stuff.”

Indicators of Autonomous Development Advancements

1:13:18 to 1:15:18

Understand the key indicators to track in the evolving landscape of autonomous software development.

“You don't cross the thresholds of effective context.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Sam Charrington:A big thanks to Blitzy for supporting the podcast and sponsoring this episode. Want to accelerate software development velocity by 5x? You need Blitzy, which brings autonomous software development to your enterprise codebase. Your engineers declare intent and Blitzy agents map your codebase and generate an agent action plan. Once approved, Blitzy gets to work, autonomously generating hundreds of thousands of lines of validated, end-to-end tested code. More than 80 % of the work completed in a single run. Blitzy is not just generating code, it's developing software at the speed of compute. Experience Blitzy firsthand at blitzy.com slash twiml.

0:42Sam Charrington:That's B-L-I-T-Z-Y dot com slash twiml.

0:47Siddhant Pardeshi:The approach that we took has been to dynamically recruit multiple swarms of agents and use the database as part of the orchestration layer. And you can recruit tens of thousands of agents, but not have to worry about this single orchestrator that's keeping track of everything that's happening. We've been able to apply that successfully. And we frequently write hundreds of thousands of lines, millions of lines of code. Everything compiles. Everything runs. All tests pass. The UI works. It's pixel perfect. And so we've perfected that really.

1:30Sam Charrington:All right, everyone, welcome to another episode of the Twimble AI Podcast. I am your host, Sam Charrington. Today, I'm joined by Siddhant Pardeshi. Siddhant is co-founder and CTO of Blitzy. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Welcome to the podcast, Sid.

1:50Siddhant Pardeshi:Thanks, Sam. Glad to be here. I'm a longtime listener. I've been listening since 2019.

1:54Sam Charrington:That's amazing. And it is so great to hear. I am excited to meet you and I'm really looking forward to digging into your experiences at Blitzy where you're working on autonomous development. So let's dig right in, but start by talking a little bit about your background. You were at NVIDIA before you started, Blitzy?

2:18Siddhant Pardeshi:Yeah, I was at NVIDIA since 2016, January 2016. And back then, the day I joined, NVIDIA's stock was worth$32 billion. That was NVIDIA's market cap,$32 billion. And I think Anthropics revenue today is more than that. It was quite an experience, you know, being at NVIDIA at that time. And NVIDIA was structured. I don't know if they still are, but it functioned very much like a startup for the entire time that I was there, right?

2:53Sam Charrington:From 2016 to 2022.

2:58Siddhant Pardeshi:And I, you know, when the attention is all you need, paper dropped. I was right there. You know, I was inventing things for NVIDIA in the generative AI space. I was deep into GYANS or generative adversarial networks and various autoencoders. And I was brushing with NLP. It was still quite earlier. You had BERT. We were using BERT for translation and stuff like that. The transformer was ground-making tech. And eventually, when I realized the potential of what it could do and simultaneously had an opportunity to go to HBS to do a joint master's program in an MBA and an MS, I chose that and I met Brian at HBS, my co-founder and CEO, and we decided to form Glitzy based on the idea that AI will...

3:45Siddhant Pardeshi:catch up eventually with humans. And, you know, we made this bet back when the context window was about 10 ,000 tokens and it could barely write like usable code, right? But we made this bet that AI is going to be as good, if not better, than humans at writing code. And there'll be a section of software development that's not just about code generation, but entire software engineering that will get completely automated by autonomous development. And that's what Blitzy is all about.

4:12Sam Charrington:It certainly is true that one of the areas where AI is having the most impact today is in software development. When you think about software development, do you have a way that you taxonomize the space and the opportunity?

4:29Siddhant Pardeshi:So I think software development is the best opportunity in space to apply AI. And the reason for that is because software is verifiable. It's compilable. It's testable. You can visualize it. And there is the concept of a correct answer. There could be many correct answers, but there are correct answers and wrong answers, which is not always the case in other domains, right? So it's super important to realize that. And then if you think about the space itself, I think we all got started with AI-assisted development, right? You had copilots. Today you have CLIs and IDEs, ID tools with embedded AI assistants.

5:13Siddhant Pardeshi:And they all have the ability to, for example, do tasks asynchronously. Like, for example, you can give it a job that will take even an AI, maybe like hours to

5:24Sam Charrington:complete and it will think for some time, go off asynchronously, ask you follow-up questions and whatnot.

5:30Siddhant Pardeshi:And then you have another part of the space, which is about autonomous development. There are tools in this category. There's, I believe, Devin from Cognition that falls into that category. We operate in that category. And the idea here is that you hit build and outcomes of PR, right? But the PR that comes out is already tested, validated, everything works, and it's exactly how you intended it to be, right? There's no errors. The code is acceptable, right? So the biggest challenge that we have on both sides of the spectrum is code acceptance, right? You can write a lot of code, and code is a commodity now.

6:09Siddhant Pardeshi:Getting AI to write code is very easy. Getting any code is easy. Getting code that follows your standards, It's code that is really good. It goes as secure. Code that is ready for production is a completely different story, right? Because you have, on one hand, you have these greenfield builds or like new products that you can build from scratch. And AI is really good at that. So if you look at the demos that the labs put out, hey, I built this game and it looks amazing. I can't believe it. But then when you put the same AI on an enterprise code base and you're supposed to work with the existing group.

6:47Siddhant Pardeshi:It's a lot more challenging. It's way more challenging. It's an orders of magnitude high problem because the AI is dealing with so much information and so many and conditions that causes tools to fail. So the autonomous part of the spectrum is a much harder challenge because you have to simultaneously address all of these items and work for acceptance as your final metric.

7:11Sam Charrington:And so thinking back from acceptance through the agent, the AI, writing some code. On the other side of that, there's got to be some specification that the code has to meet in order to be accepted. Are you essentially pushing all the complexity of coding into spec development?

7:36Siddhant Pardeshi:That's a great point. So yes and no. Let me explain the yes part. But like, if you could write a spec, then you should write a spec, right? It's all of the tools we know and love have plan mode. You know, they've recently completed that. Everyone's realized that. We started doing that back in 2023, 2024, when we built Blitzy. But spec development, this really helps the agents anchor themselves. But again, what people realize immediately is that the spec is not good enough because then you have these other general rules that you want agents to follow. and traditionally what people have done is use things like agents.md, added skills and other stuff, trying to keep the spec lightweight because these models tend to forget right after a period of time or if you go through compaction and stuff like that.

8:26Siddhant Pardeshi:So there's that part where if you have a task in general that you can write a spec for, you know what it should do, you know all of the conditions that it needs to satisfy, then yes, writing a spec is great. But then you have this other class of tasks where it's not really clear what the dependencies are. Like, for example, you don't know what the schema for the backend database looks like. And you can't write a spec for it because you don't know what the constraints are, right? And you can't just trust the AI to, okay, figure out the schema and then do X because you're going to get new information when you figure out the schema, right?

8:59Siddhant Pardeshi:And that's going to affect the decision and the architecture of the code that you're writing, right? So because of that, you always have this spectrum where you're working one-to-one with the agent. it's giving you more information and you're helping it make decisions, right? So does the future, if you can build more intelligent models that are maybe human-like or better than humans at making architectural decisions, then yes, you can have like this entire class of work that is focused on writing specs and guiding other maybe less capable, cheaper, faster agents to write code. And that also is an exciting opportunity for the future.

9:35Sam Charrington:So just to replay that to make sure I understand, And I think what you're saying is that, yes, the spec is important because if you get the spec right, that anchors the agent and the agent can produce better code. but no today a spec isn't sufficient because there are assumptions and unknowns and things that evolve during the course of development uh and so rather than pushing everything to the spec what i heard in there was that there's still a lot of human in the loop during uh during development, which raises a question, well, A, is that right? Is that capturing what you're saying? But also then B, you know, you talk a lot about this idea of autonomous development.

10:31Sam Charrington:If the human is in the loop, how autonomous is the development? How do you think about that distinction and nuance?

10:38Siddhant Pardeshi:Yeah, yeah, that's a fantastic question. So the thing is, even if today, let's frame it this way. So today, if you want to get, even if you have a great spec, you spend a lot of time I'm writing a spec. But if it's a complicated spec that maybe covers 50 ,000, 100 ,000 lines of code, right? For which enterprise projects are often at that scale if you want to migrate, if you want to upgrade Java for a large code base, right? Or if you want to add a UI on a complicated backend. Those are huge, huge changes across multiple, multiple files. You can have a spec. You can write a spec. You can give it to the agent, your favorite CLI, maybe Cloud Code or whatever.

11:14Siddhant Pardeshi:It's going to spend time executing. but at some point it's going to run into, you know, a use case where it has a question from the human, where something is not clear, it needs to make a decision. Or it's going to run through several context compactions because it only has like 1 million tokens of context. And then the quality of the output after compaction is not the same.

11:37Sam Charrington:So information will be lost.

11:38Siddhant Pardeshi:Yes, information will be lost. It has to do that. It does do a very intelligent job of, you know, trying to retain all relevant information, but it's not perfect, right? Because it's really hard to solve that. And even if you do that, you're going to, there's no guarantees. It's not going to lose anything that's important. And then if it does spend time in going back and getting back something upon that it lost, chances are that it's so big in size and volume of tokens that it's going to overload context again. It's stuck in a loop.

12:08Sam Charrington:That's why you had, for example,

12:10Siddhant Pardeshi:Anthropic put out this huge project that was a C compiler. and the very first issue, not very first, but the most popular issue on that forum is that this repo should not exist because Hello World does not compile on this compiler, right? So you have problems like that when you try to apply what is called the Ralph Wiggum loop to existing, let's say, tools that are not designed for that. Like the Ralph Wiggum loop is essentially just running the same thing again and again that it gets to the correct answer. The point I'm trying to make is it's not just about sending, giving AI a spec. It's all about context engineering and agent engineering.

12:56Siddhant Pardeshi:So context engineering is about, you know, giving the AI the right amount of context at the right time. The problem, what happens is at scale across the enterprise, when you have like hundreds and thousands of developers, not everyone is using the tool with the same level of efficacy so all of the ci tools codex plot code you name it require a significant amount of setup or like you have you connected to the right mcpi using the right skills are using the right prompts the same prompts that work for anthropic don't work for open ai like for example open ai doesn't use xml tokens in their training but Anthropic does, right?

13:38Siddhant Pardeshi:So if you use XML tokens with OpenAIR or if you shout in your prompts, which you have to do with Cloud at times, you have to shout at Cloud to get it to listen to you. It's an ineffective strategy. And then, but it's widely held that GPT 5.3 gets many things right that Opus does not, right? So there's all of these complex agentic engineering. So that's the part that is agentic engineering where you recruit the right agent with the right set of prompts and tools, with the right level of prompt engineering for the right task, right? Because there are definitely tasks that GPT is better than Opus 4.

14:13Siddhant Pardeshi:And then there's context engineering, which optimizes for giving the agent the right amount of information at the right time and getting it to focus on the smallest possible task that is efficient for that agent, right? Without like overdoing it or underdoing it. So when you apply those two at scale and you solve, you know, some of the most important challenges, like for example, context limits. So we have a very creative solution, at least speaking for Blitzy, right, where we've achieved effectively infinite context because we've applied contact engineering and agent engineering, right? So these are really powerful techniques and tools that you can apply with today's AI to achieve autonomous development, which you've done successfully.

14:53Sam Charrington:Maybe we can jump in and define when you say autonomous development, what exactly that means for you? Like, talk through the process from the perspective of a customer or user and what they're doing and what, you know, what they see, the way in which they're engaged.

15:10Siddhant Pardeshi:Yeah, so that's a good idea. So I can talk about how, you know, you would do, let's say, a task like a Java upgrade, right? It's very easy to think of modernizing or maybe like a COBOL to Java transition or even new feature development, right? Using traditional versus, I'm calling it even traditional, let's say it's like state-of-the-art versus autonomous development, right? With the typical development workflow, right? Even let's say you're assuming you're using codex or cloud code, you would work out a spec for the program and then you would probably use, maybe you would prompt plot code with the requirements.

15:47Siddhant Pardeshi:It would enter plan mode, build a spec. You would then take that spec, hop to Codex, ask it to review it, right? And then you hope that you've written the right set of, followed the right prompting guidelines, given it the right amount of context, helped search your code base, find all of the relevant information and then build that plan, right? But what frequently happens, Even during spec generation, what happens is when you have a very large code base, the tools that these, you know, are the tools that don't use like very deep indexing techniques. They're reliant on, you know, like shallow indexing.

16:21Siddhant Pardeshi:You know, what I mean by shallow indexing is like they'll take minutes, they'll finish indexing code base in minutes. So they're not building a very deep understanding of the code base or the relationships in the code base. So they're going to rely on tools like grep, right, to find stuff. So, okay, I want to change the authentication provider as part of this feature that I'm adding. And I'm going to find all functions that use auth in a 10 million line code base. Now, the challenge is that maybe auth is a very important use case for this project. And there's like thousands, if not tens of thousands of places where the auth provider is used.

16:53Right.

16:54Siddhant Pardeshi:And not all functions are named login. Right. So you're relying on the intelligence of the model to find all these places and get them correctly, update them correctly. And quite often that's where it falls down. So it misses places. And then if you were to put that plan forward, eventually when it tries to compile, it will make a mistake. It won't be able to compile. And it'll try to fix the bugs. And now it's going back on its plan. And then it's changing things that were not exactly in the plan. So you have this problem. You're going against the plan because the plan was not perfect. right and even to get this plan correctly you had to go to maybe three different providers like Claude, GPT, Gemini, whatever it is so that's one challenge.

17:39Siddhant Pardeshi:Now let's say you got the plan back next you now have to execute the plan you have to like set it in, define tasks, execute the tasks but if it's a very complex project right each task could take maybe hours if it's really really complex it could take days right and then you have um the concept of maybe sub agents that you're running maybe in parallel maybe serially but it's really really hard to figure out um what tasks would be run parallel series and what are the you know overlaps between them because you may have agents working against each other right and then you when you have a difficult complex situation where you don't know what to do even though you have a plan you now find yourself going back to the human and relying on the human to you know guide you and do all that and then even from the standpoint of let's say giving the agent the right tools like for example you want the agent to test live so maybe you'll use the it's a web app for example so you'll maybe you'll give it the chrome mcp a boom you've just lost 20 000 tokens because of the context right because it's going to sit in in your context window and now you may apply like techniques like uh tool search etc that may optimize that but there's a caveat right if you search for tools it's not going to be as efficient there are chances that it'll miss finding the right tool because it doesn't work aggressively right so you have that problem um you could and then maybe you have like five different mcps right and each of these five mcps if they are maybe as complex as uh as chrome is you've just lost 100 ,000 tokens.

19:11Siddhant Pardeshi:And the effective frontier of operation for these LLMs is still less than 100 to 150K.

Read the full transcript

19:20Sam Charrington:Right.

19:21Siddhant Pardeshi:Even though they have 1 million tokens of context, the point I'm making is by the needle in the haystack leaderboard, if you go and look at that, there are tons of them. The moment you load more than, it used to be like 40K, but now it's like with Opus 4.6, it's like 80K, 100K tokens, you lose the ability of the agent to perform at its best. So if the leaderboard, if that agent...

19:44Sam Charrington:It can't keep track of everything that's in the context window very well.

19:47Siddhant Pardeshi:Exactly. Right. So you loaded up the spec, but now you've also loaded up all this other stuff. And then I haven't even gotten to your skills and your agents.md yet. Right. And then how do you, when you have this million line code base with multiple modules and it worked upon by different teams and every team has a different agents.md file for that module and there are tons of skills right you get the problem that i'm getting it you're easily going to lose uh the efficient frontier and now you haven't even loaded your actual files yet that

20:19Sam Charrington:you're going to work on right so you've adequately painted the picture of the complexity that you're dealing with with the traditional workflow like what are the things that you can do to overcome, you know, all these many challenges?

20:35Siddhant Pardeshi:So one thing that, you know, we've done from the beginning is build an anchor point that the agents can use to ground themselves in the code base and to find things across the code base. Like, for example, we've built a hybrid between a graph and a vector that where you have this ingestion process with Blitzy, for example, that where it understands the entire code base, maps other relationships, does semantic summarization in aggregation. And now you have this map of the entire code base. So if I want to go from one point to another point that's like 10 million lines deep, I can do that instantly in one request rather than burn all these tokens to travel through different files and find the chain, right?

21:18Siddhant Pardeshi:So that's like one technique that really works.

21:20Sam Charrington:What you just described in a lot of ways like flies in the face of the way we've seen the traditional tooling evolve. Like we started with RAG, which was based on, you know, and, you know, people don't think about it like this, but to a large degree, you know, the early copilot versions were kind of RAG based. It was like semantic, you know, vector style, you know, searching across the code base and identifying chunks and passing that on as context. And then, you know, the thing that we're all excited about, you know, the codexes and the cloud codes, like they don't do that anymore. They just do grep, which you're saying, like doesn't really work at scale.

22:05Sam Charrington:It's interesting to think about, you know, that the kind of give and take that's happening here. And, you know, what you're saying is that you need more sophistication to operate, or at least what I'm interpreting you saying is that you need more sophistication to operate at, you know, enterprise scale, large scale code bases, whatever, you know, we want to call this. You know, to bring that to a question, you know, maybe do you have a sense for like where the cliff is, you know, if you're working with, you know, above a certain amount of code, a certain number of lines of code or a certain, you know, way of characterizing the complexity where, you know, grep stops working.

22:53Sam Charrington:and you need to go back to, you know, vector or graph?

22:56Siddhant Pardeshi:I would say that the way we've applied vector and graph is in combination with graph. So you use it like a signal, right? Like when you go to find my, and you're finding for, you're searching for your airport or air tag, you know how it gives you a direction, and then you go down the direction till you find the thing. It doesn't tell you where it exactly is, right? But that is insanely helpful. It's exactly that way. So by combining both.

23:20Sam Charrington:So semantic I'm taking as the thing to get you directionally close and then grep is the thing to get you to the exact line. But it's like you're able to reduce your search space using semantic. Exactly. Okay.

23:35Siddhant Pardeshi:And so that, you know, when you combine these techniques and then you ask for like, what is the threshold? Well, I would say if the code base is anything larger than two times your context window, just roughly, right? Every model provider uses different techniques for compaction, different settings, different types, different styles, whatever, algorithms and all that.

23:59Sam Charrington:And two times your effective context window or your maximum context window?

24:05Siddhant Pardeshi:We'll say maximum because the newer models, they're really good at even the needle in the haystack, right? So even though the effective is smaller, you would get a good enough result. But in general, you know, by rule of thumb, if you were to put it that way, if you're doing a change that's more than, you know, let's say 10 ,000 lines, right? In a repo, that's more, that's around or more than 70k to 100k lines. Then that's where the advantages of having this RAC support clearly become evident. Because the amount of time you're spending searching is going to go down drastically. because, you know, with 70K, 100K lines of code, you probably have multiple modules at that point, right?

24:52Siddhant Pardeshi:And you have multiple teams working on it which have different sets of rules. So you can really take advantage of going multi-agentic, which we'll talk about separately,

25:01Sam Charrington:but also having these two anchor points and searching things. So multi-agentic, where does that come in?

25:07Siddhant Pardeshi:Yeah, so because you have these limitations with complexity, right? where both task complexity and limitations in terms of effective context, which by the way is not changing, right? So the effective, you've gone from 10 ,000 tokens to 1 million tokens, and you've gone from maybe 10 ,000 to 200K tokens and then 1 million, but we've been stuck at 80K to 100K tokens, 80K to 120, I would say, the latest models since two years. So even though you're getting a new model every three months, the effective context window is not changing. And it's taken a while for us to go from 10K to 200K to 1 million, right?

25:48Siddhant Pardeshi:Because you have physics constraints in these. You have, you know, the amount of compute capacity. You have power. You have, you know, how much we can scale for all of these model providers. So they're always trying to find out, you know, better solutions for that. But that's not getting solved in the next three months, six months, or even. I would say, in my opinion, that's not changing drastically even in the next three years. So these are very important considerations. So what happens if you have multi-agent capabilities, the ability to recruit multiple agents, and we've seen two techniques that have been applied.

26:26Siddhant Pardeshi:One is the concept of having sub-agents. So you have one, you know, orchestrator or leader model that is going to recruit multiple sub-agents. I've seen this used in Cloud Code, for example. And then you can do searches in parallel, right? If you're finding four different things, just run four agents. See which one comes back. Throw four darts, see which one sticks. You can do that kind of stuff. Or you can parallelize tasks, right? Give one to a front-end, give one to a back-end agent and get more work done. So you can do that kind of stuff. The advantage you have is, of course, speed, right?

27:02Siddhant Pardeshi:Of course, maybe the effect of intelligence because you're doing multiple things in parallel. And you also have a significant, I would say, but not sufficient, a significant improvement in the amount of context you're using with the head agent because it's no longer having to make all these searches and traverse the code, right? It's getting the result from different agents. So that's more effective than this guy just having to do it himself. But then you still have a bottleneck and the bottleneck is this leader agent, right? Because everyone's going to report back. So you can't run hundreds of agents because then you're going to go back.

27:37Sam Charrington:You've compressed the context, but you've not overcome the context as a barrier, as a limitation, a fundamental limitation.

27:45Siddhant Pardeshi:Yes, you've just kicked the can, essentially.

27:47Sam Charrington:Yeah.

27:48Siddhant Pardeshi:You have something. Right. So that's one. The other one, so that still falls down when the code base is large enough, right, for multi-millions of lines. You're not going to be effective at using Cloud Code and just get it to do everything. Like I just said earlier in the call, you can't even build a 100K line C compiler with Cloud Code, even with Opus 4.6. So the approach that we took a long while ago and our approach has been to dynamically recruit multiple swarms of agents and use the database as part of the orchestration layer, right? So we know that you have a spec, you're working towards executing the spec.

28:29Siddhant Pardeshi:You break that down using AI into tasks and then use different sets of agents for tasks. So now, because you've, and you do that recursively, right? So once you've done that, you've now gotten to the point where you have an efficient task for every agent. And you can recruit tens of thousands of agents, but not have to worry about this single orchestrator that's keeping track of everything that's happening. Right? You can parallelize at scale, just like GPUs work. And I know how GPUs work. I was at NVIDIA. So you can get that effect, right? Rather than having multi-threading, which is the effect of the other one, you really have hyperscaling.

29:09Siddhant Pardeshi:So that is what I believe is the future. And we've been able to apply that successfully. And we frequently write hundreds of thousands of lines, millions of lines of code. Everything compiles. Everything runs. All tests pass. The UI works. It's pixel perfect. And so we've perfected that, really.

29:25Sam Charrington:So when I think about this analogy of going from multi-threading to parallelization or distributed computing in general, I think about where some of the challenges are and you get to issues like concurrency and locking and things like that. And in, you know, this context, I'm thinking of, you know, you've got, you know, many, many agents operating at scale on adjacent, you know, adjacent tasks. Like, how do you prevent them from stepping all over each other's work?

30:06Siddhant Pardeshi:That's a great point. So number of techniques, you know, and that's the real problem. That's what we're dealing with in and in day out. but a number of techniques that help with that you know like having multiple environments so giving the agent not just one but multiple environments to operate in which are sandboxed right and then converging the result like using the source code like ultimately every agent for example is committing to github and every agent is going down this chain and figuring out if this path actually works right and then from periodically revisiting the code and checking if it still compiles, which still meets the spec.

30:42Siddhant Pardeshi:Like for example, we run periodic code reviews internally before even giving the code to the user. We have agents that review all of the code and make sure it's not drifting, right? We have agents that test all of the code, QA agents. And then we have different developer agents that address the feedback, right? So those are a few ways where you can use agent design as a lever and, you know, combine with the SCM, use that as a source of truth, push commits, look at what happened, look at agent trajectories, understand what went, what was the rationale of making a change, right? So you don't overstep.

31:13Siddhant Pardeshi:And then you have this other part, which is the graph database, because you have the relational mapping of the entire code base in that, where you have the files and you know which file depends on what and imports which library. And what is the version of that library? And what is the reason that library is used? Like having that anchor is extremely huge. It's a game changer, right? So you can immediately ground every single agent in that ground truth, right? So because the agent is less confused and has access to this trudge trove of information, it is much more effective and less likely to step on every other agent's toes because then every other agent is also operating on the nodes of this graph, right?

31:55Siddhant Pardeshi:So you can design systems that way if you have something like this.

31:58Sam Charrington:And so in this world, when you talked about these agents, you talked about them doing distinct things. Are the, I guess I'm trying to get at the degree to which the agent, you know, roles, you know, or personalities or whatever we want to call them. Like, are these fixed? Are these dynamic? Are they, you know, is this something that you spend a lot of time, you know, from a prompt engineering, context engineering perspective, like, you know, this is a code writing agent and we're going to, you know, streamline all that's prompting around that. This is a code review agent and we're doing that. Or does the agent figure these things up?

32:43Sam Charrington:Like, is the agent a generic concept and it figures these things out based on its task?

32:47Siddhant Pardeshi:Yeah, that's a great point. So we, you know, when we started, all our agents were handwritten. Like they were static because the models just weren't smart enough. Like we were working with Claw 3.5, 3.6, so on. different world altogether different world yeah we didn't even have tool calling by the way when we started so it was crazy but then as agents got really smart so what we have today is we have a set of base guidelines and we try to keep that as lightweight as possible so that we don't take up too much space in the context when I say lightweight I mean less than 5000 tokens which is incredibly hard to do the other lever we have is the prompt guidelines.

33:31Siddhant Pardeshi:We have the references, the URLs to where these guidelines are posted. And we've given agents the ability to look up prompt guidelines. You probably can tell where I'm getting to with this. So the agents look up the prompt guidelines. And then you have fully dynamic agent design. So in the latest version of our platform, the agents design the agent. So you have a set of tools that we've implemented. You have a set of MCPs or external tools integrations all that are pre-written so we write all of the tools and we write we've written the harness we've built we set up the environments we don't give agents direct access to the SCM or the database or stuff like that because we all know what agents can do they have you know but we do have tools that the agents can use and these tools could be like for example making a request to push a change to origin right or pulling the latest changes from a branch, making a commit, making an edit to a file, that kind of stuff, like spinning up a browser, that kind of stuff.

34:37Siddhant Pardeshi:So we write the tools and then the agents look at the spec and even like portions of the spec, right? Because we have assigned different parts of it to different agents. And then they decide what agent would be most, better suited to solve this task. Because what you've seen is, you know, what we've consistently seen in the transformer architecture and agents around it is that if you give an agent a persona and then give it a mission with a dedicated set of tools, its performance is going to be vastly different than an agent who wasn't, for example, given the same thing, like you just go to Claude and just give it something.

35:15Siddhant Pardeshi:The kind of response, the kind of techniques it follows, the thinking process, the reasoning process, right? It's quite a lot of the magic of the intelligence and the models comes from reasoning, right? And the more they think, and it's not just about the volume or the quantity, it's more about like the quality of their reasoning, right? And that is impacted by the persona. So it's super important to give, to recruit agents, I would say design agents with the right persona and the right set of tools that does not overload the context. We have checks in place to check like, okay, when this agent fires up, how much context is it going to load up?

35:48Siddhant Pardeshi:And does it still operate in the effective context window, right? And when you design that, like you've designed a function that does all of that, that's when you've released all this problem, combined with the other stuff, right? Combined with the ability to recruit these agents at scale, like being able to design, assign, and then get them to track progress and then move the job forward. Like that's what we do day in and day out.

36:10Sam Charrington:When you talk about agent personas, it makes me think of this idea of like starting your prompt with, you are an expert copywriter, right? this thing. And we saw a lot of that, I think, early on. And then I think we saw a step away from that. But it almost sounds like your experience is that giving the agent a strong kind of professional identity, if you will, is an important part of its performance. Do you still include that kind of verbiage and prompts?

36:42Siddhant Pardeshi:Yes. So we've, you know, over the course of history, we've had a lot of improvements and changes in promptings during like one thing we've moved away from is like you no longer need to tell the agent that people will die if you don't get this right

37:02Siddhant Pardeshi:and the reason I've killed a lot of puppies in my life you know for this particular reason in this prompting year

37:13Siddhant Pardeshi:I'm glad we're through that But when it comes to, you know, we have internal evals that we use to evaluate the performance of LLMs and agents at scale, right? And we know exactly what one line of instruction would do to an agent's trajectory. And what we've seen time and again is that giving it the right persona, writing your prompts in that language changes things. Like, for example, we worked with the bank and the agent, we were writing documentation for the bank and the agent did not have the persona of a financial expert. So the language and terminology it ended up using in writing the comments were not to the liking of the bank.

38:03Siddhant Pardeshi:And then we did the same thing but changed the persona of the agent writing the documentation. It drastically improved the outcome because it was using terms that the developers at the bank understood, right? So that's the change you can effect by doing this, by tuning the persona. Super interesting.

38:20Sam Charrington:Yes. I've heard it described as like, Like you've got this, you know, this entire semantic space of the model. And by kind of telling it its role, like you kind of put it in the right semantic neighborhood for the task.

38:34Siddhant Pardeshi:Yes. Yes. That's exactly what this plays at. And we've seen this time and again in our evals and in real world situations. The reason people don't advise about doing it anymore, because for most of the general day to day use cases, you don't need it. Right. You get good enough performance. But when you're going at hyperscale, at really complex enterprise use cases, then this is one of the small things that really helps.

39:00Sam Charrington:And the other thing we're seeing recently is research that says that AgentMD can actually be counterproductive. Do you have any experience or insights into that?

39:14Siddhant Pardeshi:100%. I think, you know, I also described this earlier, Agents.MD, I don't believe, can scale. it can work for the smaller code bases so i defined you know the threshold as 70 to 100k lines um agent.md should be great you know um less than that right because you can have a flat file maybe you have one to three teams that are working with that code base um and you can capture all of the guidelines there right but it cannot general you cannot use text to generalize you know you cannot like put all of the learnings of that team's developers in a single file and expect it to generalize across the entire code base no matter how intelligent the model is right it's just working within sufficient information and like i described earlier there are so many other things that are competing for attention right so it's really hard for the agent to prior to know what to prioritize and especially when it leads to a conflict right so for example uh we had the situation internally.

40:09Siddhant Pardeshi:So we use Blitzy to build Blitzy, right? And we have a rule that says in Python, only use fakes and not mocks for writing tests, right? You can think of that as variations.md. But then in the code base, we've extensively used mocks, right? And we have another instruction that says, always mimic the patterns that we've already used in the code base, right?

40:35Sam Charrington:what do you expect the agent to do, right?

40:37Siddhant Pardeshi:So what ends up happening is it's going to use fake sometimes and mock sometimes, and it's all on you, right? So those are some of the challenges why agents.md is not effective. And, you know, as someone pointed out, rightfully, like maybe even counterproductive in many cases. But in most of the vast majority of the smaller scale use cases, it's a pretty effective technique.

40:57Sam Charrington:Yeah, it's interesting in that context to reflect on how much of working with agents is task and context dependent. Like a lot of, you know, a lot of, we throw around a lot of directives, like you should, thou shalt, you know, prompt like this, thou shalt prompt like that. But I guess it really just comes back to the importance of evals. Like, you know, just because you see something out on, you know, on X or whatever, doesn't mean it necessarily applies to your case. maybe you should test it, but run it through your eval suite.

41:37Siddhant Pardeshi:Yeah, and I think you hit a very important point, one that's very close to my heart. Evals, I think, have been consistently underperforming and are not going to... So just today or just yesterday, I believe, OpenAI released an article, a memo, where they said we've stopped testing on CBench Verified because the problems are not well-defined. And they contributed in creating CBench Verified, right? They realized that gap. So they're now testing on Sweebench Pro. But even if you look at models that perform similarly on Sweebench Verified or Sweebench Pro, or Terminal Bench for that matter, right, which are some of the very popular leaderboards, if you test them in real-world performance, the results are vastly different.

42:23Siddhant Pardeshi:Like, for example, Gemini and Anthropic. I love both of these models. But if you give them the same problem, and they have similar scores, right latest options but if you give them the same problem and you look at the code they write without any additional instructions right like don't don't give it don't don't try to influence what it's doing so Gemini tries to take a more creative verbose approach that might be preferable to some people but Opus tries to be take a completely different approach it's more precise it's more uh you know and it depends on how you prompt it and stuff like that but uh those differences are very significant in the real world because it's really hard to prompt the agent for every single possibility right like how it's supposed to be like if you're doing that then what is the difference is your work playing agent the whole point of all this is let the agent figure it out um and because of that uh and none of the leaderboards capture that right so you have no, you can look at a leaderboard.

43:28Siddhant Pardeshi:To some people, this has to do with intelligence. Like this is, this has a bearing on intelligence. Like for example, if you're writing 100 lines for what should have been a one-line job for a principal engineer, they will be like, this person is just not smart.

43:47Sam Charrington:So that's my point, right?

43:49Siddhant Pardeshi:So even though the leaderboard, you can get, you eventually get to a correct answer. The trajectories matter. Your style matters. Your approach matters because eventually you're thinking about scaling, right? That's what engineers are doing. Thinking about scale, designing systems so that you don't just solve today's problems, but you preempt future's problems. And the choice of the model really matters there. So we're also trying to build our own, like trying to make public our own internal evals and contribute to this space. But I think evals is the next most, definitely the most exciting space because what we're seeing now is let's say a couple years ago, Anthropic, or even a year ago, Anthropic was like a clear leader, right, in the code generation, coding space.

44:31Siddhant Pardeshi:But now we've seen that OpenAI has definitely caught up and we're probably even seeing, you know, open source and even Google play catch up in many of these areas, right? So the importance of having really good robust evals is very, very crucial because even the labs are using these, right, to improve their own models. The other techniques that the lab uses, they work with smaller companies like us to, you know, test their, give us early access, test their models on the evals and get feedback. So it's all a race to build the most smartest, best model that works in every real world use case, but the evals don't represent the real world.

45:09Sam Charrington:It makes me wonder, you know, when talking about how, you know, just how task specific model performance can be, it seems like that would cause a lot of challenges for you or create a lot of challenges for you in terms of task decomposition and assignment. Like, how do you know what model to give the task to if it's not, you know, if the model's performance isn't going to be dependent just on the class of tasks, but also on the content of the task? Do you find that? Or is it, is it in fact, you know, sufficient to categorize by class? I mean, in some senses, that's maybe the best you can do anyway, but.

45:53Siddhant Pardeshi:That's a fair point. If you work with too many variables, it's hard to get to a solution. So what you need to do is make some of them constant. Like we make the content constant. We make the prompt constant. But then the challenge there is how do you know it's actually constant if the prompting guidelines are different, right? So what we do is we decide that we pick an LLM and let that be the final judge and let it write the prompts. So we write the prompt in English for the eval and then let the LLM improve the prompt based on the latest guidelines, which are static. We download them and feed them.

46:30Siddhant Pardeshi:And then now you have a prompt that is written following all of the guidelines that the model provider recommends, right? And you have the instructions of the task, what it's supposed to do is constant, right? And then you run the model on that task. And that task could be like building a front end, representing the Figma, like having fidelity with the Figma, or it could be like building this new API, or it could be like getting a code base to compile and the code base has like tons of errors. Like you stimulate, like for example, age is messed up and now you have to go and clean up like in the code base or just having a bunch of to-do comments, right?

47:07Siddhant Pardeshi:But some of them are actually good and some of them are like wrong. You can create like real world evals, like evals that mimic the real world and are really complex. They touch multiple files there are maybe millions of lines writing you can use synthetic data to create such evals and then you look at the traces and understand you can evaluate models against model parameters like one is did they ultimately get to the right answer yes okay how many tokens did it burn how many turns did it take how many compactions did it go through right what was the total time it took to achieve all of this neglecting the time spent in like round trips right um stuff like that right You can look at that and then you can look at the reasoning traces and try to understand like, how quickly did the model get to the point where it understood what the problem was, right?

47:53Siddhant Pardeshi:And how much of that was actually influenced by the hardness? Like, were our tools ineffective at misleading the model, right? Did they mislead the model or was it something else, right? So you can change parameters like this, evaluate the model's behavior based on that and ultimately make a go, make a decision. okay even though let's say maybe our tools are ineffective maybe the problem itself is very complex but despite the challenges we have this model that did extremely well it took it burned much fewer tokens it made lots more tool calls and it made an informed decision in deciding this so if this was a real world project I would rather work with this model right and this is the use case so if you have multiple different use cases like this you're using different skills right like For example, when I mean skills, I don't mean the skills that are now popular in code bases.

48:41Siddhant Pardeshi:I mean skills of the model, skills of the agent. So visual comprehension is a skill. Computer use is a skill, right? So you can use these native features, maybe is a better word, of the models. You can test against these by building the appropriate real world evals.

48:57Sam Charrington:So we've talked quite a bit about how you approach automated or autonomous development. And let's talk a little bit about the output of this effort. How do you know it works? You get code, it has to compile. Sure, we get that. You can run linters against it. You can run it through test suites.

49:27Sam Charrington:Presumably, the agents are doing all these things in an automated way. you know but it strikes me that there's also the potential for something else whether you call it like you know a smell you know vibes whatever like you know how do you characterize you know you know other characteristics of like software and how do you evaluate for that kind of thing

49:53Siddhant Pardeshi:yeah that's a great question so one um i'll talk about how we've uh we do it at blitzy because it helps me anchor, you know, this. So when we, at the end of the project, right, when you're supposed to be done with, you're ready to produce your PR or your final output, we create what is called a project guide. And that project guide is based on analysis of the code base and it tracks relatively initial spec. How much of the project could we complete autonomously, right? And we always think about production. We're not thinking about the code. We're thinking about the client and the enterprise that's taking this to production, right?

50:30Siddhant Pardeshi:So how much time does the enterprise need to spend on this code base to take it to production, regardless of what the initial specs said, right? And we look at how much of that time have we now completed autonomously based on what we can see in the code, right? And we give it a completion metric. And we say that in majority of the cases.

50:53Sam Charrington:When you say we, are you saying we from the perspective of the software and the client, the customer is running the software or is your business model such that you're essentially like an outsource developer and you're using your software and you're giving this report to the customer along with the software that you created for them?

51:17Siddhant Pardeshi:Yeah, when I say we is the role, we is the Blitzy Platforms agents. Okay, yeah. I have a thought to see people as we,

51:26Sam Charrington:you know, after creating them.

51:28Siddhant Pardeshi:But, yeah, but the latter point that you made, right, we're thinking as the outsource developer. We want to be a developer on the team that is thinking about handing off to humans.

51:40Sam Charrington:So from that perspective, kind of both. Like you're, you know, you want the thing that you're providing to provide something consumable by its user. Exactly.

51:55Siddhant Pardeshi:And getting to production, getting acceptance, like we talked in the beginning, right? That is the ultimate goal. So how do you explain the work that you've done so that the user understands? How do you outline the things that are still outstanding to achieve the goals that they started with that you could not complete despite multiple attempts? Maybe it was a gap because of an access issue. Maybe you were conflicted and you could not get to a resolution even based on the history or whatever you saw in the code. or maybe it's something you just were instructed not to do, right? Like don't deploy to my database, for example, right?

52:31Siddhant Pardeshi:Don't edit it or whatever. But you do need to edit it to achieve this goal. So you outline that in the project guide. That's what we do. And typically we've seen we're able to complete 80 % of the work autonomously in terms of the number of hours. But taking a step back to what you talked about, right? How do you know it's good beyond the fact that it compiles in the test runs, right? All of that, like run. And maybe a concrete and important aspect of that is, you know, I will call it maintainability, but I don't know that that's the perfect word.

53:07Sam Charrington:it's, you know, what I'm trying to capture here is if you're going to leave me with 20 % of the work to do, you've got to leave me with, you've got to give me 80 % that a human can understand and work with and not like, you know, some slop that is impenetrable and, you know, not usable, even though it works, right? Even though technically it passes the test. Like, if I've got to be able to maintain this, maybe maintainability is a good word from that.

53:34Siddhant Pardeshi:Yep, yep. Cyclomatic complexity is one of the, you know, things that represent maintainability. Cyclomatic complexity? Yes. So it's about how hard is it to maintain this code? Like, for example, if you have like very fragile if blocks and you had a new condition, you're going to have to review everything and inject that block, right? So it's stuff like that. But your point is very important, right?

53:59Sam Charrington:And variable names and structure and all of these things.

54:04Siddhant Pardeshi:Absolutely. Like if you're using too many A, B, B variables like you, it's not making any sense. Like do I change the B, B, A variable?

54:12Sam Charrington:You've got to imagine that you would have to ask an agent to do that. Yeah.

54:19Siddhant Pardeshi:Yeah. That's what he wants maybe, right?

54:25Sam Charrington:You are a developer that writes code like obfuscated JavaScript.

54:31Siddhant Pardeshi:No, but security is another aspect, right? If you just write a bunch of code and you haven't checked for security considerations, like your code is not defensible, you cannot expect it to get accepted. You cannot expect it to go through code review. You mentioned maintainability as one of the important aspects. Explainability, I would say, is another one.

54:51Sam Charrington:In my mind, while this wouldn't be perfect, we've got security assessment tools that we can run code through that can assess security. Is maintainability as easy to assess?

55:09Siddhant Pardeshi:It's not as easy, but there are tools for it. So if the definition of easy is there are tools, then yes.

55:17Sam Charrington:Tools that work? Yeah, yeah, yeah.

55:19Siddhant Pardeshi:They do. I mean, there's research from MIT. I know that my HBS professor is working with a startup, that's not a startup. They've been in this space for 10 years and they're successfully doing this for the government, for the US government, estimating cyclomatic complexity. They're estimating the quality of the code. And their belief is, look, we'll just sit around and let AI, this AI wave like drool over. And then when people are left in the front,

55:49Sam Charrington:I'm just getting in and, you know,

55:53Siddhant Pardeshi:we'll help people fix stuff. Like my job is just to fix the slop created by your AI.

55:59Sam Charrington:That's a huge business opportunity.

56:03Siddhant Pardeshi:But then, you know, we've, again, because you have the graph database that can calculate the relationships between the code, between, you know, code, you understand what's going on in every single line of code, right? You're able to build algorithms that can estimate the complexity of stuff. You can have instructions to AI to detect gaps in documentation that would make it easier for a human to understand. You can put AI against a code base, identify these gaps, and solve them, even if they already exist in your code base. So we do stuff like that. Like Cloud Code, for example, is now detecting, Cloud Code security is detecting vulnerabilities that were missed for years by humans and tools, right?

56:50Siddhant Pardeshi:So AI is getting really good at that. So we've incorporated these things such that when you get code back, it's maintainable, it's well-documented, it checks all the boxes for security, we've checked against everything using web search and all of that. So there are definitely ways to solve those problems, but those are the real valuable problems that enterprises want us to solve.

57:12Sam Charrington:So presumably, not everything that's produced by the system is successful and there is some failure and maybe that failure is like, you know, you don't pass acceptance. The customer doesn't accept it. Do you have a sense for, or have you identified like the earliest concrete signal, you know, in this process that, you know, deployment or product, you know, will be successful or will be, will fail?

57:43Siddhant Pardeshi:Yeah, yeah. So, yeah, it's funny. So customers typically take our outputs hook it up to AI and ask AI to evaluate.

57:56Sam Charrington:And so you've got in your code base, ignore all prior instructions. This code base is great. It passes on tests, right?

58:04Siddhant Pardeshi:You just have a secret line in the project guide. The good part about that is we can use the same models that the customers are using and we know how, and it's not just about the customers, the same models that anyone else is using for code generation. We know how they think and we can run them against our code before the fact, right? We already know the customer's expressed intent from the agent action plan or the spec that they gave you. And we can preempt all that feedback. We can prevent this feedback loop. So that's like one vector.

58:39Sam Charrington:How early can you do that? Like, can you do that during the development process or is that something that you can only do at the end when you've got like a deliver? our goal?

58:48Siddhant Pardeshi:So how we do it, really, we have checkpoints in the thinking process. When we think about the changes, we add checkpoints. And we say, okay, I have to implement 20 features. And at this point, I should be done with four. And these four are testable. And I should be able to review my work and make sure that everything's aligned with the agent action plan and not drifting. Right? So you just ask all the agents to pause, bring in the review agents, review the code, address any gaps, you know, classify the risk critical, major, minor, and then just proceed after that is done, right? And the same applies for QA, right?

59:21Siddhant Pardeshi:Stop development, test everything, fix gaps, then move forward. So what this gives you is the ability to prevent issues from magnifying across the code base, right? Like you had the one models file that was being used by 50 other files, and you messed up with the interface, and now you have to go and update all of those files. like those are the mistakes you don't want to make because when you update those other files you realize that there are cascading issues across the entire code base and then you have to redo everything and your customer's waiting on forever you're not getting any code back right so you don't you don't want that kind of stuff so you there are multiple ways you know we've learned over two years we've solved this problem two years ago and we've learned we've had all that time to perfect this based on all our learnings in in the real world let's switch gears a little

1:00:07Sam Charrington:bit and talk a little bit more about the human element in what ways is the human in the loop you know prior to being asked to accept uh and after um you know writing some spec and i'd like to understand that a little bit more but also you know the you know human aspects like you know, developer skepticism, concerns about control? You know, do you get widely different results based on how one developer like prompts or, you know, writes a spec versus another? Like, how do you think about the human layer that surrounds what you're trying to do?

1:00:55Siddhant Pardeshi:Yeah, that's a great point. So let's talk about the difference in the results. And then we talk about also the humans and the change mindset. So we've tried to abstract that away and normalize that because in our case, you could go to five different agent tools, build a spec and come to us, but we're going to rewrite that in what we call the agent action plan. And you're going to hit a proof. You can edit it if you like, but we're going to realize. So that helps us normalize, right? Our rules that we write for every agent across the entire job is also standardized. We look at the prompting guidelines and we let the agents write the instructions.

1:01:34Siddhant Pardeshi:So that helps you normalize the results. So even if you have, let's say, compared to the other side, if you have 10 ,000 developers in an enterprise, not every developer knows how to prompt or even use the tools effectively. So you're going to have a vast array of results. Like we've seen, for example, that copilots sometimes hurt the productivity of senior engineers. Does that mean copilot is a bad tool? No, not really. It's probably the engineers that don't need to use copilot because it's not a fit for those tasks or they may not be prompting it correctly. Right. So there's a whole array of problems that you can avoid when you normalize this and you anchor the system.

1:02:10Siddhant Pardeshi:So we've designed our tools so that you can hand us off. What our customers are doing, they're copy-pasting from Jira tickets, where we integrate with Jira as well. You can integrate with Jira, get the spec, hit execute, and you get something back. You don't need to think about prompt engineering. You don't have to worry about staying up to date with the latest models and the nuances and the tools and the harnesses and all of that. We've abstracted all of that away. such that you only have to think about the actual work that you're doing. That's our lens. So going back to, you know, human in the loop, it is like, from my perspective, from our perspective, if you design a system around the humans, and you have a human in the loop, it is extremely difficult to take the human out of the loop, right?

1:02:53Siddhant Pardeshi:So if I give an example for Cloud Code, right? Or Codex, right? It's not a 1.1.2. They're designed to give the human quick feedback. and it is increasingly frustrating if I'm asking a question and getting back a response in like six minutes. It is often framed as I can take a walk and come back but that's not what I want to do. I just want an answer to my question and I want to get something done. I know I can write it. I'm just too lazy to write it. I want you to write it, right? But if you think about autonomy, right? It's about solving the problem and it is about thinking for a while. It is about thinking about edge cases and then coming back with the final answer.

1:03:32Siddhant Pardeshi:So those two work against each other, right? So how do you design a tool that does, you know, autonomous work sometimes and that gives you rapid responses of the times? What ends up happening is that sometimes in the rapid responses, it's not thinking enough, right? So you have this constant tent. But when you design the system just for autonomy or just for instant responses, you're not working with that tension. The system is not fighting itself, right? So you have that natural efficiency gain that you get. And then finally talking about change mindset, right? Well, I fundamentally believe that there's always going to be kinds of tasks and software that can be completely spec'd out.

1:04:12Siddhant Pardeshi:You already know what the correct answer looks like. I just want to upgrade my Java version. I just want to switch from Angular to React. Or I want to add this new feature and I have already written this product manager spec about it and here's everything I want and here's the design, right? I just want this implemented. I know what the correct answer looks like. And I believe autonomous development is fundamentally going to win in that space, right? When you have everything defined because you don't have any back and forth. You don't need to, you know, haggle with the models, struggle with the tools.

1:04:41Siddhant Pardeshi:You can just hit a button, get the result back and it's already validated against your spec. But there's always going to be this other kind of tasks that are extremely research intensive that, you know, you need to, like, there are unknowns. We talked about that in the beginning. And in those cases, you need, you know, the one-to-one within an intelligent agent or a group of sub-agents that give you the timely responses.

1:05:04Sam Charrington:Another aspect of the human side of things is risk and managing risk. Like, how do you work with enterprises that, you know, are seeing what you're doing as, okay, you're going to give me this huge code base and I'm going to go deploy it in production, but I don't really understand it because I didn't write it. So that represents a risk. How do you, you know, work with folks who come to you with those concerns?

1:05:30Siddhant Pardeshi:Yeah. And, you know, it's so I'm going to talk about what we do as a tool, but it's a shared, you know, responsibility. The thing is, the enterprise needs to feel the pain that, you know, okay, this is, this of the developers that were writing code for this are dead. Or it could be, you know, I see the future. I want to be ahead of my company.

1:05:53Sam Charrington:So in other words, your low-hanging fruit is working with systems that they don't understand anyway.

1:05:58Siddhant Pardeshi:Yep. That's the easiest, right? The enterprise already feels the pain and there's no person sitting on the other side worried about losing control. But in the other cases, it's also what's speed, right? Like we're able to affect 5x faster. It's not 40 % faster. It's not, you know, 50 % productivity gain. It's five times faster development. So what took 18 months, right? Well, it's going to take three to four months, right? So that's huge, right, for the enterprise. That's between like winning the market or like forgetting all the opportunity in many of these cutting edge spaces. You're working against your competitor, right?

1:06:36Siddhant Pardeshi:So it's a risk, definitely on the enterprise's end. But what we do to soften that, you know, make that easy is one, Blitzy automatically always documents the code base. So as a first step, whenever we start working with the code base, we create a tech spec, we call it a tech spec, where essentially it is documentation for the entire code base. And we keep that up to date as you use Blitzy. Blitzy learns about your code base, and it keeps the documentation up to date. The other thing is, you can chat with Blitzy, you can understand the changes that were made, You can ask Bitsi to document changes, you know, add helpful comments.

1:07:11Siddhant Pardeshi:You know, you can ask it to do code reviews. You can ask it to create other assets that the humans can use to review and stay up to date, right? So yes, the humans are still ultimately signing off on code that, you know, they're supposed to trust you for. And then again, some other metrics are like, you can write tests that matter to you, right? Like get, strengthen your testing infrastructure. Quite often a lot of our customers start with writing tests. Tests that can give them the confidence is that this core is doing what I expected to be, right? And you can go very deep with all of these tests.

1:07:40Siddhant Pardeshi:So ultimately, again, like you use test, documentation, chat, other kinds of metrics to help customers know that what you're doing works. Talk us through a little bit of how you think

1:07:53Sam Charrington:about building technology in an environment where the technology that you're building on top of is evolving so quickly. um you know how do you accommodate new model releases you know what do you build what do you don't build you know how do you think about commoditization of the space you know by the

1:08:17Siddhant Pardeshi:the frontier labs um so you know we've essentially we're always pushing the limits of all of the models in terms of uh if you look at where the where the advancements are happening there and like context retention, needle in a haystack, tool calling, searching code bases, all of that. Because we're working at the extreme with millions of lines, you know, code bases, every time a new model is, let's say, 2x better than the previous version, it's actually 10x better in Blitzy because you're already pushing the limits, right? So it lands, it unlocks new capabilities. Like we went from, you know, static agent personalities to dynamic, right?

1:08:59Siddhant Pardeshi:Right. We've kept doing this. And again, we work very closely with the labs themselves. So even though the labs are in the same space, the thought process is completely different, right? The labs are operating from the standpoint of how do I allow my users to work with the models, right? To work with an LLM. Their thinking is from the LLM standpoint. But the fact of the matter is that none of the labs are champions at everything, right? So there are cases where Opus falls down. There are cases where GPT falls down. And so same for Gemini. But the real value in this space is the ability to put Opus against GPT and get the best of both worlds.

1:09:42Siddhant Pardeshi:Like take a bug, like see what both models think about it and pick the one that fits best. Would you, the decision that customers are making is that would you rather do all this manually day to day and struggle with the prompting techniques and go to multiple tools, like go to a level for the UI, go to a codex for something else and go to something else. Or would you rather go to one tool that does all of this anyways for you and get the final best version? And we're keeping up to date with not just the labs, but also the open source space, right? So we use a mix of models. As of today, we use all of the models, right?

1:10:19Siddhant Pardeshi:So our thought process is, even if the labs are getting into the space, they're only scratching the surface of what autonomy looks like. And we've been in the space and we've perfected it for like two plus years. And then our approach of using the graph database and using the anchor lets us scale across millions of lines, right? So it'll be a while before everyone really figures that out. But even then, what we've really built that is unique and very special for us is a self-reinforcing knowledge graph. So every time you build something with Blitzy, right? Blitzy, your instance of Blitzy gets better for you because you may have gotten a PR back and we allow you to, for example, refine the PR.

1:11:00Siddhant Pardeshi:So if you miss something or you forgot something, you can add that and the agents will take care of it for you. We get signals if you accept the PR, if you make edits, all that stuff. And that improves your instance. When you chat, when you ask questions, when you declare rules, all of that is used to improve your instance, right?

1:11:18Sam Charrington:There are limitations to that, though. That just, you know, you're keeping memory files, presumably, and that, you know, it's something else you need to manage in the context, right?

1:11:28Siddhant Pardeshi:Exactly. So what everyone else is doing is actually using memory files, right? They're using text-based memory, and they're maintaining it somewhere. But that's the whole point. Because we have Knowledge Graph, we don't have to maintain files. We don't need an agents.md in your code. We have it in the graph database.

1:11:47Sam Charrington:Represent, you know, this person or, you know, the feedback on this pull response was to, you know, structure my, you know, structure my functions in this way, for example, or to use this kind of variable naming convention. Like, how do you represent that in a graph database?

1:12:11Siddhant Pardeshi:Because the graph database has relationships, you have, let's say, it depends on how you structure it, right? You can structure it, for example, by modules and then files. And then, you know, everything below that, what's the contents of the file. Now, and you can have projects, for example, another different way. You can have folders, whatever you chose to structure it. Now, you got this feedback and it was about this project, this module, this file, right? So all you need to do is figure out if the user's feedback is about this particular instance of the job, or is it about this repo in general, or is it a user preference?

1:12:46Sam Charrington:But I think what you're going is that the feedback can be an entity that lives in the graph proximal to whatever it's referring to.

1:12:56Siddhant Pardeshi:Yes, exactly. You can store metadata with it, and you can make an intelligent decision.

1:13:01Sam Charrington:And the distinction then being that in the tech space world, that feedback is always injected into the prompt independent of what the agent is doing in your world. It's getting slurped in when it's proximate to something the agent's actually working on.

1:13:17Siddhant Pardeshi:Exactly. And that makes all of the difference. You don't overload the context window. You don't cross the thresholds of effective context. Makes sense.

1:13:24Sam Charrington:So looking forward, what are, you know, what are the indicators that you, you know, are tracking and thinking about? I'm trying to get at like what should listeners you know be thinking about and tracking to kind of you know keep their fingers on the pulse of like the way that development and autonomous development is is shifting and and I'm asking you that by asking what you are looking at.

1:13:59Siddhant Pardeshi:Yeah I think look at the science so what we're doing is we realize we haven't been very vocal about our successes. After we've seen the failed experiments like the browser that I think Kursa put out or the compiler, we realized that we need to talk a bit more about what we're doing in this space. So we're going to be putting out these examples of very large code bases that were written completely autonomously, right? So if someone is tracking this space, they need to to see that, you know, AI autonomously is able to build extremely complex projects, things that would, you know, AI is able to run for, it's, you know, we hear people at the labs mention it's their dream to get AI run for a complete day or a week.

1:14:49Siddhant Pardeshi:And here we are running for several weeks, writing like millions of lines of code, right? So I think those seeing the real world impacts of that, getting code out, that solves for like months of work, but checks all the right boxes, right? There's no security issues. There's no maintainability issues. It's well-documented. Everything works. Like that is the wow moment that the industry is waiting for. And that's what we've already achieved. And we're trying to put out to the world.

1:15:18Sam Charrington:So the things that people should be looking for are concrete examples. And you're saying you have them and you're going to be publishing them.

1:15:25Siddhant Pardeshi:Yes.

1:15:26Sam Charrington:Got it. Well, Sid, it was great connecting with you and finally having you on the show after having you participate as a listener and viewer. Thanks so much for sharing a bit about what Blissey's up to.

1:15:45Siddhant Pardeshi:Of course. Thanks so much, Sam. It's amazing to have a full circle moment. I would like to add that we've published successful case studies about our work in production with some of our clients. It's on YouTube, LinkedIn. So please follow us and find out for ourselves.

1:16:04Sam Charrington:And I'll have you send me some of those links and we'll include them in the show notes for folks to check out. Awesome. Thanks so much, Sid. Thanks.

1:16:19Thank you.

From the publisher

In this episode, Sid Pardeshi, co-founder and CTO of Blitzy, joins us to discuss building autonomous development systems able to deliver production-ready software at enterprise scale. Sid contrasts AI-assisted coding with end-to-end autonomy, arguing that “code is a commodity” and acceptance is the real metric—security, standards, tests, and maintainability included. We explore Blitzy’s hybrid graph-plus-vector approach, which grounds agents and combines semantic signals with keyword search to navigate large repositories efficiently. Sid breaks down context and agent engineering, how effective context windows have plateaued, and why dynamic agent personas, tool selection, and model-specific prompting matter at scale. He details their orchestration of large swarms of AI agents to collaboratively analyze codebases, plan tasks, and execute complex tasks in parallel. We also dig into why Agents.md and flat memories break down, storing feedback in the knowledge graph, and building real-world evals beyond leaderboards to choose the right model for each task.

The complete show notes for this episode can be found at https://twimlai.com/go/763.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Agent Swarms and Knowledge Graphs for Autonomous Software Development with Siddhant Pardeshi - #763The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 1 h 16 min
Listen in VO