SED News: OpenCode, AI Code vs. Shipped Code, and the LiteLLM Breach

2 Apr 2026 · 57 min · 16 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

SED News headlines on ARM CPUs returning, the LiteLLM API-key breach, and OpenCode’s open-source agentic coding; main deep dive on “writing code vs shipping code” with CircleCI’s 2026 State of Software Delivery Report.

Guests

No guest interviewees; only hosts Gregor Vand and Sean Falconer.

Guest backgrounds

N/A (hosts discuss their own work: Sean mentions work at Confluent acquired by IBM; Gregor plugs Superbase and Stripe-related tooling).

Key claims

  • CPUs are gaining attention as agentic workloads run locally; ARM is moving toward branding/manufacturing its own chips (via major fabs like TSMC/Intel/Samsung).
  • LiteLLM’s dependency/API-key exposure enabled attackers to steal credentials and run up bills; compliance (SOC2) doesn’t guarantee security.
  • LLM coding boosts “writing” throughput, but “shipping” throughput often lags due to verification bottlenecks (PR/security reviews).

Notable examples

  • LiteLLM breach documented by Callum McMahon (Future Search); his machine was shut down.
  • CircleCI stats: 28M CI/CD workflows; top 5% throughput nearly doubled, median +4%; feature branches +50% while main branch throughput ~+1% and feature-branch throughput +15% vs main branch -7%.
  • “Prompt-to-prototype vs prompt-to-production” (e.g., VS Code moving to weekly releases with strong automated checks).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Personal Updates from the Hosts

0:45 to 2:36

Gregor and Sean share personal updates about their lives and work.

“of catch up on what we've been doing what has been in your sphere sean since we last did a scd news last month?”

Tech Headlines Overview

2:36 to 3:59

The hosts discuss recent tech headlines and transitions into the main topics.

The Resurgence of ARM Technology

3:59 to 9:26

A detailed discussion on the return of ARM in the CPU market and its significance.

“But I think the headline here is ARM is back in a big way.”

LightLLM Hack and Security Concerns

9:26 to 14:02

An analysis of the LightLLM security breach and its implications for the industry.

“So there is always like a tradeoff, right?”

Understanding Modern Cybersecurity Threats

14:02 to 17:08

Explains the evolution of cybersecurity threats focusing on new targets like API keys.

“I also think the interesting hearing here is where a lot of times when we think about hacks historically or breaches, it's like, what is it that people are trying to go after?”

The Rise of Open Source Alternatives in Coding

17:08 to 20:06

Discusses the emergence of open source coding tools and their impact on developers.

“Yeah, I think we're going to, of course, see more and more of these open source alternatives.”

Market Dynamics in AI Tools

21:47 to 25:58

Analyzes the competitive landscape among various AI tools and their sustainability.

“You know, there's a lot of these tools now available.”

Ethical Considerations in AI Development

25:58 to 28:00

Explores the ethical implications of AI companies partnering with government entities.

“And it's not like we've seen things like this probably appear in the past where a company potentially anchors itself a bit closer to a government or not.”

Consumer Perspectives on AI Ethics

28:00 to 29:30

Explore how consumer convenience can overshadow ethical concerns in AI.

“You know, you can take any major retailer in the world that's probably manufacturing goods in developing countries where maybe the standards of work are not necessarily great.”

Understanding Code Shipping vs. Writing

29:30 to 31:30

Discuss the implications of AI on code writing and shipping in software development.

“And today we are looking at effectively writing code versus shipping code, and especially the emphasis on shipping code.”
Show all 16 chapters

Analyzing Throughput in Software Teams

31:30 to 36:25

Delve into the statistics on code throughput and team performance with LLMs.

“One is that, you know, not everyone has reached the level of the most high performing teams.”

Challenges of AI in Production

36:25 to 41:30

Examine the challenges and potential errors generated by AI in production environments.

“Yeah, I mean, I think that the reality, at least at the moment, is like organizationally, you have to probably change.”

The Value of AI in Software Development

41:30 to 42:00

Assess the benefits and limitations of AI-generated code in software engineering.

The Impact of AI on Software Development

42:00 to 47:50

Exploration of how AI influences software creation and the importance of human oversight.

“Like, I think one of the things that this gives us is very interesting.”

Interesting Finds from Hacker News

47:50 to 56:00

Discussion of intriguing projects and articles from Hacker News, including creative uses of technology.

“So yeah, just some, hopefully some interesting thoughts there on writing code versus actually shipping code.”

The Rise of CLI in Software Development

56:00 to 57:28

Exploring the resurgence of Command Line Interfaces (CLIs) in the tech landscape.

“in someone's office and they had one it was kind of cool to just use it and know that this was the most powerful Mac on the planet at that point.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:12Gregor Vand:Hello and welcome to SED News. We are your hosts. I'm Gregor Vand. And I'm Sean Falconer. And as I think a lot of you do know already, this is a slightly different format of Software Engineering Daily, where we pick off some of the big tech headlines that you might have seen in sort of more mainstream news we're going to dive into a bigger topic in the middle and then we just look at some of our favorites from hacker news towards the end which usually ends up in sort of weird and wonderful rabbit hole type projects that developers have been working on so as usual we just like to kind of catch up on what we've been doing what has been in your sphere sean since we last did a scd news last month?

0:55There's been a lot in the last month. So I moved into a new house, which has been great. But while I also moved in a couple of weeks ago, I actually have traveled each week of this new adventure in this new house. So I was in DC and then I was in Seattle for our data streaming world tour last week. And then my wife was away this week. So we've been here almost three weeks and there's never been a completely full week where everyone as a family has been in the house yet so i'm looking forward for that happening and then confluent also the company i work for was officially acquired by ibm about a week and a half ago so that's exciting news that's all completed then yeah yeah everything's been completed so lots and lots of conversations now going on with those folks over the ibm side figuring out how we're going to work together which is exciting lots of stuff happening and then you're looking ahead to april my kids are going to be on spring break which means that everybody's going to be on spring break for vacation.

1:50So we have some travel coming up and then I have some work travel to Cloud Next. And then I'm also going to India. So lots and lots of stuff happening over the next little while. But what's going on in your world? You're in a different location than usual.

2:04Gregor Vand:Yeah, well, I'm in Scotland. So that's where I'm from originally. I try and come over here for a month or two of the year. When I do, I'm up in the far north in the highlands of Scotland. So it's always kind of fun. That's what technology makes possible. you can still do a podcast from all the way up here where there's i actually looked it up the town i come to is a population of 80 80 people so i thought i was from a small place yeah so i didn't actually grow up here but it's sort of where i've decided is where i come back to you know when i come back to scotland so yeah it's always quite a change from singapore yeah but that's i think by design but yeah no i mean work wise has been kind of exciting actually obviously small plug here for super base of course but yeah we've just launched so stripe we've been doing this interesting project called well called stripe projects which is a cli tool effectively so you install it through their cli and basically there's a bunch of partners it's not just super base you'll find like all the usual names on there but it's been really great i've been kind of overseeing that one and it got out yesterday so we're recording this on a friday and this became kind of public in preview yesterday it's just awesome when you see something like actually hit the public sphere that you've been working with engineering on for the last like we've really been crunching on this last couple months and nice to see that go out so yeah check it out and yeah i'll be in sf later this month as well that's for sessions that's stripe sessions so for obvious reasons as i've now just disclosed is like we're working a lot with stripe and stuff so yeah if anyone is at stripe sessions do look me up and say hello always great to meet new faces as well yeah that's exciting that'll be our first opportunity to meet in actual person in real person yeah in real life yeah yeah sure yeah for anyone listening sean i have never actually met in person so yeah this is gonna be cool and you're coming to singapore at some point as well yeah so we have two opportunities there we go yeah whether we'll do an in-person episode that's that's a different thing because actually recording in person means like a whole studio setup i think or something whereas people are surprised i've done a few in person it's there's a lot more technical complexity with doing in person recording on your computer remotely exactly yeah i've had a couple of guests on se daily just say oh can we go to a studio they're like well we don't have a studio se daily is a remote podcast that's how we do it so right moving on to the headlines so some of you that do listen to this sed news a lot we have talked about arm in the past, but more so the lack of ARM.

4:35Gregor Vand:But I think the headline here is ARM is back in a big way. And this all kind of focuses around CPUs, because the headline stories of any of this compute inference, the story has really been around GPUs. And we did talk a little bit about TPUs, which are Google's derivative of that. But we are really going back to CPUs. So yeah, Sean, what did you pick up on this one? I mean, it's like everything in text, Like the old thing is now the new thing again. Like it's always cyclical. I think part of it, the reason we have GPUs, we have TPUs and now, you know, CPUs are back. Like it's just, you need a lot of compute to run all this stuff.

5:12If you're running models, you want probably GPUs, TPUs to be doing sort of the matrix multiplication there. They're really good at that. But CPUs are also good at certain things. You know, they're good at logic and branching and task switching. And I think given now that people are running these agents on their desktop and not just running one, like running multiple agents. I talked to a startup that I worked with recently, and I was like, how many instances of quad are you running simultaneously? And they're like, well, I start to lose track of things when I get up to like six. So that's a lot of compute that you're running locally.

5:43And I think essentially that's what's kind of driving this is that people are running these agents locally. I don't know that that was something that we necessarily foresaw in the industry a year or two ago, that so many people would be kind of running these agent experiences on their local computer versus the cloud or something like that, where you maybe you're using some other type of compute but if you're going to be running these locally for things like open claw or some other version of claw that now exists or you're doing it with the agentic engineering tools that are out there you're just going to be consuming up a lot of cpu and a lot of memory so there's now a lot of i don't know appetite i guess in the industry to find

6:20Gregor Vand:more of this compute yeah and i think one of the interesting pieces here is that arm is actually going into more of the actually manufacturing their own chips not fabbing exactly because you fab being effectively the sort of manufacturing plants that can do this and those are incredibly there's only three really players that have those i believe is tsmc the taiwanese is the main player intel do have that capability and samsung is actually the other huge one that has that capability not entirely clear who's fabbing these for ARM but the difference being that ARM traditionally was supplying the designs of chips like that's their whole thing virtually every chip that is created has some design piece in it that comes from ARM and so they basically take a royalty if you want to call it that off almost every single chip that is created in the world but they hadn't really gone full on into hey this is like an arm chip with an arm logo on it for example and this is kind of where they're going and yeah as you've called out sean it's the rise of agentic tasks and it really has just accelerated as we touched on i think in the last scd news and in the last like two three months that suddenly everyone's doing agentic stuff not just engineers like anyone seems to be quite interested now in getting an open clock running and like leaving it on overnight on their mac mini or something so there's clearly just a lot of appetite now for having these independent environments that have their own cpu to run a specific agent and i think that's kind of what we're seeing here yeah and i think that's only going to going to continue like these companies whether it's anthropic or others are verticalizing a lot of their success around cloud code cloud desktop to be geared towards finance and geared towards lawyers and whatever profession that you have.

8:15So every sort of white collar job eventually is going to have people who are running these things in some sort of environment. So you need compute. You probably need some sort of containers and security parameters around this. So there's not just, there's, I guess, like a ton of people's startups all the way to large companies that are really trying to meet the needs of the market right now. And it's something that like everything now, when it comes to AI, it's exponential growth. it's not this kind of like you know nicely growing linear thing it's like as soon as something takes off it's like extremely viral and companies have to move really fast to essentially meet that demand

8:50Gregor Vand:yeah i mean nvidia are like are in this space as well you know they offer cpu racks but i guess they're still known i guess on the gpu side but yeah it's kind of i would say like quote great to see cpus coming back i think the fact that they've powered so much to this point in time from a technology standpoint and then everyone kind of thought oh they're like yesterday's technology almost but they're incredibly powerful but for specific tasks and you don't always need like a incredibly expensive gpu running every single task as i think a lot of people have decided just for the sake of if they can afford it that's what they're going to do but it isn't actually needed yeah when i did an scd episode unfortunately i can't remember the person's name but he was a turing award winner in supercomputing one of the topics of conversation that we had was how The industry's focused so much on workloads catered towards is essentially the math behind model crunching and inference has led to essentially poor performance in supercomputers for certain types of tasks.

9:47So there is always like a tradeoff, right? Like if you overly specialize in something, GPUs are really good at a certain type of math. TPUs are even more biased towards being really good at that particular math. you're also giving up something that might be important for other types of logic or compute that doesn't essentially fit the profile that you're going to run on one of those other pieces of hardware like a GPU. And when you're running agents, it's not like it's just model crunching. There's all kinds of other things that they're doing where you want to be able to ideally use the right compute depending on what the profile of the task is.

10:22Gregor Vand:Exactly. And it's just the volume here. like there isn't really a sort of realistic world for like an average person at the moment where every agent has its own GPU. So even for most people in tech, like the numbers don't add up if that's like what it takes. So if you want the number of agents that you think would get your tasks done, you're probably looking at CPUs to run most of those logic tasks. So yeah, very cool to see ARM back in the fold. And yeah, I read a book about ARM kind of just before, I think it was just before they IPO'd or something like that. And I just, I really fully at that point and understood like, wow, they have just behind the scenes powered so much of what we take for granted with all these designs, but never really got the recognition for it.

11:07Gregor Vand:So it's kind of, yeah, it's fun to see them actually stepping into branding their own CPUs as well. So exciting. moving on to this does sort of hit the main tech headlines in the sense like this was a lot of reporting by tech crunch on this one and i think especially anyone in software engineering has maybe seen it on hacker news as well but i do think this kind of really crosses into main news in terms of our world the light llm hack now this was a takeover of a dependency effectively that then extracted the credentials of many services so light llm basically gave nice api access to a multitude of llms and helped with like things like metering and so on so forth so it was a pretty used tool by so many developers out there so that kind of makes it unfortunately incredibly ripe as a target for someone to think about what can we do if we can get the credentials to that.

12:05Gregor Vand:And I think it's just the focus here, LightLM themselves were a ex-Y Combinator. And then yeah, the kicker here is the fact that another Y Combinator company, Delve, they were already in a bit of hot water for potentially fabricating. So they're a compliance startup. Sock2 reports giving, you go through the checklists and you get the, Sock2 is still a funny industry where it's actually accountants that do the final stamp on things. That's a whole other story but they apparently auditors as well yeah exactly they're your auditors and it just it still seems like kind of strange that you've got effectively a financial company saying that your security of your technology is up to scratch but that's how that system was put together and yeah there's been allegations that delve have been fabricating the sort of the rubber stamping of these reports unfortunately delve had given one of their reports with a clean bill of health to LiteLM before this happens.

13:05Gregor Vand:And it doesn't, I think that people are pains to point out, getting stock 2 does not mean they have checked every single process that is going on in the developer's workflow within LiteLM, for example, but it should be a pretty good indicator that they're following best practice. And while something kind of wasn't going by best practice here and it's had a catastrophic effect on a lot of developers. But yeah, that's, I guess, the sort of slight security-minded focus for me, but what did you make of this one sean yeah i mean i think it's just another example of where like compliance isn't necessarily security compliance is really a series of kind of check marks where you're saying yes we are following the best practices but every company pretty much that has some sort of data breach has some sort of hack is in compliant they pretty much all have like sock two or whatever various isos like whatever sort of security check marks that they need.

13:57And I think in a lot of ways, like compliance is really about insurance, while security is actually about trying to stop the attacks. So there's a very big difference between, I would say, like the companies are really like secure by default, and they take a hard stance on this versus people who are, hey, we have to make sure that we have these badges, so that when we're going through procurement with a vendor, they're not going to immediately just say no. I also think the interesting hearing here is where a lot of times when we think about hacks historically or breaches, it's like, what is it that people are trying to go after?

14:31Well, they might be trying to go after someone's password. They might be going after credit card information, banking information, personal information for identity theft. And here, I think there's now this new industry that's going to be the value is, hey, I want your open AI API key or I want your Anthropic API key. And then not only can they run up bills on you, but they could be leveraging those models to be doing something malicious against a bunch of other people too. So they can spin these things up using your precious tokens, run up a big bill for you while also using that thing to attack somebody else.

15:04Gregor Vand:Yeah. So this was discovered by a guy called Callum McMahon of Future Search, which is a company that offers AI agents apparently for web research. but he like documented how this all unfolded basically ended up with his machine being shut down so that's a pretty clear indicator that something has really gone wrong here in terms of how deep this sort of hack has gone it's something we've covered on se daily with guests obviously of companies that help to mitigate this the one that comes to mind is whiz we had rami mccarthy from whiz and we actually covered a case that he'd worked on that is exactly this it's like a dependency has a sort of leak within the maintenance of it and that just cascades through so many other packages and this one particularly just the number of developers right now using light lm is just shown to be such a problem i think these kind of like dependency supply chain attacks are hard to catch and it's easy thing for attackers to kind of exploit and then in a world and we're going to talk more about this during sort of the main topic, but in a world where we're using AI more and more to generate code and maybe not always looking at that code that closely and is picking up packages in order to do some type of work on our behalf, we might not have as clear a handle on what those dependencies are.

16:28We're kind of just looking initially at the outcomes of that. Does it pass the tests? does it kind of look correct if it generates some sort of ui based on what it is that we're trying to accomplish we commit that thing and then we get it running production and then boom

16:41Gregor Vand:we have some sort of problem yeah so the next headline here i guess again this might have crossed more slightly into the hacker news from a headline perspective but i think it was interesting that yeah open code really made a lot of noise in the community here and this really goes to like the fact that agentic coding has just taken off but most of the time we were talking about either cloud code or we're talking about codecs from effectively chat gbt and that means you're buying into again one of the big names and here was a product that's able to say hey we're actually fully open source you've got access to local models you've got access to free models or if you really want you can hook it back up to opus or something like that but i think that's just the amount of opinions that were then given about it is kind of really interesting.

17:31Yeah, I think we're going to, of course, see more and more of these open source alternatives. In a lot of ways, it reminds me of, you know, if you looked at the IDE market, maybe 15, 20 years ago, it was very similar to where open source IDs started cropping up as alternatives to the mainstream ones. And the thing that I've always thought about, or since people start, even with GitHub Copilot was, I think we've always gotten to a place where developers don't really like paying for their tools. And then you end up with like open source alternatives. And then with the IDE market in particular, the IDE very much became like a commodity.

18:06And there are certain IDs that survived that and people pay for, but predominantly I would say most became open source or were given away. Even things like Visual Studio that historically had been like a big moneymaker for Microsoft. So today, people are able to monetize these agentic engineering tools because they are incredibly valuable. If someone took away my quad code, I'd be very upset about that. So there is a lot of value there. But as you have more and more of these open source equivalents that are maybe, even if they're not as good, they might be as good. But even if they're not as good, maybe they're 80 % as good.

18:40People are happy to use something that's 80 % as good if they don't have to pay x dollars a month to use some other tool so then you get into a world where people are willing to pay for the tokens but people aren't willing to essentially pay for subscriptions at scale to have those tools so if you want to monetize that market like how do you create a big enough value and a big enough moat around what it is that you're providing so that people don't just go to the open source equivalent yeah and it was kind of interesting just to look at

19:09Gregor Vand:some of the conversation around it in hacker news as well and i guess what maybe even i hadn't sort of fully appreciated is the fact that like clot code okay it runs in your terminal but it is actually an electron app i didn't even realize this and like that's how you kind of achieve all these amazing visual niceties effectively like yes it's in a terminal but there's a ton of stuff going on from a rendering perspective like terminals are yeah being hacked a bit in this way I love the visuals, but then I was kind of starting to think, how is this even possible inside a terminal? Codex apparently is done in Rust, and people say that that is a tangible speed difference.

19:49Gregor Vand:And one of the things that has been questioned about OpenCode is the fact that basically you're using up a gig of RAM basically for a terminal, and that it's a pretty chunky TypeScript code base that's running this thing. So this is it. When you're designing these kind of tools for developers, well, guess what? You're going to have some pretty, I think, spicy opinions on like, how did you put this tool together in the first place? So I think that's going to be an interesting adoption data point that they have to think about. In mobile application security, good enough is a risk. GuardSquare uses advanced, multi-layered code hardening techniques and automated runtime application self-protection and mobile application security testing, combined with real-time threat monitoring to deliver the highest level of mobile app security.

20:40Discover how GuardSquare brings all these together to provide mobile app security for your Android and iOS apps without compromise at www.guardsquare.com. Today's episode of Software Engineering Daily is brought to you by Unblocked. Your coding agents have access to your code base. Maybe you've even connected other tools via MCPs. But access doesn't mean context. Agents can't reason across MCPs. They don't know your architectural decisions, your team's patterns, or why the API was shaped the way it is. So agents look in the wrong place and deliver bad outputs. Then you spend time correcting, turn after turn.

21:19Unblocked is the context layer your agents are missing. It synthesizes your PRs, docs, slack, and tickets into organizational context that agents actually understand. So they make better plans, write higher quality code, use fewer tokens, and require fewer correction loops. If you are running Claude Code, Cursor, or any agentic workflow, Unblocked is worth a look. Get a free three-week trial at getunblocked.com slash sedaily. You know, there's a lot of these tools now available. You have Claude, you have the mainstream sort of commercial tools from OpenAI, from Anthropic, and from others. You have the open source equivalents, like you have Klein, you have OpenCode.

21:59There's a whole plethora of these. How many can the industry really support? I feel like sometimes one of these tools crops up, kind of catches fire for a period of time, gets a lot of GitHub stars and love and people start using it. And then it kind of like tails off as the new hotness hits the market as well. So I don't know that that's like sustainable in the long, at some point, I think some of this stuff starts to level out. And they'll emerge like a couple of main players that dominate the market. And maybe there's a couple alternatives, but I don't think you can have like 100 alternatives.

22:29Yeah.

22:30Gregor Vand:And just to round out the headlines, this sort of, I guess, broke cover maybe a couple of weeks ago, but I think it just sort of is useful to touch on a slightly wider story that we're seeing play out, which is how OpenAI and Anthropic are kind of, how are they playing their markets now? and there was obviously a lot of noise around how the pentagon they wanted to the department of war or justice or whatever they're called these days wanted anthropic to allow clods for what they call all awful purposes and that it does include things like surveillance and autonomous weapons and anthropic refused and they didn't get the contract and then apparently open ai quotes swooped in and signed a deal and there's a lot of i think user backlash around this and it did make Anthropic look like the real human winner there if you want to use one of these technologies like why would you use the one that really wanted to get in with the government the US government and if we kind of then just like play that through to like well where have we seen OpenAI versus sort of Anthropic like where are they positioning themselves like it does to me at least look like OpenAI started to becoming this kind of super app type construct which which we've touched on before, Sean, like in sort of other deeper topics and Anthropic, Cloud, they're really going like full enterprise and in a sort of say in a positive way, but they're really going and saying, we can be the enterprise tools that you need.

24:00Gregor Vand:If you look this back to the whole Pentagon thing, well, a lot of companies really don't want to be associated with a company that they think is like aiding the government with say surveillance or as we speak the iran conflict is on right now and we don't know how the technology is being used there so just kind of staying away from that i think is what companies are looking at and it's just a very interesting development i think in this sort of the landscape of the two companies yeah absolutely i mean i think that it ends up kind of highlighting i think something where those two companies are already somewhat have like opposing views of what is the right thing to focus on when you're building you're sort of responsible for building these models and even this Dario and some of the other founders at Anthropic were open AI at one point and they kind of split off because at least the story goes that they didn't agree essentially with the philosophy that open AI was taking towards like building models or releasing models.

24:56They wanted to make sure that as a company that before they released the model, it was really well tested and really like security and trust was sort of the number one value. So they splintered built Anthropic. And I think this whole thing that's happening politically has been probably some of the best marketing that's ever happened to Anthropic. Like you saw people leaving or canceling their open AI subscriptions. I think even under the best of circumstances, like with politics, people, companies, individuals might not be that comfortable with a company supporting something like surveillance of individuals or automated warfare or anything like that.

25:34But on top of that, we're also in a place in the United States where politically things are very polarizing right now. And there's a large chunk of the country that doesn't want to be associated with the government at all. So this is a very, very polarizing issue. And it's kind of Anthropics kind of chosen to take one stance on this. OpenAI has taken a different one. And as a result, I think it just highlights the differences between the two companies even more. Yeah.

25:59Gregor Vand:And it's not like we've seen things like this probably appear in the past where a company potentially anchors itself a bit closer to a government or not. I'm not going to bring up any specific examples. But what I will say is like, in the past, okay, would one have refused to use this? I'm just taking pure examples that might not be true here. But would I refuse to use Google, if I knew that they were aligned with a certain government? I'm not sure because like okay it's a lot of your search history and that kind of thing but now that we're talking about what LLMs are used for well it's quite often people are divulging a lot of their personal information beyond just oh this is what I purchase or this is where I'm located right now which is you know it's personal information of course but people using LLMs they're kind of writing about their deepest darkest thoughts and in companies especially people are just literally hooking up whether through official integrations or full-on just pasting in slack conversations and documents and that's it like where does that information end up then ultimately and i think that's why this is such a concern for companies especially yeah absolutely there's something i think more personal about chat than necessarily search i think people have kind of exactly like taught themselves that okay well maybe i shouldn't put my social security number into like Google search.

27:18I don't know that people think of that necessarily with chat as much, or they, you know, paste documents in that are incredibly sensitive, not really thinking that through because of the value. There's just so much value with using these tools that people kind of throw out maybe their reasoning skills sometimes with like, should I actually be sharing this? What is the consequences of sharing this information?

27:38Gregor Vand:Yeah, I've definitely noticed that people that previously would have sort of said, I will never divulge certain information to a tech company. And then the next day, oh i've had this amazing conversation with claude and told them all these things and i don't comment on it i just think that's really i'm just observing it's very interesting that this is someone for example that i know would not have done this in the past but actually just they've felt some trust there that's a whole topic for a different day but i think it just sort of underscores like why who you give your information to especially between the you can call like the two main players here anthropic and open ai as soon as one then anchors himself to a entity that people disagree with that's like adds a whole extra layer so we'll watch how this plays out i guess and yeah we'll sort of see does this affect open ai enough that they have to walk anything back or are they kind of going more like palantir style they actually think there's too much value to be had by being very close to the government and helping them achieve things like i guess we'll see how that happens yeah and i think from like a consumer standpoint a lot of times when there's so much value behind whatever the product is, consumers end up turning off a blind eye to maybe things that they wouldn't normally agree with.

28:48You know, you can take any major retailer in the world that's probably manufacturing goods in developing countries where maybe the standards of work are not necessarily great. People kind of sort of know that, but they also like their inexpensive, you know, t-shirts or whatever, you know, and they're kind of willing to turn that part of their brain off because they're focused on like their own personal outcomes and i think you could potentially see something similar with you know even if you don't in conversation agree with what open ai or any other company is doing there's so much value behind their tools you kind of turn it off when it comes to your own personal objectives yeah so it's that time within sed news

29:29Gregor Vand:we're going to move to the sort of main topic, deep dive. And today we are looking at effectively writing code versus shipping code, and especially the emphasis on shipping code. How does this look in the age, if you want to call that, of LLMs now that, you know, probably at least 50 % of developers in a given company are going to be using LLMs to write code? What are the kind of maybe cascading effects on how that code actually then ends up to be frank on main branch you know what does main branch look like from a speed perspective and actually getting things merged in and we kind of anchored this around there was a circle ci as you know a company that many people know you know in the cicd space and they did a 2026 state of software delivery report so that was just back in february we're recording this just at the end of march they said that they analyzed this is gonna be quite data heavy this main topic so get ready for some numbers because thinking about when we talk about writing code versus shipping code it's basically all numbers like lines of code and how many mergers are made and so on and so forth so they said that they analyzed 28 million ci cd workflows across 22 000 orgs in 149 countries so that's a pretty good spread this is the interesting anchor points here that the daily workflow throughput was up 59 9%.

30:53Gregor Vand:But it's kind of misleading that the top 5 % of teams nearly doubled the throughput, whilst the median team increased only 4%. So that basically suggests that super high performing teams could super accelerate their throughput, i.e. things that got merged into main, again, just to kind of bring it back there, they could double their throughput, the top 5 % of teams. But the median, just call it like, you know, the average, I know average and median is different, but when people think about it that to only increase four percent on throughput for most for like just call it the majority that's pretty interesting because you know most developers are saying oh i'm so productive i can you know i can do so much more per day now if i use an llm but if actually from a team perspective and a shipping perspective the majority are actually only increasing by four percent that looks way off what we might have expected i think from this yeah You know, there could be a couple of things going on.

31:52One is that, you know, not everyone has reached the level of the most high performing teams. Adopting some of these tools is still relatively new to many parts of the industry. So you're going to have people who are kind of trailing behind, maybe they're kind of as a company, maybe just, you know, dabbling, or this is like fairly new efforts. So they're kind of getting used to how do we get value out of this. but I think the other thing too you could say is like well maybe the teams that are shipping these like you know top performance maybe they're shipping a lot but are they shipping also a lot of bugs and they're just you know okay with that whereas other people are you know trying to control that so there's a bunch of stuff that it would be nice to kind of be able to dig into the details of like why is it that some people are at you know four percent increased throughput whereas you have some people who are you know double this and there could be many many different reasons for that to be the case.

Read the full transcript

32:42The other thing too, is that, you know, just because you speed up writing of code doesn't mean you speed up the entire software development lifecycle. So, you know, you have this software development lifecycle of, you know, just a very simplistic one of like discovery, sort of requirements, coding, testing, putting it into production, and then running that thing in a loop. Well, if you speed up that middle part of, you know, writing the code, it doesn't necessarily mean all the other parts of it has sped up as well. So those can become choke points in the company. And on top of that, if you speed up the lines of code, those choke points might become even slower because the thing like, let's say it's your security team that has to review these things.

33:23The fact that they were somewhat slow in the past was maybe okay because the part of writing code was really, really slow in comparison, but now they suddenly have like a thousand times more things to do. So what happens as a company, either you cut corners and you start releasing products that don't necessarily go through that full security review, which leads to negative consequences. Or you slow down the entire release cycle so that you only get 4 % increased throughput because the poor security team is completely slammed and overwhelmed. I mean, it's the same thing even with reviewing PRs.

33:57If we're generating more code, that means there's more PRs to review, which means that that has to be paid for from somewhere. So either you have your engineer or spend all their time essentially reviewing PRs, which slows down, of course, things getting to production, or you leverage LLMs to also review the PRs, which then also potentially has certain consequences where less and less people within the company really sort of deeply understand what is happening in the code base, what the dependencies are, and then you end up with some sort of issue on production. And then the people who are debugging it are also not that familiar with it because AI wrote it.

34:33There's all these kind of like consequences of what's happening in the industry right now. Whether it's like, hey, we're not really seeing the full throughput value of the LM because we're generating code a thousand times faster, but these other things are getting completely bombarded and slowing things down. Or we're generating code a thousand times faster and we're pushing it out of production. And then we're also generating bugs a thousand times faster. Yeah.

34:58Gregor Vand:The kind of stat that does really make this, you can really visualize it, is that feature branch, like this report that came from CircleCI, they saw the feature branches, the throughput was up 15%. So, okay, great, 15%, generally more feature branches are being created. But the main branch throughput itself was down 7%. So basically, but actually, sorry, that's throughput and the actual creation of branches, feature branches are up 50%. Whilst the main branch, again, only like sitting only at 1%. So it basically just means that a ton of code is being generated and a very, very, very small fraction of that code is ever making it onto main branch.

35:41Gregor Vand:And I think to kind of go exactly to what you were just saying there, Sean, you know, that middle piece, the human in the loop there is still obviously like the biggest question mark. You could call it a bottleneck. But I think most companies are saying it's still the most critical thing that they believe needs to be there. like it's all very well generating a ton more code or a lot faster at least and having all these supposed features or or so on like all queued up and waiting to be potentially merged in but the actual review piece well that's where anyone really is going to be saying well hey we're not comfortable with these just like being whisked through by like an lm that sort of knows how we do things and i think that's the piece that's going to take like a lot longer to really kind of make it through to say like production production apparently there's an example from the new stack back in march we're in we're still in march so i mean coming out in april but we are a flat circle yeah well it is an ai terms yes but or ai age but vs codes uh so they've actually moved to weekly releases after 10 years of monthly and they say ai is actually making this possible but they say that only works because they say i mean they obviously i think bs code specifically probably have you know reasons that they at least want to push the idea that this is all possible but you know they say it only works because they have strong automated checks you know and a team that's invested in this review pipeline so yeah i think it's just like what does it actually look like to try and get your team around this paradigm i think is probably what a lot of people are wondering about?

37:24Yeah, I mean, I think that the reality, at least at the moment, is like organizationally, you have to probably change. Because if suddenly the resources that are scarce is not necessarily the production of code, but the resources are for that scarce are essentially verifying the code is correct. Like Addy Osmani has written about this, and you know, generation is not the bottleneck anymore, verification is. So if it's all about PR reviews, or in the world of like infrastructure, The core piece that a lot of times really matters is like networking, security, how you do authentication. And if you mess that up, you've basically, you might destroy the core infrastructure of thousands of different customers.

38:07You can't mess that piece up or it's going to be dramatically negative consequences to your business. So that's where you really need to take care and maybe you have to deploy more resources there. And you can kind of cut back resources in other places where suddenly people are, you don't need five people to maybe generate as much code, but you do need five people to verify the code is correct. So it kind of, it's almost like we're shifting the problem where the problem historically has been, it takes a long time to write the code. That's been the sort of hard, long cycle. and then now I don't know that we've really sped up everything.

38:43Like I was saying, like we've kind of shifted where the problem occurs essentially. I think the really negative part of this is that in the sort of media, we focus so much on these headlines of like, you know, such and such a company replaced a team of engineers with one engineer and, you know, six open claws or whatever it is, or people were going from idea to app with one prompt and built an application in 20 minutes. And I think the thing that we're confusing is that, you know, when the demo comes together in afternoon, I think it ends up recalibrating what the stakeholders think possible. Like executives who watch some AI generated prototype naturally ask like, hey, if we built this in a day, why does production take six months?

39:24And it's this whole like 90-10 rule that has always existed in software where first 90 % of the work takes like 10 % of the time and the last 10 % takes the other 90 % of the time. because there's all this other stuff that has to happen around productionization. And we're really confusing like prompt to prototype with prompt to production. And I think that is going to have massive consequences of the industry. And I know engineering teams are under tremendous pressure right now to be shipping at the speed of essentially prototyping. And you either end up not being able to do that, which could be negative for your career, or you end up in a situation where you bypass the normal checks and balances.

40:03and then that's going to, I think, lead to more and more of these outages and failures. I mean, even if you just statistically take, like assume that with human-generated lines of code to the number of errors that you have, if that statistically stays the same for AI-generated code, then if you're generating more code, you're going to have more errors and you're going to have more production outages. So I don't think it's a particularly bold statement to say that we can expect more outages from companies. over the next year easily because of the fact that if companies are in the place where they are actually able to put this into production faster and there is all this pressure to do that, we're going to end up just generating more code that generates more bugs at some point.

40:49Gregor Vand:Yeah, and as you touched on this, it's going to be a complete, it's a sort of a culture competition at this point because the company that already wasn't putting a ton of thought into the pre-writing code and then the post writing code piece. Well, they never saw that anyway. So like, well, awesome. We can just accelerate the bit that we thought about most, which was just generate code and generate stuff and ship it. And yeah, I think they're starting to get sort of a bad surprise in many areas when things don't suddenly like accelerate 10x because the problems have accelerated 10x. And it's the companies that already have these, literally from writing whatever people call them, like PRDs or RFCs before something gets shipped.

41:31Gregor Vand:again use if you augment that process with llms well actually that's very interesting because a ton more thought and discussion can happen around something before it's even thought to be getting code written against it and then the code being written against it potentially gets sort of accelerated but i can't say that i think the best features are getting accelerated by llms it's actually i would almost argue it is that maybe that first piece where the lm has done a ton more thinking behind the why's and the how's that everyone can then read over and comment on and then you know a classic pizza team works on the feature itself and that's still exactly what should be happening not to make me sound anti-llm on code but because i'm very much pro it's just it's interesting to see where does bringing lm into the loop actually affect the the final outcome and the quality of the outcome including security and robustness and all and all these kind of things yeah i agree i mean i think that i'm certainly don't want to come out sounding like i'm anti ai generated code it's absolutely not the case i think these tools are tremendously valuable and as i said earlier you know i'd be very upset if cloud code went away like very dependent on these things at this point.

42:50But I think that the problem here with sort of sensationalizing the fact that you can go from prompt to, you know, building a compelling demo, and then sort of confusing what that really means, hides some of the real value here. Like, I think one of the things that this gives us is very interesting. And you kind of touched on it there, where, in many ways, like software has become a thinking tool, because the cost of essentially generating that idea is so low now. So as a product manager like myself, I could spin up a working demo of an idea instead of writing a one pager and then hoping my stakeholders can kind of visualize it.

43:28The demo isn't the product, but it's a better way to have a conversation essentially about the product that we want to build. And in many ways, I think we're entering this era of where AI is kind of making software ephemeral. And that's a real shift in how we can do work because we can, anybody, not just people who are technically adept, like suddenly you have people who, you know, with a little bit of training can have the ability to convey their ideas through software. And I think that's very, very interesting because, you know, if software can only be created by a small set of people in the world, you're of course going to have certain biases that come with that or certain, you know, worldview.

44:08If suddenly you open that up where it doesn't mean that you're building production software, but if anybody can kind of prototype something, then it's a way, I think, for a completely generation or essentially a completely different set of people to be able to convey their ideas and software. And I think that could lead to some really, really interesting outcomes that we haven't seen before. And that's always the case when you kind of bring new worldviews, new ideas, new cultures to something that they didn't have it before. So I'm excited about that. And I think that there's tremendous value here.

44:40I just think that a lot of the real value gets kind of lost in the conversation. And by overemphasizing the idea that, you know, prompt leads to faster productionization, it's going to have negative consequences that companies don't really necessarily foresee right now. And it's hurting engineering in general by sort of misleading where the real value is right now.

45:03Gregor Vand:Yeah. So kind of the, I guess the sort of TLDR here is just effectively that your sort of validation layer if you think that llm coding like the point of it basically is to increase throughput generally and by throughput here we would define that at this point is as merging domain in a state that you are that meets the bar of quality that you as a company think is acceptable then you know the validation layer effectively needs to keep pace with that and at least yeah i mean where i sort of sit i'm not talking specifically about where i work but like you know you can look at a bunch of code bases open source etc it does not look so far like that is quite keeping pace if that's your kind of metric because rightfully probably at this point the human in the loop is still very much there and still a lot of like post discussion happens and a lot of the volume of code needs to then be reviewed by a human and And that's a very labor intensive task because I get the impression that a lot of people aren't then comfortable saying, well, I'll just use an LLM to review the code that was generated by an LLM.

46:09Gregor Vand:I mean, okay, sure. Like I know that people use models to rank outputs of models, but like this is precisely the piece. I don't think anyone's gonna still have a job at the end of the day if they told their engineering manager, well, the code that was written by an LLM, all I did was have it reviewed by another LLM and then I pressed merge. Like in professional software production for big companies that, for example, that we work at that just is not going to have you a job at the end of the day if something if something goes wrong it will just be you on the end of the line saying like well that was not the way to achieve that yeah i mean we both work for companies where if we mess up something in our core infrastructure that means we mess up the core infrastructure of potentially thousands of customers oh yeah like six figures of people could get messed up easy yeah by one problem yeah exactly so like there's like massive massively negative consequences to something like this I think part of it also is companies kind of have to think about, you know, where if you want to accelerate things, are there certain parts of the stack that you could accelerate?

47:09Like, are there places where, you know, a bug is more forgiving than other places? You know, maybe it's not networking, but could it be part of the UI layer where you can move faster? Because, you know, if you end up messing up the UI, like some way, like you can, you know, fix it immediately, push out a change is maybe less catastrophic, depending on what you know what the company provides. I think you kind of have to think of this as like layers of an onion, like which layers of that onion do you need to be really, really scrutinized and make sure are up to, you know, whatever your power is.

47:41And then are there other places where you can be a little, you know, looser on that and the consequences are not going to be as catastrophic.

47:48Gregor Vand:I think that's a great place to leave it. So yeah, just some, hopefully some interesting thoughts there on writing code versus actually shipping code. You probably got a few anecdotes, listeners thinking about how you work in your own companies now and just like, what does that look like in your own flows? So, and we're always happy to hear about those as well. So moving on to our final piece of the show, usually our kind of favorite piece of the show. This is Hacker News, but nothing too serious, just interesting things that have popped up on Hacker News. I think I could scan to see what Sean's is and I want him to go first because I definitely saw this as well.

48:27Gregor Vand:And I was like, I guess we're going to cover this because spoiler, it's to do with Doom. So, you know, it's going to be Doom on something. What is it this month? Yeah, it's so funny to me, probably every week, certainly every month, there's always going to be a Doom on something project on Hacker News that, you know, catches our eye. But it's also funny to me how like doom has just become that thing that everybody wants to like jam into anything we're going to run watch we're going to you put it in your wi-fi enabled earbud like and now this is doom on over dns because why not so really really fun project i think then one was the the person who posted it there's a github you can take a look at but essentially the project It compresses the entirety of the shareware Doom, splits it into nearly 2000 DNS text records and across a single CloudFlare zone and then plays it back at runtime using nothing but a PowerShell script and public DNS queries.

49:23It's just like amazingly absurd, but incredibly awesome.

49:27Gregor Vand:I do think despite the Doom in TypeScript types, which was an incredible feat, and we did have an episode on that with one of our other hosts, I got to say this to me is like creatively just one step above because like I don't think anyone thought that writing software via DNS was a thing or could be a thing. no and as usual like doom i think i guess why is it always doing well i mean it's just a kind of meme developer meme at this point but the fact that it's this very tangible outcome that as soon as someone's like well it's doom it's not just saying oh i created like a calculator app by from dns it's doom a game that's insane so yeah that's very right yeah i'm sure the creators of doom like back in the day had no idea like they had no idea they were gonna become like the it's almost like a some sort of benchmark standard at this point where it's like can you get doom to run on you know x device or whatever it is yeah the two johns i read that book masters of doom that's a great book if you want to ever hear about how doom actually came to be but nothing to do with all the funny ports we've been talking about but yeah on my side a couple of sort of just interesting ones one was why this is called why so many control rooms were seafoam green so i I often do bring up like some of these maybe more designy ones that hit Hacker News because it seems like Hacker News has this amazing community of people that can sort of dig into like why things were designed a certain way whether it's pure colors or so on and so forth.

50:57Gregor Vand:So this was posted by user Amori Meltzer. Thank you for that. And the TLDR is seafoam green is apparently when we talk about seafoam green we're talking like a sort of very light green color, mossy, light, very light moss in scotland right now i see it everywhere but think of a very light pale green and it's on walls in these sort of bunkers of control rooms and why is that the case well the tldr is there is a chap by the name of faber birren who was somebody at the art institute of the university of chicago and when he was in new york he actually became like a color consultant and he helped people understand well if you paint your walls this color in a certain restaurant like or a butcher for example like that might drive more sales of steak so like because the meat looks redder against a certain color for example he was then hired by the government to during world war ii to create pallets for different like buildings effectively to help people like just understand And like, oh, this area is a danger zone.

52:05Gregor Vand:Like so far red was used sort of for all emergency stop buttons, which I think we would think of as pretty obvious these days. But that was not obvious before some thought was put into it. Anyway, light green, that was simply used to reduce visual fatigue. Apparently it's seen as a color that you can sit in a bunker with all these switches and dials all day. And that is apparently the color that reduces fatigue on the eyes. So given that you've got to be staring at, back then, it's certainly staring at all these other machines. So very just interesting, funny bit of trivia there. I've heard sometimes certain fast food chains choose a particular color in order to encourage people to not stick around because they want to turn over the seats.

52:48I don't know how much truth there is to it, but there's definitely something where certain colors, I think, do affect people in different ways.

52:54Gregor Vand:The one that I never understood, and it's probably why I have not flown this airline for maybe, or I can't even remember, since I was a teenager, literally. Ryanair, I'm sure some people know about Ryanair. It's a very low-cost airline in Europe. But they used to put a very bright yellow plastic thing on the top of their seats. And it was just like headache-inducing. And the problem is you can't leave the plane. You're on the plane. You can't get off for the next two hours or whatever. So I have no idea why this was, I mean, they have a very controversial CEO. But yeah, you were just talking about wanting to make people leave your establishment.

53:27Gregor Vand:Well, that makes sort of business sense to me. But why trap people in a plane where they probably want to get out? That does not make sense to me. But what else did you pick up? Well, so the other one that caught my eye here was a guy who basically took Tesla Model 3 crash car components in order to assemble the Tesla Model 3 computer on his desk. and this was motivated by tesla's like bug bounty program but he had to like salvage a bunch of this hardware from these various parts the article kind of goes into detail about like how he did this so the i guess the car computer has two main parts there's a media control unit which is like the mcu and there's an autopilot computer or ap then you need like power supply you need a touch screen but i think the the thing that they ran into he ran into the most challenges around was just kind of like connecting everything because the actual connectors in the car is kind of like this like massive cables you know like crazy amount of cables yeah yeah so then he had to figure out like how can he actually connect the mcu to the other devices in some way where he could you know provide the right cabling because essentially no one it's hard to find that part because it's either been destroyed or it's like just this mass of cable that's not like super useful yeah that does feel quite deep end bug bounty okay has provided a reason for that but yeah still that is like you got to be really into this kind of stuff to like pick through how on earth this is gonna all fit together so yeah yeah just a quick one to round out i think it's interesting yeah the mac pro has been discontinued by apple no yeah i know i can still remember yeah going to apple store you know in like 2012 maybe in new york and like the mac pro is like this like beautiful aluminium box and it's like that was if you were any kind of like serious industry especially you know video or something like if you had a mac pro like that meant you were a professional on this like but apparently yeah mac studio is kind of where their head is mac studio i actually looked this up i wasn't fully aware of it but mac studio is kind of looks like a mac mini but it's just got you know beefy internals and it looked like their kind of product lineup was a bit muddled by this point but they had the mac studio and they still had the mac pro they've decided to officially kill the mac pro and just focus on on the mac studio line so probably quite a sensible move but yeah a little bit sad for some of us that still remember like all the that was the sort of if you ever wanted in someone's office and they had one it was kind of cool to just use it and know that this was the most powerful Mac on the planet at that point.

56:10Yeah.

56:11Gregor Vand:I guess briefly looking ahead, anything you're thinking might pop up when we're talking next? I kind of already stated my point of view on prediction, but I think we're probably looking at more outages. With this extra code and code and validation problems. Yeah. Yeah. For sure. Definitely something to keep an eye on. Yeah. Yeah. Absolutely. My side, I think, again, not to plug here, but it has been interesting, the Stripe projects thing, people were asking, does that make them like a marketplace in a CLI? I can't give the answer to if it's yes, definitely yes or no on that one. But I think what we are seeing, and I think we're going to see more of this, are other players coming into this CLI driven market.

56:53Gregor Vand:I think it was sort of surprising that to some people that Stripe had decided to branch over here. Perception was quite positive. So I think we're going to see more of that i think we're going to see more companies that never would have thought to like go back to their cli and maybe think that that's where they now need to target some product releases for example clis are kind of getting a moment again because of especially because of you know mcp versus cli that's all that's actually a whole other topic we should cover maybe next month but cli is coming back in a big way oh yeah clis it's crazy like everything i think is now about the CLI.

57:29And I wonder, you know, I think there's probably a bigger topic conversation around CLI versus MCV versus, you know, APIs and all this sort of stuff. But yeah, CLI is definitely having a moment right now, which is interesting.

57:40Gregor Vand:Yeah, so I'm sure something to that effect will pop up, I think, by our discussion next month. So thank you again, everyone tuning in to SED News. We hope you have a good April and we'll catch you next month. Yeah, thanks everyone for listening. We'll see you next month. Thank you.

From the publisher

SED News is a monthly podcast from Software Engineering Daily where hosts Gregor Vand and Sean Falconer unpack the biggest stories shaping software engineering, Silicon Valley, and the broader tech industry. In this episode, they cover the resurgence of ARM and CPUs as serious compute infrastructure for running local AI agents, a supply chain attack

The post SED News: OpenCode, AI Code vs. Shipped Code, and the LiteLLM Breach appeared first on Software Engineering Daily.

More from Software Engineering Daily

All 195 episodes
SED News: OpenCode, AI Code vs. Shipped Code, and the LiteLLM BreachSoftware Engineering Daily · 57 min
Listen in VO