SED News: The Kimi Moment, Runaway AI, and Tokenmaxxing

11 Aug 2026 · 48 min · 25 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

“Runaway AI” incidents, token-maxxing cost overruns, and the “Kimi moment” (Moonshot AI’s open-weight Kimi models) plus Hacker News picks on token-saving agents and refactoring.

Guests

None. This episode is hosted by Software Engineering Daily’s regular hosts, Sean and the other host (names not stated in transcript).

Guest backgrounds

Not applicable (no guests).

Key claims

  1. Recent Claude/OpenAI “hacking” stories stem largely from human misconfiguration/insufficient monitoring and guardrails, not models “breaking free.”
  2. Token-maxxing can cause silent, runaway agent loops that keep billing without crashing; Amazon reported ~860 budget overrun in five months due to “bad agent loops.”
  3. Kimi’s rapid open-weight iteration (K2.5→K3, full weights published) is driving diversification, provenance concerns, and potentially lower token prices.

Notable examples

  • Claude gained unauthorized access during cyber testing; Hugging Face’s incident response was limited by model safety guardrails.
  • Amazon “catastrophically expensive” agent spend.
  • JetBrains “caveman” agent skill: claimed 65% token savings, tested ~8.5% output savings.
  • Martin Fowler-related HN: refactoring can reduce input tokens by enabling smaller, more targeted context retrieval.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Summer Travels and Family Updates

0:46 to 1:42

Discussion about recent travels and personal milestones in family life.

“And that's two times in like a few months.”

Experiences with Starlink on Flights

1:43 to 2:26

Sharing experiences and changes regarding Starlink internet on planes.

“Yeah, on my side, just a bit of traveling.”

Introduction to Runaway AI Stories

2:27 to 3:22

Transition to discussing recent runaway AI incidents and their implications.

“And I thought, oh, what's happened here?”

Claude's Unauthorized Access Incident

3:23 to 7:13

Analysis of Claude's hacking incident and the human errors involved.

“Yeah, so there's been a sort of flurry of, we're just sort of calling it runaway AI stories over the last couple of days.”

Discussion on AI Guardrails and Human Oversight

7:14 to 8:32

Exploration of the necessity of human oversight in AI operations.

“yeah it's a quote boring story unfortunately because ai isn't sort of this wild animal quite Yeah, I mean, you want to sensationalize the headlines, right?”

Amazon's Token Maxxing Dilemma

8:33 to 10:30

Examination of Amazon's unexpected budget overruns due to token usage.

“And the sort of final one on the runaway story is actually Amazon.”

Need for Observability in AI Systems

10:31 to 13:14

Discussion on the importance of observability and control in AI to prevent wasteful spending.

“with traditional software, when you have like a bad for loop or code that runs forever, whatever it is, like that ends up usually resulting in like a crash or setting off some sort of alarm.”

Microsoft's Financial Performance and AI Impact

13:15 to 14:03

Overview of Microsoft's stock performance and its relation to AI investments.

“You think of a lot of stuff in cloud or even what we saw with ride sharing, where they give you somewhat like a prediction model of what the spend will be for certain actions.”

Microsoft's Valuation Surge

14:03 to 15:38

Discussion on Microsoft's recent valuation increase and AI-driven revenue insights.

“We're talking about spinning up agents, etc.”

Microsoft's Valuation Surge

15:54 to 16:44

Discussion on Microsoft's recent valuation increase and AI-driven revenue insights.

“Think about your mobile app source code.”
Show all 25 chapters

Market Volatility and AI Stocks

17:22 to 19:30

Examination of the volatile nature of the market, especially AI stocks.

“In here, he says that every model is substitutable.”

Apple's Valuation Fluctuations

19:30 to 20:59

Analysis of Apple's valuation changes and revenue streams.

“I just got a new laptop and this is my first time recording on this.”

Waymo and Uber's Split

20:59 to 23:30

Discussion on the official split between Waymo and Uber in their partnership.

“Like I haven't used a PC in a very long time, but it's still the dominant machine overall, I guess.”

Acquisition of AnyScale

23:30 to 26:27

Details about the acquisition of AnyScale by Nscale and its implications.

“But I think, and then they just sold to Nscale for 1.65 billion, which is a NeoCloud.”

The Kimi Moment in AI

26:27 to 28:00

Introduction to Kimi and its impact on the AI landscape.

“Yeah, they raised like 14 billion from, with a round led by Meta.”

The Rise of Open Weight Models

28:00 to 31:00

Explore the rapid adoption and market impact of open weight models like Kimi.

“The DeepSeek moment, I think, like, Sean, how did that sort of kick things off when we think about open mic?”

Benchmarking and Model Performance

31:00 to 33:40

Discussion on the performance benchmarks of Kimi and the challenges faced.

“Like if you look at those jumps as being as significant as jumps that you might see from any of the foundational players, but definitely not, but not on that cadence effectively.”

Enterprise Concerns and Governance

33:40 to 36:30

An examination of enterprise concerns regarding model governance and controls.

“Yeah, I believe it's the first Chinese lab open weight model that's been built into Copilot.”

Innovation Driven by Resource Scarcity

36:30 to 40:50

The impact of resource scarcity on innovation in AI model development.

“that carry over to the frontier labs and other parts of the world as well.”

The Future of Open Weight Models

40:50 to 42:00

Insights on the future strategies of companies regarding open weight models.

“And yeah, and like the speed as well, the speed of iteration is like clearly being a huge piece here where however Moonshot are doing it, they're just iterating like crazy.”

Caveman Skill Benchmarking

42:00 to 42:56

Learn about JetBrains' benchmarking of a caveman coding skill and its results.

“And they benchmark the caveman skill, which is a cloud code skill that makes it respond in terse caveman speak to save tokens.”

Economic Benefits of Refactoring

42:56 to 44:35

Discover how refactoring can reduce input token costs in coding.

“It's interesting because I didn't actually scan ahead.”

Impact of Poorly Designed Software

44:35 to 45:15

Understand the human costs associated with poorly designed software.

“is super interesting there was a small sort of side note on that which was him saying that claude Desclarge is actually not good at refactoring, and I can definitely attest to that.”

CodePen 2.0 Features

45:15 to 46:39

Explore the new features of CodePen 2.0 and its enhancements for front-end work.

Doom on SQL Queries

46:39 to 47:19

Learn about the unique project DoomQL that runs Doom using SQL queries.

“Yeah, I think that's probably the closest I can think of that it rivals.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:11Gregor Vand:Hello and welcome to SED News. As I think many of you know by now, this is the monthly format of Software Engineering Daily where we dive into the tech headlines, we go into a deeper topic in the middle. And then we just do a spin around our favorites from Hacker News highlights. And spoiler, there's also a fun one that didn't appear in Hacker News, but we'll get to that at the end as well. But yeah, as usual, I think I saw you last, Sean, as opposed to spoke to you. I saw you in Singapore, which was fun. Yeah, that's right. You've been traveling, as you often are, but traveling in my neck of the woods was fun to see you.

0:48Yeah, that was great. It was my first trip there. So it was great to hang out. And that's two times in like a few months. This might become a regular thing we might just have to start doing these in person yeah and spoiler on coming back to sf in a couple months so yeah it's good and i'm off to not quite singapore but i am off to

1:04Gregor Vand:australia here in the next week nice but yeah what else has been keeping you busy over the month i mean a summer i feel like has just flown by like i feel like my kids were out of school and then suddenly it's like august and they're going to be back to school in a few weeks so things have become really hot and fast this summer i think we've done a lot of traveling and then just stuff moves quickly. But it's been a big summer for my, I usually don't talk that much about my family, but a big summer for my son who learned to swim, learned to ride a bike, and also now has gotten significantly better at reading.

1:36So I'm very proud of the amount of work that he's put into the summer to learning new skills. How about you?

1:42Gregor Vand:That's like a whole model change there for your son. Yeah, it's the new Kimmy model. Yeah, on my side, just a bit of traveling. I'm recording this on my side from Scotland. I try and come back here a couple times a year so yeah nice to get out of the city I'm up in the the highlands of Scotland so lots of nature around yeah as I was flying just back to something we've talked about a few times Starlink on planes but yeah they used to be super seamless with like no login screens and that was like I believe dictated by Starlink but they've apparently had to cave in to airlines wanting to put like a login interstitial for that so yeah I mean it's such a trivial thing but like expecting Starlink to just connect.

2:23Gregor Vand:And then suddenly I found the airline saying, hey, if you need to log in with your membership number. And I thought, oh, what's happened here? And I read up on it. Yeah. Do you still like, do you have to be up in the air or does it work immediately? That's a good question. For some reason, I don't think it was working kind of gate to gate on the flight I was on. This was Qatar Airlines again. Yeah. But I think it's still supposed to, but yeah, still you just have to log in with your Qatar membership number now, which you don't need to be a specialty or anything. You just literally have a number.

2:50It seems like United is doing some experimentation, but I think I've only had like one or two times I got like a text message saying that there's Starlink on this flight. And then I think both times they ended up having problems with the plane and having to switch planes and then we lost Starlink.

3:07Gregor Vand:Yeah, that's the one thing about Kassar, they really rolled out like big time across most of their fleet now. So if you're on a Airbus A350, you know you'll have it, which is always nice. But right, out of plain chat, which I can easily head into at any time, and on to software. So on to the headlines. Yeah, so there's been a sort of flurry of, we're just sort of calling it runaway AI stories over the last couple of days. And the first one is actually the most recent. So that's just disclosed today or yesterday, which was Claude. And this is actually in reaction to OpenAI, which we'll also talk about.

3:41Gregor Vand:But basically in reaction to something that OpenAI disclosed, Anthropic has now also disclosed that Claude hacked into three organizations whilst they were just testing their cyber capabilities, or at least that's what they said. And they basically said that Claude had gained unauthorized access to outside companies during an evaluation of its cyber offensive tasks. And it was a misunderstanding, apparently, that Claude had access to the internet in its testing environment. We're going to see this as a kind of theme here of like, is this really a runaway model or is it just a human who like forgot to do something but yeah i mean this one feels like human just forgot to place it under strict no internet access controls but yeah what did you think of this one yeah i mean i think that all this really ends up coming back to some human decision making right like that perhaps the intention was not to have the model like hack something but essentially something got lost on the human side of we made a mistake here and gave it access to the internet or we made a mistake here or whatever the issue is.

4:46It really comes back to some human level of control and what guardrails are in place. And then the thing, though, with all this that has me thinking from both the open AI attack on hugging face and then also this latest one with Claude is that if you can just apologize and say, like, oh, it was unintentional, it was an accident and stuff like that. But when this stuff happens where there was intention for abuse or misuse or something like that, like, does it give people basically license to excuse the fact that they are doing something potentially malicious and just say, oh, it was an accident?

5:21I mean, we've seen that with viruses, too. I remember one of the viruses, I can't remember if it was the one of the back in the early 2000s, there was all these viruses that were like these scripts that went around in emails and then you click on it and we copy your contacts and then fire off to them and it would take down people's mail servers. and I believe like one of those started as a student who was like testing something in like an MIT lab and then it accidentally got out and I don't know what happened to that student but like this is certainly not something new but I just wonder like where does it kind of stop with being able to provide an excuse of like this was an accident versus there was intent behind it.

5:55Gregor Vand:Yeah and like just to sort of give context for anyone who hasn't been keeping up with the news the opening eye one was the fact that basically yeah it hacked hugging face they set it off to do something and it ended up repeatedly trying to get into hugging face but what was the secondary part to that story courtesy of tech crunch was the fact that actually hugging face could have detected this like earlier but they had a sort of human failure on the security end that yeah they caught it but they just they didn't escalate it to a human fast enough so right exactly so So their systems actually work the way it is.

6:28So in some ways, there's a part of this that is not really an AI problem.

6:32Gregor Vand:It's kind of the boring problem of a human wasn't in the loop at the right time. Yeah, and I think that's the piece. It's hit mainstream headlines. This is exactly what we all were saying. AI can be catastrophic. It can go off. It's got a mind of its own. It's hacking things left, right, center. But especially in both cases, the open AI and anthropic cases, a human did set it up to do something along these lines, but the human did not a keep a watch on just this repeated actions it was taking and clearly there was no fail safe for it to stop at any time again these are all human determined things that can be set but it's not like it broke out out of the cage exactly and it seems there's some quite fundamental things that could have been done by a human that just weren't and that's actually yeah it's a quote boring story unfortunately because ai isn't sort of this wild animal quite Yeah, I mean, you want to sensationalize the headlines, right?

7:25Like even my sister, who's not in tech at all, sent me the headline and was like, this is scary. So that you get the attention there. But I think one of the things that was interesting about the hugging face attack was when they tried to investigate, they couldn't actually use Claude or GPT because those models have safety guardrails in place where you can't tell whether, are you an incident responder or an attacker? So they have essentially mechanisms in place that, like, I can't go to Claude and say, like, help me break into the Pentagon or something like that. Like, it's going to prevent me from doing that.

7:57That means that you also are limited in using those models to investigate. So they ended up using an open weight model out of China for the forensics. And then I think that also we're going to talk a lot about the open weight models from China later in the episode. But that's interesting. a consequence of this is that we have these guardrails in place, but there's always ways to, I think, manipulate the models, even with the guardrails in place to do this stuff. But then when you want to use it for something intentional, like incident response, you might not be able to do that because the protections are in place there in the first place.

8:31And then you have to circumvent it by going to a model where maybe there's less safety guardrails in place. Exactly.

8:37Gregor Vand:And the sort of final one on the runaway story is actually Amazon. Not security related, but cost related. We have touched on this the last couple of SED news episodes, just where's the tipping point of cost overruns, making it not viable for businesses to be allowing employees to sort of quote token max and this kind of thing. but yeah amazon basically said that they had a ton of unplanned spend they called it catastrophically expensive this is all courtesy of the ft financial times apparently they had about 860 budget overrun over five months and this was basically in their word caused by bad agent loops just didn't crash loudly enough and they've just kept being billed so yeah i mean it's pretty interesting for amazon to come out and actually say that so yeah i just don't understand how they could be surprised by this.

9:30I feel like we've been beating this drum for months now. And I think this is just the beginning of these kind of stories that we see. But if you build a leaderboard to encourage people to use AI, and that's the metric you're optimizing for, but there's no connection to the value of the use of that AI. Like, what do you think is going to happen? Like, this is really good heart's law, essentially showing up on some sort of schedule. Like, you're rewarding people for tokens consumption. So it's like you're giving people a license to be wasteful and not looking at like the productivity metrics of those.

10:02So that's like the epitome of token maxing.

10:04Gregor Vand:So yeah, they kind of have themselves to blame in that sense. But yeah, it's kind of similar to like the initial stories we were talking about with like hugging faces. Like it's still not AI necessarily just like running rampant. It's a human decision that in both cases, even though like one is this attack vector and the other is spend, but it's still like a person making the decision at Amazon to say like, hey, we're going to just have a KPI where we're just going to reward people for maxing out tokens. The other thing that's interesting about this that people have to think about is that with traditional software, when you have like a bad for loop or code that runs forever, whatever it is, like that ends up usually resulting in like a crash or setting off some sort of alarm.

10:48And you typically know fairly immediately that something's gone wrong. And I think the challenge with things like agents and so forth is that you can have a bad agentic loop, but it doesn't result in a crash. What it really results in just keep calling that model over and over again and trying to make adjustments and then giving you sort of plausible outputs or updates, but the whole time you're getting billed. So even outside of the wasteful token use of maybe me using my company's token budget to do my grocery shop list for the next one.

11:21Gregor Vand:Your laundry, that'd be great. Yeah, that'd be great to do my laundry. But then there's also like legitimate excess spend where you might just end up having your agentic harness spin for some period of time where it's just churning against tokens. And you don't even know that something wrong is happening. And that kind of goes back to the earlier stories as well of just, we ultimately need a lot more observability into like what is happening. Like the presumably someone at OpenAI, if they hadn't been really paying attention to what was going on in this experiment, would have saw that this agent with the right observability tools in place is hammering, hugging face and trying the same.

12:00And that should set off certain alarms. And I think it's similar in this case, where if you have an agent loop that's out of control and spending excess tokens, there should reasonably be some guardrails. I mean, you have that with other services. If you spin up a particular Elastic instance or something in the cloud, you're typically setting your top line provisioning of those types of things and you have some controls over it it just feels like we haven't thought through all those controls for ai right now and i guess part of it's just this race to try to out compete everybody and everybody kind of feeling like

12:34Gregor Vand:they're behind yeah i'm going to just jump ahead for a second on i won't say what it is because that's a spoiler but and one of the things i'm going to bring up on hacker news highlights it's a very reliable source as you'll find out at the end but basically this was someone who'd done a bunch of stuff with Claude and we'll get to that but he points out towards the end that Claude doesn't provide reliable methods of counting tokens despite live showing token counts reporting token consumed for accessions and billing for tokens and he said but I'm sure this is temporary and this will be fixed it's just crazy that we do actually have a system at the moment where you literally just don't know what is happening and like exactly what it's going to cost and why And as we're going to get into an open weight side of things, like this is really feeding into the rise of open weight models as well.

13:19You think of a lot of stuff in cloud or even what we saw with ride sharing, where they give you somewhat like a prediction model of what the spend will be for certain actions. So it's like, OK, well, I want to go from here to here in Uber or Lyft. And I'll be like, oh, that's going to probably cost you X number of dollars. So you have some visibility into like what the cost would be. And you can do similar things with certain cloud calculators and stuff that it can get kind of gone. Do we need that for AI? If I'm saying like, create a engineering plan for some sort of feature, can I get an estimate of the budget required to do that?

13:52And then based on what that budget is, maybe try adjusting the plan or something like that to try to optimize it down.

13:57Gregor Vand:Yeah, that's interesting thinking sort of, could you effectively put in your ask or prompt? And I say prompt, I mean, that almost sounds like a year ago or something. We're talking about spinning up agents, etc. But yeah, try and get some kind of estimate before it sets off. slightly digressing so let me get us back to the headlines which the next one we have is just the fact that microsoft has sort of it's gone back on a bit of a tear when it comes to its valuation which is interesting it's like one of the largest jump off shares or sorry for their fourth biggest jump on record so it's up 16 and you know we don't often cover just like pure financial news of tech companies but to see microsoft making these strides is pretty interesting some people might then think, oh, well, this is like partly to do with OpenAI, but it does own still a quarter stake of OpenAI and claimed that that had contributed 24 billion of revenue, which was about 7 % of the 332 billion in sales it reported.

14:55Gregor Vand:But a lot of it was really just AI driven revenue and like massive investment in data centers, but actually that's been completely in theory vindicated by the amount of revenue they're also making. you're building agents that can write code summarize documents and automate workflows but they're missing one thing awareness of the world around them x weather combines enterprise grade weather intelligence with agent ready apis natural language capabilities and an mcp server built for tools like claude codex copilot and modern idees so your agents can adapt workflows automate responses and make better decisions based on real world conditions backed by visilla whose instruments fly on NASA missions to Mars, Xweather delivers trusted data and unique insights that go beyond conditions to actual impact, from real-time lightning strikes to road surface forecasts.

15:43Start with 15 ,000 free API calls every month and pay only for what you use as you grow. Your full weather stack for developers by developers. Start building for free today at xweather.com. Think about your mobile app source code. Once it hits the app store, it's out in the wild. And without the right protection, decompiling is easy for malicious actors looking to steal your IP or tamper with your software. That's where GuardSquare comes in. GuardSquare provides the highest level of mobile app security for Android and iOS applications and SDKs. Their advanced tools integrate seamlessly into your CICD pipeline.

16:23We're talking polymorphic multi-layered code hardening techniques and automated runtime application self-protection. paired with mobile application security testing and real-time threat monitoring to deliver the highest level of mobile app security without compromise. Don't leave your hard work exposed. Secure your mobile applications today. Go to guardsquare.com to learn more. You're shipping faster than ever with AI coding agents, but those agents don't vet the packages they pull in, and they don't have security contacts built in. Ori by Endor Labs fixes that. It plugs directly into your editor via MCP, catching vulnerabilities, blocking malicious packages, and flagging exposed secrets in real time.

17:05No separate tool to switch to, no dashboard to babysit, security that fits how you actually build. Teams using Ori see 10 times fewer security tickets and six times faster fixes. Free for developers. Get started at www.endorlabs.com slash A-U-R-I. I think if you look at also Nadala's quotes related to the announcement of their quarter performance and so forth, and also the recent thing that we covered also in the last episode where on X he had written about how the AI companies or the model companies are charging you twice and so forth. In here, he says that every model is substitutable. He's telling essentially investors that Microsoft is deliberately building its infrastructure so it can swap out things like OpenAI for its own models or for Anthropic.

17:53And I think that it seems like they're moving towards a view where they're trying to allow, essentially, their customers to be very flexible and adapt, which I think makes sense. I think the average enterprise now is using at least five different models. And you probably want to be able to do that so that it's a little bit like being hybrid cloud, although it's easier to be hybrid model, where you get power essentially in the negotiations if you're not wholly dependent on a single vendor. And it seems like, I think Microsoft's kind of weaning in that way. But I remember back in March, Microsoft stock dipped and then everyone was like freaking out and essentially calling for the death of Microsoft.

18:31And now it's back. And I just think that the overall, the market is like extremely volatile right now. there's these huge swings constantly from i mean ibm had its biggest drop recently biggest single day drop in like 50 years or something like that recently and that was coming off like a huge pop just two months earlier so yeah i don't know what goes up must come down so who knows like we might be talking about microsoft in another quarter or two how they had the largest single day loss in

18:59Gregor Vand:a day or something like that yeah for sure it is quite hard to predict as market should be i guess But yeah, we're seeing just huge swings when it comes to AI-related stocks when one minute chipmakers up, one minute chipmakers down. And just on a tangent there, yeah, Apple briefly hit 5 trillion valuation, which is pretty insane. But just before recording, I double-checked and they had actually reported numbers very recently in the last, I think, couple of hours. And they've gone back to 4.9 trillion. So I don't feel too sorry for them. Only 4.9? Yeah, only 4.9. but interesting to see that they still notched above five and yeah what's driving their revenue well still they've got very good iphone sales and they've got still very impressive services revenue and even greater china revenue as well but both of those services in greater china were a little bit less than what was expected which is again just what's kind of driven that but very interesting to see that they can still we've talked about this on previous sed news like still managed to stick on these like lines of business that are not that ai driven or even really like ai adjacent to be honest very interesting but they are having issues with memory chips which is another sort of topic but even in uh super base we've like been told by one of our suppliers like we have different suppliers depending on where you live to get laptops and some of our new employees are being told like weeks before their macbook will arrive because we do custom specs, not just like off the shelf, but yeah, now we're being told weeks, which is quite exceptional.

20:29Oh, wow. Yeah. I just got a new laptop and this is my first time recording on this. But it's interesting, like with Apple too, a lot of their lines of business, which are like incredibly successful, there's still like minority in the particular vertical. Like iPhone is wildly successful, but it's not the most dominant phone, I guess maybe from a single vendor, but Samsung might actually be bigger. But then obviously from an operating system standpoint, like Android is, there's more Android devices than there are iOS devices, similar with computers as well. Like I haven't used a PC in a very long time, but it's still the dominant machine overall, I guess.

21:08But I think it's, they have incredible like brand loyalty and they do make fantastic machines overall. So clearly the things that they're doing is working for them.

21:19Gregor Vand:Yeah, absolutely. So then moving on to, we've also covered Waymo and driverless, like from a few different angles. But this one's interesting, because we did talk about Waymo and Uber, like being partners in ways, and then obviously, frenemies in other ways. And yeah, we're now actually seeing an official split of that partnership. So if you use one in San Francisco, this might sound confusing, because it's a Waymo app, it's a Waymo car. So where does Uber figure in that but actually this is for some of their other u.s territories so waymo had like first partnered with uber apparently in may 2023 and that was to launch in phoenix then it was followed by austin and atlanta and that's like so waymo cars available through the uber app and then uber managed the vehicle fleet apparently with another partner called avomo but then in may Uber and Waymo parted Waze in Phoenix and then the two companies have also clashed over the quality of Waymo services apparently in Austin and Atlanta so yeah it's kind of interesting to see they tried but I think it was always going to be challenging to see how Waymo being owned by Google like how is this actually gonna now at the end could these two actually be true partners or was this just always going to be again a tipping point of where like that partnership just kind of had to end Yeah, I mean, they say all partners are meant to be broken at some point.

22:43I mean, especially where they're both in ride sharing, like clearly at some point, their mutual interests are going to become too competitive to each other, essentially. in a lot of ways it's i think very similar to how a lot of partnerships work where uber was essentially supplying demand and operations while waymo was weak on those particular spots and then as waymo has grown and raised more money essentially they don't need those training wheels anymore they can build out their own and build their own network and own it end to end so i feel like this was probably always going to be the ultimate end of that relationship yeah so yeah

23:20Gregor Vand:it's sort of not officially over yet but just people familiar with the matter apparently again via financial times but yeah it seems pretty unsurprising really that this was gonna probably break apart at some point and then yeah just to wrap up on the headlines there was an acquisition and we were just talking before we started recording so let me try and get this right i believe it's any scale was acquired by n scale and when we were talking earlier i said is that scale and it's like no that's not scale.ai is different again so we've got amazing naming these days but yeah what what's this one about yeah so any scale which is known they were the creators of ray which a lot of inference infrastructure and fine-tuning and model infrastructure runs on a very well established open source project any scale was the company that tried to build or built essentially like a managed version of ray around that and they raised like a billion dollars or so in 2022.

Read the full transcript

24:19But I think, and then they just sold to Nscale for 1.65 billion, which is a NeoCloud. So for those that aren't familiar with NeoCloud, NeoCloud is essentially what the term is used for companies that offer primarily GPU as a service versus kind of general computing. So there's all kinds of these NeoCloud companies are now available. I had actually talked to any scale at a variety of different times. Like I think this was probably like a decent outcome for them because from, I think it's probably hard to really grow that managed Ray as an independent company into like a really big company. But as part of a like GPU infrastructure company, it's probably a good like pairing.

25:01So I think that it makes a lot of sense. But I do like, I always confuse any scale with scale AI and then the event scale. And I don't know the history of how those company names came together. But generally, the thing that people a lot of times strive for with naming companies is you want a name that you can say and people can remember. And I'm not sure they hit the mark there. It's like nscale, scale, any scale. It's hard to remember who's who in that Venn diagram.

25:29Gregor Vand:For sure. I mean, scale.ai, okay, that's a great name to have,.ai. you can basically have anything that is your company and it's like going to sound good and it's you know five letters but then any scale and then n scale that's pretty confusing but and then yeah just a sidebar piece of news scale.ai now have a new ceo as well which is interesting because that was founded led by alexander wang not the fashion designer if anyone knows that one but this is Alexander without an E at the end of the Alexander. And Meta had taken us virtually just under 50 % stake. That was kind of the point. So Scale and Theory are still running their own show.

26:11Gregor Vand:But when you've got, I think, 49 % stake from Meta, you're quite beholden to them. But yeah, in that transaction, Alexander went to head up the AI side of Meta. Total, now Scale have a new CEO. So that'll be interesting to see how that all nets out. Yeah, they raised like 14 billion from, with a round led by Meta. So pretty significant. Yeah, I mean, actually just the new CEO of Scale is actually, he has a former Google Cloud executive called Francis de Souza. So definitely interested to follow along with that one and see sort of how that all nets out. Yeah, I saw also like Fireworks, who just raised a huge round and made a lot of news.

26:52Their new head of engineering just came over from Google. So I think you're starting to see, it's a common pattern though. You have people who reach executive positions at large companies like Google and, you know, they get maybe a little bored with that and then want to go back to, you know, building and moving faster. Yeah, for sure.

27:10Gregor Vand:Yeah. Fireworks, super interesting. We do have an episode with Fireworks with one of the co-founders, Benny Chen. So yeah, go check that out. I think that came out around March this year. So well before this fundraising was confirmed, but yeah, they're definitely having a moment. They are, you know, sort of the infra for open weight models. And that's becoming increasingly interesting to many companies for many reasons. But yeah, that's actually quite a nice segue into the main topic for today, which is we're kind of calling it the Kimi moment. You know, Kimi, which is an open weight model. The Arch company is called Moonshot AI.

27:45Gregor Vand:So if you've heard of Moonshot AI, that's Kimi and vice versa. From a Chinese, Moonshot AI is a Chinese original company. They do have offices, I believe, in the Valley and in Singapore and that kind of thing, but very much seen as a Chinese company, which sort of frames a lot of why this is quite interesting. but yeah like when we think of open weight models there was this the deep seek moment first which we can sort of touch on and then really in the last like almost like three months we've just seen this sort of especially rapid adoption by many companies of Kimi and really forcing companies to sort of think differently about using you know foundational models from the big players or the big names rather and getting some quite interesting and very high quality results out of Kimmy.

28:29Gregor Vand:The DeepSeek moment, I think, like, Sean, how did that sort of kick things off when we think about open mic? Yeah, I mean, I think the DeepSeek moment ended up having sort of more, like, market impact in some sense, because I think prior to that, the sense was that you couldn't really get these really powerful models, except from these, like, frontier labs like Anthropic and Google, OpenAI and so on. And so it was kind of shocking when DeepSeek came out. And NVIDIA's market cap cratered as a result of that. And of course, it's certainly bounced back since then. And then when you look at the Kimi launch, which is the largest open-weight model ever released, beating some of the top-close models on particular benchmarks, it seemed like nobody really panicked.

29:18And I think part of that is not because it's less impressive, but it's because we've kind of gotten used to the idea that these Chinese lab overweight models are competitive. And so it's less of a shock, essentially. We've been desensitized to it. But I do think that what the market reaction has been more around, and I think this is something we're already coming to, is where this was somewhat on the back of what happened with the mythos and fable. and the reaction to that where people got scared that they could invest in a model and then have potentially like a foreign government say like you can't use this model anymore then i think that has increased the interest in diversification of models and then also having an open weight model strategy along with having a sort of closed weight model strategy yeah and we'll probably touch on

30:06Gregor Vand:this sort of the banning of models because of course it comes into this as well if you want to I say ironically, because yeah, when you ban, at least pause, ban a foundational model, a closed source model, if you like, then open weight becomes very interesting, but then a lot of open weight comes from China. So we're going to kind of be interested in that. Why sort of the last couple of months has Kimi really just like exploded, I think in interest and popularity. It's kind of moved from, oh, it's good enough for professionals to this kind of actually competes with GPT and Claude, you know, And it's kind of closed that gap in basically weeks, at least from what has been sort of released.

30:43Yeah, and they also released like a ton of models on like back to back. Essentially, they're moving super, super fast.

30:49Gregor Vand:Yeah, exactly. So, I mean, yeah, if you sort of look at the actual lineage, I guess, like they've been doing sort of quarterly drops. If you like, you know, July 2025 was K2. And then within 12 months, you've gone through 2.5, 2.6, 2.7, and now K3. That's an incredibly rapid cadence. Like if you look at those jumps as being as significant as jumps that you might see from any of the foundational players, but definitely not, but not on that cadence effectively. Yeah. Yeah. I mean, I think the model distillation practice has really sped up how quickly people are coming up with models and models that are comparable in performance, which was a big part of the conversation around the initial deep seek launch.

31:28I do think that some of the criticism on some of the benchmark results that we've seen from Kimmy is that, in particular, there was a lot of headlines around their MCP tool calling performance, but that beating Opus, for example. but in particular that benchmark is relatively new and there is i don't know if this is more just sort of like jealousy in the dialogue or whether there's some truth to this but you can essentially sort of bias towards performing really well on certain benchmarks whether it's mcp1 or it's you know the ml lu style benchmarks as well and then you can have like really good benchmark performance but it might not actually match like reality that's where some of the criticism has been on even some of the tests of you testing a model against lstat and stuff like that and then saying oh it's you know outperformed lawyers on the lsat but then like the lsat's actually not necessarily a good indication of what a lawyer does on a day-to-day basis and stuff so these are some of the the nuance and some of this is like general you know model criticism but this is some of the dialogue that i've heard around kimmy in particular is like are they building to optimize for the metric kind of like what we talked about in the amazon on story is like you know whatever the kpi is people are going to try to optimize for that kpi so you always have to be careful essentially what the metric is that you're measuring people against

32:50Gregor Vand:yeah and on june 12th uh k2.7 code was released you know and this is obviously very much a software engineering tuned model and as you were saying yeah that became kind of the benchmark moment especially through mcp and that was a correct tool invocation is sort of how that's how that's measured apparently like scoring in theory you know 81 and that was versus say opus 4.8 at 76 which is a pretty if you believe in these benchmarks that's a pretty meaningful jump but then the distribution piece was kind of interesting like github actually made 2.7 generally available via copilot but there was kind of like an interesting piece there where for enterprise teams on copilot business and enterprise like it was actually this model was off by default and admins had to explicitly enable it and github's changelog flagged this and sort of said it may be less aligned than other co-pilot models which is interesting because obviously they have a their own interests being you know part of the microsoft open ai ecosystem but you know obviously they couldn't miss not providing that to users given that it's on azure and they still get inference from it but yeah very interesting that this sort of slightly odd warning came with it that wasn't really clear what that was about.

34:06Yeah, I believe it's the first Chinese lab open weight model that's been built into Copilot. I do get like having, you know, working with a lot of enterprise customers, they do are really careful about which models they use. So I can kind of understand having some level of governance there. But of course, there's a certain biased interest from Microsoft point of view of how do you craft the language around this particular model. But there is, I think, that fear with the enterprise. So I understand some level of controls there, but obviously there's also a certain bias that can be put into it. I think one of the things that's interesting about this is Moonshot, as well as, you know, this goes for all the Chinese labs, they don't have these like top tier NVIDIA chips available to them, like the US labs use because of the various chip embargoes.

34:50Gregor Vand:We don't think they do, but yeah. Even if they did have some, probably not at the scale. So there's a certain scarcity of resources that they have to deal with. And I think that in some ways, like that is forcing them to innovate in a way that maybe the team, the US-based companies or the Western-based companies where they don't have that scarcity of resource aren't necessarily forced to do that. So some ways, like maybe we're creating the thing that we fear, but good news of that is it's probably forcing the other companies like the Western-based companies to react to this to lower token prices and maybe also think about how they keep costs down.

35:31It's a little bit, if you look at Google's beginnings, Google started as a research project at Stanford, and you have students essentially didn't have a lot of resources. So they had to be very, very creative about how they scaled Google. And then even that carried over to when Google did initially have some funding. But if that project had started out of like a bigger, more established, well-funded company, they probably wouldn't have been ripping apart like cheap machines and like wiring them using Legos and like stuffing them together as quickly to compact the size within the data center as much as possible.

36:09And they would have just bought like, you know, super beefy servers. And the downside of that would have been like a lot of the innovation that we've seen since then in the, in cloud infrastructure kind of came from some of those, like forcing yourself to deal with these scarcity resources. And my hope with all this is we see a similar thing of innovation driven from the scarcity resources that carry over to the frontier labs and other parts of the world as well.

36:35Gregor Vand:Yeah, I mean, it's interesting sort of just sidebarring to like, she was even sort of behind the founder, Yang Zilin, he sort of probably was known as Yang the Genius. There was a very good profile on him in the Financial Times actually last weekend. He's 34 years old. He was known at university for being like just as much into music. and having academic sort of soirees, if you like. And actually Kimi, when he first released it, you know, as a chatbot, it didn't do very well. It was sort of outages and this kind of thing. And so it's clearly, as you're calling out, it's like managed to catch up despite not having the same access or we don't think, couldn't possibly have the same kind of access despite what it still may have.

37:14Gregor Vand:Back to sort of then the turning point here, like July 16th, which is for sort of history, that's then 34 days after 2.7 had dropped. K3 dropped and that was a model with 2.8 trillion parameters, which is, I believe, the largest open weight model released to date. And I think what's interesting here is then on the July 27th, the full model weights were actually published. And this sort of then leads into this whole like providence piece, which is when companies are starting to be asked like, well, you know, you're using AI to generate so much information within your own company. or analyze data, you want to actually know how this was determined and so on and so forth.

37:58Gregor Vand:And here we are, we actually have a competitive model with the closed foundational models with the full weights published on Hugging Face, which is like massively impressive. To me, that's like a huge turning point when you've actually got like effectively all the weights right there. And this isn't just, you know, a model that takes up some tiny little specific task. It's really competing now. I think overall, this is whether Mooshaw and Kimi kind of went out or not. I think overall, this is good for consumers of these models, because it will force the other companies to essentially react to this.

38:37And hopefully, actually, I saw a headline today that OpenAI was reducing some of their token costs by up to 80%. And we saw a similar thing after DeepSeek, it drove down token costs as well. And then all the stuff we were talking about earlier of like token maxing and companies kind of getting sensitive to how much they're spending. Like overall, I think it'll be good for consumers of these things. You know, one thing I was thinking about with this story, too, you know, we've covered a lot of Meta over the last year and their decisions in terms of like hiring and what they're trying to do with their AI lab.

39:10But, you know, Meta had such a head start in this like open weight model space with the llama models. and it's been a long time since i've heard anything about the llama models funny you said

39:21Gregor Vand:that because yeah there was something i didn't dig into it but i just saw the headline which was that zuckerberg is basically lobbying to ensure that chinese models don't get banned and to me that only meant one thing which is like well they have to keep that door open because like llama's not maybe going where he hoped it was i think like his instinct of trying to own the open weight model space was probably right. The execution was bad. I don't know what happened internally there, but they were clearly on a good path, especially with what we're seeing now between the Chinese labs, between what's happening with Fireworks and the other inference providers that have focused on OpenWeight.

39:59There's clearly a huge TAM available for OpenWeight to own a big part of the business. And I think realistically, especially in the enterprise, most enterprise businesses will probably have a mixture of both open weight investments that they've done and as well as the closed source models. And that'll probably be the strategy that many, many companies take for some period in the future. I don't know how Meta ended up, and maybe they'll bounce back, but they were onto something, but they haven't been able to execute.

40:30Gregor Vand:I think that's really good analysis. It is the right strategy, wrong time, like too early effectively perhaps, or just like too early in the, given they were trying to do it from the US, being asked to kind of produce too much too soon and like that can distort things so who knows maybe like a rogue or meta with llama they'll kind of adopt more of a kimmy approach to make this succeed who knows but yeah i mean just to kind of wrap up on kimmy and so this open weight obviously resurgence but like yeah maybe search model provenance is a real thing that's being sort of talked about now when it comes to you know it's a business risk right like you're putting money and basing your company on whether it's like for coding or for other tasks but you are now having to really decide like it's a procurement question on what are you buying effectively and like what are the risks with that like could it get too expensive or could you move it onto your own infra if you really needed to or not your own infra but exactly via fireworks you could still own that sort of end-to-end there that's interesting pricing we've talked about that like pricing is just kind of getting out of control and it was really a case of like when And not if is this going to like stop being possible for, you know, Anthropic and OpenAI to kind of like charge this way because it's just not sustainable for many businesses.

41:45Gregor Vand:And yeah, and like the speed as well, the speed of iteration is like clearly being a huge piece here where however Moonshot are doing it, they're just iterating like crazy. So on to our favorite parts, Hacker News highlights. Do you want to go first, Sean? Sure. So this headline really caught my eye and made me laugh, which was the speaking agents like cavemen save 65 % of tokens we test. So this was by JetBrains. And they benchmark the caveman skill, which is a cloud code skill that makes it respond in terse caveman speak to save tokens. And they claim 65 % token saving. So they essentially put that to the test.

42:22It was focused on agentic coding tasks. And they found actually in reality, it was about 8.5 % output token savings versus 65%. But a big part of that is they were focused on coding tasks where a lot of it's going to be code. You can't turn that into caveman speak. You got diffs, you got tool calls. So the caveman skill might not be best for that versus some sort of more conversational chat. But I just really loved it. It reminded me of like some ignoble prize research out there. there. It's just kind of a ridiculous topic.

42:54Gregor Vand:That's like, yeah, super funny. It's interesting because I didn't actually scan ahead. So that is actually a little bit similar to one I picked out, which is called the economic benefits of refactoring. So this is on the fairly popular martinfowler.com website posted by Java user on Hacker News. So thanks for that. But this is not written by Martin Fowler himself, but basically this was, you know, someone else that thought works and it was can you basically decrease especially the input token cost if you refactor your code or like what are the consequences of that and i mean the tldr is yes if you refactor then basically you can dramatically reduce your input tokens and kind of the theory behind this was that the saving is because the agent has to read less code but the bit that might not be like that sounds obvious but it's actually not because there is less code to read it's actually that the overall code in a certain layer they used to help test this like it stayed constant but the fact that the agent is then able to successfully identify smaller subsets of files to read is like the key bit there so like refactoring into smaller chunks but also making sure that those chunks are very clearly dri don't repeat yourself etc etc so that the code that the agent needs to go and grab is smaller you can't just sort of chunk it up into small chunks and then go well it's smaller like it then doesn't even know it has to still get all the small chunks because it doesn't know what's most important but the refactoring is what then turns it into having that context of having like what is most important go and find that small chunk input that small chunk and the output token and output code was virtually the same in this experiment so which is super interesting there was a small sort of side note on that which was him saying that claude Desclarge is actually not good at refactoring, and I can definitely attest to that.

44:46Gregor Vand:But if you can go through the motions of get it to refactor, then this could be quite a huge saving for anyone who's trying to reduce their input token costs. Yeah, so basically, bottom line, good design leads to also optimized token costs. Funny that, yeah. We just go back to how we used to design software. I mean, it's the same if you think about the human costs. If you have poorly designed software, then there's going to be more human costs each time someone unfamiliar with it needs to like ramp up and make some sort of change exactly and then you have a another one on code pen my second one was yeah just thanks to user robin reala posting the fact that code pen 2.0 has come out this takes me back that's why i was quite interested in this i don't know if anyone else out there this makes it sound terrible so if you like code pen this is sorry about this it's just i haven't been doing a lot of front-end work for a long time but used to love code pen used to you know put up all sorts of things on there and a really useful tool as well like i did a little coding like class i went back to my old high school like a long time ago and did like a coding class and codepen was amazing because you could just get people spun up in a browser writing front-end code and just see it do its thing straight there they didn't have to build with like you didn't have to get them spun up with some sort of repo or anything so that was really helpful but yeah i mean codepen must have come out like probably 15 years ago or something like that so codepen 2.0 i'll just quickly run through a couple of things that they call out that they have files and folders now so it's interesting it's almost becoming a bit ide-esque but files and folders and then they have sort of they do you know build steps anyway but they've now like added this concept of blocks which looks quite nice where you can sort of see exactly which bits are in this build process or like the linting process so that's kind of fun yeah real-time collaboration as well so finally multiplayer on code pen so for anyone that uses code pen a lot i'm sure that's quite a huge uplift so or if you haven't checked out yeah still a fun place to go and just experiment with front-end stuff so yeah and then just a special one i guess not technically through hacker news but thanks to ilia reshetnikov for tweeting at me and sean we do often cover these doom can you run doom on something and thanks to ilia he pointed out that there's another one called DoomQL, which is basically using an SQL query is the frame buffer, which is like, I think that definitely rivals, you know, Doom on TypeScript types.

47:12Gregor Vand:Yeah, I think that's probably the closest I can think of that it rivals. So thank you, Ilya, for shouting that one out to us. That was very, very fun to read through. So yeah, that's it for another SED news. Have we got any looking ahead predictions, Sean, which we usually get wrong? Yeah, I mean, I think the safe predictions here would be that we're going to see more headlines on the token maxing issue as companies start to adjust. I think we'll see more also conversations around open weight versus closed model, diversification of models. I think those are both going to be big topics of conversation through to the end of the year.

47:47Gregor Vand:Yeah, for sure. I will then say, well, because we just talked about iteration and Kimmy, let's assume that by this time next month, we're already on like Kimmy 3.2 or if i'll just push the boat out give me 3.5 3.5 by end of august let's see if that lands so thanks everyone for for tuning in as always and we'll be back next month with another sec news thanks everyone cheers

48:27Thank you.

From the publisher

SED News is a monthly podcast from Software Engineering Daily where hosts Gregor Vand and Sean Falconer break down the biggest stories shaping software engineering, Silicon Valley, and the broader tech industry.

In this episode, Gregor and Sean dig into a wave of “runaway AI” stories, including Anthropic and OpenAI disclosing that their models accessed outside organizations during cyber evaluations, and Amazon reporting a staggering budget overrun blamed on bad agent loops. They explore why most of these incidents trace back to human decisions rather than models breaking free, and the awkward reality that today’s systems can bill you for tokens without reliably counting them.

They also talk about the “Kimi moment.” Moonshot AI‘s open weight model has closed the gap with frontier models like ChatGPT and Claude at a remarkable pace, and the hosts unpack what it means for open weight strategies and how chip scarcity is pushing Chinese labs to innovate.

As always, the episode wraps up with a few standout Hacker News threads, including a JetBrains test of a “caveman speak” skill that promised big token savings, how refactoring can cut input token costs, the release of CodePen 2.0, and a build of Doom that renders through SQL queries.

Sponsorship inquiries:
sponsor@softwareengineeringdaily.com

The post SED News: The Kimi Moment, Runaway AI, and Tokenmaxxing appeared first on Software Engineering Daily.

More from Software Engineering Daily

All 195 episodes
SED News: The Kimi Moment, Runaway AI, and TokenmaxxingSoftware Engineering Daily · 48 min
Listen in VO