In short
The episode “The Tokenpocalypse Is Here” argues that AI adoption is shifting from “unlimited, cheap, productivity” to “token economics,” where companies scramble to curb AI usage costs.
Guests
Sam Cole and Emmanuel Mayberg, both co-founders of 404 Media (journalist-founded, subscriber-supported).
Key claims
token spend is rising because non-engineers are using AI for everyday work; companies are throttling access, charging per token, and forcing “tool-like” AI behavior.
Notable examples
Accenture internal audio (Justice Kwok, agentic AI strategy lead) says “token chewing” is driven by non-engineers, including converting PDFs into Markdown; Accenture is reportedly pushing “Token IQ” and previously pressured employees to use AI for promotions. GitHub moved to token-based pricing; Uber capped employee AI tool use after exhausting its AI budget; Walmart also restricted usage. A leaked memo highlights “Caveman Claude,” a plugin that makes Claude/Codex output terse and can estimate ~65% token savings.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOFrustration with Software Errors
0:00 to 0:22
Explore the annoyance of vague software error messages.
“The most annoying thing in the world is when software is like cute.”
Setting the Stage for Discussion
0:57 to 1:35
Hosts discuss the absence of a co-host and refer to recent articles.
“I'm your host Joseph and with me are two of the 404 Media co-founders.”
Introducing the Tokenpocalypse
1:35 to 2:05
Discussion begins on the concept of the 'Tokenpocalypse' and its implications.
“Sam, do you want to take us through this one?”
Understanding Token Costs in AI
2:05 to 4:24
Exploration of how AI companies are adjusting to rising token costs.
“I feel like we have coined a few token-related words.”
Accenture's Role in the Token Economy
4:24 to 6:31
Accenture's involvement in AI token spending and consulting practices.
“I mean, it's rich because of the last few months.”
Non-Engineers Driving Token Consumption
6:31 to 10:12
Discussion on how non-engineers are responsible for significant token usage.
“I mean, we'll talk more about this in a minute.”
Corporate Contradictions in AI Usage
10:12 to 14:00
Examination of conflicting corporate messages regarding AI usage and spending.
“And then, I mean, Justice goes on and says more broadly that they're really seeing what they call a rapid escalation in AI token spend.”
Navigating Corporate AI Strategies
14:00 to 18:27
Learn about the conflicting messages companies send regarding AI tool usage and the pressure on employees to adapt.
“It's going to make you 100 times more efficient.”
Navigating Corporate AI Strategies
18:30 to 19:59
Learn about the conflicting messages companies send regarding AI tool usage and the pressure on employees to adapt.
“I've noticed that before I leave the house, there's a few things I always check for.”
Navigating Corporate AI Strategies
20:04 to 21:22
Learn about the conflicting messages companies send regarding AI tool usage and the pressure on employees to adapt.
“What's softer than cashmere and warmer than wool?”
Show all 17 chapters
Navigating Corporate AI Strategies
21:30 to 23:04
Learn about the conflicting messages companies send regarding AI tool usage and the pressure on employees to adapt.
“If you've listened to this show for more than a week, you know our beat is uncovering how your data is quietly harvested and exposed online.”
Navigating Corporate AI Strategies
23:11 to 23:21
Learn about the conflicting messages companies send regarding AI tool usage and the pressure on employees to adapt.
AI Cost Management Strategies
23:21 to 28:00
Discuss how companies are adapting their AI strategies to curb spending, including the use of the caveman plugin.
“If you sold somebody a loaded gun who you knew was in a vulnerable state and they shot themselves.”
Engaging with AI as a Tool
28:00 to 30:22
Exploration of how AI should be treated as a tool rather than a conversational partner.
“It was actually a way to scrape contracts from DHS procurement databases and that sort of thing.”
Caveman Claude's User Base
30:22 to 33:41
Discussion on the internal usage of the Caveman tool by developers in notable companies.
“But I agree with you that the most annoying thing in the world is when software is like cute.”
Token Efficiency and AI Usage
33:41 to 35:36
Insights into how Caveman Claude helps reduce token spending and its implications for businesses.
“So at a minimum, Shane Sweeney has written codes to work along with this tool as well.”
Corporate Responses to AI Usage
35:36 to 40:58
Examination of how companies like Citi and IBM are managing AI tool access and usage limits.
“Obviously, I'm being facetious, but exactly what Emmanuel was saying, all the cutesiness of it.”
Transcript
Automatic transcript. May contain errors.0:00Emanuel:The most annoying thing in the world is when software is like cute. Microsoft made this change years ago where like if you have a crash, you're like, uh-oh, something bad happened. It's like, shut the f*** up. What's the error code? What's the error code? Because I'm going to have to like plug it into Google.
0:21Hello, and welcome to the 404 Media Podcast, where we bring you unparalleled access to hidden worlds, both online and IRL. 404 Media is a journalist founded company and needs your support. To subscribe, go to 404media.co. As well as bonus content every single week, subscribers also get access to additional episodes where we respond to the best comments. Gain access to that content at 404media.co. Also, remember to subscribe to our YouTube channel where you can watch all of our episodes. Subscribe at youtube.com slash at 404media.co. I'm your host Joseph and with me are two of the 404 Media co-founders.
1:02The first being Sam Cole.
1:04Emanuel:Hello. And the other being Emmanuel Mayberg. Hello. So Jason's not here. He's probably on a private jet somewhere. If you have read the website at the time of recording, you'll know what that's referring to. I won't spoil it. Go read that really, really good article if you want and you should. But otherwise, I'm sure you all will speak about it. next week on the pod. I don't think I'm here next week. But Jason will fill us in there for sure. Sam, do you want to take us through this one? Yeah. So this first one, we have a set of stories that kind of go together. So the first one is by Joe. The Tokenocalypse.
1:48Emanuel:Tokenpocalypse? Dude, I know. Tokenopocalypse? I even considered... Well, this was the hardest bit. It's the Tokenpocalypse. Tokenpocalypse. Tokenpocalypse is here. Companies are scrambling to stop spending so much on AI. I feel like we have coined a few token-related words. Did we coin token maxing? Or was that an existing term? Emmanuel, what do you think? I believe the brain geniuses at the AI companies actually came up with that. Okay, fine. All right. So we'll give them that. On this one, I thought I coined it. I thought I coined tokenpocalypse. And then I think I Googled it after the fact.
2:29I found TechCrunch should use it in a headline a few days before. Fuck. It's in the air. Yeah.
2:39Emanuel:So we're introducing some new words here. We're tokenmaxing. We're tokenpocalypsing. So do you want to just define what we're talking about when we're talking about the tokenpocalypse? Yeah. So in my definition, after coining this term days after TechCrunch and presumably some other people, the reason I kind of use this term is that something has really, really shifted in the AI industry where we've had companies rushing to adopt AI in all sorts of forms. Maybe that's in coding. Maybe that's agentic AI in their businesses as well and whatever. And companies have sort of been doing this, maybe not regardless of the cost, but it's been presented to them as like cheap, right?
3:28Oh, well, this is cheaper than people. You do your subscription or whatever it is to an AI company every month, every year or whatever. Maybe if you're an enterprise, you get a big deal. But there's this shift happening where now there's much more emphasis on the cost of individual tokens, which is basically user usage metrics, right? So now you'll have companies like GitHub being like, we're actually going to charge you per token. And all of a sudden, companies are realizing, oh, AI actually costs us a lot of money. And we're burning through these tokens, which we'll talk about in a bit. So that's why I use that term.
4:08And I think that's why other people are using it as well. There's this big paradigm shift where, oh, this isn't actually just automatically saving us money. We now have to figure out, shit, how do we stop spending so much money on these tokens, basically?
4:24Emanuel:Yeah, which is... I mean, it's rich because of the last few months. Jason wrote a story that... And we'll get into just the ways that this has been working in the last few months. But Jason wrote a story a couple months ago about token maxing and how that was the thing in March and April to be bragging about how much token spend you have or these leaderboards that we can get into in a bit. But before we do that, this story that you wrote involves audio that you obtained from a consulting company called Accenture. And I think if you're in tech, you know what Accenture is. But I even when I was editing the story, I was like, what are we even...
5:04Emanuel:Are we talking about Accenture's token spend or Accenture is consulting on token spend? So can you just walk us through like, what... How does Accenture come into play in this particular story? What's their place in the token mania, I guess? Yeah, I mean, who knows what Accenture does? Yeah, it's one of those. I think that's an open question. Yeah, obviously, they're a consulting giant. And they've done stuff like work on content moderation for Facebook, right? Like a significant amount was outsourced to Accenture back in the day. And they will fulfill this role where they come into companies and be like, Hey, we can help you optimize this task or we can cut off expenditure on this resource or whatever it is.
5:55And in this particular case, they are talking about two things. They are talking about token spend inside Accenture itself, but also with their clients. They clearly have visibility into what their clients are doing and what their clients are saying. In this audio, it doesn't mention any specific clients. We can't play you the original audio for source protection reasons. But they do say like, well, we're seeing this sort of thing inside and outside with our clients as well. But Accenture... I mean, we'll talk more about this in a minute. but it's funny how they position themselves sort of as the problem in the first place and the cure where they've told companies, you need to get on AI, you need to do this as quickly as possible.
6:45And then it's like, oh, wait, all of these companies are spending a bunch of money on tokens and maybe we can help with that.
6:51Emanuel:Yeah, weird how that works. Weird how consultants keep getting away with it. So yeah, like you said, this is based on audio that you had obtained and it was the audio is from an internal meeting. and so obviously they're discussing things that they don't necessarily want the wider public to hear or are not ready to discuss in public but there's one part that's pretty key to the conversation and it's coming from Justice Kwok who is the agentic AI strategy lead at Accenture that's a hell of a title but in the audio Justice says we're seeing from some of the data internally at least it's actually not our engineers that are driving the token consumption.
7:35Emanuel:It's a lot of the non-engineers that are doing some of those behaviors that we're talking about in the meeting. So what are the behaviors? What are the non-engineers that Justice is talking about actually spending tokens on? Yeah, it's a funny piece of audio because it actually starts with a joke. And it's kind of hard to describe that in an article. It starts with very sarcastic comments. So Justice, who you just introduced, is preparing to present in this meeting. And then someone interrupts Stuart Henderson, who's a very senior person in Accenture. And he jokes that, oh, you have these slides, but I hope you didn't use AI to just convert a PDF into Markdown and all of that sort of thing.
8:20And then Stuart says, quote, I'm learning that's one of the big token chewers, turning PDFs into Markdown. is that right? Which obviously, as I said, is a joke, but then leads to a pretty serious question. And then that's when Justice replies with, yeah, that's actually the behavior that we've been seeing. So I found this fascinating for a few different reasons. The first, obviously, I love that term token chewing. I mean, that is a consultant-ass term if I've ever heard one. Maybe Emmanuel or Jason have heard it when they've looked into token maxing or whatever, but I never heard that before.
8:54So I liked that. But just the idea that it really undercuts that narrative that, oh, this explosive use of AI is all about supercharged 10x engineers or whatever, using AI to produce mountains of code, and then we can ship products faster or whatever. Here, it seems to be non-technical employees do non-technical tasks like converting a PDF into a PowerPoint. And it's kind of like the here-mitted meme that actually, it just isn't all that sophisticated. And to be clear, this isn't to say that AI is not being used in clawed code or codex or whatever to make a bunch of code. Obviously, it is as well.
9:40But we've entered this more mature stage of AI being in companies where the reality is setting in. It's like, oh, people are just using for a bunch of dumb fucking shit, basically. And hey, it might be a time saver. But converting PDF into a PowerPoint might save you 30 minutes, if not more, it depends how much of a perfectionist you are with your slides. But there's going to be a cost to that down the road that I think people are only just starting to realize. Yeah. And then, I mean, Justice goes on and says more broadly that they're really seeing what they call a rapid escalation in AI token spend.
10:21I think they say soaring token spend as well. Maybe stuff isn't dire, but it's to the point where obviously, Accenture sees an opportunity here and they're having internal meetings and then we're hearing about it as well.
10:35Emanuel:Yeah, it's just lazy. It's just the laziness of it all is wild. I mean, you can get AI to summarize your PDF for you, read it for you, now turn it into slides. It's like, what is your job actually? Especially for Accenture, who a lot of people made this joke in the comments, Like Essentia's whole thing, if you're being a bit mean to them or whatever, is making slides, basically. A bunch of consultants just make slides and now apparently they're not even doing that. Yeah, amazing. Incredible, incredibly lucrative career, I guess. So yeah, not to over-explain the joke of it all, but it's just so...
11:09Emanuel:We see this over and over. It's like the AI is kind of hollowing out our ability to do anything yourself. It's like, can we not make a grocery list for ourselves anymore? Can't turn... Can't read a PDF anymore. Can't read anything anymore. have to have it summarized, have to have it turn into slides. You slop those slides over to somebody else. Maybe they're using AI to read the slides, recording the meeting, summarizing for you later. I don't know. So yeah, it's trying to put some value into an already fake job, I guess. So this, like we hinted at earlier, is part of a much wider trend that we're seeing around AI.
11:47Emanuel:A couple months ago, we had token maxing, We had bragging about token use and spend. And now we're seeing companies scaling back. So just quickly, what are some of the companies that are actually talking about this outside of Accenture and talking about charging customers or scaling back employees using AI? I mean, the first big change, which I think I mentioned, was GitHub moving to a token model. And that happened, I think, at the start of June. and we're recording this at the end of June. And that's when we've seen, I think, a bunch of media coverage and also for us, like a bunch of leaks from various companies where, oh wow, that took no time at all to start impacting people where the change comes in June 1st.
12:38And then halfway through the month and then towards the end of the month, we have all of these leaks, including some of which will be in an article that we'll publish after this podcast is actually out. But sort of the biggest one, I think for me, is probably Uber, where they capped employees' use of AI tools like CloudCode and Cursor as well. That came after the Uber CTO said the company had blown through its entire AI budget in just a matter of months. Walmart as well, I think stopped people using tools recently. So a bunch of massive, massive companies are either trying to find a way to save money, and we'll touch on that in the next story, or they're just putting the brakes on it entirely.
13:31On the token maxing stuff, I think Emmanuel might be better at that. What do you think, Emmanuel, about how this sits in with that whole idea of everybody wants to use AI as much as possible?
13:45Emanuel:I think we'll probably get into this a bit more definitely in the story that we're about to publish. But when we talk about the story in the next segment, I think it's just... Employees are in a very contradictory... They're getting very mixed messages. On the one hand, it's like AI is going to change everything. You have to use it all the time. It's going to make you 100 times more efficient. And that is now overlapping with, but please don't bankrupt the company because we're spending way too much money on it. And I'm trying to think about if I've ever experienced myself such a corporate 180.
14:25Emanuel:And I'm guessing it may be. It's like if you imagine one of those pivots to video where they're like, we need everything to be video. It's video all the time. And then you wake up one morning and they're like, we don't have the money to shoot anything and get crews. But even that is more simple than this. They're just getting completely conflicting messages. I don't know how you're supposed to navigate that, to be honest. yeah that's a good analogy i think or it's the closest i can think of is you know we're we're shifting strategy this way oh wait we can't actually fund that strategy but we still expect you to like perform and do the work with these tools that you don't have which is just lovely um and just one last thing on this story because i think it is so good um a couple months ago the financial times reported that accenture started forcing people to use AI or risk missing out on promotions.
15:18Emanuel:So it's not just like, we really want you to use this. It's if you don't use it, you will not grow in this company. You will not ever get a better paycheck, get a better role. You will be stuck where you're at. Or I assume probably the implication is we're going to let you go because you don't fit in here anymore. And at the time, a spokesperson told CNBC, our strategy is to be the reinvention partner of choice for our clients and to be the is client-focused, AI-enabled, great place to work.
15:46Emanuel:And then later in the audio, Justice Kwok says, Accenture plans to formally launch a product called Token IQ, which I'm sure is like some kind of throttling mechanism for the token use for employees, but we don't know yet, right? Yeah. Yeah, because I mean, there's a bunch of other stuff in the audio that I didn't really get into because it does just like get probably a little bit too in the weeds and boring for listeners. But basically what they do go on to say is like, there's all of these different ways that we might be able to curb token spend. And that's like putting budgets in place, user controls and that sort of thing.
16:22And I don't think it's as serious as obviously like the content of like chat bots where we have this big problem where there aren't guardrails and fucking chat bots are telling people to kill themselves and all of that sort of thing. Obviously, it is not as bad as that. But there's an interesting parallel where chatbots didn't have the guardrails for that. And it did all sorts of horrible, potentially unexpected behavior. Here, AI tools have rolled out and there haven't been guardrails in place to stop companies spending millions of dollars or something on LLMs or burning through their tokens or whatever.
17:01So there's an interesting constant there where these tools are coming out and there might be a chatbot, it might be a coding tool, but you don't have these guardrails in place. And now Accenture can position itself after telling its clients to adopt AI as quickly as possible. We can help you deal with what Justice calls token economics. I think they're the start of the audio. They're introduced as token ops as well. This is a whole thing now where they're making sort of umbrella terms to talk about just this part of the AI industry. There are people now dedicated just to figuring out how do we not spend so much money on tokens and that sort of thing.
17:46Justice actually makes the point that the guardrails weren't there. That's just not me making this point. He's making that as well. And now we're in the position to be able to do that ourselves, which is incredibly convenient for a consulting giant, obviously.
18:03Emanuel:Interesting how it works. Okay, let's take a break and then we'll come back and talk about a very related story, but I don't know, a little bit even funnier, even more absurd. So we'll be back.
18:27Emanuel:This message is sponsored by Raycon. I've noticed that before I leave the house, there's a few things I always check for. Phone keys, wallet, but also lately, my Raycon Everyday Earbuds Classic. Raycon Everyday Earbuds Classic have become part of my everyday carry and there's something that I don't leave the house without. Whether I'm grabbing coffee, taking a walk, running errands, or sitting outside getting some work done, I almost always have them with me. What I like about Raycon's Everyday Earbuds Classic is the active noise cancellation. Well, that's one of my favorite features because when I want to focus on a podcast or music, I can just tune everything else out.
19:03Emanuel:But if I'm walking around town or riding my bike, I just switch to awareness mode so I can still hear what's going on around me. It makes me feel a lot safer in a very hectic LA environment. They're also packed with features like multi-point connectivity, so they stay connected to both my phone and computer. Plus, they have up to 32 hours of battery life with the case. And if I ever forget to charge them, the quick charge feature gives me about 90 minutes of listening from just the 10-minute charge. I also really like that they come in a bunch of different colors. I like these black ones. It makes them feel different from every other pair of earbuds that you walk around with and see out in the world.
19:41Emanuel:And honestly, it's very hard to argue with the value. Raycon gives you premium audio quality at about half the cost of the big brands. Plus, they have over 3 million happy customers who can't be wrong and a 30-day happiness guarantee. The Everyday Earbuds Classic are a great option for everyday listening. Go to buyraycon.com slash 404 to get 20 % off. That's buyraycon.com slash 404 for 20 % off. Thanks to Raycon for sponsoring. What's softer than cashmere and warmer than wool? It's not a riddle. It's an alpaca hoodie. And I had to check it out after hearing some of my favorite podcasters talk about paca.
20:21Emanuel:Paca hoodies are great for fall and winter, and honestly, pretty good for chilly summer nights or at the beach late at night when it gets a little bit breezy. But I also love their t-shirts and socks, which have been my go-tos during this very hot summer. I picked up paca t-shirts in several different colors, and I also think their socks look stylish and feel great whether you're hiking, traveling, or just bopping around town. Paka makes outdoor and lifestyle apparel from alpaca fiber, one of the world's most sustainable natural fibers. Their clothes are softer than cashmere, warmer than wool, and most importantly, they're really breathable.
20:58Emanuel:Their hoodies are built for real life. They're thermoregulating, odor-resistant, durable, and made to last. Over 250 ,000 people have already picked up Paka Apparel. If you've been thinking about leveling up your game, this is your sign to do it now. Again, they're not just hoodies. They make amazing t-shirts. They make boxers. I have some of those. And they also make amazing socks that I wear all the time. To grab your Paka hoodie or other apparel, go to www.pakaapparel.com. That's www.pakaapparel.com. If you've listened to this show for more than a week, you know our beat is uncovering how your data is quietly harvested and exposed online.
21:42Emanuel:And if you're like us, working remotely, constantly on the move, and running your entire life off your laptop and phone, well, you're leaving a massive digital footprint. Every time you jump between networks, check emails on the go, or browse from a new location, your device is constantly broadcasting your activity. Advertisers can track your activity, leading to targeted ads, price discrimination, and a sense of being monitored. It's the kind of systemic tracking we'd normally write an article about. That's why we recommend Surfshark. At its core, Surfshark VPN keeps your online activity private by hiding your IP address.
22:17Emanuel:It encrypts your connection, making your online activity much harder to track, so your research, your sources, and your personal data stay private no matter where you're working from. You might have noticed I'm traveling right now, and I'm using Surfshark. I've been using it for months at this point and I really like how easy it is. But Surfshark does far more than what I've just told you about. It blocks ads and trackers, which cuts down on the price discrimination game sites play based on what they know about you. It also includes Alert, which monitors your ID for data leaks and an email scam checker that protects against phishing attacks.
22:52Emanuel:And if you game, it helps defend against DDoS attacks and ISP throttling. Since you're always on the move, the best part is that one subscription secures unlimited devices so your entire household is covered. Go to surfshark.com slash 404media to get four extra months of Surfshark VPN with the reassurance of a 30-day money-back guarantee. Or just use code 404media at checkout. That's surfshark.com slash 404media. If you sold somebody a loaded gun who you knew was in a vulnerable state and they shot themselves. I think it is murder. Just because you're using the internet doesn't mean you get away with murder.
Read the full transcript
23:33Emanuel:I'm Damon Fairless, host of Hunting Warhead. This season, I take you inside the business of suicide and the places desperate people go when they can't find what they need in the real world. Hunting the Suicide Salesman, available now wherever you get your podcasts.
24:02Emanuel:So we're back with another story about AI cost, AI spend at companies now they're dealing with it. Headline is companies are making Claude and Codex talk like cavemen to stop AI's soaring costs. Joe, this is another story that you wrote. Just walk us through how you came across this, I guess, to begin with. Yeah. So Emmanuel and I are working on another story, which probably won't be in the show notes here, but we're going to publish it just around the same time this goes out to all of our free listeners. If you're a paid subscriber, you're getting this beforehand. Please don't listen to it and then go and scoop us.
24:43That would be really annoying. But we'll talk about that a little bit. So we're working on that story about the various ways that companies are actually trying to stop this. So whereas the Accenture audio was like, there's this issue and companies are spending too much on AI and we need to help them stop token spend, we looked at the ways that companies are actually doing that. And we'll talk about those in a bit. But a really, really interesting one I got was a leaked memo from a company that does sort of digital infrastructure and actually works on data centers now as well. in this memo that we talk about the various ways to curb token use.
25:25And one was use the quote unquote caveman plugin. And I never heard of this. And it was highlighted to me. And I went to look around. And yeah, it's there on GitHub. That's how I first come across it. Yeah.
25:42Emanuel:Gotcha. So yeah, you start looking into this. What does it actually do? We should have done this whole podcast in Caveman Speak, by the way. I don't think I could do it. I was already planning to do that for my behind the blog on Friday, but I did something equivalent last week. I think I kind of need to give people a real behind the blog. Yeah, this is becoming too experimental. Because now at the time, I'm just phoning in with weird bits every week. We should have done that. Why would it? No, I was going to do Cape Man voice, but I'm not going to do it. It's like a meme on TikTok and Instagram.
26:15Emanuel:Have you seen this? It's like explain things in Cape Man. Have you seen this, Emmanuel? Emmanuel is more online than... No? It's like your physical therapist explains to you, it's like, yes, protein, no, electrolytes. It's like, electrolytes good? It's like, it's a whole thing. It makes me feel like totally brain rotted. Anyway, so yeah, tell us what it does. And also you tried it. So tell us what your experience was with it. Yeah, I mean, it really does that. It changes the output of Clawed and Codex. and I think you can do Gemini and basically any sort of AI tool at this point. And it changes the output to not be the normal, verbose, flowery, over-the-top responses that you'll get from these LLMs and make it really, really to the point where it is avoiding completely unnecessary grammar to communicate its point.
27:13and I mean there are examples in the piece but I think it was you Sam that sort of gave me this idea when I was writing it but yeah it's less the sort of voice of the AI chatbot which is like oh you were right to push back I was wrong I'm so sorry here's a big explanation it just goes Hulk smash basically so if you're asking it to do code it just like here's the code it doesn't give you this over the top explanation if you're asking a question, it will just shoot out a response, which it believes is still accurate. And I guess in my tests, I would say it's still accurate. I haven't done it a whole bunch.
27:52But the tool does work as advertised. I downloaded it. I integrated it with Claude Code. I asked it to look at some code I had previously worked on. It was actually a way to scrape contracts from DHS procurement databases and that sort of thing. And I wanted to check if I ask it to look at this code that was previously written, is it going to give me a really in-depth analysis or just tell me how it works? And it just replied to something like, use this API, no scraping. I'm like, damn, that's true. That is right on point. Didn't need to provide any more information than that. So, it does what it says, which is that it doesn't change your input.
28:41Obviously, you can still be as verbose as you want. But it changes it, as the creator told me, just into a terse tool. Like you're not having a fun conversation with this thing, which frankly, if we're going to have to live with AI chatbots for a bit longer, or if I worked in a company where it was forced upon me like Accenture or whatever, I personally would much prefer engaging with it like it's a tool, not a conversation they have to have this fucking computer program every single day. So it seems like it does the job. Yeah.
29:18Emanuel:I've tried. So I'm not on a mission to make ChatGPT useful for me, certainly. Most of the time when I'm like, Oh, maybe it could do this. It doesn't do it well. So I was trying a while back to make it because I was like, maybe there is something in like, if it only spoke like a computer, If it was actually just returning tasks instead of giving me all this, it's like, shut up. Stop talking. So I basically told it. I was like, shut up. Stop talking to me. You're a person. You are a computer. Speak in computer syntax. Only speak in the necessary words and return the tasks. And it's really hard to make it do that.
29:59Emanuel:So I thought it was interesting that they created a plugin for this instead of just prompt injecting and making it like you are a caveman. Because I bet it would try to do a fucking character. Right. It would do like Ugo Haga. Yeah, yeah. It would work. With a club and shit. Like banging itself on the head. It would do a whole storyline. It'd be like, I'm the greatest caveman that existed. It's like, stop. Shut up. Stop speaking. Tangent. But I agree with you that the most annoying thing in the world is when software is like cute. Yes. And tries to talk to you cute. Like Microsoft made this change years ago where like if you have a crash, you're like, uh-oh, something bad happened.
30:38Emanuel:It's like, shut the fuck up. I'm going to shoot my screen. Shut up. What's the error code? What's the error code? Because I'm going to have to like plug it into Google. It's so annoying. And I also, I've just had an experience that was like, I forget what I was calling, but it was like an AI whatever interview thing. And it takes so much longer just because of the pageantry of trying to pretend like you're a person. And it's just like a person would actually be way more efficient than this. So just cut to the chase. It won't do it. It's so hard. I hate that stuff. I hate it. I do think it's like this is...
31:12Emanuel:I start to feel really old and grumpy. But we have strayed so far. when we started doing UI and graphical interfaces that spoke to you like a person, even just like the copywriting on the air message, like you said, the original sin. I think we should have been command line the whole way. We should have never strayed from that. But who could have predicted that this is where we would be today is trying to code using AI and making it talk like a caveman. I think just briefly on the point you made, Sam, that you could just do a prompt, right? As you say, and be like, hey, be a caveman, but there could be all sorts of side effects on that.
31:52I can't find the exact quote in front of me right now because the GitHub page is a little bit long, but the creator does mention there somewhere. This is more effective than a prompt or something, which I think exactly goes to what you're talking about. They've had to develop a skill. I think it's called a skill, right? In Claude Code parlance, but they've had to make something that goes a little bit further than could you just shut the fuck up, please? And it apparently works, yeah.
32:18Emanuel:Yeah, yeah, that makes sense. So we have some idea of who's using... who might be using this thing. Who is actually using Caveman Claude? Yeah, so the memo I got was from a company called Legrand, which I'll be honest, I had not heard of before. I think when Jason edited this piece, He was like, I've not heard of this company. You need to define it. And yeah, as I said, they're electrical and digital infrastructure. And I found it quite funny that they have moved into the data center business and that sort of thing. And they were the ones that have this internal memo that said, hey, maybe you should use Caveman.
33:02When I spoke to Caveman's creator as well, he mentioned that he's heard It's been used by individual developers in NVIDIA, OpenAI, and GitHub, and then some others as well. I'll just stress, that doesn't mean there's like a company-wide deployment of Caveman or OpenAI or NVIDIA or something like that. But that's pretty interesting. Ordinarily, it would be like, well, the creator is saying that. And maybe it's in his interest to just say that big companies are using it or people at big companies are using this tool. But what was really, really interesting is that he flagged to me, when you go through sort of the commit history of the tool, Shane Sweeney, who is the Director of Engineering at OpenAI, actually contributed code to this really, really early on when it launched and added functionality for Caveman to work with Codex, which is OpenAI's coding agent.
34:05So at a minimum, Shane Sweeney has written codes to work along with this tool as well. So I felt much more confident including the creator's assertion that people inside big companies are using this. None of those companies have got back to me, GitHub, OpenAI, or NVIDIA. None of them got back to me. But I absolutely would not be surprised if more and more companies are using this or more individual developers are using this. Because something I should have said earlier, but it does work not just in making the output more basic, but apparently it works with reducing token spend as well. So you can do a command in there, which is like caveman-stats or something like that.
34:55and you'll come out with, this is how many tokens we think you've saved. So it estimated without Caveman, it was like 8.8k or something. And then the token saved was like 5.8k, around 65 % saved. I would have to look a little bit more into the methodology to figure out how exactly are you determining that. But putting those together, yeah, this could be really, really, weirdly valuable to companies that are like, I'm burning through my token allocation or my budget or whatever. And it does make you think, maybe we should have just made the software be a fucking piece of software in the first place rather than an AI companion.
35:39Obviously, I'm being facetious, but exactly what Emmanuel was saying, all the cutesiness of it. Now, you have to install an independent tool and someone has to make an independent tool to actually get these sorts of code agents to where they want to be. And speaking of agents, it's not just like integrations with Claw code or Codex or whatever. I didn't try this, but you can download like a version for OpenClaw, which is the agentic AI, right? And you can even download an entire agent which does everything in Caveman. So like not just the outputs, but apparently a more all-encompassing approach to Caveman as well.
36:21I haven't tried that one, but I imagine that one works as well.
36:26Emanuel:Yeah. So you teased at the top, Emmanuel has a story coming. You and Emmanuel both have a story coming about companies, other companies that are dealing with this entire problem that we've been talking about for the last 30 minutes. Emmanuel, what else do we have coming down the line? Like you saw some stuff with Citi, maybe IBM. What are we looking at coming up? So, I think we want to save our best nuggets of information here for the article, which will be published tomorrow. But there are a few things that I think are not going to make it into the story that I think are pretty funny and are a good example of what people are dealing with.
37:11That is just examples of how companies are trying to limit the token spend, which is really just the amount of money that these companies are spending on AI tools.
37:23Emanuel:The long and short of it is that access used to be unlimited, and now there is a limit. This I have experienced where it's like, oh, we used to have unlimited access to Getty, and you can use all the Getty images you want. And then it's like, you need to sign four forms, and you only get one image a week. and that's kind of like what is happening at these companies and one of those companies is IBM which has its own internal coding agent called Bob and in order to use Bob you need to use Bob coins and it's I saw a screenshot on Slack and it's a bunch of IBM employees complaining that they're fresh out of Bob coins and it's like oh no my Bob points.
38:10Emanuel:I'm out of Bob points. How am I going to finish this project? And I thought that it's just such a funny dystopian snapshot of what the push and then the pull from AI tools is doing to the average employee. I think, Joe, there was like a... Did you want to talk about the Citibank? Yeah, I'll talk about the Citibank. So Citibank, the, you know, banking conglomerate that I'm sure many people will know, is that I've seen emails where Citi says, essentially, whoa, whoa, slow down. Please stop using so much AI. Obviously, I'm paraphrasing. You'll be able to see the actual quoted email in the piece. But they approach it in a couple of different ways.
38:59The first is that they really, really encourage people, hey, please use the appropriate model for the task. And I think we're seeing this across a bunch of companies where users are defaulting to the most recent and the most powerful and the most token-hungry model for Clawed or whatever. And that would be Opus 4.7 or 4.6 and then GPT 5.5 for OpenAI. and people are using like those models with reasoning as well, which is supposed to be, you know, the AI can really work through some problems and figure it out itself. They're using that for basic stuff, which I presume is stuff like making presentations.
39:44We were talking about that sort of thing. So they're encouraging, or pleading with employees, please stop using the powerful stuff for that. And then in other cases, according to the emails I've seen, Citi has just cut off access to a bunch of the more powerful models as well. So it started with please don't, and now it's like, you can't use this unless you have some very, very specific use case. I should say I went to Citi for a comment about this, and they said on background that we haven't actually closed off any access to models, and we're not doing this. I then saw a screenshot where someone is unable to access the models.
40:26So I don't think that statement is fully accurate. Or maybe you need to read between the lines there. But yeah, that piece will come out around about the time this podcast is at. And I'll try and get it in the show notes. But we got a bunch more detail from a bunch more companies about all of the different ways they're trying to curb this AI use, essentially. Because I mean, I think it's out of control for a lot of people. And I think you'll probably see that in some of the stats in the story as well. All right. Should we leave that there? If you're listening to the free version of the podcast, I will now play us out.
41:04But if you are a paying 404 Media subscriber, we're going to talk about flowers that don't even exist. You can subscribe and gain access to that content at 404media.co. As a reminder, 404 Media is a journalist founded and supported by subscribers. If you do wish to subscribe to 404 Media and directly support our work, please go to 404media.co. You'll get unlimited access to our articles and an ad-free version of this podcast. You'll also get to listen to the subscribers only section where we talk about a bonus story each week. This podcast is produced by Alyssa Midcalf. Another way to support us is by leaving a five-star rating and review for the podcast.
41:45That stuff really does help us out or just tell a friend about us too. This has been 404 Media. We'll see you again next week.
From the publisher
Go to surfshark.com/404Media to get 4 extra months of Surfshark VPN, with the reassurance of a 30-day money-back guarantee, or just use code 404MEDIA at checkout. That's surfshark.com/404Media.
We start this week with Joseph’s story about the Tokenpocalypse, which is companies scrambling to stop spending so much on AI after providers started charging per AI token. After the break, Joseph and Emanuel tell us about the ways companies are trying to do this, including using a tool to make their LLMs talk like cavemen. In the subscribers-only section, Emanuel explains how entirely fake AI-generated flowers are all over eBay, Etsy, and Amazon.
The Tokenpocalypse Is Here: Companies Are Scrambling To Stop Spending So Much on AI
Companies Are Making Claude and Codex Talk Like Cavemen to Stop
AI’s Soaring CostsScammers Sell Seeds for Exotic AI-Generated Flowers That Don’t Exist
Youtube Version: https://youtu.be/Sia4LZGNkVs
Learn more about your ad choices. Visit megaphone.fm/adchoices
