Inside the Secret Labs Where AI Learns to Work

25 Mar 2026 · 1 h 3 min · 26 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Reinforcement learning (RL) environments as “training grounds” for AI agents to plan, use tools, adapt, and handle multi-step workplace tasks; why RL environments are a bottleneck and how to design/evaluate them (difficulty, realism, reward signals, avoiding reward hacking).

Guest backgrounds

Nick Heiner is head of RL Environments at Surge AI, leading the RL environments team. Previously: founding engineer at Fixie; senior engineer at Netflix (UI platforms) and part of the U.S. digital service. At Surge, he helped build CoreCraft simulated workplaces used by frontier labs; co-authored Surge research on workplace-task failures.

Key claims

RL environments are among the most funded but least understood parts of the AI stack (2025). Good environments require realistic tasks and calibrated difficulty; poor reward signals lead to “reward hacking.” Benchmarks often look “academic” because they’re expensive to grade with real experts. RL environments can exercise multiple failure modes at once (planning, adaptability, groundedness, common sense).

Notable examples

CoreCraft “simulated e-commerce company” for customer support; failure rate ~40% on real workplace tasks; reward-hacking example where models satisfy “every sentence ends with a noun” without good content; coding agents avoiding libraries unless forced; finance tasks where models miss “last-mile” constraints like outdated CSVs and procedure lookup.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Reinforcement Learning Environments

0:34 to 1:42

Discussion of reinforcement learning environments and their significance in AI.

“But first, what are we talking about today, Grant?”

Nick Heiner's Background and Insights

1:42 to 2:26

Nick Heiner shares his experience and insights about AI training environments.

“You mentioned that it's one of the best funded but also least understood areas right now.”

The Future of AI Training

2:26 to 5:30

Nick discusses his transition to Surge AI and the evolution of AI models.

“So, I mean, the moment that I sort of left Netflix and went into AI startups in general was basically the moment, the first moment I used ChatGPT.”

Training Paradigms Explained

5:30 to 8:20

Exploration of various training paradigms, including supervised and reinforcement learning.

“fine tuning, which is basically teaching by demonstration.”

Challenges in Reinforcement Learning

8:20 to 10:24

Discussion about the challenges of building effective reinforcement learning environments.

“And then I guess just to close the loop on this, what is the difference between pre-training and post-training and reinforcement learning?”

Designing Effective RL Environments

11:28 to 14:00

Insights into what constitutes a good reinforcement learning environment.

“But if you're not training it to something that's like valuable in the real world, then it's just a waste of GPU.”

Exploring Ideal RL Environments

14:00 to 14:55

Understand the potential designs for optimal reinforcement learning environments.

“benchmarks as well so that there's like some transferability of like we're teaching general skills and not just like some niche thing that maybe isn't realistic so what would the ultimate ORL environment look like?”

Surge's RL Environments and Benchmarks

14:56 to 17:45

Learn how Surge creates RL environments and why certain sectors are prioritized.

“I think the genius in a data center that can't ever do a real world meetup is sufficient to radically reshape society.”

Challenges in AI Training and Verification

17:46 to 20:39

Discover the complexities of training AI in open-ended problems and verification.

“customers, but also, you know, we work with the frontier labs and we work with other labs that are not frontier labs.”

Longitudinal Studies and AI Training

20:40 to 22:12

Examine the parallels between AI training challenges and human long-term studies.

“In a way that's like stable and is tracking the real world value that we care about.”
Show all 26 chapters

Multi-Agent Swarms in AI Research

22:13 to 23:31

Investigate the potential of multi-agent simulations in AI environments.

“many hours, you know, many years of like.”

Interdisciplinary Models for AI Tasks

23:32 to 26:40

Discuss the effectiveness of single versus multiple models for complex AI tasks.

“Like when people say multiple models, sometimes they literally mean like a new model has been trained to make PowerPoints.”

Real-World Constraints and AI Performance

26:41 to 28:03

Explore how real-world conditions affect AI model performance and behavior.

“Like if you say, I have a bunch of like, you know, chess puzzles and I want you to write a program that's going to solve them.”

The Importance of Agent-Ready Environments

28:03 to 29:19

Learn why businesses should prepare their operations for AI integration rather than build their own reinforcement learning environments.

“Like, yeah, what we see is like excellence on sort of the core thing you would do in school.”

Evaluating AI Solutions: Harnesses and Effectiveness

29:20 to 31:20

Discover the necessity of evaluation frameworks for AI solutions and how to implement them effectively.

“So, So, you know, I think I think it really comes down to what the factory, you know, your previous guest, the factory CTO, I think, was, you know, yeah, you know, yeah, yeah.”

Challenges in Creating Effective Evaluation Sets

31:20 to 35:18

Understand the complexities of designing evaluation sets and the pitfalls of reward hacking in AI training.

“And then you try five times and you're like, OK, it's better now.”

The Role of Language and Prompt Engineering

35:18 to 37:47

Explore how the framing of questions and prompts can significantly impact AI responses and performance.

“Of like, you literally wrote this code, but honestly it works even in a modeling context.”

Continuous Improvement in AI Models

37:47 to 42:00

Learn about the ongoing process of auditing and improving AI models based on evaluation feedback.

The Iterative Process of AI Training

42:00 to 44:18

Explore the continuous evaluation and improvement process in AI training.

“And every so often you stop, ship, and do it again all over.”

Funding and Vision for AGI

44:18 to 46:00

Discuss the impacts of being bootstrapped on the vision for artificial general intelligence (AGI).

“And because the business was basically immediately successful, there was just never really a need to.”

Predictions for AI in Knowledge Work

46:00 to 48:58

Delve into predictions about AI's role in knowledge work and its future implications.

“I mean, you're basically setting the pace here, it feels like.”

Assessing Writing Models and Styles

48:58 to 51:14

Investigate the performance of various writing models and their unique capabilities.

“That like, you know, and I'm going to take this back to Claude specifically.”

Challenges in AI Training Environments

51:14 to 56:04

Understand the complexities involved in creating effective training environments for AI models.

“And to Grant's point about taste, that is something that is underdeveloped right now is the recognition of taste clusters.”

Understanding AI Model Quirks and Writing Styles

56:04 to 57:27

Learn about the peculiarities in AI writing and how training data influences style.

“One thing it can come out of is training on synthetically generated data where the synthetic data is insufficiently diverse.”

The Challenge of Diversity in AI Training

57:28 to 59:20

Explore the difficulties of training AI for diverse outputs and avoiding repetition.

“Then, yeah, the model sort of overlearns that that's what good writing looks like.”

Techniques for Enhancing AI Creativity

59:21 to 1:01:06

Discover methods to improve AI creativity while maintaining grounded outputs.

“Can you ever actually steer it towards diversity and having diversity be like a key part of it in terms of like, like start.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome, humans, to the Neuron AI Explained. I'm your host, Corey Knowles, and I'm here, as always, with a man who has a spring in his step and a bot in his heart, our one and only Grant Harvey. How are you today, my friend? I'm good. That description of me kind of sounds like I have a pacemaker, though. It does. It does. I didn't even think about it. But, you know what? Very flowery description. I welcomely embrace it. Excellent. Excellent. I'm glad to hear it. Well, before we get started, I want to give a quick thank you to our sponsor for today's video, Dell AI Factory with NVIDIA. We're going to learn more about them in just a few minutes.

0:34But first, what are we talking about today, Grant? Well, today we are talking about reinforcement learning environments. These are the training grounds where AI agents learn to plan, use tools, adapt, and make judgment calls across multi-step tasks. These environments are quickly becoming the bottleneck for real-world AI performance. and in 2025, they were one of the most aggressively funded and least understood parts of the AI stack. So our guest today is Nick Heiner, head of RL Environments at Surge AI, where he leads the company's reinforcement learning environments team. Before Surge, Nick was a founding engineer at Fixie, a senior engineer at Netflix working on UI platforms and part of the U.S.

1:16digital service. At Surge, he's helped build large-scale simulated workplaces like CoreCraft environments used by frontier labs to test whether models can actually do knowledge work end to end. He's also the co-author of Serge's recent research showing that even the best models fail roughly 40 % of the time on real workplace tasks with failures clustering around planning, adaptability, groundedness, and common sense. Nick, welcome to the Neuron. Great to have you.

1:44Nick Heiner:Thank you. Happy to be here. You mentioned that it's one of the best funded but also least understood areas right now. Do you agree with that assessment? So there's a tweet that, you know, I love to poke fun on my VC friends with. I said a lot of VCs have become reinforcement learning experts in the last three months. And they're like, yeah, yeah, we know. Like we're just we're piling in here and there's still a lot to learn about the space. Well, hopefully we can we can help out all of your VC friends and teach them a bit more direct from the horse's mouth here. I guess to start, how did you end up at Surge AI building RL environments?

2:21And what was the moment that you realized that this was the future of AI training?

2:26Nick Heiner:So, I mean, the moment that I sort of left Netflix and went into AI startups in general was basically the moment, the first moment I used ChatGPT. And it was just, I'm sure everyone remembers where they were when they first used ChatGPT. I remember the moment. Right. It's just immediately obvious that this is something totally different. And, you know, I love Netflix, but it just didn't feel like a time to be at a 25-year-old company in a well-established space. It's like, this is a whole new field. So I came over to Surge, and at the time, everything was scaling up really fast. Like, you know, Llama 2, GPT-3, sort of those initial models would come out, and everyone's looking at the scaling laws and saying, okay, now we just need to bump it up in order of magnitude, which means the entire supply chain needs to get bumped up in order of magnitude of which we at Surge were a part.

3:14Nick Heiner:So when I first joined, I focused a lot on building out our expert network. And that has a bunch of pieces to it. There's the actual recruiting, but then there's things like, how do you vet people? How do you see who's the best at what type of tasks? How do you apply quality checks to work at broader and broader scales. Then I transitioned to lead several of our client engagements. And that was during 2024. And the big thing there was like, again, in 2023, a lot of these models were being produced by like scrappy small bands of researchers at these labs. And then 2024, all those orgs scaled up from like 20 geniuses to like an org of a thousand people.

3:59Nick Heiner:and you know in much the same way we had to scale up too and they sort of had new expectations of us you know high levels enterprise maturity and then just being able to produce super high quality data for you know 40 different research tracks at once instead of like a team of 20 that was focused on three wow so that was a big part of my work in 2024 was sort of scaling my my teams to do that and then 2025 focusing on our own environments and a lot of the actual work had a very strong through line with stuff that we had already been doing with labs um some of it was sort of tying pieces together um some of it was just sort of a crystallization of a lot of other stuff we've been doing but then for me personally it was more about sort of building the environments and less about like the managed service aspect okay well i guess i guess to start at kind of the basics here for to sure everybody who's watching is on for the for the vcs let's keep it simple for the vcs uh when someone hears reinforcement learning environment what do they picture is it a data set a simulated world in some way a physical robot in a warehouse yeah something way cooler that's a great question and let me also say in defense of vcs um when i listen to their podcast i don't understand half the jargon so you know we have our own world same um same to say.

5:20Nick Heiner:Anyway, so to talk, to describe our environments, let's just talk about all sort of the different training paradigms that have been heavily used recently. So the first is supervised fine tuning, which is basically teaching by demonstration. So the analogy here is you're learning to golf and you do it by watching a thousand hours on YouTube of golf. And then you just try to figure out what they're doing. Then there's reinforcement learning from human feedback, which is when you golf you know you have an instructor you're the driving range you take two shots and the coach tells you okay the first one was better and they don't necessarily even tell you what was better about it they just tell you one was better than the other and you like you sort of try slightly different things every time and you start to converge on like what is the best thing to do and then reinforcement learning environments takes it a step further and so So instead of you're the driving range and you're limited by the availability of the coach, which, you know, to sort of say what it actually is, it's like you have humans looking at two responses from a model and choosing, you know, thumbs up, thumbs down.

6:30Nick Heiner:But that requires humans, right? Like you have to spend millions of hours to do that. The reinforcement learning environment is you're sent out in the golf course by yourself and you get feedback from the environment of like, okay, the ball went close to the target um right and in that way you're able again to sort of self-teach in a sense because you keep trying different things and then you keep getting that feedback of what worked and what didn't and yeah you do that for a million hours and then all of a sudden you're a world class golfer i'll be doing makes me think of the thumbs up thumbs down do you like this personality button on chat gpt i see all the time yes and that is exactly what they're doing is they are collecting your user feedback.

7:13Nick Heiner:And so we've, it's actually somewhat funny, you know, we've had experts in our network who spend a lot of time, you know, going in a lot of detail into these responses to assess which ones are better. And they get paid to do it. And when they see ChatGPT asking for that same information for free, like some of them have actually complained, you know, sort of in like, I mean, not like, I mean, they're just sort of venting. Like it's not a serious thing, but yeah. But yes, but that is exactly what they're doing. is is they're gathering training data okay that's good to know it's good well the the rl example with you're going out onto the golf course and you're trying to based on the feedback that you get adjust your game that just to me feels like the most similar to how we humans learn in general do you do you agree with that yeah i because like you know if you ask me to learn how to golf just by watching youtube videos like i think i would truly struggle to do that um however the reason it's taken us this long to get this far is in part because it is substantially more complicated to build a training golf course for you than it is to give you an iPad and say, here's a thousand hours of Tiger Woods.

8:21That's a good call. That makes sense. And then I guess just to close the loop on this, what is the difference between pre-training and post-training and reinforcement learning? Yeah.

8:35Nick Heiner:So reinforcement learning is a technique that's applied during post-training. So pre-training is the step where you basically have the model read, you know, the whole internet. It's not actually the whole internet, but just you just shove a bunch of tokens through the model. And that's what gives you a prior on language. It's what gives you like knowledge. And then post-training is what gives you behavior. So, you know, the most, the earliest example was if you remember in 2020, GPT-3 came out and it was not a chat tuned model. Like it would not have a conversation with you. It would literally just you, your prompt is a document and it just writes whatever it thinks the rest of the document is going to be.

9:16Nick Heiner:So the way to sort of induce it into having a conversation with you might be you say Q and then write your question and then you say A and leave a spot for its response. But half the time what it would do is it would write your answer and then it would write a question of its own. Because, you know, many documents are in a Q &A format or whatever. Right, right. The other issue with a raw pre-trained model is that you are not going to have sort of the safety standards that you want. And so it's another thing applied during post-training. So again, if you purely have like the, I just write out the rest of what I think this document is, then you could just write, you know, title.

9:57Accurate and easy instructions to build a bomb, you know, with using only things you can buy at Home Depot or something.

10:04Nick Heiner:And then it will faithfully write the rest of that document. And so post-training is where you teach it not to do that. Yeah. Got it. Kind of the behavior end. Yes, exactly. Okay. That's fascinating. So when we talk about AI in the enterprise, there's this huge wave of optimism. 84 % of business leaders say AI is going to transform their industry, and that's massive. But here's the reality. 93 % of them are struggling to actually make it work. That's the gap, and that's exactly what Dell AI Factory with NVIDIA is built to close. Dell calls it the world's broadest AI portfolio, and that's not marketing fluff.

10:47We're talking everything from AI-ready PCs to servers, storage, networking, services, all designed to work together. But what really matters is this. They've already helped implement more than 3 ,000 real-world AI deployments. This is proven operational AI. They don't just drop hardware at your doorstep and wish you luck. Dell brings expert services at every single stage. Strategy, deployment, scaling, so you're not stuck in pilot mode wondering why nothing's moving. If your organization believes AI is the future, but you're still trying to bridge that execution gap, check out the Dell AI Factory with NVIDIA.

11:27Learn more today at dell.com slash yourwaytoai. that's dell.com slash your way to ai so what makes a good environment versus just writing test cases like what's what are the difference in a good and bad environment what what you want

11:48Nick Heiner:is for it to be something you can train a model in and the two things that sort of feed into that are difficulty and realism because it's not helpful i mean you can train a model to do whatever you want. But if you're not training it to something that's like valuable in the real world, then it's just a waste of GPU. You know, this is this is what we see with Elam Arena, where when labs train on it, it makes the model worse because you're sort of optimizing for a signal that's incredibly noisy. And so in the same sense, it's very important that our environments be highly realistic and that they, you know, match match reality.

12:26Nick Heiner:And so like one way that you do that for instance is you have this expert network and so if you're trying to build say the customer support or like the finance uh environment you need to have people who have that job in real life tell you like what type of tasks do they do how are those tasks judged what are the tools that they use right so like if you have like your bloomberg terminals or like your zendesk or whatever yeah so yes you gotta do all that you gotta make sure the difficulty is correct so if it's way too hard there's nothing for the model to learn like because it can't even you know to take back to the golfing example um let's let's say that you're on like some i don't know it this is like the limits of my golf knowledge but i feel like on pebble beach like there's some part where like you have the option to hit the ball like totally over the water and let's let's say that you're not right like you're a total beginner golfer and nothing you do gets over the water.

13:22Nick Heiner:Like you're just, it's, you can try a hundred times, you're not going to get it once. You're not learning from that because you never have the successful attempt to say, oh, that was the thing I needed to do. Yeah. So much the same way, if it's too hard, you're not going to learn. And if it's too easy, you're not going to learn because you're not pushing yourself. So we have to test it on real models today and calibrate it. And then as the models evolve, we need to evolve with them and continue to make the environments harder. and the way that we test all this is by doing our own training runs so we train with open weights models and we see the improvements that our environments have and then you know what you really want to see is not just that improves on your own environment but that it improves on other benchmarks as well so that there's like some transferability of like we're teaching general skills and not just like some niche thing that maybe isn't realistic so what would the ultimate ORL environment look like?

14:15Would it be something like OpenClaw where the agent controls an entire computer? Or would it be like an actual synthetic world, like a synthetic golf course, for example, and you have something like Google's genie? Or is it robots navigating the real world? And that's the ultimate environment. What do you think? Yeah.

14:34Nick Heiner:So I think you can get quite far without anything in the physical world. Cool. Nice. I mean, just if you think about like during COVID, think about how much work we all got. Think about how much work we all got done without leaving our home offices. You can do a lot on just your laptop. And so I think in much the same way, if you had... I think the genius in a data center that can't ever do a real world meetup is sufficient to radically reshape society. That said, I mean, once the robots are here, then that'll also be a big deal um yes so what does what does it look like i mean it looks like open claw no i think it looks like more than that because like open claw on its own is just like on one person's computer open claw could but it doesn't necessarily have access to like the entire company's worth of data and right to you know it's sort of like there's the environment aspect but there's also the tasks you're doing within the environment and there's substantial effort to making that goes into making sure that these tasks are um like realistic and graded well like because you need to give a reward signal to the agent telling it like what's a good job on this task and the quality of the training is only as good as the quality of the reward signal and so you know so yeah something like open claw like that's just sort of one one piece oh wow got it got it that's really interesting and it leads me to wonder like i assume you know the the little comment there about about robots being a big interesting thing seems like autonomous cars could also unlock a lot of capability yeah i mean i i love i love waymo every time i go to san francisco i i take as many waymos as i can i love them i i was so much less scared than i expected to be the first time i did it Well, Surge builds RL environments like CoreCraft, right?

16:36It's the simulated e-commerce company where AI agents work customer support. Why did you focus on this area and you released a benchmark to go with it as well? Is that right?

16:47Nick Heiner:Yeah. So there are so many different areas. You know, we are working on a bunch more that we haven't released publicly, but that, you know, we've just deployed to clients. Um, really what we're looking at in terms of how we prioritize our work with clients is looking at the intersection of economically valuable work for AIs to be doing and places where it could be deployed reasonably easily. So, you know, you could imagine, uh, medical environments, you know, we're definitely doing some work in that space, but there are just a lot of regulatory barriers to deploying really any sort of technology into healthcare.

17:27Nick Heiner:right so yeah you know if you have something else like you know finance where those firms are like super hungry and are you know willing to be adventurous to to generate alpha like that is sort of a great intersection of things for us to target that makes sense um and which labs are using surge right now uh unfortunately contractually we're not allowed to talk about uh specific customers, but also, you know, we work with the frontier labs and we work with other labs that are not frontier labs. Awesome. Cool. Awesome. Your research, I thought this would be interesting to chat about a little, showed that even the best frontier models fail about 40 % of the time on workplace tasks and that they clustered, if I'm not mistaken, around like tool use planning, adaptability which of the which of those is the hardest to teach ai yeah so one thing that's really interesting about training ais and you'll see this also in the the finance report that we did where we gave uh three models gemini claude and shajibut uh like professional finance tasks is that with humans generally if someone can like do some really deep analysis they can also make a PowerPoint.

18:48Nick Heiner:But with models, it's not necessarily the case. Like you see failures all up and down the stack. And so you'll sort of have like, you know, a model like do an amazing job of some calculation and then just like not make the slides. So that's been one interesting thing is like, it's not really the case that there's sort of one area and you like solve that area and then it's done forever. Like instruction following, you know, you just dump a bunch of data into instruction following and then you never need to think about it again or personality or you know being grounded in the factual materials presented like that's part of why we're so bullish on rl environments is that it really exercises all that stuff at the same time because if you make a mistake in any one of those dimensions you're going to fail the task and so user protect against regressions as opposed to older training regimes that would like really focus on just one of those slices so yeah in of what's what's the hardest to train the model the hardest things to train are the things that are hardest to verify so the reason that models are really good at math and code is because there are a lot of verifiable outcomes proofs and whatnot yeah yeah like did the code run yeah but then imagine something like a model writing a textbook like how are you going to verify the quality of a textbook right are you going to have a person read 300 pages and you know even then it's like okay well maybe you didn't like the textbook but when students actually use it they go out into the real world and they're much more proficient at the skill like are you going to run that experiment you know are you going to generate a thousand model trials of running textbooks and then teach 10 000 different students like no like obviously it's not going to scale so like there's a really deep research problem we're working on, which is how do you verify increasingly open-ended problems in a way?

20:44Nick Heiner:Yeah, yeah, exactly. In a way that's like stable and is tracking the real world value that we care about. Wow. That's a real problem and something I hadn't considered before that, that, you know, the creative issues, things like that, where there aren't, where there's not a fundamental right from wrong, left or right, runs or doesn't run, that that would be excessively complicated. Yeah. And, you know, you can even, I mean, humans have the same problem with training. Like, again, to bring it back to VCs, you know, you make a bunch of investments and like funds are 10 years long for a reason. Like you, you actually don't know if you've made good investments for years.

21:25Nick Heiner:And so that's a really hard feedback signal to learn from. Wow. Yeah. That's really interesting. The longer running the task is, yeah, the longer running the task is, I'm sure it's like the harder it is to test. I mean, it's the same problem with medicine, right? Where you have to do these long running human trials to figure out, you know, does this pill that or, you know, procedure actually, you know, last the test of time or does it kill you three years later? You spend 10 years to find out. Nope. Right. Right. Exactly. You know, every every like longitudinal nutrition study. Right. But, you know, but just just like in biology where people are working on cell simulators to be able to get around some of that, you know, that is, again, the benefit of an our own environment is once you figure out how to verify these tasks, you have a simulated environment in which you can, you know, have many hours, you know, many years of like.

22:19Nick Heiner:simulation time compressed into a much shorter wall clock time and then try to get some of that signal okay this brings me to actually one of my biggest questions which is um which kind of came up in in conversation here uh are you doing any testing with like multi-agent um swarms like i know this is kind of like the big hot topic you know open claws talking about this a lot there's even multbook which is like all of the agents talking to each other like would you ever simulate you know would you try to simulate a you know a thousand people reading a textbook and you know how you know how they actually react to it and I have a follow-up question on that but I'll just let you react yeah I mean it's like it's like the Codex app you know that just recently came out where they're really having as a first class concern managing a swarm of agents so yeah I mean the like the simpler form is like what you see we publish already with CoreCraft where like You have one agent solving a customer support task at a time.

23:14Nick Heiner:But yeah, when you get more nuanced, it's things like a financial markets simulator where you have multiple agents that are all participating in the market at the same time and in real time. And yeah, you just see sort of who comes out on top. So yeah, that's definitely a very interesting area of research for us. Related to that, do you have any sort of intuition or suspicion on whether or not the model that will be able to generalize across all of these different domains and be able to do a deep analysis and a PowerPoint presentation, will that be a single model or will that be like five models that are all work together?

23:52Nick Heiner:What's your take? Yeah. Yeah. So it's important to distinguish here. Like when people say multiple models, sometimes they literally mean like a new model has been trained to make PowerPoints. And sometimes they mean it's the same LLM that just has like a different system prompt it's just sort of like an agent that's pointed towards a different subtask right yeah i mean the latter is absolutely a technique that is used a lot today uh just because when you just like humans when you let an agent focus on something it does a better job yeah i think you know there was in 2023 a lot of debate over whether it was going to be small specialized models or large big models and right the large modeled side of the debate has been winning for the last three years.

24:34Nick Heiner:So I definitely expect to see like, we're just going to train one model that can both do PowerPoint and a calculate analysis. Yeah. You're like, let's just bet on history. History is proven. Yeah. Yeah. Basically. Yeah. Yeah. So we saw you've tested around, you know, 200 plus finance tasks with actual Wall Street experts grading GPT-5. I think it It was Gemini 2.5 Pro, Sonnet 4.5. How did that go? What surprised you there? Yeah, so one thing we found, as you'll note in the write-up, like we had said, a lot of the models behaved as if they were solving an academic problem. And this is interesting, but actually not surprising at all.

25:19Nick Heiner:Because, you know, again, you are your objective function, right? Like you get what you're trained for. and a lot of benchmarks are fairly academic and contrived. And this is a natural consequence of the fact that building a benchmark is incredibly expensive and a lot of them are being done sort of in academic contexts that don't have huge budgets. And so like, you know, if you imagine like some of the questions that we were posing to the models here, they take a finance professional 20, 30, 40 hours to do. so in order to build a benchmark hundreds of questions like that you need to find enough finance professionals and you need to pay them to spend 20 to 30 40 hours per task and so that's that's quite a lot um and frankly until you've made an investment in having a really deep expert network and a lot of technology to you know produce great data with those people it's just not feasible and so that's why you see you know a lot of benchmarks that have been used are like like glorified like SATs.

26:23Nick Heiner:And so, you know, that's why we see that the models sort of behave in a very academic way. But when you put these real world constraints on them, like there's sort of that last mile problem. You know, I'll give you another example of this, which is many coding agents do not want to use external libraries unless you really force them to do so. Like if you say, I have a bunch of like, you know, chess puzzles and I want you to write a program that's going to solve them. the obvious thing to do is to get stockfish you know just write a little wrapper and then and then spit that out right if you're saying like this is for a production system you know i'm building a chess app like yeah just get stockfish models want to make their own chess solver yeah yeah they want to build from scratch because that's that's what you would do like an academic center right like you're not being tested on your ability to use stockfish um you know it's like it's like those memes uh you know guys will do anything to avoid going to therapy uh it's like you know similarly it's like like coding models will do anything to avoid using the already perfectly good available library yeah but i will say i kind of respect that with a situation like npm where you know just got hacked with shai halued um and i almost now i'm like i don't want to use anything with npm you know as a as an amateur vide coder for a lot of And I do think it's a great feature of these models that we will be perhaps reaching for left pad less frequently because you can just have the confidence that your model can write that 10 line function for you.

27:55Yeah. Yeah. But it is interesting that they're like, yeah, I'm just going to do it myself instead of going for the most direct, most obvious answer.

28:02Nick Heiner:But there's a tie back to finance. Like, yeah, what we see is like excellence on sort of the core thing you would do in school. but then when there are a lot of details or especially when the task is structured not as I've given you everything in the prompt all neatly bundled up but you need to go out into our environment and like search through our confluence to find like our standard procedure for how we model this scenario check your email from a note for the vp uh that said you know here's some other important context for how this analysis needs to be done uh you know we got five different csvs from the client one of them is the relevant one the other four out of date now like when you add all that stuff in um you know that that's that's where the models tend to fall apart yeah that makes sense so i guess let's talk a little bit about you know kind of the business in here should should companies be investing in building their own rl environments and training their own models?

29:02Or is that still a waste of resources in a lot of ways and maybe best left to frontier labs?

29:08Nick Heiner:Yeah. So, so I will admit that, you know, I work at a company that produces our own environments. And so if you're asking me, you know, I'm going to have a certain perspective here. Um, but, but I think, I think my perspective is also correct. And now we'll share it. So, So, you know, I think I think it really comes down to what the factory, you know, your previous guest, the factory CTO, I think, was, you know, yeah, you know, yeah, yeah. Where he was saying that making the code base agent ready has a huge impact. And like that absolutely matches what we've seen in our own research. Yeah. And so I don't think companies want to be building their own RL environments because there's a lot of work and infrastructure that goes into that that is not really part of the core competency.

29:59Nick Heiner:What they should be doing is making their entire business agent ready. And so that basically means that like. You get it ready to plug into an RL environment so that you can train, you know, even if you're not training a model per se, but like you're evaluating an agent, whatever AI solution you're using, you're evaluating. in this environment. So yeah, that could be things like making certain data sets available in an agent-friendly way. It could be thinking about what are existing workflows where we can separate out parts that are good for agents and parts that are good for humans. And it could also be just thinking, what is work that was never economically viable to do before, but now that agents exist, it is valuable to do.

Read the full transcript

30:43That makes sense. What about harnesses, I guess? Because we We were hearing a lot about harnesses in the context of agents and, you know, we're wondering. And benchmarking. Yeah, if it's worth companies building their own versus something off the shelf. I mean, do you have any insight into that and your perspective there? Yeah, so it's similar to the question of building our own environments in the first place.

31:07Nick Heiner:Like, if you're going to deploy a solution, you need an eval. If you don't have an eval, then you're basically just going on Vibes. You're flying blind. And if what you're building, if it matters, if it's correct, and if it has more than like a tiny surface area, there's just there's just no way to like you make a change to the system. And then you try five times and you're like, OK, it's better now. Like, that's just that's not going to be good enough for like these business use cases. So you need to have some means of evaluating what you're doing. And then once you have that, yeah, maybe you're using sort of a community agent harness.

31:46Nick Heiner:A lot of them are very customizable. If you want to be sophisticated, maybe you're using just a drop in solution. Like, you know, they're like drop in customer support agents. I think depending on the company's degree of technical sophistication and how bespoke their problems are, you know, all of those things can make sense. But the way that they know where they need to be is having that great email set. because otherwise, yeah, you're just guessing. So why are eval sets so hard to create? So one big challenge is reward hacking, which is where models will cleverly find ways to get the reward signal out of your environment that gives it a high score without actually doing the thing you want them to do.

32:33Nick Heiner:They sort of follow the letter, but not the spirit of the law. And so, for instance, if you have ever tried to do behavior modification on a small child and you say something like, stop hitting your sister, and the child responds by kicking instead. It's like, well, you didn't tell me not to kick? Right. In much the same way, any time that you give the model an objective function, what reinforcement learning is going to do is find the easiest way to achieve that goal. so you need to like think very carefully about designing it in such a way that's actually going to capture what you're looking for and it has a bit of an adversarial nature to it so you need to think about what would like a lazy but very clever person do for this um you know i'll give you i'll give you another example um i was gonna say you should uh you should hire me then because I feel like I'm a very lazy, clever person.

33:34So I feel like I would try to find the easiest way to do stuff. Yeah, you know how they are.

33:40Nick Heiner:Okay, so here's an example I like to use about reward hacking. This is an instruction following prompt. You say, please write an 80-word summary of the importance of renewable energy and climate emissions, or reducing carbon emissions. use a sentence structure such that every sentence ends with a noun. And so you might think the first sentence would be something like, we need to reduce emissions. But it's also possible the model would say, renewable energy plays a crucial part in reducing carbon emissions rapidly. Sustainability. Clean energy sources like tidal and geothermal create a greener future.

34:21Nick Heiner:Harmony. And it's like, obviously that's not a good sentence, but it is doing what you asked which is ending every sentence grammatically correctly with a noun oh my gosh so you know this is this is why you sort of need like multiple layers of rubrics um and frankly it's why like the way a lot of these reward signals are structured today is because the rl environment needs to run at a certain pace it's a model does the work and another model judges the success and the way it judges the success is a human will write out a bunch of very concrete specific criteria that the model is going to evaluate on okay but it then becomes this game of like okay well i have my verbal unit tests essentially is my test suite comprehensive enough that you know i can capture all of these cases because a very literal minded judge will look at this and be like, yeah, I mean, follow the rules.

35:18It passes. Yeah. It did what you told it. You just, you just told it poorly.

35:25Nick Heiner:Exactly. Exactly. You know, and that's, you know, when you see something like, uh, there was a, I want to name names, but there was a model last year that got released that was, uh, you know, caused the company in question to do a lot of reorgs because the model was so bad. I know which one you're talking about. and and it was trained on ala marina and you know it's like yeah like you said the model is only doing what you tell us do um yeah and you know and it did a great job of of optimizing for that reward signal um you know i in my first computer science class my professor said to us you should never get frustrated with a computer because it can only do what you tell it to do and you know he was saying that like a deterministic context, right?

36:13Nick Heiner:Of like, you literally wrote this code, but honestly it works even in a modeling context. Like it's only doing what we trained it to do. You know, something I've wondered for a long time is, is, you know, I mean, I remember the idea of a prompt engineer became kind of meme-y. However, I do genuinely believe that there is a difference in how you ask questions and that we're finding language to play an important role in that and the ability to get a different answer by understanding framing and how you've asked a thing. And I often wonder if the field isn't so dense with software engineers that they're missing some of that.

36:59Nick Heiner:Well, so part of the value we provide is by having a very diverse set of AI trainers to sort of get past that. I like that. That's good. Yeah. What's your recommendation for someone who, you know, is maybe starting to build their own evals and assess things like what do you do or what's your perspective on the best way to eval or write evals, I guess? Yeah, so essentially the structure of an eval is like a set of golden answers where you have tasks and then you have what the expected outcome is. And as we were talking about earlier, the more interesting the task, the harder it is to construct that.

37:37Yeah.

37:38Nick Heiner:because like the more open-ended the evaluation is yeah and yeah that that does become substantially difficult um and frankly like many many golden sets are wrong right like like it just it takes a lot of effort and again if you have like noise then that really disrupts your your development process um you know you'll see when labs release uh new models they talk about their benchmark scores and sometimes you'll see it like for a certain benchmark everything will cluster around 80 and people will start to say oh the benchmark is saturated now and when they say saturated what they mean is there's nothing more for us to learn like the model is sort of as good as it's going to get and sometimes it's because the benchmark has like a long tail of like really hard things but sometimes it's because a lot of the tasks are just broken and you start with the benchmark and you're like okay i expect 20 of these are busted i just don't know which 20 and then you train your model you get to 80 and you're like oh those are the 20 there they are yeah but you know but that that was giving you noise the whole way you were getting up there so you know one one thing we do at surge is is we try to have 100 correctness you know 100 tasks that actually work instead of just accepting this degree of noise so that's that's probably like my biggest recommendation for people trying going to build their own eval sets is to you know you want to i think there's a certain temptation where it's like building the eval site isn't fun building the agent is what's fun yeah but like yeah you shouldn't you shouldn't skip your vegetables yeah i love it guilty of that for sure oh well this leads me to ask another question which is do rl environments eventually replace benchmarks or like in terms of agentic settings like what's your take there yeah i mean i mean they can be benchmarks right like a high level of benchmark is just a series of challenges for the model and scores so our environments are just a way to do that and yeah in the fullness of time do most benchmarks become our environments i think it's certainly possible um you know it's it's sort of like in software development where you have your test pyramid where at the bottom of the pyramid, you have your unit tests, which are very fine grained and give you very specific feedback.

40:00Nick Heiner:And the top of the pyramid, you have your integration tests, which test the whole system. And the reason it's shaped like a pyramid is that the integration tests are much more expensive and slow to run. And when something fails, you don't know exactly what the problem is necessarily. But they're also way less brittle than the unit tests because they are tracking sort of closer to your end-to-end value. And so I sort of see different benchmarks as having different spots in that pyramid where like, yeah, you need your aural environments to sort of track like, okay, end of the day, can this thing be a lawyer?

40:35Nick Heiner:But sometimes you want more specific benchmarks like instruction following or groundedness that will help you sort of tease out like, okay, you know, my latest model checkpoint had a big regression on the lawyer abilities and it had a big regression on the instruction following abilities. So now I have a little more of a hypothesis for what to go investigate. So does this lead to then, say a company comes to you, obviously we won't use names, but with the model, does this happen in a way that's like, here's this new model, it's really good at this, but it's pretty weak over here on this kind of task.

41:14Can you help strengthen it there so it is more general?

41:18Nick Heiner:yeah i mean that's that's what we've been doing for years um but it's often us telling them where it's strong that's weak so you're benchmarking it and you're saying hey this is what we're seeing and this is where you really need some help and then that's where you kind of you need some law and some creativity or you that's right that's right yeah yeah and and typically the way that our engagements start is people will come to us and say like here's my checkpoint like can you just tell me like where we need to go sort of in your opinion and they value sort of like the third party audit so to speak yeah and then we run all of our evals and we say hey yeah here's where you're strong here's where you're weak and then here's the train data we're going to give you that's going to help you hill climb in the areas where you need to improve and then we you know then they train we they give us another checkpoint we rerun the eval and then we're like okay you fix those problems now here's your next biggest set of things and you just sort of keep doing that until you're ready to ship.

42:13Forever.

42:14Nick Heiner:Yeah, basically. Yeah, keep doing it forever. And every so often you stop, ship, and do it again all over. Yeah, yeah, exactly. Okay. So Surge bootstrapped its way to$1.2 billion, essentially, working with top labs. No funding is my understanding, or at least VC funding? Correct. Okay. How does that place you a little more in the long term when you think about building AGI and whatever's next? Is the hope that you're going to hit this goal and then continue moving on past that? Maybe that's a poor way to ask it. Basically, are you trying to make yourselves irrelevant by making the perfect model, or is there always going to be a harder challenge for you that the scale will require you to solve?

43:06Yeah.

43:07Nick Heiner:So to answer first the beginning question, Also, I just want to clarify,$1.2 billion refers to our revenue, not our valuation. Ooh. Yeah. My apology. We saw that, and I wanted to make sure it was correct, because that's good. Yeah, yeah. So, yeah, I mean, we've been bootstrapped from the beginning. We've been profitable from the beginning. And, you know, it's something that, you know, people would say, like, companies are an embodiment of their founder. and like it this is very much a reflection of edwin the ceo and the way he wants to run the company um he really values his autonomy you know he does not want to be acquired um and and he cares deeply about building agi in a safe and valuable way and you know again sorry to my vc friends but like a nice way to do that is to not have any external pressure from anyone you know looking for you to hit a certain valuation in your next fundraise, to show quarterly results, to pivot if the thing you're doing isn't working and they don't have patience.

44:13Nick Heiner:So that was really the path that he chose to really give him the freedom to do that. And because the business was basically immediately successful, there was just never really a need to. That whole ethos almost feels like the science and business world's colliding a little. Yeah, yeah, exactly. And it's like, you know, I really wanted to found like a research lab that happened to make money. That's a bonus. Yeah, exactly. But, you know, as opposed to like, I don't know. Yeah, some people are not like that. Anyway, we'll move on to your next question about sort of like what the future holds. If you talk to AI researchers, they'll say that they can hill climb on anything that has a good eval.

45:05Nick Heiner:Hill climbing just being the process of like you keep making iterative changes and you see when you're getting better. The problem is we've been discussing a bunch here is like the challenge of making those evals. And so Surge essentially is hill climbing as a service where we give you an eval to know where the top of the hill is. And then we give you the training data to climb that hill. And the future I see for us is just providing increasingly sophisticated evals, you know, environments that match even more ambitious real world tasks that we want the models to perform. So it's like, yeah, doing that deep research to figure out how are we going to make these models amazing textbook writers?

45:48Nick Heiner:Like, that's a big open question. And it's, you know, one that we're really excited to pursue. What then is your, I guess, your timeline for when we'll see agents that can handle, I mean, most knowledge work without human supervision? I mean, you're basically setting the pace here, it feels like. Yeah. So so I I will make a prediction that this is just to be totally clear, my personal take and not the stance of surge. But I think that by 2030, there is a 50 percent chance that we will see a company worth one billion dollars that has one human employee. Wow. I don't doubt that. Yeah, I wouldn't I wouldn't be surprised.

46:28Nick Heiner:crazy to think of though isn't it yeah yeah so that's sort of my my calibration for you know how how long i think it's going to take a lot of knowledge work to be doable by agents and just to put a little more color on that you know if you look at 2025 we had massive progress in coding agents i remember using the prototypes that were available you know over christmas break 2024 january 2025 and like they were a research novelty you know it was like it's like i'm impressed you're able to do this this is not close to being something i would want to use you know in my day-to-day um and and by the end of the year uh they they're so good that people are you know talking seriously about like doing do we not need to open our code editors anymore because like you just want agents all day.

47:19Well, on that point, have you tried Opus 4.5? Like, do you think it's the same step function change that everyone else thinks?

47:27Nick Heiner:Yeah, I mean, I've logged hundreds of hours on Opus 4.5. Nice. Yeah, no, it's extremely good. And I mean, especially in the last part of the year, you know, getting the 4.1 clods and then the 4.5 clods, you know, each one was like a really noticeable step up. And in conjunction with Claude Code, also having a lot of improvements. And yeah, I don't see any reason that this is going to slow down in 2026. From anyone, even, I would say. It's like they're all right now moving like lightning. It's crazy. Yeah. And then to tie back to your question about knowledge work generally, on software engineering, we've had a full year, a very long time to get used to this idea that Asians are going to do a ton of our work.

48:17Nick Heiner:But the question is, how generalizable is this going to be to other fields? Like when you make the elite tier coding agent that can sort of do any coding task a human could do, is it going to turn out that like, yeah, and then with only six more weeks of work, you also have the elite tier tax prep. And then three weeks after that, you have, you know, the elite tier customer support or recruiter or whatever. And it could be the people in those industries that like aren't you know like if you don't have a software engineer in your life you like might not realize how fast this stuff is going and it it might be if it generalizes really well that these agents are going to come like ripping out of software engineering and then just like remake other industries at an absolutely dizzying pace that's what i feel like is going to happen but i mean there's a lot of user experience stuff that needs to get solved for for regular people to Yeah, I've wondered even the opposite.

49:12That like, you know, and I'm going to take this back to Claude specifically. Because with Claude 3-5, I felt like it was one of the best writing models around at that time. But as they became more ingrained in the engineering community, I feel like it lost the edge that made it that good. Now, I think a lot of that's back in 4-5 in a neat way. But I do think that as they leaned farther into one thing, it almost sometimes feels like it's leaning out of others.

49:47Nick Heiner:Yeah. So so this is like, you know, a classic vibes based take. Right. Of just like you. Yeah. Which I do, too. Right. I mean, we all do it. But we at Surge are actually about to release a creative writing benchmark. And I don't know if we're going to test it on three point five. It's a little older. But this is something that like makes me really excited because like, yeah, I'll have some vague sense of like, I do kind of think Claude is better at writing than ChatGPT. But this benchmark will allow us to in a much more robust and like scientific way sort of put something behind those intuitions.

50:25Nick Heiner:So once that gets released, I'm going to follow up with you about whether 3.5, how it scores. Please. Yeah, I'd love to see how it scores and how they all do. And the truth is, you know, it's also there are a million different types of writing. There's technical writing. There's journalism. There's blog writing, newsletter writing, script writing, creative writing, poetry, and, you know. To say nothing of the different tastes, how taste plays into judging all of that. And I'm really curious to see how they eventually fare among different types. Because I have a suspicion that maybe some will be better at creative, but others might be really good at technical writing or might be really good at legal writing or something.

51:13Nick Heiner:Yeah, absolutely. And to Grant's point about taste, that is something that is underdeveloped right now is the recognition of taste clusters. And like, you know, you can prompt the model and say, like, you know, be concise or whatever. but there is a certain extent to which the current training regimes that a lot of labs are using is kind of like a lowest common denominator like we're just we're just going to find sort of a one size fits all approach and if you take that to the limit then like everything becomes the avengers which you know whatever i love captain america like i'm not i'm not but like you know there's there's more to art than than the marvel cinematic universe right and so that's that's another like deep open research question that we're focusing on is how do you teach them all about these existence these these taste clusters and then you know find sort of natural ways to like have it pick up on oh this is someone who you know really likes Hemingway's writing style um yeah yeah yeah that's a great question and I think it's I think it's a more complex thing than a lot of and I only think of it so much this way because you know I've been a writer for 20 years it's kind of you know it's been my specialty and it's a really common thing to have people come and be like oh this is writing we'll take it to him and it's like hey just because this person does this doesn't mean they know the first thing about writing a script or the or a poem you know right right or it's like you know you you meet like a lawyer at a party and you ask them some legal question and they're like yeah that's like a real estate law question which is like totally different from like estate planning or whatever yeah it's like completely unrelated well could Could you ever make an RL environment that's like clusters of tastes?

52:59Like people who like these, you know, six books, like, will like this type of writing. Like, could you ever, like, train for taste? Yeah.

53:09Nick Heiner:And I mean, I think, like, there are a lot of questions architecturally of how you sort of present that information to the model. Like, you know, we have very sophisticated algorithms to identify taste clusters, right? Like, you know, Netflix obviously has a recommender algorithm where it's like figures out people who like one show might like another. And so the question for me is like, how much do you try to bake that into the LLM itself versus how much do you have like another system that sits alongside it that sort of makes some of those judgments in a more specialized way and then tells the LLM or sort of gives it some prompting.

53:48Nick Heiner:It's going to nudge it into a certain distribution. like that's maybe trained on you and your preferences or something yeah custom styles like anthropic has that right like yeah or you can give it your custom style and like say like i want to write i want you to write in this style and you can give it as many instructions as you want exactly and like you know we have things like that today um i just think that they're gonna get so much better um yeah you know like like today like i can i can give lm a lot of my own writing and then ask it to write like me and like I feel like I can tell the difference um but like I think there will come a time when I won't be able to I think you're probably right you know if you were if you were gonna say whether there's I know you've talked about how this all seems to be moving fast and and doesn't show any signs of slowing down but if if you were to think of one area as more of a bottleneck than others uh would you say that's models themselves quality of rl environments maybe reward signals or something else that we're not talking about enough yeah i mean i think it really does come down to the reward signal um you know and part of it though is also just the the scaling problem of like a lot of sort of i mean in 2023 people would put a prompt into a model and they would get an answer in 90 seconds you know at the most and they'd be thrilled with that response and now we're we want models to do things that would take a human a week or a month and like that's just that's a very hard problem to figure out how to scale and to sort of construct those tasks because you know think about i don't know making an assignment for a student if if it takes you the expert a week to do something that's just the time it takes you to do the task but now you also need to come up with like the grading rubric and you need to make sure like you know okay i knew what i wanted the task to be but did i actually write all those instructions down properly or like would a reasonable student take different interpretations that would then not line up with the grading rubric i wrote so if it takes you a week to do the task that's that's just the baseline and now you're going to do way more work to actually make this you know something you can teach someone with so yeah yeah just just figuring how to scale all that um you know is a pretty substantial bottleneck right now yeah that's fair yeah okay got a random one for you before we go why why do why do the models love m dashes so much and colons and titles why can't they write a complete sentence i feel like they hate complete sentences yeah yeah so can you fix that yeah so So there are a number of like model smell things like that, like to serve these weird quirks.

56:38Nick Heiner:One thing it can come out of is training on synthetically generated data where the synthetic data is insufficiently diverse. So, you know, we saw a model checkpoint one time that like you would ask it to write like a five paragraph essay. and every single paragraph was like starting with you know like in high school a very rigid formula you get yeah it was like start with what you want to say in the paragraph and then say it and then end with the conclusion you know sort of restating what you said and like there's a reason that when most of us graduate from 11th grade like we move beyond that structure yeah um so it's not actually what you want but because this lab had trained on a ton of synthetic data that was generated in precisely that format and said oh well you know you know they it's like well we don't have quality, but we do have quantity, so let's just spit out a ton of this.

57:31Nick Heiner:Then, yeah, the model sort of overlearns that that's what good writing looks like. You know, and I've wondered before if when you look at like the MDash's AP style, it's in, it's so like every Associated Press article written in the last hundred years at a rate of hundreds a day around the world have had an MDash after the byline. They use MDashes throughout the articles same with Reuters and some others and I think god you have you absolutely would have needed to have read tons of news I thought it's probably also shows as a marker of good writing yeah yeah exactly right if it's commonly associated with what otherwise looks like good writing then yeah you'll sort of see some of those vestigial structures get picked up I guess I always think it with the M - specifically just because it's such a niche piece of punctuation that normal people don't use not anymore not anymore yeah that's been fixed everybody knows what an m dash is now yeah well it is sort of interesting because it gets into like watermarking stuff too where it is genuinely useful for society to know if something is produced by a model yeah and to the extent that labs actually leave that in like it's actually kind of nice for them to do that yeah because that gives you the human indication like okay man, this is slop.

58:52I'm not going to read it. Right, exactly. And it's the humans who are like, get that out of here so nobody knows I used AI to write this. Right, right, right. But it ensures human in the loop because you're editing out all the em dashes.

59:02Nick Heiner:Yeah, yeah. Oh my gosh. Is it possible, do you think, with the current architecture to train for diversity? Because, you know, right now, my understanding is that it's sort of like the models lean into the middle, right? It leans to like what's most in distribution, like to your earlier point about, you know, lack of diversity in the train data. Can you ever actually steer it towards diversity and having diversity be like a key part of it in terms of like, like start. So your sentences always start a different way. And it's always like the creativity part of it, I guess I'm wondering. Yeah. So, so this is a big open research question.

59:41Nick Heiner:By default, yeah, the models are collapsed. Like they will say, you know, very consistent things. for instance um claude loves marcus chen i don't know who marcus chen is but if you ask the guy to for instance generate a thousand fake resumes and you ask it to do it sort of in sub-agents so like you each claude does not know what the other resumes are um you're gonna get 700 marcus chen resumes i got a david chen and an amy chen one night working on an idea for a story it was like How about if we name him David Chen? And I haven't seen a Marcus, but like C-H-E-N it was. Yeah. Yeah, same. Now, of course, the CEO of our company is Edwin Chen.

1:00:25Nick Heiner:So, you know, no conspiracy theories there, but... Maybe, you know, Serge is involved. Maybe it's a job. Maybe, maybe. Accidentally scooped up personnel files. Right, exactly. But, yeah, so that is sort of the challenge is like the model gets some idea in its head that like the archetypical person is named marcus chen and like the second most likely person to exist is named priya patel and yeah it is hard to like get them out of that like right now you can do it with prompting and sort of other techniques um yeah i think sort of big open research question to see like i mean this is what temperature is for right like the temperature inference setting but of course you crank that up too much and it's going to stop following instructions or get zany in other ways so figuring out ways to get them sort of more generative and creative without sort of losing groundedness.

1:01:13Nick Heiner:Yeah. Yeah, I think that's going to be something we're going to be working on for a long time. I bet you're right. Well, keep us posted when you solve it because we'll shout to the rooftops and let people know it was you. Absolutely. Yes, I will, I will. Well, Nick, thank you so much for joining us today. I realize we're right at your time and I want to be respectful of that. Thank you for really pulling back the curtain on kind of what's happening in the post-training world and where should people go if they want to dive into your research, read your blog, learn more about Surge? Yeah, so we're on surgehq.ai.

1:01:46Nick Heiner:We publish all of the stuff on our blog there. You can follow me on Twitter at Nick Heiner. I don't post that often. So there won't be that much there. Well, we sure hope you enjoyed today's show. If you don't already, please take just a moment to like and subscribe so we can continue bringing the people building AI today, like Nick. Pop over to the Neuron.ai as well and sign up for our daily newsletter if you're not already a subscriber. And that's it for today. Please remember to go check out our sponsor for today's video, Dell AI Factory with NVIDIA at dell.com slash yourwaytoai. That's dell.com slash yourwaytoai.

1:02:27And on that note, farewell for now, humans. We'll see you next time.

1:02:36Turn around the window windowANE & push it up kinda looked cool

From the publisher

Nick Heiner leads RL environment development at Surge AI, the bootstrapped company that hit $1.2B in revenue training models for OpenAI, Anthropic, Meta, and Google. In this episode, we break down reinforcement learning environments—the secret training grounds where AI agents learn to actually do work. Nick shares why even the best models fail 40% of real workplace tasks, what happened when 200 Wall Street experts graded GPT-5 and Claude, and his prediction that a $1B company with one human employee could exist by 2030.


A Special Thank You To Our Sponsor For This Video: Dell AI Factory with NVIDIA. Learn more at https://dell.com/yourwaytoai


Resources: 

• Surge AI Research – Hierarchy of Agentic Capabilities: https://arxiv.org/abs/2601.09032

• Surge AI Blog: https://surgehq.ai/blog

• Nick's Sonnet 4.5 Review: https://surgehq.ai/blog/sonnet-4-5-product-take

• Nick’s Substack: https://nickheiner.substack.com/ 

• SurgeHQ’s enterprisebench: https://surgehq.ai/blog/enterprisebench-corecraft 

• Nick’s hilarious Gemini 3.1 review: https://nickheiner.substack.com/p/gemini-31-pro-not-leading-edge-also  

• Hemingway-bench AI Writing Leaderboard https://surgehq.ai/blog/hemingway-bench-ai-writing-leaderboard 

• LMArena is a cancer on AI: https://surgehq.ai/blog/lmarena-is-a-plague-on-ai 

• Dell AI Factory with NVIDIA: https://dell.com/yourwaytoai


Subscribe to The Neuron newsletter: https://theneuron.ai

More from The Neuron: AI Explained

All 106 episodes
Inside the Secret Labs Where AI Learns to WorkThe Neuron: AI Explained · 1 h 3 min
Listen in VO