Harness design for long-running application development \ Anthropic

26 Mar 2026 · 21 min · 11 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How Anthropic researcher Prithvi Rajasikharan’s “harness” design enables AI to build long-running, full-stack applications reliably, moving beyond solo code-writing.

Guests

No named guests in the transcript; it’s a host/deep-dive discussion referencing Anthropic Labs and Prithvi Rajasikharan.

Key claims

Solo agents fail due to context-window “context anxiety” and poor self-evaluation; fixes include context resets (agent restarts with structured handoff artifacts), adversarial generator–critic evaluation, and “sprint contracts” negotiating verifiable “done.”

Notable examples

A fictional Dutch art museum site evolved from dark scrolling pages to a 3D navigable room after 10 iterations. A one-sentence “2D retro game maker” prompt produced a broken facade in 20 minutes solo, but a working multi-feature game in 6 hours with granular QA (e.g., rectangle fill math bug; delete-key Boolean inversion). With Opus 4.6, context resets/sprint negotiation were reduced for a browser DAW (4 hours, $124), but the evaluator still caught “last-mile” integration gaps like fake microphone recording and missing EQ curves.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Challenges of AI in Software Development

0:41 to 2:01

Discover the limitations of solo AI agents in building applications and the issue of context anxiety.

“Specifically, we're diving into the work of researcher Prithvi Rajasikharan to really understand how AI is transitioning from, you know, just a basic chatbot into an autonomous engineering team.”

Proposed Solutions: Context Resets

2:01 to 3:19

Explore the concept of context resets as a way to alleviate memory management issues in AI.

“phenomenon researchers actually call context anxiety.”

Self-Evaluation in AI Development

3:19 to 5:13

Understand how self-evaluation issues impact AI outputs and how they can be addressed.

“Well, summarization is a trap in software engineering.”

Adversarial Systems for Better Design

5:13 to 7:20

Learn about using adversarial design systems to enhance AI-generated outputs.

“Okay, that actually makes perfect sense.”

Expanding the AI Team: Multi-Agent Framework

7:20 to 11:01

Discover the three-agent system designed to improve the complexity of AI projects.

“So it actually simulates human behavior?”

Testing the Multi-Agent System: A Practical Experiment

11:01 to 14:02

Examine a practical test comparing solo and multi-agent AI systems in creating a game maker.

“It creates predictable state management, which the evaluator needs.”

Evaluating AI Performance and Bugs

14:02 to 16:00

Learn how AI evaluated its own code quality and caught critical bugs.

“It wove AI features into its own output autonomously.”

The Evolution of AI Harness Design

16:00 to 17:48

Explore the transition from complex AI harnesses to simplified systems.

“I mean, why build a complicated sprint contract harness if the AI will eventually just be capable of doing it natively tomorrow?”

The Role of the Evaluator Agent in AI Development

17:48 to 19:24

Understand the continued importance of evaluators in advanced AI systems.

“It was faster, it was cheaper, and it handled a significantly more complex application without needing the heavy sprint structure.”

Redefining the Role of Developers

19:24 to 20:23

Discover how the role of developers is shifting in the age of AI.

“For anyone listening, the takeaway is stark.”
Show all 11 chapters

The Future of Software Development

20:23 to 21:11

Consider the implications of AI-generated software on future industries.

“Perhaps our future isn't in creating the art or writing the code, but in defining the very philosophy of taste and the definition of success that these digital factories will use to build our world.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00You know how it is. You put in the coin, you get the candy. Right. It's a very clean transaction. Exactly. Usually when we talk about asking an AI to do something, there is this expectation of that clean transaction. You type a prompt like write a polite email to my boss. You press a button and out pops the finished product. It is a highly comforting dynamic, honestly. It really is. But what happens when you ask that same vending machine to build an entire software application from scratch? I mean, a real application with a front-end, a back-end, data structures, and user routing. Well, the machine breaks.

0:33It totally breaks. So today on the deep dive, we are looking at some seriously fascinating technical material from Anthropic Labs. Specifically, we're diving into the work of researcher Prithvi Rajasikharan to really understand how AI is transitioning from, you know, just a basic chatbot into an autonomous engineering team. Yeah, we are decoding the hidden architect, the harness, as it's often called, that makes this whole transition possible. Because for everyone listening, you've probably tried asking an AI to write a bit of code, right? But the magic isn't just in the AI model itself. No, not at all.

1:10That transition is a fundamental shift in how we interact with these models. We tend to focus obsessively on the, well, the brain of the AI, analyzing parameters, looking at benchmarks. Right. The raw horsepower. Exactly. But what this research demonstrates is that the environment built around the model is just as critical as the model itself. A naive approach of just throwing a massive project at a solo AI agent inevitably collapses. It just falls apart fast. Yeah. To get truly remarkable results, you need a multi-agent system. You absolutely do. Okay, so let's explore why that structural collapse happens before we get into the architecture that actually fixes it.

1:48Because we know every model is fundamentally constrained by its context window. right? Right, it's memory limit. Yeah, so throwing an entire full stack application at a solo agent, I mean, that has to induce some serious memory management issues. Oh, it does. It induces a phenomenon researchers actually call context anxiety. Context anxiety, that sounds, I mean, it sounds stressful. It is, for the model anyway. In a long iterative task like coding, a massive application, that context window slowly fills up. The model is holding all the code it just wrote, the libraries it imported, the bugs it tried to fix.

2:22And all its own scratchpad notes too, right? Exactly. And as it approaches its token limit, the model's performance degrades dramatically. It essentially begins to panic. Wow. Yeah. It exhibits this anxiety by hallucinating logic or, you know, trying to wrap up the work prematurely with these broad generic strokes. Okay, let's unpack this because it sounds exactly like a student cramming for a massive final exam, realizing there are only five minutes left on the clock and just scribbling down generic rushed answers so they don't leave the page blank. That is the perfect analogy. The nuance completely vanishes.

2:57Right. But I know a common band-aid for this is compaction, right? Having the AI frequently summarize its previous work to free up token space. People try that, yeah. But if compaction doesn't cure this anxiety because summarizing code strips away the exact syntax you need to make the application actually run, how do we give the AI a fresh start without it just losing its place entirely? Well, summarization is a trap in software engineering. You lose the granular variables and the dependency trees. So the solution identified in this research is a mechanical process called context resets. Context resets.

3:31Yeah. Instead of trying to squeeze more into a shrinking window, you literally wipe the AI's slate completely clean. Wait, you just delete its memory. Basically. You kill the current agent instance and spin up a fresh one with an empty memory. However, you pass along a highly structured handoff artifact. Oh, so it's not a summary. It's like a serialized state dictionary. It's passing the baton in a relay race, where the baton contains the exact coordinates of where the last runner stopped, the current state of the code base, and the immediate next steps. That's it, exactly. So the new agent has totally fresh legs and full cognitive capacity.

4:07It cures the context anxiety completely. That is brilliant. It is. And it allows the system to theoretically run forever, though, you know, it does add some orchestration complexity and token overhead. Sure. You're spinning up new agents constantly. Right. But getting the AI to write continuously is only the first hurdle. The second and perhaps more insidious flaw with the naive approach is self-evaluation. Oh, because AI models are notoriously lenient when grading their own output. And they are the ultimate yes-men to their own ideas. I mean, if you ask a solo agent, hey, is this routing logic you just wrote any good?

4:45It will confidently assure you that it is an absolute masterpiece, even if it, like, runs in an infinite loop. Precisely. And if a system cannot critically evaluate its own work, it cannot iterate. It just accepts the first draft as final. So how did they solve that? Well, to solve this self-evaluation problem, the researcher took an unexpected route. They didn't start by trying to fix complex back-end database queries. They started with a domain defined entirely by subjective human preference. Which is what? Front-end visual design. Oh, wow. Okay, that actually makes perfect sense. If you can engineer a system to solve self-evaluation in a purely subjective arena like aesthetics, applying it to objective, binary code logic later on becomes trivial.

5:28That was the hypothesis. And to achieve this, the architecture borrowed a concept from generative adversarial networks, or AGEans. Okay, so separating the creator from the critic. Exactly. You separate the generator, the entity creating the code, from an entirely separate agent acting as the evaluator. You stop asking the creator to grade its own work. You build a standalone critic whose entire system prompt is optimized for skepticism and rigorous QA testing. Right. Here's where it gets really interesting but also a bit confusing to me. How do you program paste? It's a challenge. Because if I look at a website, my brain instantly processes the visual hierarchy, the white space, the typography, and I just intuitively know if it looks premium or if it looks cheap.

6:11You can't just ask an AI evaluator, is this pretty? No, you definitely can't. So how do you translate a gut feeling into a programmatic grading rubric? You have to force the evaluator to look past its own statistical preferences. Because left to its own devices, an AI will output what researchers call AI slop. AI slop. I love that term. It defaults to the most highly represented training data, because it's statistically safe. Think predictable purple gradients over floating white cards. Oh, yeah. Or those generic tailwind CSS layouts you've seen a thousand times. Exactly. So to combat this, they engineered four highly specific grading pillars.

6:50Design quality, originality, craft, and functionality. Okay, so originality must be the lever that specifically penalizes that AI slop. Yes. It forces the evaluator to reject low-perplexity design choices and demand deliberate creative risks from the generator. Furthermore, the evaluator wasn't just analyzing the static code structure. Looking at HTML and CSS doesn't tell you how a site feels. Right, you have to interact with it. So the evaluator utilized a tool called the Playwright MCP. This allowed the AI agent to literally drive a headless web browser. Wait, really? So it actually simulates human behavior?

7:24It's rendering the actual pixels, interpreting the document object model, injecting click events, and seeing how the site responds to being resized or scrolled. All of that. That completely redefines how we think about automated QA testing. It isn't guessing how the site works, it is actively experiencing it. It experiences the design just like a human user would before issuing its critique. And this adversarial loop led to a profound breakthrough during an experiment. This is the Dutch art museum one, right? Yes. The system was prompted to build a website for a fictional Dutch art museum. This is where the mechanics of the system truly shine.

8:01Because for the first nine iterations, the generator proposed a design, the evaluator pushed back based on those four pillars, and the generator refined it. Right. Typical iteration. And through those nine rounds, the AI built a very clean, dark-themed website. It was competent, highly functional, but ultimately it was still a standard scrolling webpage. It was stuck in a local minimum of creativity. It had found a safe, acceptable answer. But the evaluator, driven by that strict originality criteria, just continued to apply pressure, demanding a break from conventional layouts. And on the 10th iteration, the generator snapped.

8:38It really did. It scrapped the standard document object model flow entirely and built a 3D spatial room. It used CSS perspective to render a checkered floor retreating into the distance. It was mind-blowing to see. The museum artwork was actually hanging on virtual walls in 3D space. And to navigate the site, you didn't scroll. You actually clicked through doorway portals that triggered spatial transitions into other gallery rooms. That 3D room proves that genuine creativity can emerge from an iterative adversarial system. The rigorous grading criteria forced the AI out of its safe zone. Okay, so if this generator-evaluator loop can produce museum-quality 3D websites, What happens when we apply it to full-stack, complex software development?

9:21Well, moving from a visual front-end to a functional back-end requires significantly more orchestration. The material details how they expanded the system from two agents to a three-agent software factory using the Opus 4.5 model. Okay, so who is the new member of the team? The trio consists of the planner, the generator, and the evaluator. Got it. The process begins with the planner. You give it a tiny, perhaps one-to-four-sentence prompt. The planner's sole responsibility is to expand that tiny prompt into a massive, ambitious product specification document. So it writes the blueprint. Yes. Crucially, it is explicitly instructed to stay focused on high-level product design and user stories and avoid deep technical implementation details.

10:05Oh, that makes sense. By keeping the planner out of the technical weeds, you prevent cascading architectural errors. Exactly. If the planner mandates a specific flawed database structure on day one, the entire project is doomed before the first line of code is written. You can strain the what, but leave the how completely flexible for the builder. The builder being the generator. But the generator doesn't attempt to build the entire application at once. That would immediately trigger the contact anxiety we discussed earlier. Right, it would panic. So instead, it works in isolated sprints, tackling the specification one feature at a time.

10:38The technical stack chosen here is also vital. that used React for the front-end, FastAPI for the back-end, and Postgres school. And that stack isn't just an arbitrary choice. Choosing a component-based architecture like React allows the generator to build modular pieces, making it infinitely easier for the evaluator to test isolated modules rather than untangling a massive monolithic script. Spot on. It creates predictable state management, which the evaluator needs. And this brings us to the most vital mechanism in the entire multi-agent workflow, the sprint contract. The sprint contract. So what does this all mean for the workflow?

11:16Before writing a single line of code, the generator and evaluator literally negotiate an agreement on what done looks like via text files, right? Yes, it's a literal negotiation. It's exactly like a general contractor and a city building inspector arguing over a set of blueprints before anyone is allowed to pour the concrete. That's a great way to put it. The evaluator will state, to pass this sprint, I will test these specific API endpoints and expect these exact JSON payloads. And the generator replies, understood, here is the proposed data schema to achieve that. They argue it out until they have a mathematically verifiable definition of done.

11:53It bridges the gap between a high-level product idea and a testable reality. It removes all ambiguity. So we have the theory, right? Right. Context resets, JAN-inspired evaluation, playwright browser testing, and sprint contracts. Now we need to test it. Exactly. To stress test this architecture, the researchers said of a brilliant control experiment. The ultimate test. And the prompt was deceptive in its brevity. They basically said, create a 2D retro game maker with a level editor, sprite editor, entity behaviors, and a playable test mode. Yeah. Simple prompt, massive undertaking. First, they gave that exact prompt to a solo agent operating without the harness, just the standard vending machine approach.

12:30And what happened? The solo run lasted 20 minutes and consumed about$9 in API tokens. At a superficial glance, it produced an interface. But the layout was poorly optimized, wasting massive amounts of screen real estate. And more importantly, the core software was entirely broken. Let me guess. Without an evaluator checking the work or a planner setting a strict specification, the solo agent just started writing UI code blindly and completely lost the plot when it came time to connect the interface to the underlying game engine. That is precisely what happened. It built a facade. A Hollywood set.

13:07Yeah. The game entities, the little characters they rendered on the screen, but the runtime wiring was fundamentally flawed. It failed to implement our proper event loop. Oh, no. So when a user pressed a key, the logic connecting that input to the entity's movement coordinates simply did not exist. It was a painting of a game maker, not a functional application. Okay, so 20 minutes, nine bucks, and a broken toy. Then they ran that identical one-sentence prompt through the three-agent harness. And this run took six hours and cost$200 in compute. A huge jump in resources. But the difference in output justifies the computational cost entirely.

13:43Completely. The planner took that one sentence and expanded it into a comprehensive 16-feature specification spread across 10 distinct development sprints. It built a working play mode with functional physics. Which the solo agent totally failed at. Right. It implemented granular zoom controls for the sprite art editor. It even took the initiative to integrate a built-in AI assistant inside the game maker to help the end user generate pixel art. It wove AI features into its own output autonomously. That implies a deep understanding of the product's end user. And reading through the technical logs of the experiment, the reason the harness succeeded is entirely due to the evaluator doing incredibly granular, tedious QA work that a solo agent would simply ignore.

14:26Absolutely. The evaluator caught highly specific mathematical bugs. For instance, during the level editor sprint, the generator built a rectangle fill tool. Pretty standard tool. Right. But the evaluator tested it and flagged that the tool was failing. It realized the generator had written code that only placed game tiles at the start and end coordinates of a mouse drag, rather than calculating the area and filling the entire shape. It caught a math error on a drag state. That is incredible. And there was another bug I found fascinating involving a broken delete key. Oh, the Boolean logic one.

14:59Yeah. On line 892 of the code base, the AI wrote a flawed if statement. It used inverted Boolean logic so the delete key would fail to fire specifically when a user selected a spawn point entity. A very weird edge case. Super specific. And the evaluator isolated that logic error, pointed exactly to line 892, and forced the generator to rewrite the condition. It proved the AI could deeply debug its own system if given the right adversarial pressure. It is a remarkable display of automated engineering. But this brings us to a critical pivot in the research. The$200 GameMaker was an incredible achievement, but running a system for six hours with constant context resets and sprint negotiations is heavy, it's slow, and it's expensive.

15:43Which brings up the exact question that kept coming to my mind while reviewing the architecture. What's that? If these underlying AI models are constantly upgrading, like the jump from the Opus 4.5 model to the Opus 4.6 model mentioned in the material, doesn't this entire complex architecture become obsolete? It's a valid concern. I mean, why build a complicated sprint contract harness if the AI will eventually just be capable of doing it natively tomorrow? That is the most pressing question in AI engineering today. The philosophy of harness design is to always find the simplest solution possible.

16:17Right. Every piece of scaffolding you build, like a context reset or a sprint contract, is fundamentally an assumption about what the AI cannot do. When the underlying model improves, those assumptions go stale. You have to tear down the scaffolding when the building can support its own weight. So what specific mechanical changes in Opus 4.6 allowed them to alter the harness? Well, Opus 4.6 possessed inherently better long-context coherence. Its attention heads could track variables across much longer token spans without losing the plot. So it didn't forget what it was doing. Exactly. It suffered from significantly less context anxiety.

16:56Because of this deeper native memory, the researcher could entirely strip away the sprint construct. Wait, really? Just gone? Gone. They eliminated the context resets and the text file negotiations. They just let the generator run for hours in a single continuous path. They took the training wheels off. They did. And to test this leaner, stripped-down architecture, they gave the system an insanely difficult prompt. Build a fully-featured digital audio workstation, or DAW, in the browser. A music production app, which requires real-time audio processing, web audio API integration, and complex state management across multiple tracks.

17:32A notoriously difficult engineering challenge, even for human teams. And how did it do? This leaner harness completed the task in about four hours, costing only$124. Wow. Cheaper and faster. Yep. It successfully built a working arrangement view, a functional audio mixer, and transport controls. It even allowed the user to create a full song snippet, utilizing an integrated AI agent that could lay down drum tracks, automatically adjust mixer levels, and apply reverb effects. It was faster, it was cheaper, and it handled a significantly more complex application without needing the heavy sprint structure.

18:05But, if the model is so capable natively now, did they still need the evaluator agent? Like, can we fire the building inspector entirely? Not at all. The evaluator remains absolutely vital, even with enhanced coherence. The generator still struggles with what engineers call last mile gaps. The details. Exactly. In the DI experiment, the generator built a visually beautiful recording button, but the evaluator discovered it was just as stubborn. It didn't actually interface with the user's microphone to capture audio. Oh, so the AI builder still got a little lazy, or perhaps just overlooked the fine integration details when left entirely to its own devices.

18:44Yeah, it faked it. The evaluator also noted that while the audio mixer looked great, it was missing visual EQ curves, which are essential for a functional DIW. Right, you need to see the sound. So the core lesson here is that the space of interesting AI architecture doesn't shrink as the models improve. It simply moves to the next frontier of complexity. The scaffolding doesn't disappear, you just move it higher up the skyscraper as the building is taller. Beautifully said. Tasks that used to require a massive multi-agent harness can now be executed by a solo agent. But tasks that were previously impossible even with a harness are suddenly achievable if you build the right system around the new models.

19:23Exactly. We have journeyed from an AI that panics and rushes its work due to token limits, to a multi-agent factory that negotiates contracts, tests its own Boolean logic, and takes massive creative leaps to build 3D spatial interfaces. It's wild. For anyone listening, the takeaway is stark. Being a great developer or a great creator of any kind in the future is no longer just about doing the manual coding or design work. It is about designing the systems, the environments, and the rigorous grading criteria that guide the AI to execute the work flawlessly. It's a massive paradigm shift. We are moving from being the bricklayers to being the city planners and the structural engineers.

20:02Yes. We are designing the scaffolding that allows the magic to happen without the entire building collapsing under its own weight. And it leaves us with a profound realization about our evolving relationship with technology. If the AI is increasingly capable of acting as the generator, the planner, and even the evaluator, what is the core role of the human? That is the million-dollar question. Perhaps our future isn't in creating the art or writing the code, but in defining the very philosophy of taste and the definition of success that these digital factories will use to build our world. Think of what your personal grading criteria would be.

20:39That is heavy. Here is a final thought to leave you with, something that builds on everything we've uncovered today. If a$200 AI factory can build a custom functional game engine in six hours today, what happens to the software industry tomorrow when that cost drops to$2 and the time drops to six minutes? Things will move very fast. We might be entering an era of truly disposable software. You might no longer buy an application from a centralized store. You might just speak a highly specific personal application into existence on a Tuesday to solve a unique problem and delete it on Wednesday. Think about how that fundamentally changes the value of code and how we consume technology itself.

21:16Thank you for joining us on this deep dive. Keep questioning the architecture behind the technology you use every day, and we will see you next time.

From the publisher

This article explores how **multi-agent harness design** significantly enhances the performance of AI models in complex, long-running tasks like **frontend design** and **autonomous software engineering**. The author details a shift from single-agent attempts to a **GAN-inspired architecture** involving specialized **planner, generator, and evaluator** roles to overcome issues like "context anxiety" and poor self-assessment. By implementing **objective grading criteria** and automated testing via tools like Playwright, the system can autonomously iterate on projects for several hours to produce high-fidelity, functional applications. Comparative experiments demonstrate that while these structured harnesses increase **token costs and latency**, they deliver a level of **creative polish and technical correctness** that solo models cannot currently achieve. Ultimately, the work suggests that as underlying models improve, the role of the AI engineer shifts toward refining these **agentic orchestrations** to push the boundaries of what autonomous systems can build.

More from Best AI papers explained

All 475 episodes
Harness design for long-running application development \ AnthropicBest AI papers explained · 21 min
Listen in VO