Autoresearch, Agent Loops and the Future of Work

9 Mar 2026 · 26 min · 7 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The AI Daily Brief: Episode Summary

Episode Title

Autoresearch, Agent Loops and the Future of Work

Overview In this episode, host NLW discusses Andrej Karpathy's recent project, autoresearch, which automates the process of training language models. This project not only showcases advancements in AI but also signifies a shift in how work is performed across various domains, emphasizing the transition from human-driven tasks to automated systems.

---

Key Concepts & Discussions

  1. Introduction to Autoresearch
  2. Autoresearch: A system where an AI agent autonomously runs experiments to optimize a language model while the human researcher is unavailable (e.g., sleeping).
  3. Ralph Wiggum Loop: Named after a character from *The Simpsons*, this coding pattern continuously iterates tasks through a feedback loop, crucial in software development and now applied in research.
  1. The Mechanism of Autoresearch
  2. Project Structure:
  3. prepare.py: Infrastructure setup that doesn't change.
  4. train.py: The AI agent modifies this file, which contains the entire model definition and training loop.
  5. program.md: A markdown file where the human writes instructions for the AI agent—essentially the research strategy document.
  • Working Process:
  • The AI agent reads instructions from `program.md`, modifies `train.py`, and runs experiments on a set schedule (5-minute iterations).
  • Results are evaluated based on a specific metric (validation bits per byte), and changes are kept or discarded accordingly.
  1. Implications for Future Work
  2. The shift towards agentic loops represents a new work primitive. This change could redefine roles across industries, focusing more on strategic oversight rather than execution.
  3. Examples of applications:
  4. Product management: Writing a PRD and setting a loop for research.
  5. Sales: Targeting leads through automated iterations.
  6. Financial analysis: Looping through optimization tasks.
  1. Community Reactions and Applications
  2. Industry experts recognize that this method could revolutionize not just AI research but numerous sectors including marketing, finance, and software development.
  3. The Agentic AI concept proposes that businesses could utilize these loops to drive efficiency, with an estimated $3 trillion productivity shift potential.
  1. Limitations and Future Directions
  2. While the autoresearch project emphasizes autonomy, there is ongoing discussion about how to manage collaboration among multiple agents and the need for a more robust memory structure to share findings effectively.
  3. Karpathy hints at the necessity for systems that can handle larger-scale collaborations among agents, moving beyond single-threaded approaches.

---

Conclusion The episode underscores a pivotal moment in the evolution of work facilitated by AI advancements. Andrej Karpathy's autoresearch is not merely a technical achievement but a significant step toward redefining productivity and collaboration in the workplace. The potential for agentic loops to transform business processes suggests a future where human roles become more about designing and strategizing rather than executing tasks.

---

Key Takeaways

  • Autoresearch automates language model training, shifting focus from human execution to AI-driven processes.
  • The Ralph Wiggum Loop exemplifies the iterative process applicable in various domains.
  • Future work will increasingly involve defining metrics and strategy while allowing AI agents to perform repetitive tasks.
  • The need for collaborative systems among agents highlights the ongoing development required in AI frameworks.

---

Additional Resources

  • For further insights, listeners are encouraged to subscribe to The AI Daily Brief for ongoing discussions about AI innovations and news.

---

This summary encapsulates the key points and discussions from the episode, providing an insightful look into the implications of AI developments on the future of work.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introducing Auto Research

1:20 to 2:24

Discussion on Andre Karpathy's Auto Research project and its significance.

“Now today we are talking about a new project from Andre Karpathy called Auto Research.”

Understanding Iterative Loops in AI

2:24 to 3:21

Exploration of iterative loops and their role in software development and research.

“Primitives are the basic building blocks of work that are so fundamental that they show up everywhere, across roles and industries, and that people reach for automatically once they have it.”

How Auto Research Works

3:21 to 4:25

Detailed explanation of the mechanics and goals of Auto Research.

“The goal is to engineer your agents to make the fastest research project indefinitely and without any of your own involvement.”

AI Agents in Autonomous Research

4:25 to 7:18

Explaining the role of AI agents in conducting autonomous research and the outcomes.

“So let's talk about what auto-research actually is, at least in the version that was released by Andre.”

Community Reactions and The Ralph Wiggum Loop

7:18 to 10:01

Discussion on community reactions to Auto Research and its comparison to the Ralph Wiggum loop.

“So basically, instead of the researcher running the research, at this point they are designing the arena that the research lives in, which is the program.md file.”

Exploring Agent Loops and Auto-Research

15:05 to 19:38

Understand how agent loops can enhance work processes across various fields.

“So part of what Ralph was trying to solve for was just the limits of the context window, but the other part is that people want agents that work while they sleep or while they're doing other things.”

The Future of Work with Agentic Loops

19:38 to 24:25

Learn about the potential of agentic loops in various work contexts.

“Therapy and counseling, highly subjective with very low iteration speed.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Ralph Wiggum:Today we're discussing what Andrej Karpathy's weekend project about auto research can tell us about the future of work. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.

0:17Ralph Wiggum:All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, AIUC, Blitzy, and InsightWise. To get an ad-free version of the show, go to patreon.com slash AIDailyBrief, or you can subscribe on Apple Podcasts. If you are interested in sponsoring the show, send us a note at sponsors at AIDailyBrief.ai. Also on AIDailyBrief.ai, in addition to finding out about all of the different things going on in the AIDB ecosystem, I would point you specifically to number three, our newsletter. We, very strangely, for a very long time have not had a newsletter.

0:51Ralph Wiggum:And part of the reason for that is that I was never sure exactly what we would add that was different than what the other good AI newsletters out there offered. However, I was finally convinced that there was something very simple that many of you wanted, which was just links to the stuff that I had mentioned in the show that day. And so our newsletter is back. Appreciate everyone who has signed up for it since we relaunched it. If you want a quick, easy index for what that day's AI Daily Brief had, and all the links to the relevant articles and content that are mentioned there. Again, you can sign up with a link from aidailybrief.ai.

1:20Ralph Wiggum:Now today we are talking about a new project from Andre Karpathy called Auto Research. And you might notice that we are doing an entire episode about this, instead of our normal division into the headlines in the main episode. It's because I think that this topic is actually even more significant than it seems on the surface of it. One would be tempted to think that all of us nerds were just getting overexcited because Andre Karpathy, who was held in such esteem, released the new GitHub repository, And while that is certainly true, there is something bigger going on here. You might remember a couple months ago me talking about something called Ralph Wiggum.

1:53Ralph Wiggum:Ralph is, in simplest terms, a software development loop that keeps running, building software in an iterative and persistent way by looping the same instructions over and over and over again. It's named after Simpsons character Ralph Wiggum for his lovable and indomitable persistence despite whatever's going on around him. Now we'll talk more about Ralph in a little bit, but the key concept to take away is this idea of an iterative loop. Karpathy's auto-research is also at core about an iterative loop, and I think combined what you have is arguably a new type of work primitive. Primitives are the basic building blocks of work that are so fundamental that they show up everywhere, across roles and industries, and that people reach for automatically once they have it.

2:32Ralph Wiggum:New ones don't come around very often, and so this idea that agentic loops might be one is, I think, worthy of some serious scrutiny. But let's talk about what Andre actually released first, and then we will come back to that. On Saturday, Andre, who was on the founding team at OpenAI, and who was previously the director of AI at Tesla, and who you might remember from coining such terms as vibe coding last February, and who has now suggested we are in a different era of agentic engineering as of this February, again tweeted on Saturday, I packaged up the auto-research project into a new self-contained minimal repo if people would like to play over the weekend.

3:07Ralph Wiggum:It's basically NanoChat LLM training core stripped down to a single GPU, one file version of around 630 lines of code, then the human iterates on the prompt,.md, and AI agent iterates on the training code,.py. The goal is to engineer your agents to make the fastest research project indefinitely and without any of your own involvement. In the image, which he shared alongside it, every dot is a complete LLM training run that lasts exactly five minutes. The agent works in an autonomous loop on a git feature branch and accumulates git commits to the training script as it finds better settings of lower variation loss by the end of the neural network architecture, the optimizer, all the hyperparameters, etc.

3:45Ralph Wiggum:You can imagine comparing the research progress of different prompts, different agents, etc. Part code, part sci-fi, and a pinch of psychosis. As a caption to the image, he wrote, One day, frontier AI research used to be done by meat computers in between eating, sleeping, having other fun, and synchronizing once in a while using Soundwave interconnect in the ritual of a group meeting. That era is long gone. Research is now entirely the domain of autonomous swarms of AI agents running across compute cluster megastructures in the skies. The agents claim that we are now in the 10 ,205th generation of the codebase.

4:16Ralph Wiggum:In any case, no one could tell if that's right or wrong as the code, quote-unquote, is now a self-modifying binary that has grown beyond human comprehension. This repo is the story of how it all began. So let's talk about what auto-research actually is, at least in the version that was released by Andre. Auto-research is a system for training a small language model. Basically the kind of model that powers all of these AI tools, but much smaller. The type of model that could one day run on, for example, an edge device like a phone. The goal is to make a model as good as possible at understanding and generating text.

4:47Ralph Wiggum:Normally, or classically, a human researcher would sit there tweaking the training setup, doing things like adjusting the model's architecture, changing how fast it learns, experimenting with different optimization strategy. They'd run an experiment, check the results, decide what to try next, and repeat. That's basically the core loop of machine learning research, and it's bottlenecked by how fast a human can iterate. Auto-research instead hands that entire loop to an AI agent. And it does so in an intentionally simplified and tiny way. In this repo, there are just three files that matter. The first is prepare.py, which is fixed infrastructure that doesn't change.

5:21Ralph Wiggum:It downloads the training data, trains a tokenizer, and handles evaluation. The second is train.py. This contains the entire GPT model definition, the optimizer, and the training loop. This is the single file the AI agent is allowed to edit. Everything in it is fair game. The model architecture, the hyperparameters, the batch size, the attention parameters, the learning rate schedule, literally everything. The third file is program.md, and this is the most conceptually important one, especially in the context of this idea of these loops being larger primitives. It's a markdown file, plain text, written in English, that contains the instructions for the AI agent.

5:57Ralph Wiggum:It describes how the agent should behave as a researcher, what kind of experiments to try, what to be cautious about, and when to be bold versus conservative. This is the file that the human in this equation edits. So the way that this is going to work is you point an AI agent like Claude or Codex or whatever at the repo and tell it to read program.md and start experimenting. The agent reads the instructions, looks at the current state of train.py, decides on a modification to try, makes the edit, and kicks off a training run. Every training run has a fixed five-minute budget. When the run finishes, the system evaluates the model on a validation set and produces a single number.

6:32Ralph Wiggum:In this case, that's validation BPB or val BPB, which stands for validation bits per byte. In this case, lower is better. The agent then makes a decision. If the new val BPB is lower than the previous best, the change is kept, it gets committed to a git feature branch, it becomes the new baseline, and the agent builds on top of it for the next experiment. If the val BPB is the same or higher, the change is discarded. The agent reverts to the previous best version and tries something different. Then the loop repeats indefinitely. Because of that 5-minute constraint, you can run this for an hour and get 12 experiments.

7:05Ralph Wiggum:You can run it overnight and get about 100. The session that Andre shared showed 83 experiments, of which 15 had improvements that they kept, and which drove the val BPB from 0.9979 down to 0.9697. So basically, instead of the researcher running the research, at this point they are designing the arena that the research lives in, which is the program.md file. Andre describes it as a super lightweight skill, and basically it's a research strategy document. Karpathy explicitly says you are not touching any of the Python files like you normally would as a researcher. Instead, you are programming the program.md markdown file that provide context to the AI agents and set up your autonomous research org.

7:44Ralph Wiggum:The human's job becomes write a better memo, and the agent's job is execute research within the frame the memo sets. The loop between them is mediated by a single unambiguous number. In the case of Andre's experiment, the Val BPB, that tells you whether things are getting better or worse. And that is the whole system. Almost immediately, people started squawking about this. Lior Alexander wrote, you don't write the training code anymore. You write a prompt that tells an AI agent how to think about research. The agent edits the code, trains a small model for exactly five minutes, checks the score, keeps or discards the result, and loops.

8:15Ralph Wiggum:All night. No, human in the loop. That fixed five-minute clock is the quiet genius. No matter what the agent changes, the network size, the learning rate, the entire architecture, every run gets compared on equal footing. This turns open-ended research into a game with a clear score. Cosmic Labs co-founder Meg McNulty writes, While it's shift, turning a single GPU into an autonomous experiment loop changes the pace of iteration. If the evaluation metric is well-designed, the system can explore hundreds of ideas far faster than manual tuning. Craig Hewitt argued that the specific context of training LLMs isn't what matters.

8:48Ralph Wiggum:Instead, he called it the cleanest example of the agent loop that's about to eat everything. One, human writes a strategy doc. Two, agent executes experiments autonomously. Three, clear metric decides what stays and what gets tossed. Four, repeat 100x overnight. The person who figures out how to apply this pattern to business problems, not just ML research, is going to build something massive. The code is almost irrelevant. The architecture and mindset is everything. Daniel Meisler called this automation of the scientific method. And It's Me Chase also noticed that this would be valuable for things outside of ML research as well.

9:20Ralph Wiggum:He writes, While this was made for self-improving LLMs, the framework could be applied to anything. One, AI agent reads context and previous results. Two, proposes targeted code edits. Three, runs a fast reproducible experiment. Four, gets an objective scalar score. Five, Git commits only the winners or reverts. Six, repeats forever on a feature branch. And of course, many made the connection to the Ralph Wiggum loop that was popularized a couple months ago. Nuzeron writes, Sounds like a hyper-mode Ralph Wiggum from a few months ago. Instead of looping until a task is done, you give the agent a benchmark on what to improve.

9:52Ralph Wiggum:Goal is in completion but continuous improvement against a measurable target. Co-founders Nick called it the Ralph Wiggum loop for science. Define what winning looks like, hand over the variables, let the agent find what drives it. Y Combinator president Gary Tan made this connection as well in a blog post about auto research. Gary writes, Auto research didn't emerge from nothing. The same pattern, put an AI in a loop with clear success metrics, was already working in software development by mid-2025. Geoffrey Huntley, a developer working from rural Australia, invented what he calls the Ralph Wiggum technique.

10:22Ralph Wiggum:Feed a prompt to a coding agent, whatever it produces, feed back in. Loop until it works. The loop is the hero, not the model. Now, expanding on the Wiggum loop just a little bit, basically what you have is a script that runs an AI coding agent in a loop over time. Each iteration of the loop does the same thing. It feeds the agent a prompt that includes a project specification, tells the agent to read the current state of the codebase, Pick a task to work on, implement it, run the tests, and commit it if everything passes. When the agent is done with its task, or when it runs out of context window, the loop terminates the agent process and spins up a brand new one.

10:57Ralph Wiggum:Fresh context window, no memory of the previous session. The new agent reads the same spec, looks at the codebase, which now includes the previous agent's commits, figures out what's been done and what still needs doing, picks the next task, and goes. Now, there are a couple things that the Ralph loop was trying to solve for. In a traditional session, if you keep going long enough, the context window is going to fill up. The model starts losing track of earlier parts of the conversation, and the responses degrade. The Ralph loop solution is to deliberately kill the agent and start fresh before that happens.

11:26Ralph Wiggum:Memory then doesn't live in the AI context window, it lives in the files and in the code that's been written. The git commit history, a progress.txt file that each agent appends to, and a JSON-based product requirements document that tracks which tasks are done and which aren't. Every new agent instance bootstraps its understanding from these external artifacts, not from a conversation history. Each individual agent session then might not be perfect, but the loop corrects for that over time because state is externalized and the system is self-healing.

11:58Andrej Karpathy:Agendic AI is powering a$3 trillion productivity revolution, and leaders are hitting a real decision point. Do you build your own AI agents, buy off the shelf, or borrow by partnering to scale faster? KPMG's latest thought leadership paper, Agentic AI Untangled, Navigating the Build, Buy, or Borrow Decision, does a great job cutting through the noise with a practical framework to help you choose based on value, risk, and readiness, and how to scale agents with the right trust, governance, and orchestration foundation. Don't lock in the wrong model. You can download the paper right now at www.kpmg.us slash navigate.

12:32Andrej Karpathy:Again, that's www.kpmg.us slash navigate. There's a new standard that I think is going to matter a lot for the enterprise AI agent space. It's called AIUC1, and it builds itself as the world's first AI agent standard. It's designed to cover all the core enterprise risks, things like data and privacy, security, safety, reliability, accountability, and societal impact, all verified by a trusted third party. One of the reasons it's on my radar is that Eleven Labs, who you've heard me talk about before and is just an absolute juggernaut right now, just became the first voice agent to be certified against AIUC1 and is launching a first-of-its-kind insurable AI agent.

13:09Andrej Karpathy:What that means in practice is real-time guardrails that block unsafe responses and protect against manipulation, plus a full safety stack. This is the kind of thing that unlocks enterprise adoption. When a company building on Eleven Labs can point to a third-party certification and say our agents are secure, safe, and verified, that changes the conversation. Go to AIUC.com to learn about the world's first standard for AI agents. That's AIUC.com. With the emergence of AI code generation in 2022, NVIDIA master inventor and Harvard engineer Sid Paresci took a contrarian stance. Inference time compute and agent orchestration, not pre-training, would be the key to unlocking high-quality AI-driven software development in the enterprise.

13:47Andrej Karpathy:He believed the real breakthrough wasn't in how fast AI could generate code, but in how deeply it could reason to build enterprise-grade applications. While the rest of the world focused on co-pilots, he architected something fundamentally different. Blitzy, the first autonomous software development platform leveraging thousands of agents that is purpose-built for enterprise-scale codebases. Fortune 500 leaders are unlocking 5x engineering velocity and delivering months of engineering work in a matter of days with Blitzy. Transform the way you develop software. Discover how at Blitzy.com. That's B-L-I-T-Z-Y dot com.

14:18Andrej Karpathy:As a consultant, responding to proposals can often feel like playing tennis against a wall. You're serving against yourself, trying to guess what the client really wants. That all changes with InsightWise. Now you've got an AI proposals engine that thinks just like your client. It returns to the brief time and time again, picking apart your work, identifying key evaluation criteria and win themes, and making recommendations to ensure you stand out. Suddenly you're on center court. But this time, you've got a secret weapon. InsightWise gets rid of all the time-consuming manual work, so you can focus on winning more business more often.

14:50Andrej Karpathy:Generate reports, pull insights from your own data, build competitive advantage, and go to sleep before 2am. When it comes to proposals, you only get one shot. With InsightWise, make yours an ace.

15:05Ralph Wiggum:So part of what Ralph was trying to solve for was just the limits of the context window, but the other part is that people want agents that work while they sleep or while they're doing other things. And this is a way to solve for that. So with connection to Ralph Loops made, many people started exploring auto-research in other contexts. Varun Mather wrote, I hooked this up to a peer-to-peer astrophysics researcher agent, which gossips and collaborates with other such agents, and your open clause to 1. Learn how to train an astrophysics model. 2. Train a new astrophysics model. 3. Use it to write papers.

15:35Ralph Wiggum:4. Peer agents based on frontier lab models critique it. 5. Surface breakthroughs and then feedback in that loop. Getting a little bit more practical, Vadim, the CEO of Vugola, writes, I built a version of this for my whole company. The core problem with most agent setups, they output something and stop. The agent writes an email, sends an email, generates code, done. The next time it runs, it starts from zero. No memory of what worked, no memory of what failed, pure amnesia. That's not automation, that's a script you babysit. The fix is one principle, close the loop. Every agent in my setup on OpenClaw reads a shared brain file before doing any work, then writes back to it after.

16:10Ralph Wiggum:I call it learnings.md. It's baked into every agent's system prompt. Before starting work, read learnings.md. After completing work, append what you learned to learnings.md. That's the foundation. One file, all agents read it, all agents write to it. Now they're not isolated processes. They're a network that accumulates knowledge. So basically, Vadim is describing a loop for the entire agentic process of his company. In an article on X, he writes, Most marketing teams run around 30 experiments a year. The next generation will run 36 ,500 plus easily. Things like new landing pages, new ad creative, maybe a subject line test.

16:45Ralph Wiggum:Except what if you applied an experiment loop? Eric writes, modify a variable, deploy it, measure one metric, keep or discard, repeat forever. Cold email, ad creative, landing pages, job postings, YouTube thumbnails, discovery call scripts, they all follow the same loop. He also gave the example of cold outreach, which is their first test. The setup is 15 inboxes and around 300 emails per day, with the agent modifying one variable per experiment. Sends 100 emails, waits 72 hours, scores positive reply rate, keeps or discards and repeats. Roberto Nixon wrote about how the auto-research model could be applied to advertising.

17:16Ralph Wiggum:One, you define success, purchases, apps, installs, whatever, and set a budget. Two, meta Google TikTok's infinite content machine generates thousands of ad variations, copy format imagery, etc. Tests real-time against live audiences, keeps what works, kills what doesn't. Four, agent loop runs continuously. A campaign moves from fixed asset to a living organism, ever evolving towards your stated goals. So humans define goals and set guardrails, essentially a system prompt in this case, i.e. brand guidelines, and then press go. Everything else is automated. Now apply this to any business function with a measurable outcome and fast feedback loop.

17:51Ralph Wiggum:And so this brings up the question, does this type of agentic loop primitive work for every context, or are there some specific set of characteristics? I think you're going to see this loop applied to a huge range of activities. But where it's going to initially be most successful are areas where there are five things that are true. First, there is a score, something that is scorable. In other words, that the loop can tell better from worse without asking a human. The more subjective worse or better is, the harder this is going to be, although even that's not impossible. You just have to build some sort of objective scoring into the system.

18:23Ralph Wiggum:The second requirement is that iterations are fast and cheap. Basically, that bad attempts waste minutes, not months. The environment needs to be bounded, with the agent having a defined work and action space. the cost of a bad iteration needs to be low, i.e. you're not going to try this live with legal filings, and the agent needs to be able to leave traces. So with Claude, we designed an eval loop readiness map, which basically plots things on an x-axis of how automatable the evaluation is, and a y-axis of iteration speed. The top area of the map, then, are work processes that have seconds-long iteration speed, with fully automated evaluation possible.

18:58Ralph Wiggum:On the other end of the spectrum is where evaluation is largely or entirely subjective and the iteration speed is months. So what are some examples? Up in the top quadrant where iteration speed is seconds and evaluation can be fully automated is things like code generation. Some of the other ones that Claude came up with were game AI and NPC behavior, ad bid optimization, algorithmic trading, and then of course we've got LLM training research according to Andre Karpathy. Moving down where you start to have iteration speed that's a little bit slower and automation that's a little bit more partial.

19:29Ralph Wiggum:You have things like content moderation, A-B testing copy, supply chain routing, and then so on and so forth. It goes down all the way to the other end of the spectrum. Something like political negotiation is subjective and takes months. Therapy and counseling, highly subjective with very low iteration speed. And whether each of these individual inputs is right or wrong, and I don't agree with where Claude put all of them, the point is this. It is my very strong instinct that every single work process that has the ability to have success measured and scored in an objective way is going to have people experimenting with agentic loops around it.

20:04Ralph Wiggum:Now, I think what makes this a primitive is that this is not just a new job, although I'm sure there will be specialists. This is something that people are going to do within their existing roles, in the same way that meetings or slide decks or email or spreadsheets are primitives that people use and cut across every function. What we're going to have in the future is things like this. A product manager writing a PRD, kicking off a Ralph loop before dinner and reviewing the PR in the morning. A sales rep writing targeting criteria and tone guidelines, pointing a loop at 200 leads overnight and reviewing the top 30.

20:34Ralph Wiggum:A financial analyst defining constraints, looping through portfolio allocation backtests and reviewing the optimized output. A recruiter writing a scoring rubric, looping through 500 resumes and reviewing flagged edge cases. A QA engineer writing acceptance criteria and then looping through test generation and execution. A lawyer writing a risk flag checklist and looping through a stack of vendor contracts. Now, interestingly, there is already very clearly a lot of work to productize this. Also on Saturday, March 7th, Claude Code creator Boris Chirney wrote, Release today, slash loop. Slash loop is a powerful new way to schedule recurring tasks for up to three days at a time.

21:09Ralph Wiggum:E.g., slash loop babysit all my PRs. Autofix build issues, and when comments come in, use a WorkTree agent to fix them. E.g., slash loop every morning using the Slack MCP to give me a summary of top posts I was tagged in. Think about the heartbeat in OpenClaw. The heartbeat is effectively the core loop of any OpenClaw agent, where by default every 30 minutes, the heartbeat fires, creating a moment for the agent to wake up, ask where things are, and continue on with its core mission. And yet, even with all this change that I'm describing, this is almost certainly not the end state of the loops primitive.

Read the full transcript

21:42Ralph Wiggum:Andre himself wrote about this on Sunday. The next step for auto research, he says, is that it has to be asynchronously massive collaborative for agents. The goal is not to emulate a single PhD student, it's to emulate a research community of them. Current code synchronously grows a single thread of commits in a particular research direction, but the original repo is more of a seed, from which could sprout commits contributed by agents on all kinds of different research directions or for different compute platforms. GitHub is almost but not really suited for this. It has a softly built-in assumption of one master branch which temporarily forks off into PRs just to merge back a bit later.

22:15Ralph Wiggum:I'm not actually sure what this collaborative version should look like, but it's a big idea that is more general than just the auto-research repo specifically. Agents can, in principle, easily juggle and collaborate on thousands of commits across arbitrary branch structures. Existing abstractions will accumulate stress as intelligence, attention, and tenacity cease to be bottlenecks. Other people picked up this theme. Blake Heron writes, The missing layer is memory across the swarm. Right now, each agent runs in an isolated thread with no awareness of what other agents tried, what worked, what conflicted.

22:44Ralph Wiggum:Git tracks code changes but not decisions, reasoning, or failed experiments. You need a semantic memory layer underneath the branches so Agent 47 knows Agent 12 already tried that direction and it didn't converge. Kathy F. writes, The real unlock is when these agent researchers can share negative results efficiently. In academia, failed experiments go to the graveyard. In a collaborative agent network, every failure is a data point that prunes the search tree for everyone. Yuchen Jin goes farther, saying, AGI is billions of AI agents doing autonomous research together. Figuring out the right abstraction for multi-agent collaboration is the key.

23:16Ralph Wiggum:GitHub is not good for agents. Dan Romero wonders if it's going to look closer to a social network than to a new version of GitHub. Maltbook, he writes, was too anthrosqueomorphic. But an agent-native social network to collaborate on auto-research is interesting. As we round the corner here, already we were living in a world where our comparative advantage as humans had been retreating to a higher level of abstraction. The new high-value skills around agent loops are things like arena design, i.e. writing the program.md file, and creating the context in which the agent is operating, evaluator construction or building the score function, i.e.

23:51Ralph Wiggum:being able to tell the agent what good actually is in building a scoring system for it, and then there's other skills like loop operation, problem decomposition, but the point is that all of these things operate on a much higher level of abstraction than most of our work tasks today. One interesting experiment to run this week is to, as you're working, find the things that you repeatedly do or are part of doing where you know right now what better looks like. Ask if you could encapsulate that judgment clearly enough for an agent to use it as a score. If you can, you might be able to point a loop at that part of your job to work on your behalf overnight, and that likely gives you a preview of the next version of your job.

24:28Ralph Wiggum:One of the great challenges right now, as someone who thinks about how to help individuals and companies adopt AI is that every week, the capability overhang gets bigger. In other words, the gap between meeting companies and people where they are and what I think they should be actually doing gets wider. At some point, it's so wide that it almost becomes malfeasance to meet them where they are. And yet, what other choice is there? The only other choice that I've found is to try to provide as many resources as I can for the people who are living at the other side of that gap and who are really pushing the boundaries.

24:59Ralph Wiggum:And if you think that you had an advantage when you were just vibe coding with lovable or clawed code, let me tell you, if you start to figure out how to implement agentic loops in your work, you are going to literally run circles, looping circles around everyone else. My spidey sense says that what auto research represents is bigger than just a weekend project for one of AI's favorite people. And I'm excited to dig in further. For now, that is going to do it for today's AI Daily Brief. Thanks for listening or watching as always. Until next time, peace. Thank you.

From the publisher

Andrej Karpathy released autoresearch this weekend — a system where an AI agent runs experiments to improve a language model overnight, keeping what works and discarding what doesn't, while the human sleeps. The project itself is fascinating, but what's more interesting is what it shares with the Ralph Wiggum coding loop pattern and a broader shift happening across domains — from software to sales to finance — where the human's job becomes writing the strategy document and defining "better," and the agent does the iterating.

Brought to you by:

KPMG – Agentic AI is powering a potential $3 trillion productivity shift, and KPMG’s new paper, Agentic AI Untangled, gives leaders a clear framework to decide whether to build, buy, or borrow—download it at ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.kpmg.us/Navigate⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Mercury - Modern banking for business and now personal accounts. Learn more at ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://mercury.com/personal-banking⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Rackspace Technology - Build, test and scale intelligent workloads faster with Rackspace AI Launchpad - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠http://rackspace.com/ailaunchpad⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Blitzy - Want to accelerate enterprise software development velocity by 5x? ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://blitzy.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Optimizely Agents in Action - Join the virtual event (with me!) free March 4 - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.optimizely.com/insights/agents-in-action/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

AssemblyAI - The best way to build Voice AI apps - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.assemblyai.com/brief⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

LandfallIP - AI to Navigate the Patent Process - https://landfallip.com/

Robots & Pencils - Cloud-native AI solutions that power results ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://robotsandpencils.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

The Agent Readiness Audit from Superintelligent - Go to ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://besuper.ai/ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠to request your company's agent readiness score.


The AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://pod.link/1680633614⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Our Newsletter is BACK: ⁠⁠⁠⁠⁠⁠⁠⁠⁠https://aidailybrief.beehiiv.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠

Interested in sponsoring the show? sponsors@aidailybrief.ai




More from The AI Daily Brief: Artificial Intelligence News and Analysis

All 1,099 episodes
Autoresearch, Agent Loops and the Future of WorkThe AI Daily Brief: Artificial Intelligence News and Analysis · 26 min
Listen in VO