The Era of Agentic Organization: Learning to Organize with Language Models

15 Nov 2025 · 11 min · 5 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Agentic organization using “async think” to coordinate multiple LLM agents dynamically (not a single model or fixed parallel workflow).

Guests

No guests mentioned; it’s a solo “Deep Dive” episode with the hosts speaking.

Key claims

AsyncThink uses an Organizer-Worker Protocol with the same LLM backbone for both roles; organizer manages via fork/join tags and can pause/resume on join results, enabling adaptive concurrency. Training uses two stages: cold-start format fine-tuning (GPT-40-generated manager examples) and reinforcement learning to optimize fork/join timing with rewards for accuracy, format validity, and agent-pool utilization.

Notable examples

AIME24 math benchmarks (28% lower inference latency; ~1468 vs ~2048 baseline) and multi-solution countdown (89% all-correct vs ~70 sequential, ~68 fixed parallel). Zero-shot generalization to 4x4 Sudoku (89.4% accuracy) and a tetrahedron geometry task using three distinct worker perspectives.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Agentic Organization

0:45 to 2:28

Discussion on the shift from single LLMs to collaborative agentic organizations.

“So to ground this whole thing, the source material gives us a pretty great analogy from computer science.”

Roles and Functions in AsyncThink

2:28 to 4:16

Explaining the roles of organizers and workers in the AsyncThink model.

“So the organizer is the project manager.”

Dynamic Management with Fork and Join

4:16 to 6:14

Describing the fork and join mechanism that enables dynamic management in AI.

“Since the data didn't exist, they had to create it.”

Training AsyncThink for Success

6:14 to 10:12

Overview of the two-stage training process used to teach the AsyncThink model.

“It only counts the moments where the organizer absolutely had to pause because it was waiting on a join.”

Future Implications of Agentic Organization

10:12 to 10:47

Exploring the future of human and AI collaboration in agentic organizations.

“And it issues a fork, maybe with a tag like F-O-R-K human, and dispatches that specific piece of the problem to a human expert for review.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive. Today our mission is to really get our heads around a new breakthrough in AI. It's called agentic organization. Right. We're moving past this idea of a single, brilliant LLM just thinking by itself. Yeah, that's really the old way. We're diving into a world where these models can stop thinking step by step and, you know, start working together like a really efficient team. It's a huge shift. And the paradigm behind it is called asynchronous thinking or async think. It lets AI agents solve incredibly complex problems, but collaboratively at the same time. And for you, the listener, this is probably the best way to understand how these systems are going to scale their intelligence.

0:44Exactly. It's a peek into the future. So to ground this whole thing, the source material gives us a pretty great analogy from computer science. It does. It helps to think about the different parts. So you have an individual agent. And that agent is just an LLM running one thought process. Right. And that's like a single CPU core. And in the agent pool, that's all your available agents. That's your multi-core CPU. Exactly. And the most important part, the organization policy. That's like the program running on the CPU. It's the brain. It's the set of rules deciding what runs, where it runs, and when to get the most out of all those cores.

1:17And this whole new structure was built to solve some really big problems with the older methods for AI reasoning. Right. I mean, historically, we had two main ways of doing this. First was just sequential thinking. The classic step-by-step. Simple, reliable, but just agonizingly slow for big problems. And the other one was parallel thinking, which sounds good, right? You run a bunch of independent thought processes at the same time. Yeah, like three different agents all trying to find a solution, and then you just take a majority vote at the end. But that had two massive drawbacks. The first one being you're only as fast as your slowest agent.

1:52Precisely. A huge bottleneck. But the second one, and this is really the key thing to take away, that whole process was based on a fixed, manually designed workflow. It's just a recipe, a hard-coded recipe. A hard-coded recipe. It can't change its mind. It can't adapt if it hits a dead end or finds something unexpected. It just, it has no management skill. And that's exactly what AsyncThink brings to the table, that dynamic management. And it does it through something called the Organizer Worker Protocol. What's so interesting to me is that it's the same LLM backbone for both roles. Right. It's not a different model.

2:27The same LLM can be an organizer or a worker. It's all about the actions it takes. So the organizer is the project manager. It takes the user's query. Uh-huh. The initial prompt. And it maps out the whole thinking process dynamically. And then at the end, it merges everything back together. While the workers, they just do the jobs they're given. They run their subqueries concurrently, get the job done, and report back. And this whole dance is choreographed by two really simple essential actions, fork and join. I love that they borrowed those terms. It makes it so clear. So the organizer uses fork to start a concurrent process.

3:03It just writes something like F-O-R-K-I with the subquery. And that assigns that specific job to a worker that's free in the pool. And then when the organizer needs a result back to continue its own thinking, to check if They agree or combine two pieces of information. It issues a join tag like J-O-I-N-I. And that join tag is the absolute key. It's the scrunchionization point. If that worker isn't finished yet, the organizer literally pauses its own thinking. It just waits. It waits. And once the result comes in, the worker's text gets added to the organizer's context and it just resumes. That ability to pause and resume is what gives it genuine adaptivity.

3:42It's what those fixed parallel methods could never, ever do. Okay, so the model has this new language, this fork and join syntax. But the next big problem is how do you teach it when to use it? I mean, this kind of organizational trace data doesn't just exist on the internet. Exactly. So how do they even get the training data? You can't just scrape the web, for examples, of LLMs managing agent pools. That is precisely the challenge. And they solved it with a two-stage training process. Stage one is all about learning the language. They call it cold start format fine-tuning. Learning the grammar of being a manager.

4:16That's a perfect way to put it. Since the data didn't exist, they had to create it. They used a powerful model like a GPT-40 to generate thousands of examples. To simulate what a good manager would do. Right. Forcing it to follow that fork-join format perfectly. And the ablation study showed that if you skip this step, the whole thing just falls apart. Accuracy plummets. So stage one teaches it how to talk like an organizer, which means stage two must be about teaching it strategy, how to be a good organizer. And that's where reinforcement learning comes in. This is where it optimizes its policy.

4:51It learns when to fork, when to join, using a really clever reward system. Okay, so you have the obvious stuff, like the accuracy reward. Did you get the right answer? Of course. And then you have the penalty for messing up the organization. Right. The format reward. Exactly. It gets penalized if it tries to use more workers than are in the pool or joins a task that doesn't exist. You know, basic project management mistakes. But the real kicker, the thing that drives the efficiency, has to be this thinking concurrency reward. It's the most important one, I think, because it explicitly encourages the model to find ways to break a problem into parts that can be solved at the same time.

5:29So it's not just about getting it right. It's about maximizing how many workers are active at any given moment. Yes. The reward is literally tied to the utilization ratio of the agent pool. It forces the AI to think in parallel structured blocks to save time. So the model learns that good organization isn't just about correctness. It's about minimizing the time everyone has to wait around. Precisely. So does all this complexity actually pay off? We need to measure the results. But if things are running in parallel, you can't just use a stopwatch. That's a great point. And that's why they use a specific metric here called critical path latency.

6:05Can you unpack that a little for us? Sure. It's basically the theoretical minimum time the task could take. It measures the longest necessary chain of steps. So it ignores the time that workers were busy in the background on other tasks. Exactly. It only counts the moments where the organizer absolutely had to pause because it was waiting on a join. It's a really honest way to measure how much true concurrency was achieved. That makes a lot of sense. So looking at the results on tough math benchmarks like AIME24, the games are pretty wild. They are. AsyncThink got a 28 % lower inference latency compared to the standard parallel thinking methods.

6:42And it did that while being just as accurate or even more accurate. Let's put some numbers on that. On AIME24, AsyncThink's latency was around 1468. Right. Whereas a typical parallel thinking baseline was up around 2048. That's a massive speed up. And it proved something important. The individual workers were actually limited to shorter responses. Only 512 tokens. Oh, that's interesting. But these short, focused, well-managed fragments, when put together by a smart organizer, outperformed one long, rambling thought process. Good management beats long hours. And it wasn't just about speed. The reasoning itself got better.

7:17We see this on tasks where you have to find multiple correct answers, like the multi-solution countdown task. This is where that adaptive quality really pays off. On the hardest metric, all correct, where it had to find all four solutions. Let me guess, it blew the others away. It did. Async Think got 89 % accuracy, the sequential baseline was stuck at 70%, and the fixed parallel one was even worse, at around 68%. And that jump in accuracy comes from the fact that the organizer isn't following a script. It can actually, you know, divide and conquer. Precisely. The organizer can get results from two workers.

7:51See, they only found two of the four solutions. And then immediately launch two new workers. Exactly. And tell them, hey, you two try a completely different approach. It adapts on the fly until the job is done. That ability to pivot is incredible. But for me, the real proof that this is a huge deal is generalization. The source material says the organization policy itself became a learned skill it could transfer. This is maybe the most exciting result. The model was trained mostly on math than these countdown problems. But then they tested it, zero shot, on something it had never seen before, a 4x4 Sudoku puzzle.

8:24No extra training at all. None. And it just, it knew how to manage the problem. It generalized the skill of organization. So it learned how to manage resources, not just how to solve one type of problem. Exactly. It got better accuracy, 89.4%, and lower latency than all the baselines on Sudoku. But the best example was another math problem, a really tough one about a tetrahedron. It didn't just try one method and fail. Not at all. Faced with this hard geometry problem, the organizer immediately stoned three workers. And gave them different jobs. Distinct jobs. It told one, you use coordinates, the second, you use unit edge lengths, and the third, you use a different coordinate system.

9:05It assigned three different perspectives. And then it just waited for all three to report back with the join protocol. And when all three came back with the same answer, that the cosine of the angle was 13, it knew it had it. That's not just problem solving. It's managing a coordinated scientific investigation. So for you, the learner, what this all really boils down to is that async think is a structured, adaptive, and just incredibly efficient way for AIs to tackle complexity. It turns one giant problem into a well-managed team effort. You get speed and accuracy at the same time. And it's only going to get deeper.

9:41The researchers are already talking about recursive agentic organization. Where any worker can become a sub-organizer and create its own little team. A fractal of intelligence. It is. But the real frontier, and this is the final thought we want to leave you with, is human AI agentic organization. Okay. Just think about the implications when a human can get directly involved in this process. Imagine a human acting as the main organizer, forking complex analytical tasks to a team of AI workers. Or the other way around. The AI is the organizer and it hits a problem that needs, you know, human intuition or ethical judgment.

10:16Right. And it issues a fork, maybe with a tag like F-O-R-K human, and dispatches that specific piece of the problem to a human expert for review. Wow. What happens when the AI can dynamically decide when to bring a person into its own concurrent managed thinking process? That fusion of human and machine planning that completely redefines collaboration. It does. And it's a huge question for the future, one that's built directly on the ideas we've talked about today. Something for you to ponder long after this deep drive is over.

From the publisher

This paper introduces **Asynchronous Thinking (AsyncThink)**, a novel paradigm for large language model (LLM) reasoning designed to enable **agentic organization** and collaborative problem-solving. AsyncThink employs an **organizer-worker thinking protocol** where an LLM acts as an organizer that dynamically structures concurrent processes using **Fork and Join actions**, while workers execute sub-queries. The authors compare AsyncThink favorably to traditional sequential and parallel thinking approaches, demonstrating that it achieves **higher accuracy and reduced critical-path latency** across complex tasks like multi-solution countdown and mathematical reasoning. Training is accomplished through a two-stage process involving **cold-start format fine-tuning** followed by **reinforcement learning (RL)**, which optimizes the model for correctness, format compliance, and thinking concurrency. Furthermore, the results show that AsyncThink's capability for organizing thought processes **generalizes well** to previously unseen domains and problem types.

More from Best AI papers explained

All 475 episodes
The Era of Agentic Organization: Learning to Organize with Language ModelsBest AI papers explained · 11 min
Listen in VO