In short
Y Combinator Startup Podcast: The Fastest Path To Super Intelligence
Episode Overview In this episode, Ian Fischer, co-founder and CEO of Poetiq, discusses the groundbreaking advancements in AI through recursive self-improvement systems. Poetiq, founded by former DeepMind researchers, aims to create reasoning harnesses that enhance the capabilities of existing large language models (LLMs) while sidestepping the high costs associated with traditional model fine-tuning.
Key Concepts Discussed
Introduction to Poetiq
- What is Poetiq?
- A startup focused on building recursively self-improving AI systems.
- Aims to create reasoning harnesses for LLMs that enhance performance without traditional fine-tuning costs.
Recursive Self-Improvement
- Definition:
- The process by which AI systems improve themselves autonomously.
- Importance:
- Represents a significant advancement in AI, offering a more efficient and cost-effective alternative to training new models from scratch.
Advantages Over Traditional Methods
- Fine-Tuning Trap:
- Traditional fine-tuning of models is costly and often outdated by the time a model is optimized.
- Recursive self-improvement offers continuous enhancement without the need for extensive re-training.
- Stilts for LLMs:
- Poetiq’s approach allows startups to build on top of existing models effectively, increasing their performance without the risk of becoming obsolete with each new model release.
Performance Metrics
- ARC-AGI and Humanity's Last Exam:
- Poetiq recently topped these important benchmarks, demonstrating their system's capability to outperform existing models at lower costs.
- Achieved a score of 55% on Humanity's Last Exam, surpassing the previous state-of-the-art at 53.1%.
Meta-System Functionality
- How the Meta-System Works:
- The Poetiq meta-system automates the generation of systems designed to solve complex problems.
- It optimizes prompts, reasoning strategies, and overall performance tailored to specific tasks.
Future of AI Development
- S-Curve of Improvement:
- As models and the Poetiq system improve, they create a continual upward trajectory in performance.
- AI Paradigms:
- The emergence of recursive self-improvement marks a new phase in AI development, differing significantly from reinforcement learning methodologies.
Key Takeaways
- Startups and AI:
- Poetiq offers a unique opportunity for startups to leverage advanced AI capabilities without the prohibitive costs of traditional model training.
- Automated Optimization:
- The system allows for rapid iteration and optimization over time, adapting seamlessly to new models.
- Advice for Founders:
- Founders are encouraged to experiment with AI technologies, pushing the boundaries of what is possible and continuously iterating on their ideas.
Conclusion Ian Fischer emphasizes the potential of Poetiq to revolutionize how startups interact with AI. By allowing them to build upon existing models in a cost-effective manner, Poetiq aims to democratize access to high-end AI capabilities.
Call to Action
- Interested startups can sign up for early access at [Poetiq.ai](https://poetiq.ai) and explore how their systems can be integrated to enhance performance.
This episode showcases not just the technological advancements but also a mindset for innovation in AI, encouraging entrepreneurs to embrace and explore the capabilities of AI systems actively.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIan Fisher and Poetic
0:45 to 2:00
Discussion about Ian Fisher's background and the mission of Poetic.
“which is building recursively self-improving AI reasoning harnesses for LLMs.”
Recursive Self-Improvement in AI
2:00 to 4:00
Exploring the concept of recursive self-improvement in AI systems.
“I mean, that seems like actually the, like, defining thing that a startup really, really wants.”
Challenges with Traditional AI Models
4:00 to 6:00
Ian discusses the challenges of training traditional AI models from scratch.
“And it's better than the thing that you fine-tuned.”
The Value of Poetic's Approach
6:00 to 8:00
How Poetic's methods provide advantages over traditional AI approaches.
“So they come out with soda and then you come in right above them every single time.”
Performance Benchmarks and Achievements
8:00 to 10:00
Ian shares Poetic's achievements on benchmarks and comparisons with competitors.
“I think a lot of founders that get very good results for agents, they treat the underlying model as a common layer that you can switch in between.”
Humanities Last Exam Results
10:00 to 12:00
Discussion of Poetic's results on the challenging Humanities Last Exam.
“So we could optimize just the prompts, just the reasoning strategies.”
Optimizing AI Systems with Poetic
12:00 to 14:00
How Poetic's system optimizes AI agents for better performance.
“And, you know, there's some unexpected stuff.”
Advancements in AI Performance
14:02 to 15:46
Learn about the journey from 5% to 95% performance in AI with reasoning strategies.
“And that got us a little bit of the way.”
Transitioning to AI at Google
16:18 to 18:16
Explore the speaker's journey from mobile apps to AI and robotics at Google.
“So you arrived at Google over a decade ago when they acquired your first YC startup, A Portable.”
Advice for Aspiring AI Engineers
18:19 to 19:19
Get insights on how to explore AI and innovate within the field.
“for engineers who want to get into more of the AI side, probably the applied AI and build startups around AI?”
Transcript
Automatic transcript. May contain errors.0:00Ian Fischer:The world is changing so quickly. This is probably a little bit obvious, but you should just try things and like every day do something with AI. Last summer, I took a weekend and used GPT-5 to help me build an iPhone app. I hadn't done that in a decade.
0:17Y Combinator Startup Podcast Host:So fast.
0:18Ian Fischer:Yeah, it's so fast and so easy. And that was, you know, an age ago. That was like eight months ago. Now it's even faster and easier. Don't limit yourself. Like anything that you imagine, you should just try to use AI and see how far you can get with it. and you'll be making the world better.
0:39Y Combinator Startup Podcast Host:Welcome to another episode of The Light Cone. Ian Fisher is the co-founder and co-CEO of Poetic, which is building recursively self-improving AI reasoning harnesses for LLMs. Previously, he spent a decade as a researcher at Google DeepMind and founded a mobile dev tools company through YC years ago. Welcome, Ian. Thank you. I'm so happy to be here. What is Poetic? How's it different than RL? How's it different than context engineering?
1:07Ian Fischer:At Poetic, what we're building is a recursively self-improving system. And so recursive self-improvement is this kind of the holy grail of AI, where the AI is making itself smarter. The core insight that we had is that we could do recursive self-improvement far faster and cheaper than all of the other ways that people had been proposing to do this. So obviously, I can't go into details about what that is, what our particular approach is. But most of the approaches out there involve, you know, they require you to train a new LLM from scratch. And training LLMs from scratch costs, you know, hundreds of millions of dollars and takes months of effort.
1:49Y Combinator Startup Podcast Host:And then Anthropic or OpenAI will come along and just eat your lunch in the next model release.
1:53Ian Fischer:Right, right. Right. And, you know, of course, Enthropic and OpenAI and Google, they're exploring recursive self-improvement, but typically at that level of having the, you know, having to train a new model for every step of self-improvement that they do.
2:07Y Combinator Startup Podcast Host:I mean, that seems like actually the, like, defining thing that a startup really, really wants. Like, I know that I want to take advantage of whatever the next model is, but the second you're in fine-tuning land, I'm spending, you know, millions to hundreds of millions of dollars. And then guess what? I just lit it on fire because the next version of the Frontier model comes out and I'll never catch up. Whereas working with your systems means that I will always have the thing that is better than the thing that's out of box. And that's sort of like the holy grail.
2:38Ian Fischer:Yeah, we think that this is incredibly valuable to anybody who's building on top of large language models. And we don't view the frontier models as competitors. They're the ones that were using the stilts, building stilts to stand on top of. But if we didn't have that foundational layer, then Poetic couldn't exist.
2:59Y Combinator Startup Podcast Host:Yeah. I mean, being the smartest model, it's a game of inches, actually. And so those inches matter a lot. Right, right. How do we actually get started? I mean, you've built something that basically any startup could use that it's sort of like stilts, really.
3:14Ian Fischer:We have built a system that can automatically generate systems for your particular problem that will always outperform the underlying language models. And without kind of the massive expense, as you're saying about the bitter lesson, where, you know, what would you what would you have done without Poetic? Like you probably would have said, okay, we're going to first collect a large data set, you know, like tens of thousands of examples for a particular problem that we're working on. And we're going to fine tune, you know, the best model we can put, get our hands on. Maybe that's, you know, one of the frontier models, or maybe it's an open weights model.
3:48Ian Fischer:It doesn't particularly matter. You're going to spend a lot of money on that fine tuning. The compute is so expensive. And then at the end of it, you have something that, you know, works better than the thing that you fine tuned on top of. But by then, a new model has come out. And it's better than the thing that you fine-tuned. You fine-tuned like three years ago on top of GPT-3.5 or whatever, and then GPT-4 comes out and it just blows you out of the water. And so are you going to do that again, or are you going to go out of business? And like in some cases, the latter. With Poetic, what we end up giving you is a, you know, people are calling these things harnesses now, but, you know, or a Gentic system or whatever you want to call it that sits on top of one or more language models.
4:31Ian Fischer:and it just performs better than them. And when the new model comes out, that same harness is perfectly compatible with it. And you don't need to change anything to get an even bigger performance bump. Additionally, we can continue to optimize for this new model, whatever the new model is that you want to use and make it even better. But you don't lose out on hundreds of millions of dollars. In fact, we do this so much more cheaply than fine-tuning would cost as well.
5:05Y Combinator Startup Podcast Host:And you've done this actually a bunch of times, right? Like, I remember when you first came out with your paper in December of last year, you shot to the top of Arc AGI V2. And then you've done this a bunch of times for other benchmarks, too. You know, what was that like?
5:19Ian Fischer:Arc AGI V2 was, this was kind of us coming out of stealth, letting people know that we could tackle these really hard problems. And in particular, we wanted to show that our system could generate these, what we call our system like the poetic meta system, can generate reasoning systems that are highly effective. Gemini 3 Deep Think had just come out, and they were really quite dramatically at the top of the leaderboard at 45%. And two days later, we released our results where we were showing that we could get a lot higher than that.
6:00Y Combinator Startup Podcast Host:So they come out with soda and then you come in right above them every single time. Yeah. Like wild to see, honestly. That's what it's like to have stilts. You know, like whatever model comes out, you can be taller than that one with Poetic, which is like, that's so awesome.
6:13Ian Fischer:Yeah, so the interesting thing is that we were half the cost of Gemini 3 DeepThink because we were building on top of Gemini 3 Pro, which is a much cheaper model. But we still got, in the end, a 9 % point improvement on the official verification. So they were at 45 % and like$70-something, and we were at 54 % and$32 per problem.
6:37Y Combinator Startup Podcast Host:So recently you guys just announced some incredible results for Humanities Last Exam. Can you tell us more about those?
6:44Ian Fischer:Humanities Last Exam is a set of 2 ,500 really, really hard questions written by experts in many different domains. They're meant to be challenging even for PhDs in those fields. AI hasn't passed it yet, but we got to 55%, which is almost two percentage points higher than the previous state-of-the-art, which came out just last week from Anthropic with Claude Opus 4.6. They got 53.1%, and we got 55 % on it.
7:18Y Combinator Startup Podcast Host:And one thing that Humanity Last Exempt doesn't publish is the cost of getting those results. In your case, this run was done with less than around six figure. How much was it?
7:31Ian Fischer:We didn't publish any cost for this, but I can say that the optimization costs us less than$100K, yeah. Which is impressive because each of these big foundation model train runs are in the hundreds of millions of dollars.
7:45Y Combinator Startup Podcast Host:And you guys, as a company, you're only seven people?
7:49Ian Fischer:That's right, yeah. Seven research scientists and research engineers, yeah.
7:53Y Combinator Startup Podcast Host:That's impressive. And I think the thing that's very interesting about your approach is sort of taking a very scientific approach to the emergent behaviors that a lot of the best founders are doing with models. I think a lot of founders that get very good results for agents, they treat the underlying model as a common layer that you can switch in between. And there's a certain task, for example, for GPT 5.2, like very hard to verify bugs gets sent to that versus architecture that gets sent to CLOT 4.6. But you're kind of doing this automatically instead of having a human conducting is very impressive.
8:32Y Combinator Startup Podcast Host:I think there's something more special going on underneath. Can you tell us a bit about how it works? Yeah, it sounds magical. So what can you tell us?
8:40Ian Fischer:Right. So you're getting at a really core thing. You know, these harnesses, they are code, prompts, data, you know, built on top of one or more language models. Right. And so this is something that in principle you can build by hand or with like cloud code or whatever. But in practice, it takes a lot of work to do these, to, you know, have all the insights to make this to make these work well. And so the core technology that we've developed at Poetic is recursive self-improvement. So we have a recursively self-improving system, which we call the Poetic Metasystem. The output of that system is systems that solve hard problems, where a hard problem is something that if you gave it to GPT-5-2, it would struggle to give you a reliable, robust result, just to use an example.
9:32Ian Fischer:So this is a very big advantage for us. We can generate these systems in a much more automated manner, which means that we can do it much more quickly and much more cheaply than if you hired a team yourself to try to make your own agent to solve your particular task. But not only that, since this is really an automated optimization process, if you already have done that work, you're a startup that's going after a particular vertical and you think you understand your problem pretty well, you've put together your agent, and maybe it's working pretty well, but you know you can get something better or you really need something better, then you can bring that to us and we can optimize that entire agent or pieces of that agent.
10:19Ian Fischer:So we could optimize just the prompts, just the reasoning strategies. There's a lot of different things that we can do depending on your particular needs.
10:26Y Combinator Startup Podcast Host:It sounds like this is a complete different paradigm than RL because we went through the S-curve of regular pre-training, RL when OpenAI released 01. And now this feels like a new one. It sounds special. It rhymes a lot with RNNs, which is a whole different paradigm than RL, right?
10:46Ian Fischer:It's going to depend on the particular task, the particular type of problem that we're going after, that we're trying to solve, and the underlying models that we're working with. But effectively, you could say like each model or each set of models that we're working with will have their own S-curve. The poetic system, the poetic meta system itself is also going to have its own S-curve. And so as the Poetic meta-system gets better and as the underlying models get better, you'll find that the S-curve that you're dealing with keeps shifting higher and higher until ultimately either you saturate or like…
11:19Ian Fischer:Reach AGI? Yeah, reach AGI, reach super intelligences, yeah.
11:24Y Combinator Startup Podcast Host:Given its stilts, you might like hit the ceiling first then. That's the goal, right?
11:28Ian Fischer:Yeah. You want to hit the ceiling first with Poetic.
11:31Y Combinator Startup Podcast Host:I think a lot of startups that we work with, and then in my spare time, I do a bunch of context engineering. And then the thing is, we're sort of like tuning it, tuning evals, tuning, like we're context stuffing ourselves. What does that even feel like to have a recursively self-improving version of like prompt engineering, context engineering?
11:52Ian Fischer:We don't spend a lot of time looking at the particular data that we're working with. instead we're letting the poetic meta system look at that data and and so like the meta system you know if it if it thinks that it needs to put more things into context do more context stuffing or whatever it'll do that if it needs to like generate a bunch of examples to get the get better performance it'll do that for you right it was pretty interesting to look at the prompt outputs in particular, I'd say, for RKGI in that, you know, I think you can read those and say, well, that's not what a human would have written pretty clearly.
12:34Ian Fischer:And, you know, there's some unexpected stuff. And, you know, it made some really simple examples. And one of the examples is actually wrong. But we didn't change it. We're like, well, this is, you know, this is the thing that it output, we'll just leave it be. You know, we don't want to go in and monkey around with things. And so historically in machine learning, the rule was you have to know your data set really well. But now we're kind of outsourcing that to the AI itself, where it's the AI's job to understand the data set and figure out where are the failure modes and where are the kind of robust reasoning strategies that the agent could use to get better performance.
13:19Y Combinator Startup Podcast Host:How much of it is like much, the output is much better prompts, and then how much of it is like the harness itself, context stuffing or summarizing in the right way or re-ranking in the right way so that like you have some number of like mega LLM calls. And then how do you get the most out of each of those calls?
13:37Ian Fischer:Yeah, and so that definitely varies per problem. But what we've seen, in fact, our last paper at DeepMind was not doing this recursive self-improving stuff, but we were showing that you could build these harnesses manually to solve really hard problems. And what we saw there is that we manually optimized the prompts really hard for these very hard problems. And that got us a little bit of the way. In this particular case, the hardest task we were working on, we got to 5 % performance with Gemini 1.5 Flash. This was a while ago. And then when we added on the reasoning strategies, we went from 5 % to 95%.
14:21Ian Fischer:And so this is typically what we see. you know, like everybody's out there kind of doing some amount, I wouldn't say everybody, but many people are out there kind of doing some amount of automated prompt optimization. You know, JEPA is this very popular paper. Everybody's kind of re-implementing that. That will get you some performance improvements, but it's very far from everything that you can get if you actually think about these reasoning strategies that are really going to be written in code rather than in just better prompts.
14:50Y Combinator Startup Podcast Host:So if startups want to use Poetic to put their agent on stilts, what should they do?
Read the full transcript
14:56Ian Fischer:Yeah, so right now, we haven't released anything yet. But if you go to poedic.ai, there is a button you can click to sign up for early access. And if you're a startup or a company who has a really hard problem, and you've tried everything that you can to make it reliable and robust, and you just can't get all the way there, you need something more, then let us know. We're looking for problems like that. So just tell us what it is that you're working on, and we'll reach out. You'll be the first to know when we're ready to work with you.
15:30Y Combinator Startup Podcast Host:I mean, if you're at the top of Humanities' last exam, then, I mean, that's pretty big. So you're already all the way out there at SOTA, and then I guess the stilts basically let any agentic company become SOTA.
15:44Ian Fischer:That's the idea, yeah, yeah. Yeah. And, you know, we view the ArcGGI results and the Humanities Last Exam results as showing kind of two different capabilities that we have. We can really improve your reasoning and we can really improve deep knowledge extraction from these models. And then you're just totally vaccinated against the bitter lesson.
16:03Y Combinator Startup Podcast Host:Exactly. YC's next batch is now taking applications. Got a startup in you? Apply at ycombinator.com slash apply. It's never too early, and filling out the app will level up your idea. Okay, back to the video. Slight sort of change of topic, but something I was curious about. So you arrived at Google over a decade ago when they acquired your first YC startup, A Portable. A Portable was importing mobile apps cross-platform, right? Like Android or whatever.
16:32Ian Fischer:It's quite different to recursive self-improving AGI. How did you make that leap? What happened once you got to Google? What made you think that you maybe wanted to shift out and do something different? And just would love to hear that story. The acquisition was this amazing opportunity to reflect on what I really wanted to be doing next, right? Like Google itself is a place where you can do so many different things. So I spent some time thinking about where I wanted to go next in my journey. I realized that the problems that I was most excited about were really actually AI and robotics. And the best people in the world, many of them in those fields were at Google at the time.
17:27Ian Fischer:And so I went and talked to them. They let me come join a new AI robotics team in Google Research, which was this amazing opportunity for me since that wasn't my background. My background was like computer security and then this cross-platform mobile, systems building stuff. I was able to join this team. And I'll tell you the truth that I very quickly realized that hardware is hard and I didn't really want to be doing robotics. is more aspirational at that moment. But I was really passionate about machine learning. So I just made a very hard switch into just doing machine learning research and did that for about a decade at Google and then DeepMind.
18:16Ian Fischer:What's maybe some advice that you have today for engineers who want to get into more of the AI side, probably the applied AI and build startups around AI? Like how should they think about that? You know, the world is changing so quickly. This is probably a little bit obvious, but you should just try things. And like every day, do something with AI. Always try to push yourself to find the boundaries of what they're capable of and build the things that you want to build, right? Even for me, last summer, I took a weekend and used GPT-5 to help me build an iPhone app. I hadn't done that in a decade.
19:02Ian Fischer:So fast. Yeah, it's so fast and so easy. And that was an age ago. That was like eight months ago. Now it's even faster and easier. Don't limit yourself. Anything that you imagine, you should just try to use AI and see how far you can get with it. And you'll be making the world better.
19:19Y Combinator Startup Podcast Host:That's all we have time for today. But Ian, thank you so much for giving us all stilts. We can't wait to use it at YC. I can't wait to use it for Gary's List. I mean, there's just so much to do.
19:30Ian Fischer:Yeah, thank you for having me. This was a lot of fun.
From the publisher
Poetiq is a new startup founded by former DeepMind researchers that recently achieved a major jump on the ARC-AGI and Humanity's Last Exam benchmark by layering a recursive self-improvement system on top of existing models. In this episode of Lightcone, Poetiq's Founder & CEO Ian Fischer joined us to discuss how small teams can build “reasoning harnesses” that outperform base models, what that means for startups and why automating prompt engineering may be one of the most powerful levers in AI today.Chapters:00:00 – Intro00:40 – What Is Poetiq?01:07 – Recursive Self-Improvement Explained02:07 – The Fine-Tuning Trap02:59 – “Stilts” for LLMs03:14 – Recursive Self-Improvement vs. Fine-Tuning05:05 – Taking the Top Spot on ARC-AGI06:37 – Beating Claude on Humanity’s Last Exam08:40 – How the Meta-System Works10:26 – Beyond RL: A New S-Curve11:32 – Automating Prompt Engineering13:37 – From 5% to 95% Performance14:50 – Early Access & Putting Your Agent on Stilts16:17 – From YC Founder to DeepMind Researcher18:29 – Advice for Engineers in the AI EraApply to Y Combinator: https://www.ycombinator.com/applyWork at a startup: https://www.ycombinator.com/jobs




