In short
The TWIML AI Podcast Episode #713: Why Agents Are Stupid & What We Can Do About It with Dan Jeffries
Overview In this episode of The TWIML AI Podcast, host Sam Charrington engages with Dan Jeffries, founder and CEO of Kentauros AI, discussing the current challenges in developing advanced AI agents. They explore the definition of agents, their use cases, and strategies to create smarter systems. Dan shares insights on LLM reasoning, the importance of open-source in AI, and future directions in the agentic era.
Key Concepts Discussed Definition of Agents
- Dan's Definition: An agent is capable of performing open-ended, long-running work in the real world, functioning largely error-free over extended periods.
- Distinction from simpler tools like chatbots or web scrapers that cannot operate independently.
Current Challenges
- Developers face significant hurdles in creating agents that can act independently and make intelligent decisions over time.
- Many existing tools are not robust enough for complex tasks requiring sustained performance.
Dan's Approach to Solving Agent Challenges
- "Big Brain, Little Brain, Tool Brain" Approach:
- Big Brain: Strategic, long-term planning.
- Little Brain: Tactical execution of tasks.
- Tool Brain: Utilization of appropriate tools and resources for task execution.
- Model Selection: Emphasizes the need for both general-purpose and task-specific models for effective agent performance.
Use Cases for Agents
- Agents could potentially handle a wide range of tasks, from data processing to complex problem-solving in various domains such as marketing, health management, and IT operations.
- Example use case: Automating repetitive tasks like data entry and analysis across platforms.
Insights on AI Development Trade-offs in Model Selection
- Dan discusses the trade-offs between general-purpose models and specialized models tailored for specific tasks.
- He emphasizes the importance of leveraging open-source tools to innovate and advance AI capabilities.
Evolution of AI Capabilities
- Dan believes the landscape of AI is rapidly evolving but cautions that many existing models still face significant limitations.
- There is a need for innovative solutions to make agents more capable of performing complex tasks reliably.
Future Directions
- Importance of Open Source: Dan advocates for open-source contributions as vital for the advancement of AI, believing they drive innovation and create transparency in AI development.
- The future of AI agents is seen as highly promising, with the potential for significant economic impact and changes in how we interact with technology.
Conclusion
- Dan promotes the idea that the agentic era is just beginning and holds vast possibilities for technology to integrate more deeply into everyday human tasks.
- The conversation ends with an optimistic view of the future, underscoring the need for continued research, development, and exploration in creating smarter agents.
Key Takeaways
- Agents vs. Tools: Understanding the distinction is crucial for developers aiming to create independent AI systems.
- Research and Development: Continuous experimentation and adaptation are necessary to evolve AI capabilities.
- Open Source as a Catalyst: Contributing to and utilizing open-source projects can accelerate advancements in AI technology.
Show Notes
For complete show notes, visit
[TWIML AI Podcast Episode #713](https://twimlai.com/go/713)
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00You know, if you think HCI is around the corner trying to build an agent, you know, sort of quickly be disabused of that notion. You know, try to build a real agent. I don't mean a talk to your PDF agent or scrape this website with Playwright agent. Those are really useful tools. And there's a ton of very useful tools, but they are not systems that are capable of acting independently over long periods of time and thinking for themselves and making intelligent decisions over that period of time.
0:36All right, everyone. Welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Dan Jeffries. Dan is Chief Executive Officer at Contaros AI. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Dan, it's been a bit. Welcome to the pod. Thanks so much for having me on, man. Really good to see you. It's great to see you. I'm looking forward to digging into our conversation. We'll be talking about what you're up to and in particular, a little bit about your perspective on agents, a topic that I've been exploring quite a bit here on the pod recently.
1:11But I'd love to have you share a little bit of your background and what you're up to. My background is I'm a lot older than I look. And so I've been, I started, I started working the dot-com era out of college. And, you know, I've been fortunate enough to be a little bit ahead of the curve, fortunate and cursed to be ahead of the curve so like i i was i started working in the internet when people were like that's a really dumb idea that's what would you do that for you know and then uh you know i went to red hat and i had a recruiter you know tell me you know actually before that i had a recruiter tell me why would you want to work in linux all the jobs are in solaris i said you know solaris won't be here in 10 years and he looked at me like i had two heads right so we had we had a bunch of time there 10 years at red hat that i was a packeter which is an awesome mls company did the ai infrastructure alliance i was at stability ai and now we're at so now i'm at cantarles so that's kind of a the compressed version of how i got to now awesome awesome and cantarles is focused on agentic stuff yeah so basically building real world agents and the kind of infrastructure behind it.
2:22We end up building a lot of tools to make them work better. So I have this joke that, you know, we wanted to build a website, but we found out we had to build PHP. An operating system. Yeah, PHP, Linux, MySQL, Apache, and then WordPress, right? In order to build a working agent, right? So yeah, we've been mostly focused on, a lot of our research has been around doing navigation agents, which we're about two years ahead of time Again, where people are like, why would you want an agent to navigate GUIs? That's stupid. And then, of course, you know, Claude Computer Use comes out and everyone's like, that's a genius idea.
3:01It's so obvious. Why didn't everyone think of this, right? So we're like, okay, great. You know, that's for us. And is the idea to be a research company or a product company, but, you know, you need to do research in order to be a product company at the cutting edge? Like, how do you think about that? Yeah, I mean, we have to do research. We think of it as applied AI research, right? So we're not trading foundation models, but product company, right? More than anything, we're getting close to using a tool called Agent Tutor. And that's really something where you're just able to show the model how to do a bunch of stuff and it records all the data, right?
3:37So you're like, I want to click through Airbnb and find all this stuff. You give it a bunch of demos to kind of, you know, to know where to click and know what that workflow looks like. And then we have these kind of micro fine tunes that can kick off from there on the model or dump it into its memory as a synthetic memory so that it can, you know, learn how to do these types of things. But, you know, we had to, we didn't, we weren't planning on building that tool. We just had to build that tool in order to get. And so we're like, okay, well, let's productize that. And then, you know, there's a lot of other kind of products kind of coming down the pipe.
4:11But, you know, now we're up against some big heavy hitters. In fact, when we started thinking of the GUI navigation stuff, we knew Alibaba was working on it. We knew a lot of Chinese universities were working on the multimodal front. We knew Apple was working on it. But we were pretty sure that they were going to keep the Promethean fire to themselves. And then Anthropic released it. And now OpenAI is going to release it. and Google wants to release stuff. And I was like, wait a second. Actually, Antropic's the most conservative organization of the group. But I was like, man, they wouldn't take this risk.
4:44And they were the first ones to come to the table with like computer use, right? Although I was very happy to see that, you know, that$10 billion startup suffered from exactly the same problems that we had suffered with, you know, with much less debate. And so we're like, yes, all right, we're fighting it out, we're wrestling away, we're punching above our weight class but we have to be careful with how we compete with some of these things because we certainly are going to spin up our research cluster with a million H100s to compete against that. We need different solutions. Any particular insights on why now beyond just the confluence of increasing LLM capability and interest?
5:32I mean, like the computer stuff, computer use idea kind of harkens back to OpenAI's gem and all that stuff from a really long time ago. Like, why even weren't they first to it? I don't know why they weren't first to it other than they weren't generally thinking about AI as a product. They were generally, you know, Google was generally thinking about AI as something you did behind the scenes and then you put a product in front of it. Right. So it was search or it was something like that. And so OpenAI, I think, really changed a lot of that. um and and i think that was different they it's funny though a lot of the research harkens back to that q star paper right so when so when when when it leaked that the strawberry stuff was called q star and again everybody was like it's agi it's you know and i'm like that's that's a 10 year old paper on playing video games right like that they probably modernized for for math and you know reasoning and that's exactly what they did like i basically was you know well everyone's like it's agi next week i was on twitter going like i think it's going to make it better at math and science because it's a deterministic policy that it learns right and so like that makes it good at hard reasoning but not fuzzy reasoning or abstraction or whatever and that's exactly what we got right um you know those things are you know are sort of super cool but, you know, I don't know.
7:00Yeah. So you mentioned, you mentioned reasoning, reasoning abilities might be a contributor to that, you know, talk a little bit about the way you think about LLM reasoning. Yeah, look, I think, so it's a number of things. I think the way perplexity started was, you know, they just started looking at, people wanted to ask it questions. They wanted to do sort of research for them and be an intermediary to web search. And I probably used that 80 % of the time and so i think the more people chatted with these things the more they decided oh what else can we do with this and there was a ton of kind of agentic stuff early that really didn't work quite frankly in fact none of it works right it's just sort of like you know we're just are you saying that in a past tense like we're beyond that no we're not right i mean we're getting there but we're not beyond that we i mean you know we're definitely at the beginning of the reasoning era and then moving into the agentic era if you look at that open ai is kind of five stage thing um but we're really just at the beginning of that but they there was a bunch of things like oh you know auto gen and all these other things where they're like we're just gonna give it a massive plan and it's gonna go work independently for three hours or five days and i'm like no it's gonna hallucinate after 13 seconds and and get nowhere and um you know so that's where the problems when we started to realize what they were right and i think people were excited by the possibilities, but then they started to have to get back into the applied AI side of the house.
8:22They had to start to figure out, how do I make these things strategically planned over the long term? How do I make them tactically aware of what to do? And how do I give them good tooling? And a lot of the early stuff, again, was this build PHP, Linux, MySQL, Apache level stuff. They're like the RAC databases, the vector databases took off. And so initially, every developer's approach was like, I'm just going to hurl a ton of data into it and it's going to auto magically work like my mind you know like you and i are talking and you're doing associations at the subconscious level of things that you thought of before maybe questions you want to ask or something you gotta you gotta pick up dinner later whatever it is you're using you're using this memory yeah the search algorithm is actually the important part of that memory right and so there's that's where the apply day work comes in you can't just dump it in there and it's like, boom, it magically knows everything in Wikipedia.
9:15Not really. Right. And so I think now that the hard work sort of begins of how do we make these systems useful? Right. And, and I think, how do we make them useful? The way we define agents are. I was just going to ask that. Yeah. It's pretty, it's pretty, so, you know, it's pretty simple. We call it anything that is capable of doing open-ended, long-running work in the real world. So that could be the digital world, it could be robotic, in the physical world. And what we largely mean is something that can function for many minutes, many hours, or many days, or many weeks, largely error-free or without unrecoverable errors, because humans make mistakes all the time, but you don't want any fatal errors in there.
10:02A cat doesn't blue screen and not know what to do next. It always has something to try. Even if it's locked in a cage, it has some routine that it can get to. It doesn't just, you know, it knows what to do, right? And is the idea that something's not an agent if it is not robust to those kinds of errors by your definition? Okay, so... Or can it be just a bad agent? No. So I think it's sort of loosely defined for a lot of folks. And so the joke that I say in my talk is if you think we're HCI's around the corner trying to build an agent, you'll sort of quickly be disabused of that notion. And let me jump in and say that part of my inspiration for talking to you for the benefit of the audience is that part of my inspiration for reaching out after a very long time, way too long, is that I caught your recent talk, Why Agents Are Stupid and What We Can Do About It.
10:59And we will link to that in the show notes. It's a great talk, but we'll cover the highlights here. Cool. I appreciate it. And, but yeah, so, you know, in there I said, you know, try to build a real agent if you want to be disabused of this notion. And what I said specifically there is I don't mean to talk to your PDF agent or scrape this website with Playwright agent. Now, I don't mean to throw shade at those. Those are really useful tools. And there's a ton of very useful tools, but they are not systems that are capable of acting independently over long periods of time and thinking for themselves and making intelligent decisions over that period of time.
11:35right and that's so that's really where i think all this stuff is going i think that's what people want you want to be able to say look go make me an excel sheet and like you know make sure it has all these parameters and does this thing and has the correct math and and that's just simple and then on the corporate side you want to be able to say look i want to go take all these leads out of salesforce and put them or take them out of email put them in a salesforce and make sure they're correctly linked and there's valuable notes and the documents in there, you know, that's something that we do without thinking about.
12:08And we don't have anything that's even close to being able to do that currently right now. Defining it around robustness is, is kind of interesting in that, you know, I had this experience where, okay, you know, I'm talking a lot about agents, let me go build an agent. And, you know, if I start not from, hey, I'm going to start with agent and like work my way back to the problem, but start at the problem and, you know, try to get to the point where I need agenticness, I can get pretty far with, you know, just LLMs in a loop. Right. And so, you know, it prompts me to ask, okay, do we really need agents?
12:45Like, what does that really, really mean? And, you know, what is the, you know, the core use case, you know, for the, where is the additional complexity justified? Cause you just talked about like, you know, take some information from your emails and stick it in an Excel spreadsheet. Like I can get 99 % of the way there with like Zapier and like an open AI, like step in my workflow. Right. So, you know, there's, you know, part of it is like the thing like lives off, you know, So it's like ready and waiting to do work for you. But that could be event-driven and not necessarily an agent. It could be a loop.
13:29It could be a webhook. Yeah, how do you think about that? I mean, go find me a pair of shoes in the style that I love for a party next week. It's something where it's generally going to fail. You're not going to get there with Zapier. And that's just a simple kind of use case, right? Right. So what I always say is there's like, you know, there's like a little unobtainium. You can like put this stoicastic element in there for certain steps that you weren't able to do in the past. OK, so you might be able to say, like, I wrote a poem. Right. And it's you know, it's a haiku. And I need something to check that it's, you know, the correct number of stanzas and syllables.
14:07And but not just that, that's very deterministic. I also need it to, you know, to check whether it's a poll about springtime, right, and, you know, these kinds of things. That's the kind of thing that you could put that kind of reasoning element into it, right, that you're not going to get out of kind of the old determinism stuff. What we've sort of found, and initially we were thinking we're going to build a marketplace for these things because they're going to explode. There's going to be millions of them, and, like, everyone's going to want one. And the problem is all the teams, we started running into problems and all the teams that we talked to were running into problems building a long-term reliable agent to do anything, right?
14:44And so that all of them were basically doing the work and then selling it to the enterprise because it was such a heavy lift. And what they were having to do is build this huge, huge pipeline, right, where they're putting a ton of deterministic stuff and then letting it kind of like do a little bit of stochastic stuff right here, right? And then get back and like close it down, right? But this kind of open-ended reasoning, I mean, the applications are endless. I mean, what are the applications? It's like when you think about robots – I think it's Andrew Yang or somebody who's like, what's the killer app for AI?
15:15Like what's the killer app for the brain? Yeah, that's right. Yeah, like I mean if you talk – I mean Dario talked about that in the Machines of Living Grace over the 10 years, right? And I talked about it in my recent article on the middle path for AI because I felt we needed something in between dystopia and utopian that looks more realistic at it. But there's a million things. And, you know, if I've got an assistant that understands my health, right, or I can recommend, you know, something that's, you know, predict whether I'm going to go face down on my oatmeal if I, you know, if I don't do certain things or, you know, can pick out clothes, you know, for a bash that's coming up, right, or a present, right, those types of things.
15:55But a million types of things too, like checking any robot within a factory doing a bunch of work and knowing how to take this thing from over there and rivet this bolt, right? Those types of things are incredibly important. Or any open-ended task that you would do on your computer, right? And you're just like, look, I want you to go research, you know, Sam Charrington and tell me like everything about him. Find every article that he's ever written about stuff. Watch all his videos. Tell me what he likes and doesn't like. Like, give me a, give me his background. And like, you know, we've got a, I've got a podcast coming up.
16:29Give me, give me a summary of that thing. Right. That company already exists. Like we can do that with today's technology. We can kind of do it with today's technology, but not from the, you know, you would, you would have to know in advance kind of what those bits are, right? Like you'd have to sort of know where those, where all of your stuff is, but like the human element is somebody who would go out and find all your LinkedIn and your Twitter posts and like read through them and make deterministic, you know, make decisions about which of these things is relevant versus you just talking about the dinner you had, right?
17:00These are the kinds of things where it starts to be really interesting. I think part of what's interesting about this, and as someone who's like really excited about agents, I'm also, you know, just using you to poke at my own arguments a little bit. I think I find sometimes that if you think very broadly, you know and unconstrained by anything that you can you know do today and you pose the problem as i want uh essentially agi right i want something to to be able to take an open-ended request and go figure out how to do it and do it then um that you know clearly requires something different than what we have today but i think the the closer you get to a very specific use case like if you're talking to an enterprise and they say we've got this set of problems we've like decomposed our business into these core business processes and we kind of know what those are we've been applying varying degrees of automation you know over many years like um those feel like more tractable, more, uh, what am I, what's the word I'm looking for?
18:12Tractable. What the heck? Brain fart. Those feel like more tractable problems. And, um, and you know, the devil's argument is like, let's do the author LLM in the loop or put our LLM in the part of that process where we have to bump up to a human to like exercise a little bit of judgment based on a combination of rules and intuition. So there's two things in there. My friend Willem built a company called Cleric and there's agents that go and look at Kubernetes. So basically when there's errors, you can talk to it in Slack and it'll go look at all the different dashboards and it'll go look at the error logs and it'll look at its knowledge base and it'll try to figure out what the actual problems are and give you kind of an answer.
19:01That's a whole set of IT tasks. right and you know that's very sort of challenging to do but it is more tractable right there are challenges to it right so one of the reasons he and i were starting to talk for next year because there's certain world gardens they can't scale right so now you need a gui agent to basically go look at a dashboard and decide what's there versus just dealing with the api which may not expose what you want right so or then there's even more kind of open-ended problems which are a little more challenging so if you think about you know we have someone who's in marketing and we say look I want you to go come up with, you know, an outreach campaign, right?
19:36And there's a whole bunch of stuff that needs to happen around that, right? There's software that has to be discovered, has to be set up. There's emails that have to be written. There's like, do you even want to write an email or do you want to reach out on LinkedIn? Do you want to touch base with people on Twitter? How would you do that in the first place? How would you understand that this is, you know, kind of the correct vote? And so that starts to get even more slightly out of reach, you know, from where we are with things. Not according to Satya and Sundar. I mean, look, you know, never – what's the old joke about – never, you know, if a man's dinner is based on saying such a thing, you know, right?
20:15Like, that's why – you know, Ilya now basically has, you know,$10 billion to just do research with no need for a product. He can say whatever he wants. He's like, yeah, we're back in the era of experimentation, right? But if you're running a multi-billion dollar company, right, then it's like – then it's AGI next week, right? So it's kind of like the president of Chesterfield Cigarettes, cancer? What are you talking about, you know, right? So I think it's sort of the same thing. I think – yeah, so these sort of on-rails agents or agents that have like a more scoped set of problems are definitely more tractable.
20:54And I think they're doable now because it's they're sort of like write a bunch of Python and then like insert the Oracle to ask a question. So like I had a simple one that we wrote in house where I wanted to find a bunch of partners and and going and scraping a bunch of partners from the web is really, you know, is is old school Python. Right. Old school programming. But I also wanted to basically read the website. Tell me a bit about the company. Tell me whether it fit like 10 different rules that I've written. in like, you know, I don't want it to be doing, you know, these kinds of things or be involved in these industries.
21:26You know, I don't want it to be a simple wrapper around Chachy Ghee or whatever it is, right? You know, very complex rules. There would have been no way for me to just write that as code, right? It's just, and working with that, the push pull of building that little internal tool to get it down to the false positives. Actually, we do, we have an internal equation. So sometimes our firm does private equity diligence advice, right? So we talk to some larger firms and, you know, they're looking at buying a company. We don't do the whole Bain and everything we work with those folks, but they're, we'll come in and look at the team, the AI team, figure out whether they should be augmented and whether what they can kind of do over a next year, five, you know, a five-year time horizon, right?
22:13A lot of folks are looking editors like what can they do right now and one of the things we have is an equation that takes into account kind of the number of errors that that cascade down from it from a system that you're building um the severity of those errors like and what your tolerance to those things are and you know kind of the cost of throwing more people at the problem and a lot of times what you find is you probably shouldn't automate it where you're you're sort of at yet right that You might find that we can automate 60 % of this work. It's a dollar per chart to do this. But throwing people at it, 25 cents overseas, and you add a little bit of old school analytics, AI, and it's 18 cents.
22:54So that's very hard to beat. And then when you look at the – it's not just – with agents, the errors cascade down the pipeline. It's like a classic example when our GUI clicking agents are – you know, you, and this is a simple one, just at the tactical level. So it's trying to click on Google flights or whatever. And, you know, it makes the mistake of misclicking and it clicks the return date. So the departure. So it says, well, that makes sense. I'm just going to go ahead and put in the departure date. Right. And then it goes and it goes, I'm going to go put in, you know, now I'm going to go put in the departure date and it like resets.
23:29Cause no human would ever do that. Right. And it's like they do the departure and then the return, right? Like every time. And so that's kind of a simple mistake that can cascade or it's like, oh, now I'm trying to fix that. And I hit the departure date to fix that. And now that's the, you know, another I put in the wrong date there. And now I've got to backtrack. Right. Or I clicked off the screen and I have to start over again. The other thing is the severity of the errors. Right. And you look at something. It's like if you're doing open heart surgery. Right. The severity of the errors is tremendously high.
23:59You look at driving a telephone. Right. People can die, get injured. but if you're if you built an agent to go out and get a bunch of companies and partners you know if you have some false positive false negative it's not the end of the world because maybe you would have never been able to scale up to go look at a hundred thousand companies yourself read every website it would take you two months you get it done in a day with this thing and so if there's false positives false negatives on either side it's still incredibly valuable and so that's kind of the trade-off every company has to go through is figuring out severity eras the the number of cascading eras and like whether it's cheaper to just hire more people and i think that's where we are with everything you talked a little bit about the um the the experience building kind of the gui agent and google flights um yeah should maybe have you kind of bet since that's a core part of what you're focused on right now back up and talk about that challenge broadly um but the question that jumped out at me was you ran into this issue you know with the um the departure date you know etc there's some root cause there uh so you abstract a little bit to the root cause and then do you like what what's the the general approach to tackling this problem creating a rule creating some you know what would have to be like a very infinitely growing list of rules or exceptions or are you you know training you know refining fine-tuning models like how are you addressing all of the holes that um your system can run into out on the out in the real world out in the wild well a lot of ways is the short answer but like you know i would say you know in the talk i kind of break it down into um colloquially we break it down into big brain little brain tool brain so big brain being kind of your strategic long-term planning like even if If you're like, I'm thirsty, you know, and I want a bottle of water, there's a whole cascade of steps to go down, you know, to the grocery store to get one.
25:58You got to find the water. Did they move it? You know, you got to get in line, wait, make small talk. Do you want a bag or not? Take out your wallet or your phone, etc. All these kind of steps. Right. So that's the whole strategy of, you know what, maybe I'm just going to drink water from the tap. Right. Or one of these two things, right? But, you know, and then there's this tactical brain is kind of like the little brain is sort of like the cluster of sub steps, right? So, you know, if I got the water, I need to take it. I need to take off like the top. I need to put it to, you know, that's there's a whole cluster of steps or I'm waiting in line.
26:32I need to wait for X amount of time until I talk to the store owner. Then I, you know, she's going to ask me whether I want to bag or make some small talk. You know, there's a whole bunch of steps, right? and then tool brain is kind of like the the appendages or the the tools that you give it so it doesn't have access to an api motor control that kind of fine motor control right right and it's like you know it doesn't matter if i know how to pick up right this thing if my hand is broken right i'm not freaking it up and so you know so we've we've had a bunch of funny things like we had found for instance we built a like number of proprietary heuristics in the early days when we started working with this two years ago, we found that Claude and GPT and the open source models couldn't actually give us precise clicking points.
27:16They couldn't return coordinates properly. But we did find they could kind of approximately return points. So we found, oh, let's layer a grid of red dots over it with yellow numbers and say, look, I'm trying to click on the Twitter compose button. Which number is it approximately close to? Oh, it's number seven. And we could mathematically We zoom in and we could do that three times till we got to a very close spot. And then we would have a precise spot to click. The thing is, is there's a lot of like round trips to the car. That sounds slow. Yeah, it was slow. It was slow, right? It was super slow.
27:47But, you know, when you, you know, now we're at the point where we have good click models and like you can fall back to that in the event that it fails, which is cool. But we had found like a little problem was a tool brain problem. So on the calendar, sometimes when you had bunched up, you had these kind of giant red and yellow numbers that were overlay numbers right so so you're like so they can't actually you can't actually the model couldn't actually see the number it was trying to click under it so it didn't know which one and so that was tool brain problem right it's like it knew what to do it knew how to click the number 23 but it couldn't get the number 23 back because that was it so sometimes we have to address it at that level i would say the biggest thing we're doing right now is something that we call skill packs and it's probably another thing where we're ahead of the time, but now we're in a race against time at this.
28:30It's sort of what we call micro-fine teams or synthetic memories. So one thing you can do is you can take the workflow data and you can put that into its database. We have sort of multiple databases for the agent memory now, both SQL and Vector. And actually, we just want them to merge, and there's a bunch of projects moving in that direction, thankfully. but um you know you can put the kind of workflow steps into that and then load that as a synthetic synthetic memory when it's doing stuff and that kind of helps it but the micro fine tunes where you basically agent tutors designed to do that where you can go in and show it a bunch of different demos of how to do these things and the idea is that we can kind of rapidly fine tune that 10 minutes 20 minutes half hour for a particular task and then the idea is that you're just going to have all these adapters and it's going to be like neo you know loading up kung fu right you're like oh i'm doing stuff on amazon you know i know how to do that oh i'm doing like i'm clicking stuff on airbnb boom i know how to do that right and so so you're envisioning site-specific adapters or skills right and and even even um we've even been thinking we've been working on kind of a rust-based inference engine that could change like load an adapter in in in time as it's doing its, you know, token output, right.
29:50To say like, I need to do this thing. Right. It almost seems like you might, and I don't know if it's like a capable, I guess it would be a capability of your underlying foundation model, but, um, at some point, you know, e-commerce site is good enough and it doesn't need to be Amazon versus, uh, J crew or whatever. Right. They're all kind of similar. Like how far, you know, based on what you see in these interactions, like how much more intelligence or LLM capability or whatever we want to call that, how much more, you know, magical Oracle thing do we need in order to be able to get to that next level of abstraction?
Read the full transcript
30:35I mean, you can ask Anthropik and OpenAge. So they just put out, you know, Claude put out called computer use. And the first time I used it, I was like, oh, man, it's so fast. It's great. And it just ripped through one of my test prompts. And then another test prompt where I asked it to go to Google Flights, first of all, I couldn't get to Google Flights. So it decided finally to just go. It didn't realize that the screen was too small. And so the European zone pop-up button was hidden below the screen. And so it kept trying to recycle that. And it was like, OK, I can't get there. I'm going to kayak.
31:08GPR will be the undoing all agents. It's the end of mankind. It's the end of agents. It's like the EU AI Act. I just saw the EU as like doing this giant supercomputer so that everybody who trains something on there can be immediately classified as a systemic threat so it can't be deployed in the EU. Anyway, the... So...
31:34I lost my train of thought on this one, but... I think anthropic. Yeah, so the anthropic thing. So then I put it through another prompt, and it's going through that. Then it had a pop-up for a Google login. They covered part of the numbers, and they didn't know what to do there. So it just completely failed, whereas my agent kind of ripped through it. And I realized, again, they're throwing a billion dollars in computing, trying to get it to learn this, and they're going to throw a ton of more at it. And so is Google, and so is OpenAI. Great, because these equal better capabilities for us, which is wonderful.
32:12But our sense, my realistic sense is that we don't actually get to 97 % accuracy across e-commerce site, right? That what you want to get to is 97%. Like when something that's kind of the old school reliability test for software, right? If it's at 97 or above, then you're willing to deal with that 3 % of errors. And like this is, you know, perplexity feels like it's 97%. I don't, it's not actually like sometimes it doesn't, but it doesn't just error out on me. Like it returns an answer, right? And, you know, maybe it hallucinates once in a while, but in general, it feels like it's closer, but most stuff is not, is not there.
32:50I am not sure how much compute we have to pour into it. And the other problem is a lot of the scaling stuff is moving towards test time compute, right, or like thinking longer. The problem is, of course, with an agent and you're clicking around on stuff, you don't want it thinking a long time. You might want it in the planning stage, but you want it to be like motor memory, like you want it to be able to do stuff like in the way that you do stuff, right? So, you know, our feeling is that actually for a large chunk of time into the indefinite future, you're going to need like task-specific adapters or whatever in order to be able to do a lot of tasks.
33:28And I think, you know, I think my friend David Ha was talking about this, right, with Sakana, where they started like pushing towards, you know, evolving smaller LLM clusters, you know, to do a bunch of stuff. And I, you know, it's, of course, preferable to have the super generalized model that does everything. And that is just auto magically awesome at like clicking on GUIs and running around in the factory. But, you know, it's the internet is such an open ended, sparse reward system. right that it's it's incredibly challenging i think we're probably going to be in the point where you need to you need to train it give it specific insights and and for for a while in fact i'm betting the company on it so i hope i might you know but but but but but i you know and and you want them to get really good as a baseline right um but i think there's also going to come a pushback where people are going to want some of these things to be open source too Right.
34:28Because, you know, you know, a closed source proprietary model is a window into your life when you're like having it use your computer and do tasks in a way that we've never had in the past. Right. And I think you're going to inevitably see laws passed and spying and that kind of stuff and just hacking, straight up hacking. And so I think a lot of people are really going to want not just individuals, but corporations are going to want kind of open source models that they can fine tune on the test. And by the way, at a corporate level, if you're using some sort of internal application that you developed that doesn't have any kind of history, you're going to have to show it how to use that application.
35:12It's like we spent$5 million building this internal application that's like this legacy app that's been with us forever, but it does what we need to do. We're not replacing it. It's not just going to automagically do that. And there's a delta where the big labs are telling us that the models are just going to get so good that you're just going to plug them in. They're going to automagically do everything. But the reality is that in business, there's always friction. There's always friction. You put it in there. There's legacy systems. There's permissions. It's what it's allowed to do. There's culture.
35:45There's data that's sensitive and not sensitive. There's a million things that I think that's where a lot of the rubber meets the road kind of in the coming years. And I don't see any magical solution to get around it at this point. Now that, uh, that either queues up or runs ahead of my next question, which was, um, you know, when I'm talking to folks about agents, I often comment that I think a lot of agent, a lot of the things that, uh, we want agents for our political problems and not technical problems. And an example that I gave is like, for years, you couldn't get a Southwest flight to come up on Google flights.
36:23That's not because Google couldn't get the data if they wanted it. It's because Southwest would sue them and has sued people. And, you know, there are lots of examples like that. If you wanted to build an agent to give you, you know, the lowest cost for a medical procedure in the US, good luck with that. There's no transparency there. Um, and so, you know, at the risk of asking you a question that's directly tied to your dinner thesis, like, do you see, you know, how do you see when I think of this, it's, you know, there's one road that's like the momentum behind AI and agents and all this stuff like breaks down walled gardens, you know, and then there's this other more, um, you know, somber theory that's like walls are really resilient and have proven themselves to be resilient over time.
37:16Like, how do you think about, how do you think about that in such a way that you can spend so much of your time assuming that, you know, you can break down these walls, I guess? We can break down these walls. And I would say it's, it's an eternal cycle of something there is that doesn't level wall that wants it down, right? That was Robert Frost, right? And, And then the wall gets built up again. If you look at the early internet, you could search and find anything that you wanted. And then all of a sudden, more and more rules went up, more and more things that you could or couldn't do. More and more sites started to get taken down, right?
37:51These types of things start to happen. And then we're starting to see certain things in the internet as well, right? Like people used to have all open free APIs and all of a sudden, the APIs are charging. And like, maybe you can't do what you want in the API because they decided you can. You can't tag someone on Twitter or on LinkedIn, or you can't use the LinkedIn one without begging them for the possibility, right? And they're not going to give it to you because they want to protect their business. But you could hire someone for X number dollars an hour to go on LinkedIn and click to their heart's content and get that information for you.
38:24So I think it's coming to this pendulum that swings back and forth where exactly what you said, where we want to take it back. Like, I don't, you know, once you go, and I don't know what economics replaces this. I know it'll figure it out. But once you, you know, decide what you want to make for dinner, and you can just ask the system, and it's like, here's how you make sausage and peppers. I'm never going to that ad-laden and shittification level, like, website, right? Where there's literally, like, 10 pop-up ads, and it's an ad every other paragraph. and you know there's a long intro of like when my mother used to make this and i remember back in the day and i'm like i just want the thing and there wasn't there was again that goes back to the old school part of the internet too right where like they thought so much about how do you make ads tasteful where do you put them so they're not bothering people and for a long time that's the experience of the net where you would go out there and you do those types like nowadays it's just like how many can i cram in here as much as possible i don't care what it looks like i don't care about the experience the content is relevant people want to they want to take that power back right and so that's i think where a lot of the our friends at open interpreter you know working on kind of a desktop control and and what they want when they when we ask perplexity i ask perplexity eight percent of the time because mostly it just gives me the answer that i want then i can go look at the individual sites right to see what i need but it's like if i have to search myself i'm like oh my god i feel like i'm going back in time it sounds like your experience with perplexity is much better than mine well i mean do you got pro version or do you got like you just i mean i was using pro until i stopped using pro because and it wasn't the generation it was the search like they're they're they would well it was both they would come up with bad sources and then like tell stories about what those sources that were saying that if you actually clicked through and looked like, yeah, that was not in that article.
40:24It was just plausibly in that article. I think it's certainly gotten better over the last three to four months, like being able to set the model and then they kind of train their own model to be good. So, you know, from my standpoint, it's been, it's maybe have to go back to it and check it out. Yeah. I give it another shot. I talk about that as a risk of the hype cycle is that you try something and it's like, yeah, this isn't really working for me. And then you like, you know, put it in a box and don't remember to come back and check it. And it evolves quickly as a lot does in this space. It evolves fast.
41:00Right. And so like UIZR is another example. So it's a UI building system. And the first time I looked at UIZR, I was like, what? UIZR? UIZR, like UI lizard. Not a great name. Great product. Right. Sorry guys. I like your product, but it's really, challenging because this happens every time i tell someone about it anyway um so but really you guys are doing great work so but um when i first looked at it i was like wow this is what is this it's terrible um but then i got in there the other time and i was able to like you know just prompt first of all it has perfect ui design tools or gives you perfect react components they look great it's almost like using figma and then you can pop up a prompt and say like okay put in the screenshot and they'll just recreate that thing.
41:45So like in the middle of my designer working on something, I was like, okay, I want to redesign all of these screens. And like, he was asleep, right? You know, it was like, I was, cause I was traveling and I'm like, I just imported his screenshot and it like remade the screen. And I'm like, just rearranging everything. And I made like 10 different versions of it. And I was able to like, give it back to him, you know and these are actual like usable components, right? And that's super cool. high fidelity yeah like actual use them actually now as opposed to wireframes yeah yeah exactly like a full interface yeah so like i use it now a lot of like a lot of time just for prototyping out a cool interface right and you know and if i you know i can show it like another interface as inspiration and it'll like be able to kind of clone that thing up which is super cool and then i can just start pulling the stuff out and arranging pieces of the way that i like so i think that kind of remix aspect of ai and where we're getting to it's look this is the same thing we experienced in in every one of these kind of diffusion of innovation curves or in this hype cycle where it crashes and comes back out and you come up against the friction of dust and like and pebbles and everything else right it's like there's all these wonderful capabilities that we have to drag out right and yet you have to you have to learn how to build the tools for them the tools still feel the still the tools still feel primitive i mean even when i'm like you know i'm using ms swift which is alibaba's you know training framework right and like you know it's it's got a ton of stuff on the back end and you know has multiple inference engines vlm and lm deploy and because you need six of them there's no unified one anyway and there's like unsloth on the back and PyTorch and a million other things that they wrapped.
43:35And so you have this command line list that's like 400 parameters. And so then it's like this black magic of trying to figure out which of these parameters have made my training run crash, right, or sticky ball, right? And then it's like training run bingo, right? You're like, get the slot machine, wait 10 minutes and see whether it works. I thought all that would have been sort of finished. Is that part of what you're doing proprietary and and the i'm asking because i'm really curious about all those pieces and what you're doing there and if it's not proprietary in the sense that you would share it and folks are interested in the sense that they reach out to me and say so like we can come back on and talk about that yeah i mean so so in the ms the alibaba training framework stuff yeah i mean that's me getting involved in in kind of training side of the house to understand what my engineers are doing with this stuff because it's like um so for instance on the um on the gui agents we've been fine-tuning momo which came out of the amazing allen institute and they're doing the very hard work of open sourcing the data set the training scripts and everything which which the industry's gotten so hostile to this right both from a from like a proprietary corporate sort of standpoint of like keep your secrets and also from the crazy crazy kind of doomer narrative of like it's good we're and so like people are just afraid to like open source anything or like get attacked from many angles so kudos to them for doing that work but mo mo for instance is about the 7b is about uh 83 accurate on clicking on guis and the the 70b 72b is about uh 99 accurate now uh there's still room for improvement across the board and so yeah we've been doing a bunch of fine tuning with Agent Tutor.
45:24So you can basically show it a task and then we have the sort of a Rust space back in that can kind of just kick off a fine tune. And so part of it is like, we were trying to figure out, we have our kind of training system that we build in Rust, but then it proxies a lot of other stuff, right? And so one of those things is MS Swift, some of it is VLLM and some of the other open source frameworks for inference and training because you don't want to reinvent the wheel on that, right? So, yeah, the idea is to kick off these kind of rapid fine tunes and then test the model again. So basically, you know whether the model can do this task now and got better at what you want it to do.
46:06And right now, we're trying to find the balance of how much data it needs to be good at both clicking and sort of workflow. It's outside of the kind of reasoning stuff currently. we've got some additional stuff for thinking there, but we're trying to attack tool brain and little brain, kind of tactical and the clicking stuff first, because the reasoning we're hoping that the frontier models basically act like a GPU upgrade or a Linux kernel upgrade. Like you just kind of swap it out. Right, yeah, the big brain part. That said, we are thinking a lot about, and we've done a bunch of experiments around test time compute and that type of stuff that we're talking about too.
46:51I think actually the open source community desperately needs to put together an awesome reasoning corpus a la O1, right? And because I actually think that's one of the places where the open source community can get ahead of the proprietary lab by building a large enough synthetic data set or a reasoning data set, reasoning corpus. I think it's the type of thing where as a community, the open source could really I could really iterate and kind of get ahead and build benchmarks around it that start to show some of the stuff that's being hidden away, like the way that it's thinking about stuff. And you're beginning to see that with questions and things like that, but it's still going to be early for that part.
47:32So we're not tackling that yet. You said something that suggested to me like imitation learning. At least that's how we might have approached it before. Or like, you know, take a video of a user interacting with a site and then like train some model based on that video. Are you doing that kind of thing? Yeah. So there's sort of imitation learning, but there's also like we've tried a bunch of different techniques, actually. One of the funds I found that was really fun is I built kind of a prototype interface in-house that was kind of like Loom. That was sort of a precursor to AgentUner where you could just show it how to do something.
48:10And then I found if you use the Gemini long context window, one to two million, you can just upload the video and give it a JSON format for what the memory structure should look like. Not bounding boxes, because our models are better at doing that than those models were. But give it back an actual timestamp set of actions and a description of those actions in the format of the memory. Oh, wow. Which is super cool. And then essentially from there, you can insert that as a synthetic memory into the database of the system. And then what we would have the agents do is basically call up a series of successes and failures.
48:54So you would also be able to annotate it. So you'd be able to annotate to say like, hey, this was a success, and you should pay close attention to that. and then the system in memories after it had finished the task when when the agent would try to do something we could label that as a success or failure and give it a bunch of feedback that would be stored with the memory so you had kind of these synthetic memories which were like the human kind of showing you what to do and like so like and then what we're getting to now is sort of two different things so with agent tuner it's sort of an extension of the idea where you're like you're able to like show it what to do and it keeps it captures the screenshots the click data and the kind of workflow, right?
49:36And so it puts that in as a synthetic memory or as a fine-tune job. And then the next phase to get to is sort of the video, right? Where you can basically take it and kind of drag it and say like, from this here to here is this action and I'm annotating that sort of sub-action or the kind of overall thinking and pay close attention to this, you know, that kind of thing. So yeah, there's a lot of, there's a lot of, quite frankly, it's still the bleeding edge of all this stuff so we're we're figuring out as we go and like half the time we're testing the hypothesis is we're building the thing right and that's that's sort of the fun of it um and uh and in kind of the the part where i think we are in the age of discovery again and it's you just have to be willing to work with a model that came out last week and try crazy stuff and think outside the box every every single day that's really where i'm living right now.
50:29Do you have a methodology or framework and really nothing that quite that strong, but like, I'm curious how you think about, you know, there's 50 models that you can try for any given thing. Do you, you know, try three to try to get a sense for if it's feasible at all, you know, and then try the 50 or not try the 50, get close enough. Like, how do you think about, um, you know, that the next level of detail of you know dealing with the the crazy times that we have you know where there's so many options they're all evolving quickly there's like you know leaderboard leapfrog that whole thing there's not so many options this is that's a fool's game right it feels like there are um but it's just a lot of noise like let's be honest there's maybe a hundred useful models in the world um but across proprietary and closed um when a useful new model comes out it is it is painfully obvious, I think, to anyone who is actually good in the machine learning community.
51:32And like our, our team and the other teams that we respect are living and breathing that stuff. And they, they read the stuff that comes out and they're like, that's cool. Right. That's useful. That's valuable. It sounds like you're saying in practice, you are not finding a lot of like local maxima or something like that, where a specific model, you know, is really good at, you know thing x but you know maybe not necessarily noteworthy on other things and you have to find you know for each model the thing that it's really good at and plug it into your system you're more taking the approach of you know finding the best general purpose thing you know maybe scope to some degree or another but like and going with those things based no based on the task i mean so as soon as the mobile model came out we were like this is awesome this is you know this is you know because we i i immediately tried it out and i i uploaded the screenshots that i had and i was like okay find me this button find me this button and it was it had already in their demo had kind of returned a red dot where the coordinates were and i was like this already looks really promising this is our next fine tune and we had we had worked earlier with pally gemma because pally gemma was sort of built for fine tuning and had already been trained on bounding boxes we built a data set by combining data sets cleaning them and adding our own sort of data to it called the Wave UI dataset.
52:53And we had managed to fine-tune Pally Gemma to be Forex better than it was at returning bounding boxes. So we got it up to like 63 plus percent. And we released that model. So yeah, we're always looking for models that have specific capabilities. Like right now, I'm on a desperate hunt for models that store memory with much higher resolution. So everybody uses Clip. Clip is 256 by 256. 256. You lose a lot of information when 256 by 256. So now you've got some of the newer stuff coming out that's at 512 or whatever, 486 by 486. I'm waiting for the day when it's just not compressing it down, right?
53:35And that's sort of all built into the memory and you have a pure multimodal memory. So I'm always looking for that kind of stuff. The whole team's looking for that kind of stuff. So yeah, we're very much looking for specific models that do very good things and then obviously when it comes to reasoning obviously you're just looking for the models that have you know the best capabilities for us you know any multimodal reasoning model is obviously the greatest you know is the is the golden goose right that would be the one that we'd love to have but but you've got to have lots of little models to do things and like finding the right the perfect multimodal rag model or finding the perfect click model finding the perfect one that could do planning on things yeah you have to look at all that stuff and you have to be paying attention to it and there's no there's no kind of perfect source for that other than reading archive and twitter and talking to people and like you know the discord communities your your machine learning people are in and that kind of thing um but the vast majority of stuff that comes out is noise it's not you know it's like well we're like three percent better than this or like we quantized it down to like you know it's like you gave it a lobotomy it's it's useless like yeah it runs on your desktop, but who cares?
54:46The fact is there's maybe a hundred models in the world that are really valuable. And then maybe there's a couple thousand proprietary ones for individual tasks that people have trained that we don't know about or something like that. But in general, it's a lot smaller than it seems if you can see through the noise, I think. You spent some time working in MLOps and kind of thinking about the infrastructure. or to what degree are you, you know, what does that look like for you? Are you at the size where you care about that? You know, how, you know, how much automation are you trying to build around, you know, various things?
55:26Like, it's not super important for us now, other than when, like, we, when I find what I have found, it's funny, it's like, as CEO, like, half of my job is just leading the company. but like i love to keep my hands you know in the cookie jar playing with stuff and so like i was always i always wanted to be a good coder but i was pretty much a terrible coder i was a great sysadmin and so like so like you know claude and o1 and all those are really giving me a great chance to code and be like a great prototyper and i think differently than coders right so like my my genius coders would be like that can't be done i'm like i hate that word and And I'll like prototype it because I don't have any priors and 50 % of the time I'll be right.
56:11And they're like, God damn. Right. And then they'll go implement it correctly. Right. So that's fun. And then I'm spending a lot of time on the system inside of the house currently, just sort of like I said, I was playing with the MS Swift, trying to really understand the state of the art and kind of distributed training. so we can, not because I need, you know, a 10 ,000 GPU cluster, but, you know, I want to make sure that we can fine-tune the 405s and the 70 VMLs and that we can tune them fast and that they're valuable to us and, like, how can I make that a stable? And almost everybody in the company is kind of a Swiss Army knife.
56:45They can do a lot of different things. I was very surprised. I pretty much felt like when we came out of the MLOps era that MLOps was solved. i very much do not feel that way now like i i i look at this and i'm like wow we the the mlf era got truncated and rug pulled by the gen ai era and all the money went over to there but it was not finished um and and the problem is a little different it's different enough it's different enough like i think the main thesis of the mlf era was that everyone was going to be doing advanced machine learning. Big companies would have thousands of data scientists training models from scratch.
57:28That is never coming to pass now. It's the frontier models and the open source models that deliver a huge chunk of the functionality. You might train a proprietary model for a small subset of tasks or something very specific that you need to do that's not covered by those things, or you're going to be doing a hell of a lot of fine tuning and culturing in data sets so the problem shifted it just became it just became very different the thesis was wrong the thesis was just wrong and so now we're at this point where it's like wow i really want the thing where it's like auto magically take my data set analyze it and be like i have clean ditched i have i have split it i have put it into the correct formula that you need based on my deep knowledge of Lama 3.3, you know, and...
58:19Sounds like, what, Paka, Paka? Paka, Paka. What was that company called again? Packager. You can't get that from like HPE or something? I cannot. So no, those systems don't really exist in a way that I think is valuable. When you look at the open source community is proliferated, so when you start to look at inference engines, right, you look... There's four different major ones, and they all have model-specific code around them to make it run faster. They've done optimizations is the word I was looking for. So it's like, yeah, you can use a general-purpose one that'll run any kind of model, but it's just going to run slow.
59:06So it all depends on what that project is interested in. And it was very interesting to see, for instance, VLM was primarily interested just in the tech space models and not really in multimodal. And the Chinese were super interested in multimodal. And so by the time Momo comes out and they were looking to support in VLM, there was a big project internally to rewrite all the multimodal stuff. and like it kind of you know i could see in the prs it was like okay well almost almost raised for it but we're going to finish the rewrite of all that stuff right so there's kind of there's all this stuff that's sort of happening and it still feels very much like black magic right it still feels like why did my you know fully shorted you know fsdp pie torch run just fail on me the last few days and have me tearing what air i have left out you know um and it's like who knows man like you know i'm watching youtube videos on this like this is not auto magical at this point it's not doing like perfect error recovery and then spitting out a clean error that says like hey you know what you should really think about ratcheting down this this and this and i think the problem is this it's just sort of like digging through stuff uh oh i think that's the era right and like working with Claude on it to be like, and feeding it the docs and reading the stuff ourselves and watching YouTube.
1:00:28That's not where we should be, but that's where we are still. And it's frustrating, to be honest with you. Open source. Yeah. You've mentioned open source on a couple of occasions and clearly excited about it as a consumer of it. Are you contributing to it or are you planning to contribute to it with what you're building? Yeah. So we've already, we have, we've already released like 10 different repos as open source. We've got a K8s. We've got a Kubernetes-style deployer for agents. We've got a virtual desktop that it can control like a VNC, so the agent can control it like you would control a desktop via VNC, so it can spin up a container and talk to that.
1:01:07There's a tool fuse, which basically mounts that into the agent so that it can talk to it. And we've got another 15 plus repos, many of them which will be open source. and i think we have an open core philosophy 50 plus in the works mate is that i mean yeah 50 50 plus in the works you know that are sort of still behind behind the closed doors still okay um you know not all of them see the light of day obviously uh um but uh many of them will i i think we very much believe in open source open core and um it's not that nothing will be proprietary but to me open source is an unmitigated good for the world.
1:01:46I've actually been really angry that I've been fighting the open source battle again in AI. I thought that was one in the Linux era and that you'd have to be an idiot not to see that open source is the most important software in the history of mankind. It's worth trillions and trillions of dollars in value to the economy, the entire cloud ecosystem, your phone, the router in your house, supercomputers, everything. and yet here we are it's under attack again right with the same old stupid things you know like you know the bad guys can use it the bad guys can use everything right i mean this is this is ridiculous like i have a kitchen knife and we don't take the kitchen knife off the market because somebody stabbed someone we let the other 99.99999 % of people cut vegetables with it so i've been really frustrated to fight that that that battle again because i thought i thought it was over but i think you've seen this sort of paranoid uh you know sort of a dimmer crowd kind of come in and try to attack it and you know we can't do this like it'll it'll go crazy it will lose control and i'm like once again i'm like describing magical properties and stuff to me it's it's a hugely uh a huge loss for society when you go back to kind of the you know the the proprietary pure software play.
1:03:05Proprietary software always has a, I'm not, I'm actually, I'm not an open source fanatic in the way that it's like everything must be open source. You know, I'm like proprietary software is super awesome. It's super useful. I want them both to coexist. And I think that open research makes things more secure. I think it makes them more understandable, interpretable over time. Like that's where the research is going to get done. It's not going to get done when you get a paper that's published. It's like, we're not going to tell you anything about the how we restrain what the architecture is you know how you know that's you know you it doesn't matter if you have 10 000 people working there that's different from you know 10 million people looking at it in the open source way that's the reason that all of the you know encryption algorithms and stuff are open right and not proprietary right it's you've got to you've got to put it out there so to me i'm i'm an open source diehard it's one of the few hills that i'll die on i think it's absolutely essential did we cover all of the highlights from the talk anything we missed we covered a lot of good ground how about we look we covered a lot of it i think it's really i think this is just an exciting era i think there's there are ways to make the the we talked a lot about how to kind of make the agents smarter i think it's going to be a dog fight to make them smarter over the course of time i think i think a lot of the cool stuff that's been done with um reinforcement learning is going to start to really bleed into it we're starting to envision you know could we just clone up like you know 100 000 websites and let let an agent go wild in there with you know so yeah like alpha web right just going in there and just kind of playing around with it until it figures out how to use that right could that be done i think so right and i think that kind of stuff could be you know super awesome and really kind of take it you know to the next level but i think everybody building these things today i think to me the timelines are not as like compressed as people think they are like and you know dario and i are going to disagree on this where you know he's like look i've seen that movie before the scaling laws will just work and we're just going to throw you know throw more stuff at it and like it's absolutely going to get smarter i have zero doubt that they're you know produce better and better models but a lot of them are plateauing with what we we can you can see it in the charts are plateauing right it's like it's so it's i i i don't i think you need some fundamentally new stuff um to kind of make them better and i'm excited for that but i think you got to get your hands dirty and i think there's there's nothing like applied at and this isn't even pure machine learning but it's like trying to get it to do real stuff in a reliable way is is where the rubber meets the road i think it's where the next crash comes but i also think it's where i think it's where the multi-billion dollar applications come from too and the agentic era takes back some power for us like you were talking about earlier until the pendulum swings the other day again, and that gets in shitified in 10 or 20 years, right?
1:05:59But it took a while, right? It took a while for the dark side to manifest. And I think the agentic applications are almost limitless. Even at the lowest level, you think of it as RPA++, but at a high level, being able to do something super open-ended, we don't have any comparable technology. It's the most general technology you can possibly have. And why would it not be as important as the printing press or the internet or, you know, the hammer and nails, right? It's just a, it's a fundamental, and now I'm starting into the hype cycle a little bit, but I just, but I don't think it's next week.
1:06:38Self-awareness is important. I don't think it's next week. I think it's a gradual process of hard work. It's a Sisyphean process. It's really hard, you know, and the further you roll the rock up that hill, the harder it gets. And I think, you know, and you're going to have some plateaus and some peaks and some troughs. But when these systems are effective at maintaining sort of human workflow or doing human workflow in an open-ended way, I don't know what you asked, what are the applications? What are not the applications? Right? That sort of becomes the aspect of it. So I think it's just exciting to be in this age of discovery again.
1:07:14And a lot of companies are going to get whacked again when the music stops and the chairs get pulled up. But the next trillion-dollar companies, the next world beaters are going to come out of this era again. And it's going to be something really awesome. Well, Dan, as always, it was wonderful catching up. Awesome. Thanks so much. Yeah, man. Hopefully not as long this time. but you've got a lot of people to talk to. I've got to make myself relevant enough that you're going to ask me to be back on here every six months or something.
1:07:55Awesome. Thanks for having me. I appreciate it. All right. Thanks so much.
1:08:16you
From the publisher
Today, we're joined by Dan Jeffries, founder and CEO of Kentauros AI to discuss the challenges currently faced by those developing advanced AI agents. We dig into how Dan defines agents and distinguishes them from other similar uses of LLM, explore various use cases for them, and dig into ways to create smarter agentic systems. Dan shared his “big brain, little brain, tool brain” approach to tackling real-world challenges in agents, the trade-offs in leveraging general-purpose vs. task-specific models, and his take on LLM reasoning. We also cover the way he thinks about model selection for agents, along with the need for new tools and platforms for deploying them. Finally, Dan emphasizes the importance of open source in advancing AI, shares the new products they’re working on, and explores the future directions in the agentic era.
The complete show notes for this episode can be found at https://twimlai.com/go/713.




