Multi-On is What You Wanted AutoGPT to Be - Interview with Founder Div Garg

23 Jun 2023 · 24 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AI Daily Brief Podcast Summary: Episode with Div Garg

Podcast Overview Podcast Title: The AI Daily Brief (Formerly The AI Breakdown) Description: A daily news analysis show focusing on artificial intelligence, examining creativity, industry disruptions, and philosophical questions surrounding AI. Episode Title: Multi-On is What You Wanted AutoGPT to Be - Interview with Founder Div Garg Episode Description: Interview with Div Garg, founder of Multi-On, an AI personal agent that executes complex tasks using the browser.

---

Key Topics Discussed

Introduction to Div Garg

  • Background in AI for almost six years.
  • Experience includes roles at NVIDIA, Apple, Google, and Uber.
  • Current focus on Multion—an AI personal agent.

Multion Overview

  • Described as the world's first AI personal agent.
  • Aims to utilize the browser for executing complex tasks.
  • Positioned as a response to the shortcomings of existing AI assistants like AutoGPT.

The Appeal of AI Personal Agents

  • Increased interest in AI assistants amid disillusionment with previous models.
  • Multion aims to bridge the gap between theoretical AI and practical usability.
  • The focus on browser interaction allows the AI to replicate human-like control over web tasks.

Multion's Unique Features

  • Task Execution: Can automate tasks such as ordering food and scheduling meetings.
  • User Interaction:
  • Two modes: Step-by-step (user approves each action) and Auto mode (automatically executes tasks).
  • Ability to ask clarifying questions to refine user requests.
  • Safety Measures: Emphasis on trust and safety, ensuring users feel secure when using the AI to control personal data and browsers.

Real-World Applications

  • Users currently employing Multion for:
  • Research tasks.
  • Social media interactions (e.g., wishing friends on birthdays).
  • E-commerce transactions (e.g., purchasing household items).
  • Streamlining scheduling processes (automatically creating calendar invites).

---

Insights on the Future of AI Personal Assistants

  • One General Agent vs. Specialized Agents:
  • Users prefer a single AI that can handle multiple tasks rather than several specialized AIs.
  • Potential for specialized agents to emerge for complex tasks (e.g., trip planning, legal research).
  • Integration with Existing Platforms:
  • Companies may integrate AI functionalities directly into their applications, reducing the need for multiple agents.
  • Multion is exploring how to enable third-party integrations and allow developers to build applications on its platform.

---

Upcoming Developments

  • Beta Expansion: Currently in a closed beta with plans to scale to 1,000 users.
  • Hackathon: Organizing an event to explore innovative uses of the Multion agent, encouraging participants to develop applications leveraging its capabilities.

---

Conclusion Div Garg's insights into Multion highlight the evolving landscape of AI personal assistants, emphasizing usability, interaction, and the potential for broader applications. The episode provides a glimpse into how AI could reshape everyday tasks, making technology more accessible and efficient for users.

---

Additional Resources

  • Learn more about Multion: [Multion Website](https://multion.ai/)
  • Subscribe to The AI Breakdown Newsletter: [Newsletter Subscription](https://theaibreakdown.beehiiv.com/subscribe)
  • Join The AI Breakdown Community: [Community Link](bit.ly/aibreakdown)
  • Follow on YouTube: [YouTube Channel](https://www.youtube.com/@TheAIBreakdown)

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00On today's AI Breakdown, I'm speaking with Div Gargard. Div is the founder of Multion, a new AI agent that uses the browser to execute complex tasks. The AI Breakdown is a daily podcast and video about the most important news and discussions in AI. Like, subscribe, and share, and go to breakdown.network for more information.

0:21Welcome back to the AI Breakdown. Right now, as I mentioned, I am traveling in Europe. And so as something a little bit special and different, I wanted to bring you a set of interesting interviews for that time when I'm going to be away. Today, my guest is Div Garg. Div has worked on AI at numerous companies, including NVIDIA, Apple, Google, and Uber. He's an adjunct faculty member at Stanford. And he's now building Multion, which they're billing as the world's first AI personal agent. Now, if you've been listening to this show for the last couple months, you know that there has been immense interest in AI assistants and AI personal agents.

0:57As people have gotten a little bit perhaps disillusioned with things like AutoGPT, I've seen a number of people who have early access to Multion and feel like it was what they were looking for out of that project. Div and I talk a little bit about his background, about Multion's current capacities and how people are using it, and what he thinks the future of AI personal agents really is. All right, Div, welcome to the AI Breakdown. How are you doing, sir? Great, yeah, thanks for having me here. Yeah, no, I'm super excited. As I was just saying to you off air, I think you are building in one of the areas that people are most excited about, so I think it's going to be a great conversation.

1:33But before we get into Multion and what you're doing now, what's your background? How long have you been in AI? What's the perspective that you're bringing to this? I've been doing AI for the last almost six years. And I think I would say I almost started around like my freshman year. So I actually did a lot of physics in high school, did like an international physics alumni period. And when I joined undergrad, I was like, hmm, what is the right thing to do in life? Like, should I do physics? Should I do something else? and seemed like physics was saturated and you needed a different thing to make the jump to solve a lot of these problems.

2:08And it seemed like AI was the right thing to do. I had the feeling that there'll be AI that will solve physics, not just humans. So a combination of both. So I made the choice of focusing on AI since very early on when I started college. And I've been working on a lot of systems. My first internship was actually working on an autonomous driving car back at Uber, where I did a lot of benchmarking for them, trained their 3D computer vision models for driving cars on the roads and detecting all the vehicles and the pedestrians' cycles on the road and making that really safe as a system. And afterwards, did a lot of research around how do you make autonomous driving cheaper or safer.

2:47So we did this research at Cornell, where we got an autonomous driving car to just work with two cameras instead of a very expensive lighter sensor. and that time like applied ours used to be like$70 ,000 more expensive than the car itself and we showed that you can do this a very similar thing just using like$100 cameras and we got a lot of like media coverage on that got a Forbes coverage put a bunch of other media publications and just got me really excited about the potential of AI and how can you apply to the real world so that has been my focus for the like the last couple of years uh worked in like a bunch of like big companies obviously like a lot of like top secret AI projects so I was like almost like the go-to guys like oh like we have this like uh like i had an internship at google they were like oh there's this like the thing on computer vision you're working on we can't disclose you but we like your details we like your resume you want the job yes or no i was like okay sure i had a similar experience in like apple survived a lot of this like sort of like ai like like very like uh security stuff in our big companies and uh there was a lot of fun there's also a time like where ai was like research but like it doesn't actually work as a product so So people tried a lot of things.

3:53I actually worked on some AI devices, like AR kind of stuff in Google back in the day. Did a lot of interesting reinforcement learning stuff in Apple. Did some diffusion model stuff in NVIDIA. So that was fun. But there was no actual real-life applications. And after I joined Stanford, where I was focusing a lot on physical agents like robots, how can you make them more controllable? And how can you have more powerful algorithms that can learn from human data? So I would say like a lot of robotics how it currently works is you just have pre-programmed loops. You just like program like subscript and it just like goes and execute that.

4:27A lot of my research thesis during my PhD at Stanford was like, how can you learn from human data? Can you like observe videos of humans? Can you like sort of like see what humans are doing? And like teach a robot or agent to like sort of like do some other things. So I created this algorithm around like learning from human videos. We actually won like the number one prize in a Minecraft AI challenge two years ago for like the New York's conference back in 2021. So we had this agent that can watch 50 videos of human players building houses in Minecraft and use that to go itself and build a house.

4:57And so that was really fascinating. Also did a couple of projects on how can you steer a robot using natural language. So we actually did the first project around how can you combine actions together with language to teach a robot to control it using human voice. So you can tell a robot, go open this drawer or pick up a mug and go and do that. So we did a lot of like explorations around this like physical agents and then like was actually working at a robotic startup for a while in their AI efforts, building a lot of their like AI algorithms and simulations. And afterwards, like it just seemed like the right time, like sort of like after Chachy PD came out, I was like, OK, like LLM as a technology is getting really good.

5:36And it seems like we are reaching this phase where we can start communicating with AI agents. Because before it was sort of like you have everything as a number, it's like a tensor. You don't really know what's happening. But now it's like, okay, we are getting past this communication gap where as a human I can tell it something and go and understand that and communicate back to me what it's doing. And that just seemed like, okay, that was something that was missing. And we are finally getting there where this can become actually usable. and everyone can go and steer this sort of agents. And so, yeah, so the child GPT has been fascinating.

6:10As a technology revolution, I think everything was there, but just it made LLMs more mainstream, especially with GPT-4. Now you have such good reasoning capabilities. Yeah, so I've been very interested in computer interaction agents recently, which is sort of multi-on how can you actually take a language command, which could be a voice or text, and actually translate that to real-life actions by controlling a human browser. and this is similar to how a human controls a website. So if a human can go and click, type, and do everything, my thesis here is we can train an AI that can also do this very effectively at the same rate as a human.

6:46And so a lot of the current approaches you will see around plugins and APIs, which is good, but the problem there is they're very restrictive, it's hard to build APIs for everything, and it's almost like using a backdoor where someone has to give you entry and expose it so you can go and use that. but like using like the browser is sort of like a front door like it's like every human is already doing it and so if you can teach an AI to like just use the front door properly we can interact with anything on the web potentially also like anything on the desktop and so this seems like a very horizontal and powerful way where we can like reach very powerful virtual agents.

7:18It's super interesting. It sounds like one common thread you know first when you were a student and then when you started to be you know in industry and building companies is a real interest in these tools moving from theoretical to actually usable, right? Going and doing things. And to the extent that that's true, I think it's interesting that you found your way into this AI agent space, given, you know, we were talking about this as well, but there has been such excitement around AI agents. You know, AutoGPT ripped onto people's view at the beginning of April and had this real sort of, you know, deflation, I think a couple of weeks later, partially because, you know, it was people who are non-technical using very fast, non-technical implementations of it.

7:57But, you know, if I had to sum it up, It was basically what people were excited about was the idea that one, an AI could figure out the steps to achieve a goal, and then two, it could actually do the steps. And I think what a lot of people found with the first implementations is that that first part happened fine. It was a great plan for how you would go do X, Y, or Z goal, but then there was no actual connectivity to actually going and executing. And it sounds like Multion's approach to this in some ways is, like you said, walking in the front door of the browser and trying to make sure that it can interact with, you know, a thing that we all interact with, which is, you know, the web browser.

8:33Definitely. Yeah, it's actually interesting. Like for us, like, Multion was actually in a very functional state back in February. And we just decided not to release it because one was like around trust and safety. So like if you just release this to like a million people, like things are going to go haywire and like how do you control things? Another thing was like we tested with some folks and they were like just like scared, like they were very skeptical. oh, this thing can go and control my computer. That doesn't seem right. And so the interesting thing was, after AutoGPT, people became more familiar with agents because before that, people didn't know what agents were.

9:06And so if we talk to someone who's non-technical, they'll be just like, oh, what is this thing? Why is it taking control of my computer? What it's doing? But now, after AutoGPT, people are like, oh, yeah, the agents exist and everything. So we'll say it helped us build this, sort of clear out the space where people know, okay, what agents are, what can they do? And if you now give multi-on to someone, they understand the capabilities and we can make it into a trustworthy experience. So it's been an interesting journey because we basically started a lot before AutoGPT, but have been just trying to focus a lot on how can we make more reliable, make more trustable, put safety guardrails.

9:43And so that's been the focus that we have had for the last three months in terms of making better and better. And currently we are in this closed beta where we have currently around 100 beta users. It's mostly invite-only. We are trying to increase that to 1 ,000 beta users over the next three weeks. And then we have around 20 ,000 people on our wait list. So we'll be doing a lot of these launches as we iteratively make it much more safer and get all the feedback from how people are finding it, how can we improve it as an agent. Amazing. I think maybe at this point it'd be great, if you're up for it, to do a quick demo so we can see how Multion works.

10:16Sure, definitely. So I can ask it something, let's say something like if I say order a burger, from Mel and Palo Alto using Node.hash, for example. Okay, I'm gonna have to log in. So yeah, the log in is one thing we take very seriously because we want to make sure that people don't start misusing it for like, building like bots and like, spamming people and like doing like crazy things. So here the agent is like sort of thinking on like, so it's making a plan on what to do. And then once that's created the plan, it starts searching and here like you can like see what exactly it's doing and it will like start taking actions so you can see like it said like it's clicking on like the first link and then can like go and actually like start like ordering the thing for people who are listening to this because this will be on a podcast as well there's basically a multi-on window in the bottom right hand of the screen that as it sort of controls the browser is explaining what's happening right So it says, I am clicking on the link to DoorDash page for the Melt in Palo Alto to proceed.

11:23Right. And in this case, it can actually ask me a question. So it asks me, do I want the Melt burger or do I want a different one? And so this is a new feature we've been experimenting with where we gave it the ability to ask clarifying questions to a user present options. So if I say something like, I want the BBQ bacon burger. Yeah, and it's a bit on the slow side today, but usually we can work very real time. and then it's trying to find the burgers and then can like find it and add it to the cart. And then it can also like automatically do the whole checkout if I ask you to do that. So it went to the checkout screen.

11:55I'll probably pause it here, but otherwise it can go and like buy the whole thing for me. Yeah, so I'm actually used to having a random do-dash order show up at my house. That's amazing. When it's doing this, how much is it asking you to approve at different steps versus just doing it itself? Yeah, that's a good question. So we have two modes we built. So one is sort of like a step-by-step mode where whenever it does something, it will ask you, like, should I do another thing for you? And if you press a hotkey, so in this case, the right arrow on your keyboard, it will take another step. And so you can control each step.

12:28You can approve, like, okay, next step, next step, next step. And so that way it's safe. If it does anything wrong, you can stop it and give it another command. The second mode we have is auto, where it will go and do the whole thing. And it's almost like watching a movie or seeing a video on your screen where it's interactively doing things. And we have built this sort of like a pause button. So if you press the space key whenever it's running, you can pause the agent. And so it's like almost having a control to a remote where you can say, okay, do me this thing. And then you press the play button and it starts doing it.

12:56And then you can press the pause button anytime. Or you can give it a new command and then you can give it a play again. So it's a very interesting sort of way to control a computer where you can imagine in the future, you might just need something like an Apple. if you've seen like those Apple TV remotes, which are like very minimal, just has like a mic button and like a play pause. That's all you need. You might not even need like a keyboard or a mouse in the future. Yeah, it's super interesting. I can imagine, and I have no idea if your tests validate this at all, but I could totally see at the beginning people almost using the step version as like kind of personal training wheels or trust training wheels, because it's like, I want to see how this thing works a few times because I bet a lot of people, you know, it's not like they've already adapted to a new interaction mode, right?

13:47They're experimenting with it. And so they'll naturally kind of step through. But I can imagine if you've done something two or three times successfully, then you just auto post it, you know, especially once you know that there's sort of a pause button that, you know, if it goes haywire, you can stop it. Definitely. That's what we've seen. We've also have some users that are like, we just got bored and we started playing with multi on because it was like fun to see what it's doing. So we've had like a lot of those sort of things. We have also seen people like to widely share videos of it to their friends.

14:15And they're like, oh, like this computer is going and like automating itself. It's almost like a ghost in the shell sort of experience for folks like using it for the first time. Do you feel like you guys have a sense? Obviously, you're very early with the test. What are people using it for so far? And how does that compare to what you thought they might use it for? We have seen like a lot of people are actually using it for research currently, for fetching information online. I've also seen people using for like social media stuff where if you ask multi on like go wish happy birthday to my friends on Facebook It can actually find everyone has a birthday and like send them a happy birthday message for example or it can like find people on LinkedIn and Send them a message.

14:52So you've seen people like sort of those sort of workflows Also on emailing another we have seen is like around ordering So we've seen people like buying like stationery on Amazon or like toilet paper or something for example or drink chairs. So those are some interesting things you've seen so far. Also, let me also do this as a demo. So this might be interesting. So we also made a multi-unabit streamline for scheduling. So it's actually very good at creating calendar invites. And it can actually automatically include your Zoom link information. So but yeah, if I want to say, create a meeting invite, that can automatically go at my meeting details, and my personal Zoom link.

15:34And it sort of saved me maybe 15 to 20 interactions, especially if you're a power user who has to book a lot of meetings or set something up. So let me do this as a demo. So if I say something like book a 2 to 3 p.m. meeting tomorrow, and in this case that can be my co-founder, and say Teamsync. So here the agent is thinking and it will create a plan. So it went to the Google Calendar and then can start filling all the information. So it can automatically start putting everything. So it added the emails. It also is adding my Zoom link here and sharing it out. And then it can extend the whole thing.

16:16And then we're also trying to build a user verification workflow where before it sends something, it can send you a notification like, oh, I created this meeting invite. correct? Do you want to, like, does it look correct? Do you want to change something, for example? And if you approve it, then can they go and send the actual thing? So if I say like send, it will send this meeting. Super cool. Some questions on this. So there are two things that were correct that could easily be not correct that I noticed. One is it went to Google. Is that sort of, is it optimized for assuming that people are going to Google Calendar?

16:47And if not, you sort of, how do you change that? And then the second is it knew your Zoom in advance. So what is that? You How did you as a user sort of interact to teach at that? So we have built this sort of like a memory scratchpad feature where you can give it your personal details. So if you say, okay, this is my name, this is my address, stuff like that, then Multion will actually know all of that and customize it. So in this case, I've told it what my Zoom link is. I've also given some instructions like, okay, include my Zoom link in Calendar Invite so I can give it any notes. Almost like talking to maybe an assistant where like, okay, this is some do's and don't do's.

17:21If I tell it my allergies, for example, or I can tell it like, like, okay, like what seats are like on a flight, like window versus aisle. And can I take that preferences into account and like take actions accordingly. So this is like a feature we're starting to build. And so that's how it knows my Zoom link. In terms of like defaults, I think currently we have hard coding defaults where like suppose like it wants to make a calendar invite. We tell Meridian by default you should choose Google Calendar over something else. Or like if you want to search for something, choose like Google over like Bing, for example.

17:48And so this is like some choices we are made in the future. We can like allow people to customize this. There's also like interesting partnership potentials where like if someone comes to us and we're like, oh, can you make us your default provider for like, say, like travel, like for like Uber, for example, or say, can you make us your default provider for food, for example, which would be like DoorDash, and we could like use that for monetization. Totally. Now, can you make the AI breakdown your default source if you're asking a question about breakdown news? No, I think it makes tons of sense.

18:14So this kind of gets to a question that I was going to ask both broadly, but also in the context of your product specifically, which is, you know, what is your team's thesis about the future of this sort of AI personal assistants? Do you think that they're going to be super general with people using them for everything just the same way they would use a personal assistant now? Do you think that they're going to get refined into, you know, you're going to have eight or 10 of these that are sort of, you know, optimized for specific experiences? Is there some combination or is it just too early to tell?

18:47I think it's going to be a combination. What we've seen is people don't like interacting with 10 different services. So if I've had a choice and if I could just interact with one agent or AI that could do like 100 small things for me in a day, rather than having an AI that can only do one specific thing. So people really like the second modality more, especially around assistants, because they want something that can reduce the friction and a lot of their everyday small things. So we have seen there's a lot of space where there's potentially one general agent that can help a lot. It doesn't have to be really specialized, but as long as it's really helpful and a lot of really small things.

19:25So we see, I think that's where people really want an AI system. And there's also space where you can have very specialized AI agents. Example could be like, maybe you want to build a travel agent. I'm taking a one-week trip to Italy. that involves like for a human that might actually take like a one week to plan the whole trip discuss with the friends coordinate on the the hotels call the hotels schedule like ubers flights everything so that is easily more than one week to two weeks of planning and you can imagine if there was a special agent that can just go and do this in like one hour or something i think there's a space for those sort of special agents uh similar could be true for a lot of like research where like for a lot of complicated like say finance research or legal research you have to spend a lot of like months like doing all the groundwork finding everything and so i think for a lot of this like complicated specialized jobs you can have specialized agents um but for like like i would say like for everyday life i think you need like some sort of like a general agent instead of interacting with like 10 different agents you want to interact with just one it's fascinating i mean i think that the other the other sort of thing which might add some heft to that theory is you have to think that almost every company is going to experiment with retrofitting how you interact with their service with this sort of interface.

20:38And it will almost mean that you don't need to go out and seek specialized agents because they're just going to live in the apps where you already are. So if you use Instacart, you know, I mean, Sam Altman has said this about ChatGPT plugins. It's why he thinks that they don't yet have product market fit. He said something to the effect of, I think a lot of these companies that think they want to be in ChatGPT actually want ChatGPT in them, you know, which I think is an interesting insight. But that does leave this space for this sort of day in day out interactions. I think honing in on things that can be done in a browser is a really interesting insight as a way to sort of limit what the focus is while still keeping it really broad.

21:17Definitely. Definitely. We also see this as like in the future, this could become an interesting layer on a platform where we could allow people to build like more like applications on top of Multion and expose this sort of like action layer, where if you want to go build like a very powerful agent for some particular use case. You can have like multi-undu controller browser and do the heavy lifting and sort of like build like experiences around that. So we see that there's a lot of space like that, which a horizontal like agent like RSC could enable. Yeah, super, super interesting. So what is next for you guys?

21:50Well, you know, you're still in a very early beta. You said you're expanding that beta probably over the next three weeks. But, you know, what else is coming up on the horizon? Sure, sure. It's very exciting things. we're actually closing a big fun year round. So we're hiring. So that's been great. We are actually organizing a hackathon this weekend. So we are organizing this agent hackathon at the AJI house in Hillsborough. And it's almost like everyone in the AI space is there. So we have Karpathy coming to give the intro talk. And so it's going to be very exciting. And we'll be giving everyone who's attending the event access to Multion, as well as like programmatic control.

22:27So we will be giving like API access where you can like, so currently if you see like multi-on you can like a user didn't go and it couldn't like give it commands, but it will allow you to like programmatically give it commands. And then you can connect with LinkedIn, you can connect it with like other things, and then you can like build very powerful applications and use cases for the duration of the hackathon. And so we're trying to see that as an experiment on like what people will do with the sort of like agents that can actually have a lot of purchasing power, for example, can actually like do a lot of interesting things on the internet.

22:56And so we are also like trying to make sure like the event is safe and like we can moderate it. So like people don't start doing like malicious things. So we have already been like banking websites, for example, so it can work on them. And it'll be like a fun experiment. But I think it'll also be like very different from any other hackathon. Yeah, that's fascinating. I just today did a part of the show about this idea for a new Turing test that came from Mustafa from DeepMind and Inflection, where he said the new Turing test should be give an agent or an AI$100 ,000 and see if it can turn it into a million.

23:30So maybe someone will get a head start on that this weekend with the Multi-On Hackathon. I will say it's almost like a superhuman test because if an agent can do that. I know. I know, exactly, exactly. Well, listen, Dim, I really appreciate you taking some time today. I'm very excited to see what comes with this hackathon. We'll definitely share what comes out of that on the story. Yeah, that was great. Thanks a lot for inviting me.

24:13Thank you.

From the publisher

Today NLW is joined by Div Garg, the founder of Multi-On which is an AI personal agent that uses the browser to execute complex tasks.   Learn more: https://multion.ai/   The AI Breakdown helps you understand the most important news and discussions in AI. 
Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe
Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown
Join the community: bit.ly/aibreakdown
Learn more: http://breakdown.network/

More from The AI Daily Brief: Artificial Intelligence News and Analysis

All 1,099 episodes
Multi-On is What You Wanted AutoGPT to Be - Interview with Founder Div GargThe AI Daily Brief: Artificial Intelligence News and Analysis · 24 min
Listen in VO