OpenAI Agent "Operator" Coming In January?

15 Nov 2024 · 17 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The AI Daily Brief - Episode Summary

Podcast Title

The AI Daily Brief (Formerly The AI Breakdown)

Episode Title

OpenAI Agent "Operator" Coming In January?

Episode Overview

This episode discusses OpenAI's imminent release of its autonomous agent called "Operator," set for January. This agent will autonomously perform tasks like coding and booking, heralding a significant advancement in AI applications. Additionally, the episode covers OpenAI's proposal aimed at enhancing U.S. AI infrastructure, likening it to a "Manhattan Project" for AI aimed at ensuring national competitiveness.

---

Key Highlights

  1. OpenAI's Autonomous Agent: "Operator"
  2. Launch Timeline: Expected in January as a research preview.
  3. Capabilities:
  4. Autonomously execute tasks such as coding, shopping, and booking flights.
  5. Competing agents from Google and Anthropic are also in development, indicating a competitive landscape.
  6. Industry Implications:
  7. Represents a shift from AI as mere assistants to agents capable of performing complex human tasks.
  8. Potential for significant disruption in business and societal structures.
  1. AI Model Slowdown
  2. Current Situation: OpenAI and Google are experiencing a slowdown in model performance improvements.
  3. Technical Challenges:
  4. OpenAI’s Orion model has not shown expected advancements similar to the leap from GPT-3 to GPT-4.
  5. Google’s models are reported to demonstrate similar stagnation, challenging the previously held scaling laws regarding data and computational power.
  6. Responses:
  7. Both companies are exploring new methodologies for enhancing AI reasoning and fine-tuning models.
  8. Google’s initiatives include forming teams to improve performance through new architectures and methodologies.
  1. Business Model Innovations
  2. Perplexity Advertising Experiment:
  3. Introduction of ads in the format of sponsored follow-up questions for U.S. users.
  4. Aimed at generating sustainable revenue alongside existing subscription models.
  5. Salesforce's Position:
  6. CEO Mark Benioff expresses skepticism about AI negatively impacting the company's bottom line, emphasizing the firm's extensive data management capabilities.
  1. OpenAI's Infrastructure Proposal
  2. Ambitious AI Infrastructure Plan:
  3. Proposed building AI economic zones to expedite project approvals and promote energy infrastructure.
  4. Advocates for increased solar, wind, and nuclear energy investments to support AI demands.
  5. Historical Context:
  6. Compared to the 1956 National Interstate and Defense Highways Act, hinting at large-scale, transformative infrastructure projects.
  7. Geographical Focus:
  8. Emphasis on the Midwest and Southwest for infrastructure expansion to foster equitable job distribution across the U.S.
  1. Future Outlook
  2. AI as a Vertical Enterprise Tool:
  3. Predictions suggest the greatest impact of agents will be in specialized enterprise tasks, rather than general-purpose applications.
  4. Concerns on Pricing Models:
  5. Discussion around how to price AI agents as they become capable of performing extensive tasks efficiently.

---

Discussion Points

  • Agent Technology:
  • The potential transformation in how businesses operate with the introduction of AI agents.
  • The balance between hype and practical application of AI capabilities.
  • AI Model Performance:
  • Diminishing returns on AI scaling and the need for innovation beyond mere data processing.
  • The shift from scaling models to discovering new methods and architectures.
  • Economic and Policy Implications:
  • The importance of government investment in AI infrastructure for maintaining a competitive edge against China.
  • How upcoming U.S. administration policies may affect AI development and infrastructure projects.

---

Conclusion

The episode encapsulates the evolving landscape of AI with OpenAI's new agent technologies and infrastructure proposals. It highlights both the challenges and opportunities within the AI industry, pushing for a more strategic approach to innovation, infrastructure, and competitive positioning in the global arena.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Today on the AI Daily Brief, we are apparently getting an open AI agent as early as January, which is a good thing because in the headlines, we talk about how Google is also dealing with the same AI slowdown that we talked about in the context of open AI a little bit earlier this week. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. To join the conversation, follow the Discord link in our show notes. Welcome back to the AI Daily Brief Headlines Edition, all the daily AI news you need in around five minutes. One of the big discussions we've been having this week is whether there is an AI model slowdown.

0:33According to reports, OpenAI's Orion model is not showing the same jump in performance that was observed between GPT-3 and GPT-4. This apparently has led OpenAI to double down on reasoning and fine-tuning as a potential way to obtain the performance boost they expect from the next generation of frontier models. And now it appears that Google is joining OpenAI and exploring new avenues to tackle some of these challenges. According to the information sources at Google, their models are demonstrating the same lack of improvement. The information writes, past versions of Google's flagship Gemini large language model improved at a faster rate when researchers used more data and computing power to train them.

1:06Google's experience is another indication that a core assumption about how to improve models known as scaling laws is being tested. Many researchers believe that models would improve at the same rate as long as they process more data while using more specialized AI chips, but those two factors don't seem to be enough. This is particularly troubling for Google, whose models have failed to see the same level of adoption as OpenAI's. There was a belief, perhaps, that Google could leapfrog OpenAI in this generation purely due to their advantage in computing resources. That appears less and less to be likely.

1:34So, following in the footsteps of OpenAI, Google is also looking to develop new methods of improving performance. Over recent weeks, Google DeepMind has put together a team to work on the development of reasoning models. That team is being led by principal research scientist Jack Ray and former Character.ai founder Noam Shazier. To give a sense of how important they consider the work to be, other DeepMind researchers are working on making manual improvements to the model, including changing so-called hyperparameters, which are the variables that determine how the model processes information and how quickly it draws connections between different concepts.

2:03Another problem Google has run into is duplicate copies of information within training data which could hurt performance. Google has also experimented with synthetic training data, essentially feeding data generated by an LLM back into the corpus of training data. They've also added audio and video. And while it was believed that these steps would lead to significant improvements, sources at Google say they didn't make a major difference. Meta Chief AI Scientist and Turing Award winner Jan LeCun has been predicting these diminishing returns from model scaling for years. Yesterday, he posted on threads, I don't want to say I told you so, but I told you so.

2:33He referenced the statement from former OpenAI Chief Scientist Ilya Sutskever from earlier in the week, who said, The 2010s were the age of scaling. Now we're back in the age of wonder and discovery again. Everyone is looking for the next thing. Scaling the right thing matters now even more than ever. Now what makes Ilya's comments more significant is that he was basically the chief proponent of the idea that you could just add more compute and data to keep scaling higher and higher. So the fact that he is moving away from that suggests that there is a changing understanding of the technical capabilities here.

3:02Lacoon, for his part, commented, we've been working on the next thing for a while. Lacoon is referring to the fundamental AI research team at Meta who are pursuing new architectures as a path towards AGI. Their focus is currently on world models which seek to train an AI on how objects and environments interact rather than just focusing on the connection between words. Still, not everyone is convinced that this is a real thing. Bindu Reddy writes, The AI slowdown is a non-story. The biggest reason AI is slowing down is that there's nowhere else to go. If you begin to saturate on benchmarks, nothing is left to do.

3:31100 out of 100 is the highest score you can get. Now, moving over into the world of business models, Perplexity says they will begin experimenting with advertising on their platform. Starting this week, U.S. users will see ads in the format of sponsored follow-up questions. These ads will be placed to the side of generated answers and labeled as sponsored. The initial brands and agency partners for the launch include Indeed, Whole Foods, Universal Mechanon, and PMG. As an example, Perplexity showed a search for information about looking for a job with a sponsored follow-up that says, how can I use Indeed to enhance my job search?

4:04In a blog post, Perplexity explained, ad programs like this help us generate revenue to share with our publisher partners. Experience has taught us that subscriptions alone do not generate enough revenue to create a sustainable revenue sharing program. Advertising is the best way to ensure a steady and scalable revenue stream. Perplexity said the ads themselves will be generated by AI rather than pre-written or edited by sponsors. Advertisers also won't get access to users' personal information. Regarding the choice of format, Perplexity wrote, We intentionally chose these formats because it integrates advertising in a way that still protects the utility, accuracy, and objectivity of answers.

4:36These ads will not change our commitment to maintaining a trusted service that provides you with direct, unbiased answers to your questions. Obviously, right now, how perplexity-style AI summaries influence the core business model of the web, which is basically search ads on Google, is one of the big open questions. So this will be really interesting to watch these experiments. Basically, I think they're a lot more consequential than just for perplexity as a company itself. Lastly today, speaking of business models, Salesforce CEO Mark Benioff thinks it's crazy talk that AI could hurt his company's bottom line.

5:08Benioff has very publicly planted his flag on the idea of AI agents over recent months, and we're about to find out whether it will save Salesforce from disruption. During a recent appearance on TechCrunch's equity podcast, he said, what if your workforce had no limits? As far as being disrupted, Benioff believes that his moat is access to client data. He said, we manage 230 petabytes of data for our customers. You could say that might be one of the main things we do for them, and we do it with a security and a sharing model. Anyways, for me, it's interesting to note that Benioff feels like he has to justify the potential disruption to the SaaS business model from agents.

5:40He can say it's crazy talk all he wants, but there is absolutely no doubt in my experience and in my conversations with enterprises, while Benioff may be right that agents are Salesforce's future and that they're an even bigger deal and the company grows to lofty new heights, there is an incredible pressure being put on the sort of traditional per seat model of SaaS companies that will not be resolved easily or quickly. Certainly something that we're really interested in and thinking about a lot as we price super intelligent, but that is a conversation for another time and place. For now, that is going to do it for today's AI Daily Brief Headlines edition.

6:10Next up, the main episode. Today's episode is brought to you by Plum. Want to use AI to automate your work but don't know where to start? Plum lets you create AI workflows by simply describing what you want. No coding or API keys required. Imagine typing out, AI, analyze my Zoom meetings and send me your insights in Notion, and watching it come to life before your eyes. Whether you're an operations leader, marketer, or even a non-technical founder, Plum gives you the power of AI without the technical hassle. Get instant access to top models like GPT-40, Claude Sonnet 3.5, Assembly AI, and many more.

6:41Don't let technology hold you back. Check out Use Plum, that's Plum with a B, for early access to the future of workflow automation. Today's episode is brought to you by Vanta. Whether you're starting or scaling your company's security program, demonstrating top-notch security practices and establishing trust is more important than ever. Vanta automates compliance for ISO 27001, SOC 2, GDPR, and leading AI frameworks like ISO 42001 and NIST AI risk management framework, saving you time and money while helping you build customer trust. Plus, you can streamline security reviews by automating questionnaires and demonstrating your security posture with a customer-facing trust center, all powered by Vanta AI.

7:19Over 8 ,000 global companies like LangChain, Lila AI, and Factory AI use Vanta to demonstrate AI trust and prove security in real time. Learn more at vanta.com slash NLW. That's vanta.com slash NLW. Today's episode is brought to you as always by Super Intelligent. Have you ever wanted an AI daily brief but totally focused on how AI relates to your company? Is your company struggling with AI adoption either because you're getting stalled figuring out what use cases will drive value or because the AI transformation that is happening is siloed at individual teams, departments, and employees, and not able to change the company as a whole, Superintelligent has developed a new custom internal podcast product that inspires your teams by sharing the best AI use cases from inside and outside your company.

8:06Think of it as an AI daily brief, but just for your company's AI use cases. If you'd like to learn more, go to besuper.ai slash partner and fill out the information request form. I am really excited about this product, so I will personally get right back to you. Again, that's bsuper.ai slash partner. Welcome back to the AI Daily Brief. A couple interesting pieces of news out of OpenAI today, starting with an agent story. OpenAI is reportedly planning to release an autonomous agent next year. Now, if you spend any time around the AI space, you'll know that basically since ChatGPT launched, We've been on the verge of the agent era.

8:45The idea of moving from just these super powerful assistants to agents actually doing human replacement style work is something with such dramatic implications for the structure of business and society and what we can accomplish that it captures a huge amount of energy. In fact, probably a disproportionate amount of energy relative to how far the technology actually is. And yet it's clear that the major labs have been making nudges in this direction. And so what have we learned so far about this theoretical agent from OpenAI? The agent, which they've codenamed Operator, can control a computer to complete tasks independently, including coding, shopping, and booking flights.

9:20According to Bloomberg sources, staff were told in a meeting on Wednesday that the tool would be released as a research preview in January. That would mean that by early next year, we could have competing computer use agents from Anthropic, Google, and OpenAI. There are also already more limited agents available from companies like Microsoft Salesforce and a host of startups. So far, we've seen two different approaches to fully-fledged computer use. Google's agent is sandboxed in the browser window, making it more limited but potentially more performant. Anthropics agent, which is the only one generally available, was trained to control a mouse and a full computer interface, so can theoretically carry out a much broader variety of tasks.

9:56In practice, the experience is still rather limited, with the company admitting the agent is slow, cumbersome, and error-prone. Bloomberg sources said that OpenAI is working on several agent-related products, and the one that is nearest to completion is a general-purpose tool that executes tasks in a web browser. Sam Altman has been hyping agents as the next big thing over the past few months. In October during a Reddit AMA, he said, We will have better and better models, but I think the thing that will feel like the next giant breakthrough will be agents. At OpenAI's Dev Day, Chief Product Officer Kevin Wheel said, I think 2025 is going to be the year that agenting systems finally hit the mainstream.

10:32The Verge writes, AI labs face mounting pressure to monetize their costly models, especially as incremental improvements may not justify higher prices for users. The hope is that autonomous agents are the next breakthrough product, a ChatGPT-scale innovation that validates the massive investment in AI development. So what do people think about this? Well, Elvis on X writes, Computer use is the kind of capability I expected OpenAI to launch first. This time, it looks like they will be following Anthropic. I'm still hoping for a unique twist and huge improvements. I think there's a lot to learn from custom GPTs and search.

11:04One thing is certain, 2025 will be the year of AI agents. I never bet against OpenAI on these things, especially because O1 for data generation, as was used in ChatGPT search, will play a key role here. They have a huge advantage. My asks, make it easy to launch agents, make it easy to use and provide feedback, versatile with tools and integrations, reduce latency, and reduce costs. Others shared their skepticism. Ilion X writes, I hate how companies always flex their AI agents that can quote book a flight for you as if that was not the worst use case for AI automation ever. This, by the way, is something that is a personal pet peeve of mine as well.

11:39I do not need an agent to book me a flight or to order me food. I know that it is just demonstration of capabilities, but I do think it shows how early we are that those are the things that people always point to. Others are thinking about the implications for various domains of the world. Callum McClark writes, For the learning world, clicking complete means we must shift from clicking complete courses to gathering meaningful learning metrics, a shift that should have happened a long time ago. But now hopefully these agents will be the final nail in the coffin of bad L &D courses. There are also questions of the business model of agents, which are swirling.

12:11Sully Omar writes, how do we properly price agents, especially when they keep getting more capable, they do days of work in one hour, and they work 24-7? We have no market reference. Is it compute per hour per task? Adam Silverman of Agent Ops writes, I think as agents scale over the next five years, pricing will be compute plus 10 % margin for specific use cases. When OpenAI and others release agents, they will only charge compute. There's a huge opportunity for startups to capitalize on charging significantly more in the interim. Now, going back to this Verge quote, the mounting pressure to monetize costly models, and the hope that autonomous agents are the next breakthrough product, I think whereas ChatGPT and Claude and the like have been easily as much a consumer innovation as they have been an enterprise innovation, really transforming and hitting both equally, and by some measurements consumers more, I believe that where we're going to see the value from agents is absolutely in vertical, highly specific enterprise tasks.

13:04I think that getting agents very good at very specific, repetitive tasks that happen over and over and over again all the time inside specific companies is much easier than getting good at very general purpose use. And I think that that's where we're going to see a lot of benefit. Now, it is still very early. What we have available and what's production ready is still very nascent. But if I were a betting man, that is where I would be placing my chips that the impact will be in very specific verticals within the enterprise. Second OpenAI story, the company has outlined a grand policy proposal to bolster artificial intelligence in the U.S.

13:37in an effort to stay ahead of China. Yesterday at a think tank event in Washington, OpenAI's head of global affairs, Chris Lehane, presented what they're calling their official blueprint for U.S. infrastructure. The company said the plan was, quote, as ambitious as the 1956 National Interstate and Defense Highways Act. They outlined AI economic zones, co-created between state and federal governments, which would give the states an incentive to speed up permitting and approval for AI infrastructure. They envisioned constructing solar and wind power, as well as gaining clearance to restart unused nuclear plants.

14:05OpenAI wrote, States that provide subsidies or other support for companies launching infrastructure projects could require that a share of the new compute be made available to their public universities to create AI research labs and developer hubs aligned with their key commercial sectors. OpenAI also wrote a bill called the National Transmission Highway Act. The legislation would expand power, fiber, and natural gas pipeline connectivity across the nation. The company argues that we need, quote, new authority and funding to unblock the planning, permitting, and payment for transmission. And they noted that existing procedures aren't keeping up with AI-driven demand.

14:36The document noted that, quote, the government can encourage private investors to fund high-cost energy infrastructure projects by committing to purchase energy and other means that lessen credit risk. OpenAI notably views the Midwest and Southwest as key areas for infrastructure expansion given the plentiful land for construction. Focusing on these areas would also ensure that the jobs and prosperity of this wave of technology aren't just concentrated on the coasts. Now, the premise of all of this is that without government investment and the removal of red tape, the U.S. will lose its lead in AI to China.

15:03The OpenAI proposal stated, given the stakes, we need to think big, act big, and build big. These decisions determine whether a nation leads or lags in technological innovation, often with far-reaching consequences for economic competitiveness and national security. The history of the U.S., they wrote, is one of iconic infrastructure projects that move the country forward. The auto industry, the Tennessee Valley Authority, the Manhattan Project, the interstate highway system. One of the things that we've been tracking a lot recently is the growing push to nuclear driven by the AI industry. And it's notable in that regard and something that OpenAI points out, that China has built more nuclear power capacity over the last 10 years than the U.S.

15:39has built over the last 40. Speaking to that rapid deployment, OpenAI's head of global policy, Chris Lehane, said, We don't have a choice. We do have to compete with that. Now, the big remaining question is whether these kinds of policies would be adopted by the Trump administration. OpenAI says they plan to work with the Trump White House on this agenda. So far, all we know about Trump's AI policy, however, is that the president-elect has pledged to repeal the Biden AI executive order, stating that, quote, in its place, Republicans support AI development rooted in free speech and human flourishing.

16:07That's about the extent of the details we have so far. Then again, slashing red tape and building out a ton of energy production seems in line with campaign promises. During an appearance in July, Trump said, we will be creating so much electricity that you'll be saying, please, please, president, we don't want any more electricity. We can't stand it. You'll be begging me, no more electricity, sir. We have enough. We have enough. So who knows? I think what's pretty clear is that this is coming into the vacuum of whatever the repeal of the executive order looks like. As Andrew Curran points out, this blueprint is being presented to influence whatever form the new regulations will take.

16:38Still, the earnings nugget may have summed it up best when they wrote, this is the new Manhattan Project. interesting times ahead. But with that, we will wrap today's AI Daily Brief. Appreciate you listening as always. And until next time, peace.

From the publisher

OpenAI is set to release its autonomous agent "Operator" in January, a tool that aims to execute tasks like coding and booking autonomously, marking a potential breakthrough in AI's role in practical applications. Alongside this, OpenAI has unveiled an ambitious proposal to advance U.S. AI infrastructure, advocating for energy expansion and streamlined AI project approvals to maintain national competitiveness. This "Manhattan Project" for AI could reshape the future of tech and policy in the U.S.


Brought to you by:

Vanta - Simplify compliance - ⁠⁠⁠⁠⁠⁠⁠vanta.com/nlw⁠⁠⁠⁠⁠⁠⁠The AI Daily Brief helps you understand the most important news and discussions in AI.

Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614 Subscribe to the newsletter: https://aidailybrief.beehiiv.com/ Join our Discord: https://bit.ly/aibreakdown

More from The AI Daily Brief: Artificial Intelligence News and Analysis

All 1,099 episodes
OpenAI Agent "Operator" Coming In January?The AI Daily Brief: Artificial Intelligence News and Analysis · 17 min
Listen in VO