In short
AI Today Podcast Episode Summary
Episode Title
Why OpenAI’s New AI Agents Are Causing a Stir
Overview In this episode of AI Today, the host discusses OpenAI's recent release of two open-source AI models, their implications, and how they compare to previous models. The conversation touches upon the benchmarks, performance, and the ethical considerations surrounding the deployment of these AI agents.
Key Topics Discussed
- OpenAI's New Open Source Models
- First Release in Five Years: OpenAI has released two new models, marking the first open-source models since GPT-2.
- Criticism and Expectations: The release has been met with criticism due to OpenAI's past as an open-source advocate.
- Models Differentiation: The episode clarifies the distinction between "open source" and "open models".
- Model Performance and Benchmarks
- Benchmarking Performance:
- CodeForce Benchmark: The 120 billion parameter model scored an ELO score of ~2600, compared to OpenAI's O3 and O4 models which scored 2700 and 2720, respectively.
- Humanity’s Last Exam (HEL): The two models scored 19% and 17%, showcasing strong performance against tough questions.
- Hallucination Rates: The new models exhibited higher rates of hallucination (49%-53%) compared to older models (16%).
- Model Training Insights
- Training Methodology: Models were trained using high-compute reinforcement learning on NVIDIA GPUs, focusing on teaching the AI to differentiate right from wrong.
- Parameter Activation: The 120 billion model activates only 5.1 billion parameters per token.
- Licensing and Ethical Considerations
- Apache 2.0 License: The new models are released under a lenient license allowing monetization.
- Data Transparency Issues: OpenAI is withholding the training data details, likely due to legal concerns over copyrighted material.
- Microsoft's Integration
- Windows AI Foundry: Microsoft plans to make the smaller model available to Windows 11 users, positioning it for tasks such as code execution and tool use.
- Accessibility: The model is optimized for varying hardware and is designed for real-world applications.
Key Takeaways
- OpenAI's Shift: The release indicates OpenAI's attempt to address past criticisms while maintaining a competitive edge in model performance.
- Performance Nuances: While the new models show promise, issues like hallucination rates raise concerns about reliability.
- Market Implications: OpenAI's decision to allow monetization of its models could stimulate innovation in the AI space while raising ethical discussions on data usage and model training.
Conclusion The episode provides an in-depth look at OpenAI's latest developments in AI agents, balancing excitement for new capabilities against the backdrop of ethical considerations and performance reliability. The host invites listeners to explore AIbox.ai for access to a variety of models, emphasizing the growing landscape of AI tools available to developers and businesses.
Additional Resources
- [AI Box](https://aibox.ai)
- [AI Chat YouTube Channel](https://www.youtube.com/@JaedenSchafer)
- [AI Hustle Community](https://www.skool.com/aihustle)
> For a broader understanding of these advancements and their implications in the AI industry, listeners are encouraged to engage with the content and participate in discussions around ethical AI development and real-world applications.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Today on the podcast, we're talking about OpenAI, which has just dropped two open source models. Now, this is actually really big news because this is the first time in five years that they've actually dropped any open source models back to GPT-2. And this is something they've gotten a ton of criticism. Basically, all of Elon Musk's online AI beef, pretty much why he says he started XAI. And just a lot of drama and heat that has been thrown at OpenAI is basically on the fact that they started as an open source company and hadn't dropped anything. And they have now officially launched some quote unquote open models.
0:31Now, I'm going to be talking about the difference between open source and open and where these models sit. I'm also going to go through the benchmarks of basically how these models perform because a criticism that a bunch of them have gotten is like they just dropped like these models that are, you know, just to say that they're open source, but they're not actually that good. And I'm actually, I'm not going to lie, impressed by some benchmarks, but interested. And there's a couple of interesting nuances I want to go over. At the same time, Microsoft has just announced that they're going to be bringing some of their smallest open models to Windows users.
1:00So there's a ton of really interesting things that are getting rolled out right now. We'll be covering all of that on the podcast today. Before we get into it, I wanted to mention, if you want to try any of the AI models that we talk about on the show, I'd love for you to go check out my own startup, which is called AIbox.ai, where we essentially have the top 40 different AI models from Anthropic, Coher, DeepSeek, Google, OpenAI, Meta, tons of others, audio models like 11 Labs, and a bunch of really interesting image models, all for 20 bucks a month, you get access to all of them. So my hope there is not just that it'll save you some money on, you know, the absorbent amount of AI models that you can subscribe to, but really that you'll be able to find and try out a whole bunch of different AI models that you hadn't heard of or used before.
1:39I think there's a lot of really great unheard of models that can do some great things in specific tasks. We have kind of benchmark data and we break down what models are best for what on the platform. So go check it out. It's 20 bucks a month, AIbox.ai. All right, let's get into what OpenAI is doing. So the first benchmark that I want to talk about is the CodeForce benchmark. They basically ran the GPT-OSS 120 billion parameters. That's the bigger of the two open source models. They have a 120 billion parameter one, and then they have a smaller one. But basically, the bigger parameter, 120 billion parameter one, got an ELO score on CodeForce of 2600, roughly.
2:22And just to compare that with OpenAI's other tools, their 03 model got 2700 and their 04 model, their 04 mini model, sorry, got 2720. So like these things aren't very far apart. It definitely did better than the 03 mini model, which only got 2000. So it did pretty decently. Now, a lot of them are also rated with or without tools. So that 2000 benchmark that I quoted you was without tools. And what exactly does it mean tools? And is that important? Yes, I would say it's important. Tools basically mean they gave the AI model things like calculators and apps and like different tools. So like it is completing the tasks, yes, but it's able to rely on like actual hard software to get good results.
3:09And this is something that we basically use in AI because with these LLMs, because we found that they're like, perhaps not necessarily fantastic at a math problem or like some sort of really intense molecular biology question when they're kind of just guessing what should come next in the line. But they're good at figuring out what they need to do. And so we actually leverage that to find what tool to use and then bring the tool in to solve maybe a more calculated problem. Okay, so does it matter that they're giving us basically benchmarks with and without tools? Yeah, I think the big thing here is when they release the model open source or quote unquote open for everyone to download and use, they're not releasing it with the tool.
3:46So no one's getting, they give this benchmark with tools or without tools, but they don't give us the tools because they're kind of OpenAI's proprietary stack. But what I will say is it is a good benchmark because big companies or software like startups, they can build their own tools. And usually if you're taking one of these models and putting it into your startup to do a certain task, you're going to be building custom tools anyways. I think back to my first startup, which is called SelfPause. It was a no-co, or sorry, it was an AI life coach. and we basically had it so you would talk to chat GPT and we would act like a life coach and work you through different questions and we built our own custom things into there to basically instruct and guide how the AI model worked and how it would run a conversation and so I think most software startups would be similar okay humanity's last exam this is kind of a notorious benchmark it's called HEL but basically it's humanity's last exam before AGI is kind of the concept but it's got a whole bunch of really complex questions.
4:44You heard this exam quoted a lot by XAI when they released their latest version of Grok. It did really well on this. So with tools, GPTOS, so they're basically their 120 billion parameter model, and their 20 billion parameter model scored 19 % and 17%. Well, I was actually really impressed that the 20 billion parameter model got 17%. that's a very that's not very far behind 19 which is the 120 billion parameter model again this is an incredibly hard test uh test so less you're like oh 20 this is terrible this is i don't think i would be able to get any questions on this on this test right it's like super in-depth uh you know like translating ancient hebrew meaning of this hieroglyph to how does it convert to this thing it's like it's very very complicated questions okay the most many experts would not excel at or succeed in.
5:41Okay. And it's a whole bunch of different areas that it tests you. So in any case, this outperforms the O3 model, but, or sorry, it underperforms the O3 model, but it does outperform a bunch of the leading open models from DeepSeek and Quen. So the Chinese companies that are releasing their open source models, it beats the Chinese companies, but it's not better than what OpenAI has closed source, what basically they're selling. So I mean, that kind of makes sense. One thing I will say, though, is OpenAI's new model does hallucinate much more than its latest, you know, 03 or 04 mini models. So that is not a particularly fantastic statistic.
6:24So this is definitely something that's interesting, because it seems like these have actually, these hallucinations have actually been getting more severe, which is not a good thing. We don't want to hear that OpenAI's models are getting worse and worse at hallucinating. They don't even really understand why, which is another thing that's not super great. They wrote a white paper where they said, quote, basically they said that it is, quote, expected as smaller models have less world knowledge than large frontier models, and they tend to hallucinate more. So basically, they're blaming on the fact that there's less data, less parameters inside of these models.
6:55That's why they're hallucinating more. But we're, yeah, it's kind of a trend that you see in a lot of places. So it's kind of interesting. Basically, they found that their 120 billion parameter model hallucinated in response to like 49 % of a benchmark called persons QA. So it's one that is asking the AI model basically about people. So like who is Tom Cruise and give me information about him. Who is, and it goes through like tons of people, some that are famous, some that are less famous and they'll see if it hallucinates. If you've ever tried it, like if you go say, who is Jaden Schaefer to chat GPT?
7:30Like sometimes it'll grab information from, you know, the web and whatever my LinkedIn, but like half the time, I especially remember in the early days, we'd thrown tons of random stuff that weren't true. And so I think with these older models, we're kind of it's kind of like, or these open source models are kind of like some of the older models, they throw a bunch of funny things in there. So that's an area that hallucinates doesn't mean hallucinates on every topic, but on people, it definitely is doing that. So yeah, it's their their newer models that they they have that you they pay for like their 01 model had a 16 % score there so you know 16 % hallucinations about people versus these models are like 49 % and 53 % it's a lot more so anyways it's kind of funny open eye said that basically they are not going to share what data they use to train these models and basically I think this comes down to a whole bunch of lawsuits that are happening where people are saying that opening eye uses copyrighted data.
8:25I'm assuming that they did. They're just not announcing it. They're trying to probably work with some regulators in Washington to make it all kosher and okay before they officially say that they're not. So I think that this is interesting. 20i does say that the model that they have for their 120 billion parameter model, it only activates 5.1 billion parameters per token. They also say that their model was trained using something called high compute reinforcement learning. Basically, this is a process. It's a post-training process that helps to teach AI models right from wrong, like the right answer to the wrong answer.
9:04And they do this in a simulated environment. And they do all of this using basically a really big cluster of NVIDIA GPUs. So this is how they trained their O series of models, right? Their O3 and O4. And so it has basically a similar chain of thought process where they, it takes a longer, but it's saying like, okay, how would I solve this problem? It comes up with a list and then it works through that list chain of thought on solving different problems. And we typically get better answers when we do this. Now, here's what's exciting, in my opinion. They are releasing both of these models under the Apache 2.0 license.
9:37So this is really considered as one of the most, I guess, like lenient licenses. is it it will allow companies to monetize this model right so you can actually charge money for it unlike some things that meta was doing when they were releasing they were they like we're like here's our like open source models long versions of llama but like companies you got to talk to us if you want to like make any money off of it open ai is being really generous letting people make money off of it um they don't have to pay open ai and they don't have to get permission from open ai so this is really just a free gift to the world of an open source model which is my opinion really really cool i will say unlike fully open source models what's the difference They're calling it an open model, but it's not totally an open source.
10:15The difference is that companies like AI Labs or sorry, AI2 have like a fully open source model. The difference is that OpenAI is not going to release the training data that they use to create their models. I kind of already talked about it. It's basically for legal reasons. They probably have copyrighted stuff in there, which is their own choice. But basically anyone that gets the model is going to benefit from it because the model is going to answer and be way more accurate and have higher quality results. So that's kind of the outcome of that. I will say that they have delayed this model multiple times.
10:48I have personally been disappointed when I've seen Sam Altman's tweets on Twitter over the last couple of months where they keep delaying it for safety reasons. They think that it is a lot safer now. Basically, the things that they said they were concerned about was cyber attacks or the creation of biological or chemical weapons. Basically, you get information from the models that could help you do those two things. It seems like they've kind of put some guardrails and made the model better now. So it doesn't do that. They had a bunch of third-party evaluators actually test it, and they said that it marginally increases biological capabilities, but it didn't find evidence that they were going to have a very high capacity threshold for danger in these domains after fine-tuning.
11:24So I think, or sorry, even if you tried to fine-tune it to be able to do that. So I think it's going to be a much safer model. It's definitely state-of-the-art amongst other open models, right? So if we're looking at like DeepSeek and Quinn and Meta's Llama. It's definitely the top of the pack there. We're also waiting for DeepSeek R2 to release, which also should give it a run for its money. So it'll be interesting to see what's happening there. So all of this is going down, which is really, really interesting. And the last thing I wanted to bring up, though, is that Microsoft is basically bringing their smallest model, right?
12:01So they have the 20 billion parameter model. They're bringing it to a bunch of Windows users, which is pretty interesting. It's going to be for any Windows 11 users. It's via Windows AI Foundry. So this is kind of their platform that lets you use AI APIs and a bunch of popular open source models on your computer. Microsoft in a blog post said, tool savvy and lightweight optimized for agentic tasks like code execution and tool use. It runs efficiently on a range of Windows hardwares with support for more devices coming soon. It's perfect for building autonomous assistance or embedding AI into real world workflows, even in bandwidth constrained environments.
12:38So basically what you actually need if you want to run this, and this will be starting on Tuesday, but it'll be able to run on most consumer PCs and laptops, but you have to have at least 16 gigs of VRAM, which basically a modern GPU from NVIDIA or Radeon would have. OpenAI said that the model was trained using high compute reinforcement learning. So it pretty much excels at powering AI agents and a bunch of other tools. It can do web search, it can do Python code execution, and all of that. So really, really impressive. I'm excited to see where this goes. I'm excited to see that Microsoft is kind of rolling out some sort of integrations so that a lot of people can use this.
13:18This is a really cool moment. You can go download this today on Hugging Face, which is super cool. And I'm excited to see what people build with it, what companies start using it. This is just honestly a gift to the world, and I'm sure OpenAI has more exciting things up their sleeve like GPT-5 that'll probably blow this out of the water. But for what it's capable of doing, anyone gets access to a really world-class AI model and so I'm quite excited about that. All right, thanks so much for tuning into the podcast. Make sure to go check out AIbox.ai if you wanna try out a lot of the different models I talk about on the show.
13:48For 20 bucks a month, it's an amazing value and I would love to hear what you have to think about or have to say about it because it's currently in beta. We're taking feedback and adding tons of new features all the time. Thanks so much for tuning in and I will catch you in the next episode.
From the publisher
With this new release, OpenAI has opened the door to powerful automation tools. These agents can plan tasks, automate workflows, and even mimic human reasoning. This episode dives into both the promise and the pitfalls.
Try AI Box: https://aibox.ai
AI Chat YouTube Channel: https://www.youtube.com/@JaedenSchafer
Join my AI Hustle Community: https://www.skool.com/aihustle

