944: Gemini 3 Pro: Google’s Back on Top

28 Nov 2025 · 8 min · 2 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Gemini 3 Pro performance and why it signals Google is back on top in frontier AI.

Guest backgrounds

No guest appears in this episode; host Jon Krohn references Professor Joey Gonzalez (LM Arena creator) and mentions company sources (LM Arena team, Wayfair CTO, GitHub).

Key claims

Gemini 3 Pro debuted first on LM Arena across text reasoning, vision, coding, and web development; it tops the overall leaderboard and reportedly broke an ELO 1500 threshold. Benchmarks cited: Humanity’s Last Exam 38% vs GPT-5.1 27%; Math Arena Apex ~23% vs ~1% for GPT-5.1 and Claude 4.5; AIM, GPQA Diamond, and MMMU Pro set new highs; ARC-HEI-2 31% vs 18% (GPT-5.1) and 14% (Claude 4.5).

Notable examples

VendingBench 2.0 simulated profit $5,500 vs $4,000 (Claude 4.5) and $1,500 (GPT-5.1); ScreenSpot Pro 73% vs 36% (Claude 4.5). Deployment: rolled out across Google Search, Gemini app, Vertex AI, and developer tools; Wayfair pilots infographic generation from support documents; GitHub reports 35% code accuracy boost over Gemini 2.5.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Deep Dive into Gemini 3 Pro's Performance

0:28 to 4:43

An overview of Gemini 3 Pro's top evaluations and benchmarks.

“I'll fill you in on everything you need to know about Gemini 3 Pro's performance and why it matters.”

Applications and Impact of Gemini 3 Pro

4:43 to 6:40

How Gemini 3 Pro is being used in various industries and its broader significance.

“Well, Google is wasting no time deploying Gemini 3.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Jon Krohn:This is episode number 944 on Gemini 3 Pro.

0:07Jon Krohn:Welcome back to the Super Data Science Podcast. I am your host, Jon Krohn. There's been a ton of excitement around Google's new Gemini 3 Pro model. It appears that Google, after trailing behind OpenAI since the release of ChatGPT three years ago and then later lagging behind Anthropic as well, it appears Google has regained the top spot at the frontier of AI capability. In today's episode, I'll fill you in on everything you need to know about Gemini 3 Pro's performance and why it matters. First, Gemini 3 Pro is topping major evaluation leaderboards. On the popular LM Arena leaderboard, for example, which we cover in detail back in episode number 707 with Professor Joey Gonzalez, whose lab devised the unique arena.

0:47Jon Krohn:But in a nutshell, what makes LM Arena unique and hard to game relative to most benchmarks is it uses human evaluators. And Gemini 3 Pro debuted in first place across all of the key tracks. Text reasoning, vision, coding, and web development. In fact, the Elemarina team noted that Gemini 3 surpassed even brand new rivals like XAI's Grok 4.1 and OpenAI's latest GPT-5 class models on a wide range of tasks from math and long-form Q &A to creative writing. This is huge news because in recent years, OpenAI's and Anthropik's best models like GPT 5.1 and Claude Sonnet 4.5 were seen as the ones to beat.

1:29Jon Krohn:Gemini 3 Pro upended that and is currently, at the time that I'm recording this, topping the LM Arena leaderboard overall. Beyond LM Arena, Gemini 3 Pro's performance on a wide array of benchmark evaluations is also impressive. For example, on Humanity's last exam, Gemini 3 Pro scores a 38 % with no tools like web search or code execution, while its next closest competitor, GBT 5.1, scores only 27%. That's a big jump on what is still a tough benchmark for LLMs in 2025. On another also tough benchmark, a math contest benchmark called Math Arena Apex, Gemini 3 Pro solved about 23 % of the problems, which might sound low until I tell you that the next best performing models, GPT 5.1 and Cloud 4.5, managed only about 1 % on that test.

2:20Jon Krohn:So this is a jump from 1 % to 23 % with his Gemini 3 Pro release. That's pretty crazy. Gemini 3 Pro also set new high watermarks, including the AIM math benchmark, the GPQA diamond benchmark of scientific knowledge, and the triple MU Pro multimodal understanding and reasoning benchmark. It also tackled tricky visual reasoning puzzles in the Arc HEI 2 challenge quite a lot better than Rival scoring 31 % versus just 18 % and 14 % for GPT 5.1 and CLAWD 4.5 respectively. That is a remarkable jump in a notoriously hard test of abstract problem solving. For those of us using LLMs to help us with coding or for agentic AI tasks, Gemini 3 Pro made some cool leaps there too.

3:09Jon Krohn:While it performed comparably to Claude and GPT models on the popular Sweebench verified benchmark, it crushed all other models on LiveCode Bench Pro. It also shines in tool use and long horizon planning. For instance, on VendingBench 2.0, where a commercial vending machine simulation is carried out, Gemini 3 Pro made a simulated profit of over$5 ,500, while its closest competitors, Claude Sonnet 4.5 and GPT 5.1, net profits of only$4 ,000 and$1 ,500 respectively. I guess you'd actually, if you added together the performance of both Claude Sonnet 4.5 and GPT 5.1, that would still be a little bit less than Gemini 3 Pro made on its shot at VendingBench.

3:56And finally, for those of us interested

3:57Jon Krohn:in developing agents that can understand what's happening on a computer screen to be able to act autonomously using that information, Gemini 3 Pro might be the first LLM we'd actually trust with that, because it scored 73 % on the ScreenSpot Pro benchmark, while its closest rival, Sonnet 4.5, scored less than half that, only 36%. And what about the model itself? Well, Google hasn't revealed much about Gemini 3's architecture, no surprise, the details are proprietary. But one technical spec worth sharing or worth mentioning is that it boasts a big million token context window that corresponds to about eight novels worth of natural language.

4:35Jon Krohn:Not such a big deal because million token context windows are becoming more and more common with frontier models, but still worth mentioning. And so how is it being used? Well, Google is wasting no time deploying Gemini 3. The launch is one of Google's most extensive ever with Gemini 3 Pro being rolled out across many of their products simultaneously. It's powering features in Google search, the new Gemini chat bot app, Vertex AI cloud services, and developer tools. Beyond Google's own ecosystem, some early partners are already test driving it in real workflows. For example, Wayfair's CTO shared that they've been piloting Gemini 3 Pro to convert complex support documents into clear, data-accurate infographics for their field teams.

5:15Jon Krohn:That's a great illustration of Gemini's multimodal power in an industry setting. It's taking long, text-heavy manuals and automatically producing visual digestible summaries. We're also hearing about Gemini 3 being evaluated for coding copilots and enterprise knowledge assistants. GitHub, for example, noted a 35 % boost in code accuracy when they switched to Gemini 3 Pro in an early test compared to its predecessor, Gemini 2.5. The broader significance is that this release signals that Google is back on top in the AI race, at least for now. Over the past year, OpenAI's GPT models and Anthropics' Claude models often grabbed the most advanced LLM headlines, but Gemini 3 Pro has emphatically put Google's flag in the ground as the lab to now beat.

6:00Jon Krohn:It leapfrogged not only OpenAI and Anthropics latest, but also up-and-comers like XAI. In fact, the team behind LM Arena noted Gemini 3 Pro was the first model to ever break an ELO score of 1 ,500 on their platform, a testament to just how far Google pushed the envelope here. For Google, this is a strategic win. It validates their investment in combining DeepMind's research with other Google resources. For us, leveraging frontier models and bringing them to the real world within applications, as I suspect many listeners are, this shows us that the competition at the frontier of AI is alive and well.

6:39Jon Krohn:As each new model raises the stakes, we all benefit from the rapid improvements in what these systems can do. On that note, I'm not gonna cover it in any detail in today's episode, but you might also want to check out Google's brand new Nano Banana Pro. I've got a link to that for you in the show notes to see what's possible for the state of the art in text to image generation, as well as image editing. Perhaps unsurprisingly, Gemini 3 Pro is the LLM running in behind on Nano Banana Pro's backend. Well, with new capabilities at your fingertips now for image generation, image editing, text generation, coding, and especially agentic AI relevant capabilities like screen understanding and long-term reasoning, I hope your human brain is buzzing with ideas about what you could do with Gemini 3 Pro.

7:25If not, chat with whatever your favorite LLM is to get ideas on what you could do given your particular industry, network, background, etc. While being able to predict what company will be leading the AI race in any given month may be tough, one thing's for sure. Folks like us leveraging and deploying these rapidly improving LLMs are all benefiting. All right, that's it for today's episode. I'm John Krohn, and you've been listening to the Super Data Science Podcast. If you enjoyed today's episode, or know someone who might consider sharing this episode with them, leave a review of the show on your favorite podcasting platform, please.

8:00And if you aren't already, be sure to subscribe to the show. Most importantly, we hope you'll just keep on listening. Until next time, keep on rocking it out there. And I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.

8:16Thank you.

From the publisher

Google is steaming ahead with launching its top-league new Gemini 3 Pro model across their product suite, from Google Search to Vertex AI cloud services. The multinational tech company is also letting eager early adopters like Wayfair and GitHub. Get all the detailed data, its performance across hard-to-game industry benchmarks, and what this all means for the way you use generative AI, in this week’s episode.

Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/944⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.

More from Super Data Science: ML & AI Podcast with Jon Krohn

All 130 episodes
944: Gemini 3 Pro: Google’s Back on TopSuper Data Science: ML & AI Podcast with Jon Krohn · 8 min
Listen in VO