Stack Overflow Becomes a Core AI Data Source

19 Nov 2025 · 9 min · 6 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Stack Overflow’s shift from a public developer Q&A forum to an enterprise AI data provider, positioning its content for internal AI agents after ChatGPT reduced forum traffic.

Guest backgrounds

No guests are identified in the transcript. Speakers referenced include Stack Overflow CEO Parashanath and CTO Jody Bailey.

Key claims

Stack Overflow saw major usage/web-traffic declines and increased AI scraping; it responded by launching an enterprise “Stack Overflow internal” tied to Model Context Protocol (MCP), plus an API and content licensing deals with AI labs. CEO claims enterprises already use its API for training and Stack Overflow offers “blanket” licensing.

Notable examples

Wikipedia traffic decline and API response; Chegg struggles; Reddit licensing deals reportedly bringing $200M+; Stack Overflow metadata (answer timestamps, tags, answerer identity) used for reliability scoring and knowledge-graph linking; CTO highlights an agent “writing” function that can generate new Stack Overflow questions.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

AIbox.ai Promotion

2:03 to 2:24

Learn about AIbox.ai, a platform offering various AI models.

“if you want to try any of the AI models that I talk about on the show, I'd love for you to try out my startup, which is called AIbox.ai.”

Stack Overflow's New Product Direction

2:30 to 2:48

Discover Stack Overflow's new tools aimed at enterprise AI integration.

“All right, let's talk about Stack Overflow.”

Impact of AI on Q&A Platforms

2:48 to 3:36

Examine how AI is affecting platforms like Stack Overflow and Chegg.

“Like every enterprise needs to have a license to this new Stack Overflow tool.”

Stack Overflow's API and Licensing Deals

3:36 to 4:28

Understand Stack Overflow's new API and its licensing strategy for AI.

“And Reddit has went ahead and made these deals where they'll license their content to companies and they're able to just make kind of like blanket deals.”

Metadata and Data Reliability in AI

4:28 to 6:12

Learn about the importance of metadata in ensuring data reliability for AI models.

“And it is essentially an enterprise version of the web forum that they have.”

The Future of AI and Stack Overflow

6:12 to 8:42

Speculate on the future roles of Stack Overflow in AI and developer communities.

“is a layer of metadata that Stack Overflow has access to, right?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Hanging out at the pool is great. Relaxing and playing Vegas style games on my phone at the same time. drink in one hand and a blackjack in the other it's all at spin quest over a thousand games including your favorite slots and table games be cool with this summer special new players get 30 coin packs for 10 at spinquest.com spin quest is a free to play social casino void where prohibited visit spinquest.com for more details you're listening to a podcast right now driving working out walking the dog if you're into podcasts chances are you have something to say too. With RSS.com, starting your own podcast is free and easy.

0:42Upload an episode and we distribute it to Apple Podcasts, Spotify, Amazon Music, and more. Track your listeners, see where they're from, and start earning from ads just like this. If you've been thinking about starting a podcast, this is your sign. Start your new podcast for free today at RSS.com. Today on the podcast, we're talking about Stack Overflow, which is essentially recreating itself into an AI data provider. I think the reason I want to cover this is because I think we're going to see this exact same trend played out with a ton of different online companies that are struggling with, you know, lower web views, lower usage after ChatGPT and a lot of these other AI tools came out that will answer questions for you.

1:25Stack Overflow is one that has been reported on extensively and seen a dramatic drop in usage, but you can also talk about Wikipedia, you can talk about Chegg, You can talk about a lot of different companies that would do kind of questions and answers in specific niche areas. The AI models came in, scraped their whole website, have all of that baked into their models now. And now the original companies are suffering because no one really is using them. So we're going to get into the future of some of these forum type websites and specifically the deal that Stack Overflow has done, how it's similar to other players.

2:01and more on the podcast. Before we do, I just wanted to mention if you want to try any of the AI models that I talk about on the show, I'd love for you to try out my startup, which is called AIbox.ai. You get access to the top 40 different AI models, Google Gemini, OpenAI, Anthropic, Cohere, DeepSeek, everything, image models, audio models like 11 Labs, all for 20 bucks a month in one place on one platform. So if you don't want to have to get a new subscription every time you want to test out a new AI tool, go check out AIbox.ai. there is a link in the description. All right, let's talk about Stack Overflow.

2:33I think this all came out at Microsoft's Ignite conference. So Stack Overflow came out and showed a whole bunch of new products that they were going to try to essentially use to position themselves as a really useful part of the enterprise AI stack, right? Like every enterprise needs to have a license to this new Stack Overflow tool. This is kind of a new take for the company. um stack overflow definitely struggled after chat gpt came out and there's a number of articles that just said their web traffic went down significantly right this is traditionally a website where developers would go on and ask coding questions and say hey look i'm running into this issue does anyone know you know how to fix this bug in my code people would respond and help debug or work on code problems together and you saw this play out in a lot of different industries i mentioned chegg which was like for students students would ask questions and other students would respond so it's kind of like more like an education side of course we saw this with wikipedia who has recently said that they are seeing a massive drop i don't want to say massive but they are seeing a decline in web traffic uh that is from humans and an increase in web traffic that is from ai scrapers bots um and maybe even some of those are agents and wikipedia has responded by making an api um chegg has been struggling and then we also see companies like reddit who again is a forum but like for everything.

3:55And Reddit has went ahead and made these deals where they'll license their content to companies and they're able to just make kind of like blanket deals. So with all of that, a lot of the new tools that they're making are specifically at Stack Overflow. They're specifically designed to feed into internal AI agents that are using the MCP or the model context protocol. And they're using that with different variations designed specifically for Stack Overflow. It's essentially, you know, Stack Overflow internal is what the new tool is called. And it is essentially an enterprise version of the web forum that they have.

4:34But they have a bunch of additional like security and admin controls on it. So companies have that extra security and control over the content. This is their CEO, Parashanath, said, talking about all of this said that they were already seeing a whole bunch of enterprise companies using their API for training. So that's another thing that Stack Overflow did, right? They saw kind of like Wikipedia, they're like, look, our traffic has dropped a lot. And we have a lot of, you know, bots that have been scraping us, they just made an API. And they're like, if you're an AI company, you should use our API for training, or you have to as our term service, otherwise, we're going to sue you.

5:10And they saw a lot of, apparently, according to their CEO, they saw a lot of success and progress with that specific model. And so then they decided to kind of take this new product direction where they're like, well, maybe enterprises would want access to Stack Overflow tied directly into the AI models that they use in a very direct way. They already made a bunch of different content deals with a whole bunch of AI labs that essentially allow them to train their models on public Stack Overflow data. And they're doing this just for a blanket fee. So it's very similar to the Reddit deal, which happened.

5:42And the Reddit deal has brought in more than$200 million for Reddit, just, you know, kind of giving like, I think Reddit is working with OpenAI and Google specifically that I know of. And I think it's like$100 million a piece. They're like, look, you can scrape Reddit, 100 million bucks, and you can kind of have access to this. So it's like a big boost in revenue for the company. And of course, OpenAI and Google are like, well, we don't have to deal with any lawsuits. It's a great data set. And others are kind of blocked from it. So it made sense for them. A really important part of this new product is a layer of metadata that Stack Overflow has access to, right?

6:17Because you could say, well, you know, if all of their website has already been scrapped by the AI models, why does everyone want access to, you know, maybe having like this custom API into it? And the reason why is because they still have some data that others don't have. Beside the questions and answers that you see inside of Stack Overflow, the data also includes some information like who answered the question and when they answered the question. They also have content tags and a lot more complex assessments of some of the internal coherence. So what this means is you could say like, look, I'm asking a question about Java, but this question was answered like back in 2012.

6:56So is it relevant to the current version of the coding language I'm running today? Or maybe I'm running, I'm using an old version of some coding language or some tool and I need like an older answer. And so what's interesting here is because they have that date not a lot of these AI models scraped that and so they're actually able to assign this sort of like an assessment score to say how likely the it's a reliability score which will tell the AI agent how likely the answer is to be trusted right it's like well based off of what you're currently asking the question about your current stack and when this answer was created this is how likely it is to be good and in addition to this they know who answered the question so they're actually able to look at those accounts and see, you know, how legitimate the accounts are, how, you know, how good of a developer they are, how good their solutions are.

7:45And then they can use all of the data from the individual users or contributors account to determine how good the answer will be. So this is interesting, right? And I really appreciate this. They're trying to lean on some data that other people might not have that they have exclusive access to and make the product better. The CTO Jody Bailey said this about it. They said the customer can set up their own tagging system or we can dynamically create that for them. What we'll be doing in the future is really leveraging that knowledge graph to connect people and to connect concepts and pieces of information rather than requiring the AI system to do that on their own.

8:20So while Stack Overflow right now is making a whole bunch of tools for enterprise agents, it isn't building all of those agents itself. So it's kind of hard to say what their final product is actually going to look like when it rolls out. Bailey is really excited about the writing function, though. Bailey is their CTO. And Bailey said that the writing function is going to allow agents to create their own stack overflow questions. If they can't answer a specific question, or they notice there's like a knowledge gap, they're actually able to ask a question on stack overflow. I think my question is, will real humans seeing AI bots ask questions on stack overflow, feel obligated to answer a bot, right?

8:59It's not really like a human, but there's usually a human behind the bot asking the question. So maybe they'll still be helpful. I'm not sure. Or are they going to just have AI bots come in and try new ways to answer the question? It's going to be interesting. The way that Bailey sees it right now, this kind of like read write function means that as the quote is, as we continue to evolve, it will require less and less effort from developers to capture the unique information about the way they operate their business. So overall, I think this is a fantastic direction for Stack overflow. They're leveraging pieces of the data that only they have access to.

9:33And I think we're going to see a lot of other companies that have these kind of question and answer forums, which are essentially deep sources of data, will have to monetize it one way or another. The blanket deals are one thing, but I think it's great if they're actually building tools and software that people can use and add extra context and data that the scrapers don't have access to. All right. Thank you so much for tuning into the podcast today. If you enjoyed the episode, make sure to leave us a rating and review wherever you get your episodes and make sure to check out AIbox.ai for all of the best AI models in one place on one platform$20 a month there's a link in the description I will catch you guys all in the next episode hey guys lady luck here are you going on any road trips this summer I know I'm going to be going on a bunch of road trips and being that I'm going to be passenger princess I love playing on spinquest.com spinquest has all of my favorite slot games, live blackjack, live craps, head on over to SpendQuest right now and get yourself a$30 coin pack for just$10.

10:32SpendQuest is a free-to-play social casino. Avoid where prohibited. Visit SpendQuest.com for more details. You're listening to a podcast right now. Driving, working out, walking the dog. If you're into podcasts, chances are you have something to say too. With RSS.com, starting your own podcast is free and easy. Upload an episode and we distribute it to Apple Podcasts, Spotify, Amazon Music, and more. Track your listeners, see where they're from, and start earning from ads just like this. If you've been thinking about starting a podcast, this is your sign. Start your new podcast for free today at rss.com.

From the publisher

The company is now supplying structured technical answers for training. Its moderation history is being used as a quality benchmark. Experts say this gives them leverage in the AI ecosystem.


See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

More from AI Hustle: Make Money with AI

All 178 episodes
Stack Overflow Becomes a Core AI Data SourceAI Hustle: Make Money with AI · 9 min
Listen in VO