#254 Prashanth: Why Developers Still Trust Stack Overflow in the Age of AI

14 May 2025 · 49 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye On A.I. Episode #254 Summary

Episode Overview Title: Why Developers Still Trust Stack Overflow in the Age of AI Host: Craig S. Smith Guest: Prashanth Chandrasekar, CEO of Stack Overflow Sponsor: Oracle Release Date: [Insert Release Date]

This episode discusses the adaptation of Stack Overflow in the era of Generative AI and the continued reliance of developers on this platform. Prashanth shares insights on how Stack Overflow is partnering with major AI providers to enhance its offerings while maintaining the integrity of its community.

Key Topics Discussed

  1. Prashanth’s Background and Journey
  2. Transition from software development to CEO of Stack Overflow.
  3. Emphasis on understanding technology company operations and scaling businesses.
  1. Stack Overflow vs. Other Platforms
  2. Unlike GitHub, Stack Overflow focuses on authoritative Q&A for tech information.
  3. Stack Overflow's unique model aims to solve the problem of providing trustworthy guidance rather than code collaboration.
  1. Community and Human-Curated Knowledge
  2. Stack Overflow hosts over 60 million Q&A pairs, forming a crucial knowledge base for developers.
  3. The community actively contributes and curates content to ensure high-quality information.
  1. Data Strategy and AI Training
  2. Stack Overflow's data is being licensed to major AI players like OpenAI and Google to improve model accuracy.
  3. The platform aims to solve issues like hallucinations in AI-generated code by providing high-quality data.
  1. OverflowAI and Enterprise Solutions
  2. Introduction of OverflowAI and Stack Overflow for Teams, which empowers enterprises with AI-driven tools.
  3. AI assistants can leverage Stack Overflow data to enhance developer productivity.
  1. Challenges with AI and Trust Issues
  2. Developers encounter a "complexity cliff" with AI tools, often returning to Stack Overflow for complex queries.
  3. Only 40% of Stack Overflow's community trusts AI-generated responses.
  1. Protecting Integrity and Quality
  2. Stack Overflow has strict policies to prevent AI-generated content from polluting its Q&A database.
  3. Licensing agreements ensure proper attribution and usage of Stack Overflow data by AI companies.
  1. Future Directions
  2. Ongoing adaptation to the changing landscape of AI and developer tools.
  3. Exploration of new features and enhanced user experience in both public and enterprise environments.

Key Takeaways

  • Stack Overflow remains a vital resource for developers, especially as they navigate AI-generated content.
  • The platform's commitment to quality and community-driven knowledge ensures it retains its trusted status.
  • Partnerships with AI companies are not intended to compete but to enhance the quality and utility of AI tools.
  • Continuous innovation and adaptation are key strategies as the technology landscape evolves.

Conclusion This episode serves as a critical examination of how Stack Overflow is adapting to the challenges posed by AI while maintaining its foundational mission of providing high-quality, trustworthy information to the developer community. Listeners gain insights into the strategic direction of Stack Overflow and its role in shaping future developer tools.

---

Stay Updated:

  • Craig Smith on X: [@craigss](https://x.com/craigss)
  • Eye on A.I. on X: [@EyeOn_AI](https://x.com/EyeOn_AI)

Listen to the full episode for more insights on the future of AI in developer tools!

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00At the moment we are not really looking to do something equivalent to what the LLM providers are doing I think what we were trying to do is to really assess the quality of our knowledge base and what does that mean. We will incorporate AI for our users ultimately in the context of making their lives a lot easier and productive on our side. The system is truly automated, completely automate all your work, especially in more complex tasks. I can totally see the foundational work, the simple questions, the simple programs, absolutely right. It makes a lot of sense for those to be solvable and through this new mode of this user interface.

0:33But ultimately, there is still a pretty large gap, especially for more advanced programming needs where you're absolutely going to need to know how to debug it. There's a growing expense eating into your company's profits. It's your cloud computing bill. You may have gotten a deal to start, but now the spend is sky high and increasing every year. What if you could cut your cloud bill in half and improve performance at the same time? Well, if you act by May 31st, Oracle Cloud Infrastructure can help you do just that. OCI is the next generation cloud designed for every workload where you can run any application, including any AI projects, faster and more securely for less.

1:21In fact, Oracle has a special promotion where you can cut your cloud bill in half when you switch to OCI. The savings are real. On average, OCI costs 50 % less for compute, 70 % less for storage, and 80 % less for networking. Join Skydance Animation and today's innovative AI tech companies who upgraded to OCI and saved. Offer only for new U.S. customers with a minimum financial commitment. but see if you qualify for half off at oracle.com slash IonAI. That's oracle.com slash IonAI, E-Y-E-O-N-A-I, all run together. oracle.com slash IonAI. Why don't we start, Prashant, by having you introduce yourself, give a little bit of your background, how you came to Stack Overflow and some background on what Stack Overflow is and how it got started.

2:28And then we'll talk about the future. Sounds great. Yeah, well, again, thank you for having me. It's a pleasure to speak with you. So my background is fairly straightforward. I have been in technology for a couple of decades. I started out as a software developer at companies like Nashville Semiconductor and Cabletron back in college, in fact. And then quickly realized that I was quite interested in understanding how tech companies worked and how they operated. So through a sort of a multidimensional career path in consulting, in banking, all focused on technology companies or technology functions or using technology in some fashion.

3:06Then sort of landed in my first operating role at a company called Rackspace, which is in the cloud services space. Sure. I spent about seven years at that company, CUNY, Texas. That's the reason I moved from New York City back in 2012. And that was a great experience learning how to really scale fast growing businesses, building great teams, and especially entrepreneurial or intrapreneurial endeavors. And that was an exciting thing to focus on. And then in 2019, after about seven years at the company, I was approached about this role at stack and the whole mission was um you know the company is obviously well known as a and i'm happy to explain what it is uh but well known as a public platform that serves you know millions of people around the world uh how do we transform it into an enterprise focused company was the the mandate so given some of what i had done uh at my prior company at rackspace the uh even though it was in a different industry some of the fundamentals uh translated so that's the reason i joined and know very very excited about what people have accomplished over the past uh six years at this company and i've been here uh you know through a multiple uh couple different sort of interesting transitions along the way and you alluded to one so which i'm happy to talk about yeah and my understanding of sac overflow is it's kind of like quora for coders and i mean and correct me if i'm wrong but you know you're you're you're working on something and you are looking for a solution or some piece of code that you can't put your fingers on.

4:44You go to Stack Overflow and, you know, what do I do? And somebody sends you a solution. It's sort of a collaborative platform in that sense and relatively live so that people aren't waiting overnight for an answer. How did it start? And I've always wondered why that function doesn't exist, or if it does exist, why Stack Overflow is dominated, why that function doesn't exist on, say, GitHub or another platform. I think it's absolutely, you know, it's been a, it's quite a journey since the company was founded back in 2008. I have to say that I think the comparison with other sites like the ones that you mentioned, I think they oftentimes do get included in the kind of broader family site, but we're quite distinct in that.

5:35When we were formed, we were formed because the founders wanted to create a really authoritative knowledge base for technology information. So when developers and technologists were writing code, as an example, they got truly, they could really get unstuck based on trustworthy guidance from the world's experts all around the world. So this is not a forum site, never been a forum site where people just have, you know, subjective conversations. We may open up new avenues, as I've described, to do more of that in the future. But the first, let's call it 16 years, 17 years of the company have been all around creating like this really accurate knowledge base of information, which is sort of the canonical answer to a question in technology.

6:17So, you know, that is, I would say, what's distinct about it. And, you know, over the past, again, as like I said, 16, 17 years, we have accumulated our community, thanks to their goodwill, have accumulated something like 60 million questions and answers on, again, every possible technology topic. and all that is very, very useful. Not only it has been useful in the past in the context of people just going to the site through, let's say, a Google search and finding an answer to that question when they're stuck, but also now in the context of leveraging this data for LLMs and agents and agentic AI inside companies, as I'll explain.

6:54So it's just been a really, really phenomenal exercise, a community exercise to bring all this human-generated expertise together and knowledge together to move humanity forward, which is, I think, which is amazing. The other question that you asked was around why hasn't this been, you know, what's special about this? Why hasn't this been replicated by other platforms, et cetera? I think it comes down to sort of the problem that you're trying to solve, ultimately. I think every product has a kind of a core problem that they're trying to solve. I think something like GitHub, and I can't speak for them completely, but I think that, but I can, from the outside in, obviously, they came from the code sharing and collaboration space.

7:34We've never been in the code collaboration space necessarily, but we have been in the technical knowledge sharing space. And so it's really, you know, it's a great combination between something like the real code that's stored on a place like GitHub or any of the other repositories, which is phenomenal for collaborating on code as a team, but also then leveraging all the context that is on Stack Overflow because it has all the the background information on why code needs to be written a certain way, what are the right packages to be using, et cetera, et cetera. So it's a very different problem that it's solving relative to sharing code.

8:11Now, both platforms over time can do all sorts of things to expand and so on. And I'm sure both companies will approach it that way. But the original nexus of it was in the two directions I described. Yeah. How live is the platform? Again, I'm not a user, so I always imagined that it was, you know, you ran into a problem. You could post a query and somebody around the world who's online and would see it. Or is it really a, you know, a knowledge base, a massive FAQ? Yeah, it's an astute question. It is a very live platform. So we have thousands of questions being asked literally on a weekly basis, which is great because we have millions of people all around the world that are not only being able to benefit from the collective goodwill of 16 years of contributions, but also they're able to answer questions that are open because they're experts on a certain topic and they get recognized for being experts.

9:17and that's you know amazing for not only their uh primary motivation which is to help their fellow developer and technologist out but as a secondary motivation there will be obviously recognition and that translates to all sorts of things like you know you get great jobs as a result of that you know we have even a partnership with indeed on the on our website where you're able to apply for jobs and so on um but the nature of the platform is again a very very strict q a uh uh form factor where it's not a discussion. Again, it's a very specific question and a very specific answer. Now, very recently, what we have announced is that that we're going to keep absolutely because it's very important to have that accurate data set in the canonical answers.

9:57But we are now experimenting with even faster engagement modes in other parts of the site. So distinct from Q &A, we have opened up, for example, an ability for users to chat real time with each other, right? And so we have opened up that highway, if you will, which has been exciting to watch as we have as we've exposed that ironically that that feature actually has existed for 16 years but has been buried deep in the site for moderators as an example to be able to sort things out when things happen on the site so but we realized that some of the newer generation folks love to have instant access or directional guidance versus you know perhaps the the canonical answer to get formed you know in pristine fashion so so we will experiment with other content types, even maybe discussions, which is another feature that we have, which is a little bit not as strict as Q &A, let's say.

10:47But they've got their own sort of lanes, distinct lanes, so we don't pollute the site. And the idea is to have high-quality information, but very community-oriented type of interactions. Yeah. And who decides what's canonical? I mean, who's the pope, so to speak? Yeah. Indeed. So very much the community. I think this is what's so fascinating about this site is that this is a very much a self-reinforcing. The experts in the world have all sort of participated, are participating and contributing to the site. And the community votes up or votes down answers based on those actually being effective or not effective in getting people unstuck.

11:35And so, funnily enough, it works really, really well because the intent behind the site and the rule sets that have been established governs the site in a way that rewards quality. It absolutely does the opposite for maybe even a little bit more harshly than people would like. But it is definitely sort of a school of hard knocks when it comes to high-quality canonical answers to questions. and the hang on just a sec here the I can see this this is I mean two things it's the live chat would be wonderful because things are changing so fast I imagine it's hard to keep up with with the question and answer side but also as a knowledge base for training these coding agents and that sort of thing.

12:40Are you doing that? Are you building your own agent or are you licensing the content to people that are building agents? Yeah, a little bit of all of it, actually. So we back in 2023, what we started noticing is that when all the LLM providers started announcing their developments, you know, we noticed obviously our data had been used by most of them. So we, if not all of them, because, you know, we are one of the few high quality again, speaking about this canonical knowledge base. We're one of the few cases that has this high quality information set, which is, you know, quite very, very useful.

13:24and what we did was we actually first took an open source model and fine-tuned it with our own data and we noticed something like a 30 percentage point improvement in accuracy as we know llm's continue to hallucinate they hallucinate less these days you know probably because we have done all these uh licensing arrangements but but generally speaking uh they do hallucinate so but the quality goes up dramatically when you use our data so when we saw that and we of course put up things like anti-scrapers and you know we prevented people from you know grabbing the latest information or even the historical information uh we started engaging with all the LOM providers the cloud providers etc and uh over the past 18 months or so have constructed a significant number of partnerships with everybody that you can imagine in the ecosystem so everybody from the open AIs and the Googles to all the cloud providers and more uh you know we have struck partnership deals and these are not only licensing deals let's be very clear but they're also bi-directional in nature so yes we want to get paid for the data so we can invest uh in back into our community platform so that that you know the world can keep going around around that way but also because we're solving real problems for these llm companies because number one uh you have uh you know what we call lm brain brain but if you don't actually have humans creating new knowledge with you know from their minds and you know documenting it then you know how are these lm's going to obviously train uh and that's important even in a world of synthetic data you're going to need human generated uh you know knowledge that's number one number two is uh when oftentimes i'm sure craig you tried it you mentioned rabbit as an example but there is a complexity cliff with these ai uh chat bots or assistants or you know lm models at some point they're going to tap out on complexity and it's just going to say, here's the best I can do.

15:15And you know, it's on you, then you're on your own. And so when that happens, what we're trying to help, you know, Gen AI and LLM companies with is, hey, why don't you actually engage our community when in that moment for the user? Why can't the user actually ask the question straight on Stack Overflow straight from there? And then by the way, we have a legal requirement for them to attribute our content every time our content is used in a black box answer as an example, which I'll get to here in point number three. But to be able to answer that question, again, generates new content on our website, and that knowledge can then get curated and added and pruned, et cetera, and obviously community curated.

15:58And that allows them to leverage all that data for future model training needs and other agentic needs over time. So that's the second issue, this AI, you know, answers are not really knowledge all the time. And so we want to help them. And the third issue is around trust, Craig, because around the hallucinations, only about 40 % of our community trust what's coming out of AI models. So we have now made sure that all these partnerships require them to attribute our data and links to our original answers every time an answer is used in the context of an LLM answer. So those are the three big problems um along with these with these licensing agreements yeah uh on the on the um link uh well first of all on the model that you guys fine-tuned with stock overflow are you pursuing that is that a product that that's going to become better and more comprehensive?

17:01I mean, are you competing with OpenAI and some of the others on code generation? I mean, you know, Sam Altman's been out or Dario has been out talking about how, you know, these systems are going to write complete code very soon. Yeah. Yeah, no, we don't. At the moment, we are not really looking to do something equivalent to what the LLM providers are doing. I think what we were trying to do is to really assess the quality of our knowledge base and what does that mean. We will incorporate AI for our users, ultimately, in the context of making their lives a lot easier and productive on our site.

17:44Right. So as an example, you know, we've integrated Google Gemini into our public platform and in our enterprise product called Stack Overflow for Teams, which is the private version of Stack Overflow that companies use. and we have thousands of enterprises using that product, where Google Gemini powers our public platform and provides assistance when users are trying to ask questions on the site. And Gemini gives them very friendly recommendations in the privacy of their Gen AI experience versus having to do all that in public in the context of your other subject matter experts around the world, which is, by the way, addressing a historical concern that people have had about our site, where it's a little bit, you know, it can be quite an experience for a newcomer especially.

18:28So this allows, you know, a newcomer to feel a lot more welcome on the site. And then on the enterprise side, we have integrated our content into all parts of the workflow. So into the IDE, in Visual Studio Code, into Slack, into Microsoft Teams. So you can ask natural language questions at any of these places, along with, of course, on the platform, and to be able to get Gen.AI summarized answers and real assistance. And we will take both these concepts, you know, we'll keep improving these on both sides, on the public as well as the private. And the private one runs, by the way, on Microsoft Azure.

19:00So we work with everybody and, you know, we think of ourselves as sort of Switzerland. And then over time, there may be other opportunities to help our users with more assistance on the site. We even launched about in 2023, AI search and material functionality on our public website and gave the ability of what we call overflow AI, where people are able to ask a question in natural language and get answers straight from the corpus of 16 million questions, but at Gen AI summarized answer. So we will continue to experiment with things like that as the industry moves very rapidly. And again, we're watching everything, including open source, which is quite fascinating to observe.

19:40So you will see more innovations from us over time. Yeah. One of the things you mentioned, the integration with some of these code generation platforms or putting it on Visual Studio, that's for programmers as they're working, for example, in Visual Studio. And they have a question and they can, without having to go to another window or something, they can immediately tap SAC Overflow's knowledge base. But increasingly, as these agents are writing code, will the agent be able to access that knowledge? Yes. Itself. Yeah. Yeah, great question. So what is fascinating is that we almost see in the context of our data, there are two big opportunities for us.

20:44One is all the LLM pre-training examples that I just gave you with all the big LLM companies, the cloud hyperscalers and everybody that we've got a relationship with there. But more recently, what we have with the explosion, of course, agentic AI as a topic in the industry and very much the pursuit of the ecosystem to create these agents, especially in companies so that they can assist every employee to be able to do their work a lot more effectively and productively to free up a lot of their capacity to other things. we see an absolute opportunity for our content, both our public platform content, our knowledge on Stack Overflow, plus the private Stack Overflow knowledge base with our enterprise product, which is again within thousands of companies where they have a proprietary Stack Overflow within their company to share information.

21:32We are noticing, Craig, multiple examples where companies like Uber and a whole bunch of other super, you know, kind of, I would say, forward-thinking companies building AI assistants or co-pilots or you know you can call them agents as we move into this next world of you know creating really sort of performing action and they're able to tap into the knowledge store of stack overflow for teams and now the public stack overflow knowledge to be able to empower those agents to be really really effective and accurate in the moment where the developer is trying to you know let's say delegate tasks to that agent to perform things.

22:09So as we know, whether it's LLMs or AI agents, they're all very dependent on high quality knowledge to be able to either provide information, search results or perform actions on your behalf. And so that is why I think we believe this is a tremendous opportunity for us, both for our enterprise business as well as our public platform. Yeah, yeah. Just as an example, again, I'm not a programmer, But I'm eager for these multi-agent systems to work, and I've got access to a couple of them. And, you know, if I try and write a simple piece of code, you know, for myself, inevitably, when it runs, I get a bunch of errors.

23:02and uh you know the i'm sure what many other people are doing as well then you copy the error back into the system or you know into a chatbot or whatever it is and it gives you a solution you copy that back and you you know you end up uh going down a rabbit hole it's like a house of mirrors you don't you pretty soon you're lost uh could the integrations with uh multi-agent systems address that where where the system i don't i don't know i assume um they're accessing a knowledge base in the weights of an llm to solve you know to answer a question or i mean to solve an error.

23:58But I don't know whether it's just the inherent probabilistic hallucinations of the models. Anyway, it just doesn't seem to be very effective. But are you working on that where an agent, if it runs into a problem, then could refer to SAC overflow and get a solution that'll actually work. 100%. I mean, this is exactly what I meant when I said the AI, the Gen AI tools that people are beginning to leverage and testing and piloting and so on do hit a complexity cliff. Right. Like your example, I have also experimented with plenty of coding Gen AI tools. And when you do actually hit very much a sort of a cliff where it's able it's really not able to keep all the context that you have in mind and it's able to it's you know it has to do with you know a data scientist can explain all the reasons why our machine learning scientist machine learning expert but the the point being that there's still a dependence on high quality knowledge and it's not yet where the the kind of the system is truly automated but can completely automate all your work especially in more complex tasks i can totally see the foundational work, the simple questions, the simple programs, absolutely, right?

25:23It makes a lot of sense for those to be solvable and through this new mode of this user interface. But ultimately, there is still a pretty large gap, especially for more advanced programming needs, where, you know, you're absolutely going to need to know how to debug it. So oftentimes, we do see people come back to Stack Overflow after they hit the complexity cliff, which is why I was explaining previously why don't we get users to be able to ask questions straight from that that point of uh you know kind of issue that they have even in the gen ai tool straight on stack over so they complete the task and you know we we help the gen ai user so you know absolutely you know and over time things will get better there's no question because you know that's just you know progress um and you know we will just want to make sure that we're adding value to the developer what we are trying to do is focus on the developer add value to them uh and two additional concepts that are important here is one is going wherever the user is.

26:17And so our fundamental thesis is to make sure that we serve our developers wherever they are. And we call this concept knowledge as a service because we just want to provide Stackpole for knowledge, public or a private Stackpole for knowledge, wherever the developer is in their workflow, whether that's in a coding Geni tool, that's in an IDE, whether that's in Slack or Microsoft Teams or anywhere else in your company or what you're using. And then secondly, we want to make sure that we are really sort of serving these users in an authentic fashion where they are getting access to true experts, human experts, when they need them, that they can trust this information and that the world's best JavaScript programmer, either on the public platform or your best JavaScript programmer within your company is able to validate the answers that are coming out of these uh these ai ai you know answers from these lms and engineering tools so yeah just as you're talking the gemini integration where that's um if you're if you're working with a code generation system you can go to Stack Overflow and can you copy code into Gemini and ask it to query Stack Overflow about what to do in this situation or do you have to explain it in natural language what you're trying to do?

27:54You can do either. I think that if you look at the Gemini experience, so now if you go to Google Gemini Code Assist, which is their code editor you can you can type in you know how do i do xyz on kubernetes and it's going to provide a gen ai summarized answer and it's going to have links back to stack goal flow underneath it and you can click those links come to the you know the canonical answer that has been generated by an expert you know god knows when uh and uh that is then going to give you confidence that you know that this answer is in fact accurate so absolutely you could paste some things into Gemini and you know that Gemini all the gen ai experiences are similar i'm not sure about you you can probably state very specifically hey go look at stack of flow and find out but i don't think you even need to do that because it's just it's going to generate the links automatically because that's what the train the data is trained on uh the other thing i was going to mention is that in the same spirit of uh going where the developer is many years ago craig you know we always thought about google search as or you know any search engine as the user interface for stack overflow 99 of folks went to a google search when they had an issue they would type in a question and it would land on stack overflow because the answer was always accurate now the user interface has changed where increasingly people seem to leverage not only a search engine but also these gen ai tools as their primary mode of engagement onto the web and so really what's happened is that people change their mode of uh you know how they ask questions and in looking for answers just more broadly, I think search itself, I think is sort of a, some sort of a different, it is changing, obviously, as we know.

29:32And that is what we are trying to adapt to. We're trying to adapt to the fact that previously we were linked to the user interface being Google. Now, if the user interface is all these Gen AI tools, great. We will also then make sure our data and our knowledge is provided in all these Gen AI tools. So we want to just move with the times. Yeah. You, on, on stock overflow for teams,

29:59you mentioned a private stock overflow. I mean, how does that work? Is that, yeah, how does that work? Yeah, yeah, absolutely. So back in 2018, this is about six months or about a year before I joined the company, there were some early signs that large companies really loved Stack Overflow because you can obviously get this community curated knowledge, upvotes and downvotes and gamification and all these things, a really effective way to keep knowledge up to date and accurate and with a communal sense. And a lot of folks wanted to leverage this internally within their companies because they wanted all that data and knowledge to be proprietary within their own company versus it showing up on the public web.

30:44And so the company launched Stack Overflow for Teams, a private enterprise version of Stack Overflow for companies to use for knowledge sharing within their companies. And that's effectively what it was. And companies like Microsoft and Apple and all these sort of big companies that were leveraging the platform. And the goal was effectively how do we make this sort of the default knowledge base for all technologists within companies. And so over the next several years, we very much went on that drive. And so, you know, we now are used in pretty much every financial institution that you can think of.

31:21All the banks, all the retail companies, all online retailers, the big tech companies certainly all use Stack Overflow for teams as a private knowledge repository for their technologists' information. and now that is being surfaced as I mentioned previously in the case of Uber as one example where they're leveraging for their AI co-pilots or AI assistants and those co-pilots are now leveraging the data that's been created. Thousands of questions inside Uber are now being leveraged in an agentic sense inside the company. Yeah, and this private knowledge that's held within Stack Overflow for teams is you also have access to the public stack overflow.

32:10That's correct. That is correct. Absolutely. It's part of your search. When people use the product, absolutely. How do I do XYZ on Kubernetes example again? So you're going to get from the 60 million questions and answers, the answers on that topic, which are more public. And then you're going to get, in addition to that, anything that's specific to your company around Kubernetes for the new company. you're going to get both results they're very clearly delineated as this is a private versus a public answer and so on uh but really really a powerful mechanism uh for for users yeah and uh another question i had uh how is i mean the the answers that are generated or that are that are given on on stock overflow by uh experts uh are are you seeing anyone using uh an llm to write those those answers and is that a danger that that you know is the more uh gen ai content that gets into the knowledge base, it'll cause some sort of problem.

33:28I'm not sure what. Absolutely. We very explicitly prevent that from happening. And back in late 2022, our community very correctly basically said, look, we can't just generate. We saw a massive spike actually in content on the site when people just started using Gen AI tools and tools like ChatGPT. and a lot of it was incorrect because of the you know the hallucination issue and it sounds really good but actually is incorrect and given like the the goal of the q a part which is to create this canonical you know human curated knowledge base it was literally you know the opposite of that so that so we very quickly the community and the company we both together banned all that content saying look you can't just use you know chat gpt or any gen ai tools to generate answers but let's let's go speak, you know, let's go understand how we can use this phenomenal technology to solve real problems.

34:22So we've spent the next year or so talking to our community, figuring out that they don't trust what's coming out, 40 % only of them trusted, but 70 % of them wanted to use Genia tools in some fashion, which is interesting. And so we incorporated Genia into the public platform, as I described, to provide assistance on questions in the enterprise product. And now we are also experimenting with answers using AI but we're trying to make sure that we don't do it in the sense of Q &A because Q &A is a very pristine sort of knowledge base we do not want to pollute it with AI slop so to speak if you know the internet is now polluted with a lot of just you know AI generated content and we do not want to be one of those sites where that happens there may be very specific use cases in the other lanes as I was describing to you in chat or discussions or maybe even other challenges and other things we do on the site in multiple content types but outside of q a because we want to again protect the quality of that q a base uh you know very very specifically yeah um the other thing is you've in this this q a knowledge base you've got i think you said 16 million 60 zero 60 million uh question answer pairs uh i've talked to uh people about breaking data silos and you know uh rationalizing and ontology and all that stuff uh and and when you start breaking silos you find out that you've got contradictions.

36:01Have you run any Gen AI systems over the entire knowledge base to make sure that it's consistent, that there aren't contradictions between answers? Yeah, I think that this system in itself was, I think, really brilliant about what the founders of this company back in 2008 set up was the very very specific rule set to account for all of that right so duplicates is a kind of an inside joke within stack gold flow so if your question is you know it's been asked before maybe in a different form factor it's going to get called out pretty quickly right and it's it is a school of hard knocks for that reason and so the standards very very very high correct so i think i think almost to a place where it is a little bit tough for new users or has historically been until our new gen ai incorporation as i described to you uh but but it does account for uh the fact that you know within the programming knowledge base so stack oflow itself it's sort of about half of our traffic half of our content uh and that's one out of 185 sites that we have so the other 184 sites are what we call stack exchange sites uh and these are you know other sort of technology topics so it includes artificial intelligence server you know system administration sort of topics devops topics and so on security topics and so yes those are more specialized communities where there may be an intersection but we have kept them somewhat apart but SAPOIO on its own is huge because you know it's got every possible programming knowledge every cloud technology it's the massive you know a part of the technology knowledge base uh but you know the other ones i don't think there's like there is uh we're working on something called unified search across all of the stack exchange across 185 uh but they've been you know somewhat distinct over time but you know there's the core stack of the knowledge base itself is quite cross-functional obviously because given the depth of the the knowledge base there and you know i think the system that's been set up uh prevents for the issues that you're describing yeah um Something else occurred to me while you were answering that.

Read the full transcript

38:10Is there a concern that this knowledge base will be stolen by some agentic system that could be pinging? It wouldn't take very long to download 60 million answers. Yeah, I think the way we think about that is back in 2022, 2023, when we noticed that it had in fact been, let's say, leveraged by a lot of the LLM providers. We took precautions to then sort of set up all the anti-scrapers, as I described. We want to stay open and make sure our community has access to all this data, because that's always been kind of our ethos, which is we're by the community for the community, so they should get access to all the contributions that they provide uh we're governed by something called creative commons um and so in that so we want to strike this balance of where our community users you know have to sort of sign up to get access to the data and tell us who they are so you know we know they're not coming from a company as an example but if you're a company if you're a tech company ai company lllm company ai agent company and so on uh you've got to work with us officially, because I think we really can solve your problems as long as you work with us through sort of official manner and licenses content in a real way.

39:37So yeah, we'll take defensive precautions like the anti-scrapers and rule sets and all that kind of stuff. But we're really trying to say, and the good news is that most companies have embraced it. That's the good news. That's we've now got relationship pretty much with every big AI and LLM and cloud provider that you can imagine and there's plenty more that we're working on. There's a long tail of companies that are interested in engaging with us. Yeah. The licensing, this data to code generation systems, are there any metrics? Can you see the improvement in code generation once they have included stack overflow in their training data?

40:19Very much. I think that if you look at some of the, let's call them league tables or the benchmark tables that are coming out on a weekly basis, the companies that work with us are at the top of the list, especially on coding benchmarks, right? Because by the way, also as a side point, Craig, the data that we have is not only used for, is really useful for coding related AI tasks, but also non-coding related AI tasks, because the depth of the iceberg under the water, so to speak of all the comment history with the underneath an answer the curation of downloading an older answer and a new you know more elegant answer to a question all this all the things that have happened on on that very specific one primitive of q a uh in amongst 60 million of them that is really really useful for things like reasoning and really sort of being able to drive logic so um so yes we do absolutely see the the performance uh on the high end of the table associated with us being associated with them.

41:17So that's heartening to see. And we have no doubt of that, only because when we continue to run these tests, we even have a technical paper out called Stack Eval, which is like a score that we actually have around various AI models. And when they use our data versus not, and so we can see the big differences between using it or not using it. Yeah. But with these large models absorbing absorbing all of this knowledge and curating their training data more and more. Are you concerned that the day will come when Stack Overflow is no longer unique, that that data will have leaked out into the public sphere, into these models?

42:10well i think that the amount of new firstly there are two things one is to be able to get access to the data you are uh you know these are multi-year agreements that are recurring in nature so the only way you are able to use this data for future model training needs is if you work with us right so if you don't work with us then you can't use it going forward that's basically the kind of the arrangement uh and number two is that obviously we're generating new information so So if you want the new information to continuously to give you the latest and greatest of what's happening in the ecosystem, again, with human thought and judgment, then you should, there's a second reason for you to work with us.

42:49So I think both one is more of a kind of a fundamental legal requirement to say, you've got to work with us if you want ongoing access and do that on a yearly basis. And then secondly, if you need all the new information, you should also work with us for that. is it yeah uh okay we're coming up to the hour um is there anything i haven't asked that you talked about uh at human x or that you you want listeners to know yeah yeah i think i think that maybe one thing we haven't uh spent time on a little bit is if you think about uh you know the ecosystem of ai players the big llm providers and the of course they're all after the whole you know, from creating like the industry leading LLM that's specialized or just more generic.

43:38And then you've got the, you know, the long tail of AI companies that are all vying for this agentic AI future, especially within companies. And I think what we have noticed is that our, every one of our business lines, Craig, and let me just explain that, which is, we have obviously our public platform, which we've spent a bunch of time on, which is our 60 million questions and answers with millions of people asking questions and then we have the enterprise business stack over for teams like we just described it is used by thousands of enterprises for inside that company for now ai agent degrees and then we have an advertising business because we have uh you know uh the fact that we we have a lot of traffic and folks we are able to showcase their products and then we of course have our data licensing business uh that that we've already spent time on.

44:25So in every one of these, I think historically people were, you know, back in 2022, 2023, I think there's this question about, okay, we need to change our business model. The model of the internet has changed. We've got to change with that. People are no longer just relying on advertising as the sole source. So data quality, et cetera, becomes really, very important. So that's what we've been focused on. But also what's interesting is that a lot of these AI agent companies, what we noticed is they really want not only our data and to be able to build great AI agents for the enterprise, but also they want access to our distribution.

44:59You know, the fact that we have millions of people that can try out their products, including the ones that you're trying out. Imagine you have access to the developers of the world to be able to showcase your product on Stack Overflow, to say, here's the leading AI agent for XYZ. And, you know, it's a huge driver of, by the way, really, really powerful demand gen driver for cloud companies who advertise with us. All the big companies do. Microsoft and the Amazons and the Googles is an example. And so it's a great place for AI agents, agentic companies to be able to get distribution to the developer ecosystem.

45:37So that's the other aspect we didn't spend a bunch of time on. So we literally are, you know, we see this as a tremendous opportunity for the company because literally all parts of our company can be useful and moving humanity forward with, you know, being able to access all this AI power in their favor. Yeah. So the Gen AI future, you guys are embracing it. I mean, a lot of people are afraid of it because, as I said, it's going to make a lot of operations obsolete. But I think it's important what you were saying about the question and answer, that pristine knowledge base. we need human knowledge that isn't polluted by you know agentic systems or generative AI answers that that is kind of the ground truth and then indeed build on that yeah absolutely yeah and I think you know this is you know as I often say welcome to technology right there's it's nothing but change and progress so companies absolutely have to adapt you know we've had to adapt over the past couple years uh to this new future and we're still making our way and it's been fun to do all the things that we've all these partnerships that we've struck as i mentioned um so it just i think it's going to open up plenty of opportunities and i think they know kind of i would not be overly pessimistic about you know it's probably going to have some sort of uh you know i would say it's it's probably harsh in the near term and good in the long term because you know it will unlock a lot more potential for people to be a lot more productive I think more companies will get created because you know again there's an infinite amount of technology to go you know be written and innovation will keep happening so I think if if people have an open mindset about it this could be a very very good thing in the long run I think near term you're right I mean it probably does create a lot of disruption for for teams and uh and hiring plans and because you know ultimately this the CFOs of the world are going to have to ask okay I've invested a lot in this technology, what am I getting out of it?

47:46So I think that's the reckoning for that will happen pretty soon. There's a growing expense eating into your company's profits. It's your cloud computing bill. You may have gotten a deal to start, but now the spend is sky high and increasing every year. What if you could cut your cloud bill in half and improve performance at the same time? Well, if you act by May 31st, Oracle Cloud Infrastructure can help you do just that. OCI is the next generation cloud designed for every workload, where you can run any application, including any AI projects, faster and more securely for less. In fact, Oracle has a special promotion where you can cut your cloud bill in half when you switch to OCI.

48:35The savings are real. On average, OCI costs 50 % less for compute, 70 % less for storage, and 80 % less for networking. Join Skydance Animation and today's innovative AI tech companies who upgraded to OCI and saved. Offer only for new U.S. customers with a minimum financial commitment. See if you qualify for half off at oracle.com slash IonAI. That's oracle.com slash IonAI. E-Y-E-O-N-A-I. All run together. oracle.com slash IonAI.

From the publisher

This episode is sponsored by Oracle. OCI is the next-generation cloud designed for every workload – where you can run any application, including any AI projects, faster and more securely for less. On average, OCI costs 50% less for compute, 70% less for storage, and 80% less for networking. Join Modal, Skydance Animation, and today’s innovative AI tech companies who upgraded to OCI…and saved. 

 

Offer only for new US customers with a minimum financial commitment. See if you qualify for half off at http://oracle.com/eyeonai 



In this episode of Eye on AI, host Craig Smith speaks with Prashanth Chandrasekar, CEO of Stack Overflow, about how one of the internet’s most trusted platforms for developers is adapting to the era of generative AI. With over 60 million human-curated Q&A pairs, Stack Overflow is now at the center of AI development — not as a competitor to large language models like ChatGPT, but as a foundational knowledge base that powers them.

 

Prashanth breaks down how Stack Overflow is partnering with OpenAI, Google, and other LLM providers to license its data and improve AI accuracy, while also protecting the integrity of its community. He explains the rise of OverflowAI, how Stack Overflow for Teams is fueling enterprise-grade co-pilots, and why developers still rely on expert human input when AI hits its “complexity cliff.” The conversation covers everything from hallucination problems and trust issues in AI-generated code to the monetization of developer data and the evolving interface of the web.

 

If you want to understand the future of developer tools, AI coding assistants, and how human knowledge will coexist with autonomous agents, this episode is a must-listen.

 

Subscribe for more deep dives into how AI is reshaping the world of software, enterprise, and innovation.



Stay Updated:

Craig Smith on X:https://x.com/craigss

Eye on A.I. on X: https://x.com/EyeOn_AI



(00:00) Intro

(02:31) Prashanth’s Journey from Developer to CEO  

(05:18) Why Stack Overflow is Different from GitHub  

(08:51) The Power of Community and Human-Curated Knowledge  

(12:53) Stack Overflow’s Data Strategy for AI Training  

(17:26) Why Stack Overflow Isn’t Competing with OpenAI  

(20:36) How Stack Overflow Powers Enterprise AI Agents  

(26:13) OverflowAI, Gemini, and the Future of Dev Workflows  

(30:09) Inside Stack Overflow for Teams  

(33:29) Safeguarding Quality: The Fight Against AI Slop  

(38:32) Licensing, Attribution, and Protecting the Knowledge Base  

(43:19) Business Strategy in the Age of Generative AI  

More from Eye On A.I.

All 266 episodes
#254 Prashanth: Why Developers Still Trust Stack Overflow in the Age of AIEye On A.I. · 49 min
Listen in VO