In short
AI text detection reliability, why most detectors fail, and Pangram’s approach to producing a low false-positive “trusted” AI detector; how to interpret results and how it’s used in education/publishing.
Guests
Max Spiro, CEO of Pangram. Background: AI/ML founder (with co-founder) who built Pangram from research; focuses on machine-learning methods (active learning, synthetic “mirrors,” interpretability) and benchmarking with researchers.
Key claims
Pangram’s false-positive rate improved from ~1 in 1,000 to ~1 in 10,000. It uses active learning with “synthetic mirrors” (AI-written versions of human texts) to learn micro “AI decisions,” not perplexity. Old detectors using perplexity misfire on memorized text (e.g., Declaration of Independence) and English-learner writing. Pangram outputs confidence and estimates AI percentage by scoring document parts.
Notable examples
Commonwealth Prize story “The Serpent in the Grove” (Pangram said 100% AI; author claimed voice-to-text). Claude/ChatGPT self-checks described as unreliable.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOListener Call-In on AI Detectors
0:45 to 1:30
A listener shares concerns about AI detectors' reliability
“For the longest time, this has been the refrain.”
News Brief: Updates from The Verge
1:30 to 3:43
Overview of current tech news including OnePlus and EU regulations
“This is 90 seconds on The Verge for Thursday, July 16th, 2026.”
News Brief: Updates from The Verge
3:48 to 4:28
Overview of current tech news including OnePlus and EU regulations
“Support for the show comes from MongoDB.”
Interview with Max Spiro on Pangram
4:28 to 14:01
Discussion on Pangram's AI detection technology and reliability
“All right, we have Max Spiro, CEO of Pangram.”
Understanding AI Detectors and Perplexity
14:01 to 18:39
Learn how AI detectors like Pangram assess text and the limitations of perplexity as a metric.
“So you can build an AI detector that measures perplexity of a text and says if it's low perplexity, it's AI.”
Understanding AI Detectors and Perplexity
20:27 to 21:58
Learn how AI detectors like Pangram assess text and the limitations of perplexity as a metric.
“Support for the show comes from MongoDB.”
Skepticism in AI Detection
21:59 to 28:00
Discuss the controversy surrounding AI-generated content and the implications for writers and educators.
“Often the response is, well, if you change the legal foundation here, we will leave.”
Addressing AI Use in Education
28:00 to 29:25
Learn strategies for educators to handle AI-generated work from students.
“If I'm a professor, what should I do if I get a positive result?”
The Impact of AI on Professional Communication
29:25 to 30:33
Discover the implications of AI use on professional writing and communication skills.
“I see this learned helplessness, not just among students, but even among professionals who feel like, oh, AI could write a better email than me.”
Evolving AI Detection Techniques
30:33 to 32:34
Unpack the challenges and developments in AI detection methodologies.
“I'm curious whether your job gets harder as time goes on because these models change all the time.”
Show all 13 chapters
Ensuring Authentic Human Writing
32:34 to 34:34
Explore the importance of sourcing authentic human writing in AI detection.
“We've talked about having an essay contest or something, but we have to supervise you to make sure that what people are writing are legit.”
Building a Sustainable AI Detection Business
34:34 to 36:46
Understand the business aspects of developing AI detection technologies.
“We were chatting before we started recording, and you mentioned that there might be some advancements coming soon.”
The Role of AI in Writing
36:46 to 40:04
Discuss the ethical implications and the future of AI-generated content.
“Then, you know, my work here is done and the problem is solved.”
Transcript
Automatic transcript. May contain errors.0:02Hello and welcome to The Verge Cast, the flagship podcast of homegrown human writing. I'm Jake Kastronakis, Executive Editor of The Verge, and today we're talking about AI detection and the one system that might actually work. I have been on the hunt for a reliable AI text detector for a while now, and I know I'm not alone. Here's a call we recently got from a listener. Hey, my name is Aiden. What do AI plagiarism or AI text detectors measure?
0:30Nilay Patel:Are they reliable? Can they reliably detect AI-generated text? Recently, there was an article at my college newspaper, and it's 100 % AI-generated according to ChatGBT. I told the publication, and they refused to take it down, stating that AI detectors are not reliable. Thank you. Bye. For the longest time, this has been the refrain. AI detectors aren't reliable. So maybe a student's paper or an executive's LinkedIn post looked like AI, but there wasn't a surefire way of knowing. That might be changing, because now the thing I keep hearing is AI detectors aren't very reliable, but Pangram says it might be AI.
1:14So today, we're talking to Max Spiro, the CEO of Pangram, which makes what might be the first trusted AI text detector on the market. We're going to talk about how it works, how much we can trust it, and what we should do with its findings. But first, here's what's happening on The Verge today. This is 90 seconds on The Verge for Thursday, July 16th, 2026. OnePlus is exiting the US and Europe. The company made the announcement today, 12 years after first making a splash with the OnePlus One. This is a real bummer for smartphone fans. OnePlus had its ups and downs, but the company genuinely was a pioneer in low-cost, high-spec devices that could go head-to-head with the big flagships.
1:53David Amell has a great piece on The Verge today about how the U.S. carrier system is a big part of what killed OnePlus. Its prices may have been great, but they never looked that great, beside an iPhone that only cost$4 a month on contract. Next, the EU is forcing Google to make Android open up more in Europe. The European Commission said today that competing AI assistants need to get the same level of access as Gemini. That means letting them be activated by voice commands and giving them the ability to control apps. Google argues that this presents security and privacy risks, but as of now, it's on the hook to make it happen by July 2027.
2:29The EU is also updating rules requiring Google to share search data with competitors in Europe. That now has to include AI chatbots. Finally, do you want a couple companies to be able to dominate U.S. airwaves? FCC Chairman Brendan Carr does. He's planning a vote to end the national ownership cap, which currently prevents broadcasters from reaching more than 39 % of U.S. households. Instead, he wants to be able to review and approve deals that would violate the cap on a case-by-case basis, letting the FCC consider factors such as, quote, viewpoint diversity. Huh. Fun fact, the 39 % cap is, in fact, enshrined in law.
3:04So this is definitely going to end up in court. Another win from the Carr FCC. You can read more at TheVerge.com. That's 90 Seconds on The Verge for Thursday, July 16th, 2026.
3:40Nilay Patel:it's about to do so you stay in control. To put AI to work for people, visit servicenow.com.
3:50Nilay Patel:Support for the show comes from MongoDB. AI-assisted and agentic coding is helping you build faster than ever. But if your data layer is still a bottleneck, what's the point? Instead of wrestling with rigid schemas or translating data formats, MongoDB's native data model mirrors the language LLMs already speak. It ships at the speed of AI, is ACID compliant, and scales to handle massive Fortune 500 workloads. Developers have a word for that kind of reliability. Actually, five words. It's a great frickin' database. Start building at mongodb.com slash AI. All right, we have Max Spiro, CEO of Pangram.
4:34Max, thanks so much for joining us. I'm really interested in talking about what you guys have been building. Yeah, thanks so much for having me. Really excited. So there have been these AI authorship debates for a couple of years now. And earlier this year, I started to notice them playing out a little bit differently. I feel like it was always people saying AI detectors aren't reliable, they aren't reliable. And then suddenly people were saying, well, Pangram says this is AI. And this tool had emerged as a name that people actually trusted. So my question to you is, what happened there? And how did Pangram, at least from my perspective, so quickly get this reputation as a trustworthy name in AI detection?
5:15It's funny. For me, it doesn't feel at all like we gained it quickly. So I think we've been around for almost three years now. but I think for the first two-ish years of our life we were just fighting this narrative on AI detectors don't work we were publishing papers, publishing technical reports and then I think we slowly were embedded more in the research community and then we had a couple big name researchers who decided to benchmark Pangram and publish results and I think that's when people started to realize like oh these guys are legit. They're not just like lying about their accuracy. They're real.
5:59So did you start out as purely a research product before productizing it? Or at what point did you get the model that kind of cracked it and went, oh, this is reliable enough that we kind of want to brag about it? Yeah. I mean, definitely it was like just research at the start. It was just me and my co-founder, we both have like AI and machine learning backgrounds. And we're just like, this sounds like a fun problem to try and crack. It seems like nobody has really cracked it to the degree of accuracy that people need or want. And so it took us probably, I think, a bit over a year to get our first version of the model that really had this flagship false positive rate that was significantly lower than everyone else.
6:47Early on, it was a 1 in 1 ,000 false positive rate. So 0.1%. And now today it's one in 10 ,000, which is 0.01%. And I think that's like confidently, it's low enough that people are able to confidently point at PANGRAM results and be like, oh, this is what it says. So why don't we back up to that? What is the approach that you took that got you to that one in 10 ,000 false positive rate that you guys are advertising? So it's a method of machine learning called active learning. essentially what happens is we take a model that's okay, it's decent, and then we say, scan this really, really large corpus of human written text and find out which examples we have errors on.
7:36And what this essentially does is it finds examples that are close to the boundary between human and AI. And then we take those, and then we say for each of these documents, say it's like a Yelp review on Denny's. Then we'll ask AI to also create a Yelp review about Denny's in the same style. And then so now we have a human side and then an AI synthetic mirror. And so we train on these, the human example plus the AI synthetic mirror. And our model is able to learn the difference in stylistic choices between these two examples. How did you tap into that? Like where did that idea originate from?
8:19from like these core machine learning ideas where you want as large of a data set as possible and you want as diverse of a data set as possible. So I think that's sort of how we landed on synthetic mirrors because otherwise if you ask, if I just ask AI for 10 ,000 essays, I'm going to get like 9 ,000 essays that sound like very, very similar. So instead of what we have to do to diversify our data is to have the AI essays mirror a human essay. So that's really interesting. So you're getting this huge body of human work. You're then mirroring that with AI versions and training it on the trickiest subset of the ones that your detector can't always tell the difference.
9:07Is that right? And looking for distinctions between how a human writes and an AI writes? Exactly, yeah. And training on these hardest examples is a really important part of it as well. Because otherwise, there's just not enough signal. For most pieces of text, it's actually very obvious to the detector if it's human or not. Is it riddled with typos? Is it messy or personal? And so by looking for these edge cases where it maybe plausibly could have been written by AI, we get like much higher signal on what are the actual AI signals. So there's a really interesting thing there where, so Pangram itself is using an AI model to assess AI text.
9:53And I've used your service and you have this feature where Pangram can kind of flag parts of a body of work that it believes are signs that tip it off as being something that was written by AI. But at the same time, you guys have this sort of warning or this caveat saying, oh, actually, this isn't what the model is looking at. We don't really know what the model is looking at. Am I understanding that right? So the model is sort of a black box. Like you guys can't quite tell what is triggering it to tell what is AI and what is human. Yeah, I think usually it's like a more holistic story than the clean story is like, oh, this sentence tips us off that it's AI.
10:41But the actual answer is that the model is really looking holistically at the document. And there's a whole bunch of micro decisions that like each micro decision alone is a decision that AI would have made and a human wouldn't have necessarily always make that decision. But when you like aggregate all of these micro decisions together, you can have high confidence that the document was AI generated. So yes, the things in our dashboard, we have some supporting evidence. We've mostly taken these from the Wikipedia signs of AI writing. And we just kind of show these to people as a way to personally train yourself on how to detect AI writing.
11:25Is that saying that you, Pangram, despite being the flagship AI detector, you don't have your own signs of AI writing? You're relying on Wikipedia to kind of, you know, turn it into something that is like English readable to people. Yeah, in a sense, yes. I think the like English readable side is the hardest thing. We've been doing some really interesting work. This field of research is called interpretability. and so in our case we are looking at like what neurons in the like pangram neural network are activating when it sees AI and how does it cluster AI text differently from human text so so we had a pretty cool blog post that we put out recently where we found that even though we're not training the model specifically by saying this AI text is from Claude and this AI text is from ChatGPT, our model still learns what model family the text is from and is able to do a pretty good job at clustering text from different models separately.
12:32You know, one thing I'm curious about, too, is I've seen people go to ChatGPT or go to Claude and ask them to assess whether writing is AI because those things are, they are AI, they generate AI. And my impression is that those things have zero capabilities specialized for this whatsoever. But there are other specialized AI detectors out there. And those are the ones that, I guess, have not developed this reputation of reliability. I'm curious, do you have a sense of what they're doing wrong or not doing that isn't giving them this success rate? Yeah. So a lot of the early AI detectors, certainly we were not anywhere close to the first.
13:14But a lot of the early ones relied on research which said that there's this metric called perplexity, which might be a good method for detecting AI text. This is a measure of how surprising a piece of text is to a language model. So if you take the sentence like, the boy ate a bowl of soup, that's pretty low perplexity. Every word is expected. Whereas if you have a sentence, the boy ate a bowl of spiders, spiders would be a high perplexity word because that's not expected. So if you look at the way AI language models are trained, they're trained to produce low perplexity, unsurprising sentences.
13:58Because if it's surprising, it's more likely that it's wrong. So you can build an AI detector that measures perplexity of a text and says if it's low perplexity, it's AI. and if it's high perplexity, it's human because humans write in a more surprising way. This breaks down, however, in a couple of cases. So any text that is memorized by the AI is going to be low perplexity. For example, like the Declaration of Independence. So if you're wondering why you put the Declaration of Independence into some random AI detector and it says it's AI, that's why. It's because it's low perplexity. The AI model has memorized it.
14:36The other issue with it is English language learners also write in simple language, which is low perplexity, which then gets flagged as AI by these detectors. So that's sort of the problem with these early detectors. And the problem with perplexity in general is it's not really a metric that can be improved upon. Like it just is what it is. It's like measuring, I don't know, like the density of a liquid. It is just a single metric. There's no improving upon it. Got it. And so just to break those apart, the old style detector with perplexity, they're essentially looking at how homogenous a document is.
15:19Whereas Pangram, perhaps to simplify, you have trained on the specific patterns in AI text. Yeah, we're sort of like learning the micro decisions that these AI language models make consistently. So I guess big question, if I see a Pangram result, and I see a lot of them these days, should I trust it, right? How should I think about a Pangram assessment? Yeah, so I mean, I think Pangram is very accurate, of course. But what most of our benchmarks say is the false positive rate is 1 in 10 ,000. So look, there's a chance that it could be wrong. Hundreds of thousands of things are scanned by Pangram every day.
16:01So like there's going to be a few errors. But I think the longer the text is, the more confident we can be that it's correct because we just have more data. So if it's like a 50-word tweet that's flagged by Pangram as AI, like it's most likely correct. But I think there's like greater error bars on how much of it was AI. Was it actually just AI-assisted? Whereas if we're looking at like a 80 ,000 word novel and Pangram says this thing is 90 % AI, like we're very confident that it's at least 90 % AI. That's majority AI written for sure. I've noticed your tool is able to do that where it will give some, it'll tell you how confident it is in a result.
16:45I was messing around with it and kind of interspersing human written text and AI written text. And there was one big chunk that was kind of 50-50. And Pangram told me it was like low confidence, human written. Whereas there are other times I've seen it tells me, you know, it thinks that something is partly human and partly AI. So how are you coming to that decision when you are deciding, oh, I see bits of both in here? Yeah, so a lot of what we're doing is we're scanning the document, first off holistically, but also in parts. And we're saying, let's look at this part, and does this part look like AI or human or assisted?
17:31And then for each part, we kind of have a score, and then we are going to staple these together and aggregate them and say, well, we looked at 20 parts of this document, and 10 of them look like they're AI. So we can say it's about like 50 % AI.
17:51Nilay Patel:Support for the show comes from Even Realities. When you walk into a presentation, you have some options to help you remember your notes. You can bring a tablet, write all of your talking points in ink on the back of your hand, or try your best to simply memorize them. Here's a fourth, much more innovative option. You can wear them on your face with Even Realities. Even G2 are productivity smart glasses designed to keep real-time support right in view. With teleprompting, conversation support, AI assistance, and more, they help you stay on top of work and daily life. And unlike most smart glasses, they're designed to look and feel like premium eyewear.
18:30Nilay Patel:With no camera and a lightweight 36-gram design you can wear all day. The more contacts you give them, the smarter they get, adapting to how you work and what you need. To learn more about EvenG2, go to evenrealities.com and see how everyday smart glasses keep helpful information in sight so you can stay productive and hands-free throughout the day. And for our listeners, use promo code VERGE at evenrealities.com to get 10 % off Even Ring 1 and or Even Clip when you add them to your EvenG2 order. That's evenrealities.com, promo code VERGE. Support for the show comes from Even Realities. For a long time, the reality of smart glasses was that they were these bulky, inelegant pieces of tech that were much cooler in concept than in practice.
19:21Nilay Patel:A far cry from the sleek, stylish accessory of our sci-fi dreams that actually provides real functionality. Well, Even Realities has that fixed. Even G2 are productivity smart glasses designed to keep real-time support right in view. With teleprompting, conversation support, real-time translation, AI assistance, and more, they help you stay on top of work and daily life. And unlike most smart glasses, they're designed to look and feel like premium eyewear with no camera and a lightweight 36-gram design you can wear all day. The more context you give them, the smarter they get, adapting to how you work and what you need.
19:59Nilay Patel:To learn more about EvenG2, go to evenrealities.com and see how everyday smart glasses keep helpful information in sight so you can stay productive and hands-free throughout the day. And for our listeners, use promo code VERGE at evenrealities.com to get 10 % off Even Ring 1 and or Even Clip when you add them to your Even G2 order. That's evenrealities.com, promo code VERGE.
20:29Nilay Patel:Support for the show comes from MongoDB. AI-assisted and agentic coding can help you build faster than ever. But if your data layer is a bottleneck, what's the point? MongoDB actually gets out of your way. MongoDB is a unified AI-ready data platform that empowers you to build scalable, generative AI applications. It eliminates the need for separate, specialized vector databases by combining a flexible document model natively with semantic vector search, full-text search, and real-time operational data. Instead of wrestling with rigid schemas or translating data formats, MongoDB's native data model mirrors the language LLMs already speak.
21:11Nilay Patel:Plus, MongoDB gives you the flexibility to ship at the speed of AI. The ACID compliance lets you sleep soundly at night and scales to handle massive Fortune 500 workloads. Developers have a word for that kind of reliability. Actually, five words. It's a great freaking database. Start building at mongodb.com slash AI.
21:57for you.
21:59Nilay Patel:Often the response is, well, if you change the legal foundation here, we will leave. Yeah. How real is that? It's dead serious. With all due respect to Swiss authorities and everybody else, I think it would be suicidal to continue down this path. Subscribe wherever you get your podcast. This series is presented by Comcast Business. The big scandal recently around, I think AI detection was there was this Commonwealth Prize. They awarded a big short story prize to a piece called The Serpent in the Grove. A lot of critics thought it read like AI. Pangram assesses it as 100 % AI generated. But the author, Hamir Nazir, he insists he wrote it himself.
22:43He just was interviewed by The Atlantic. He said, oh, you know, I used voice to text to actually write this story. And maybe that created some linguistic oddities. I'm curious how you think about that. I am very skeptical of his claims. First off, I think like just my personal intuition, it just reads like total AI. Like it reads like you ask ChatGPT to write a literary prize winning essay. And like, this is what it would say, which is, you know, kind of vapid and has a bunch of empty metaphors and is like fake pretend deep. At least that's like my read on it. I really don't want to be like too critical of the guy, but I think there were also like a lot of inconsistencies in the interview, which also make me suspicious.
23:34For example, like he was asked what his favorite author was. He mentioned a couple authors and then the interviewer asked, Hey, okay, what's, what's your favorite piece from this author? And then And he's like, I can't actually remember a piece from this off. Which is very... I think that lights up some alarms for me. Is this person really just bullshitting? Or did he actually spend a lot of time on it? Also, I think the speech to text is kind of surprising to me. I don't see how that would trigger Pangram unless you took speech-to-text, rambled for five minutes or whatever, and then asked AI to turn that into a coherent essay.
24:28That sort of gets at something interesting where part of the evidence here is first people read it and people who are really familiar with AI writing said, I'm noticing some things here. Then people ran it through Pangram and Pangram said, yeah, we think this is AI. And then people asked him and kind of assessed the evidence that he provided. And some parties, you know, I think including the prize board, found it compelling. I think other parties perhaps, I don't know that the Atlantic issued judgment, but certainly suggested some skepticism as part of that interview, which I think is pretty well warranted.
25:12You know, it's interesting, Granta, which ran the story, said that it asked Claude to assess whether the story is AI. and I believe Claude said that it wasn't. And I think that's a very reasonable approach if you listen to all the leaders of these AI giants who are saying they've built super intelligence. But we both know this is a specialized skill to detect AI writing and that check was in all likelihood probably basically useless. I'm curious, you run this program that is supposed to be reliable. Who is using Pangram right now? A lot of industries seem to be completely unprepared or unaware at this point of how to assess AI writing.
25:59I think it's really only been in the last six months where I think the tides have really turned against AI writing. I think for a while people were holding out and saying, well, you know, maybe like AI writing is going to become more common, is going to be like generally accepted. And now we're realizing like, oh, there's actually a lot of use cases beyond our initial ones. So who's using Pangram? I think like there's a very big cohort of educators and schools and universities who use Pangram to check if a student is cheating on an assignment, if they're using AI to fully generate their essays, etc.
26:35I think we also have users in publishing. Either they have a magazine or they're an agent or something like that, and they want to use Pangram to just make sure that everything they are going to publish is above board. We also have AI companies who use Pangram to make sure their data is clean, for example. You don't want to pay an expert to write up some data or write an explanation or solve a problem. but actually instead of doing that, they're just feeding it into ChatGPT and then sending it back to you. So I think that's been a growing business as well. That is a wild problem of their own creation.
Read the full transcript
27:17It's true. I'm curious, what do these deals look like for you? Do you have a specific product for these companies or for educators or is it just come on in, buy a bunch of Pangram credits and scan away? Yeah, we try to bring Pangram to where people are at. So like with higher ed, they use a learning management system called Canvas typically. And so this is where students will submit assignments and receive grades. And so we integrate directly into Canvas to automatically score assignments with Pangram and then just show that to the instructor. And so kind of similarly, we work within people's like content management systems.
27:59And wherever they work, we're trying to bring Pangram there. If I'm a professor, what should I do if I get a positive result? Do you think that that's enough to fail a student? I mean, I think the first step is always talk to the student. I think the reason that somebody would use AI, it's like a symptom of a deeper problem. It's either the student didn't have enough time to complete the assignment, they didn't feel like they had the understanding to complete the assignment, or they don't care about your class enough. And so I think all three of these are like something that you probably want to dig into deeper rather than just like giving a zero.
28:35As a very polite way of saying, though, you think the assessment is correct and they should chase that down. I mean, I see this all the time where like educators, they get better results if they do something beyond giving a zero when they suspect AI use. because I think that doesn't really solve the core of the problem, which is that this student feels incapable of producing an assignment. Well, and I think this has been a big problem in education, and there are a lot of professors who are very unhappy about just how much AI has invaded college campuses, and it's a very, very tricky thing to combat.
29:21There's just no world in which people are going to stop using it wholesale. I see this learned helplessness, not just among students, but even among professionals who feel like, oh, AI could write a better email than me. So why would I ever write an email myself? And I think this kind of comes from a place of insecurity. But yeah, also it's just like, it's kind of concerning. Like if you have ChatGPT write all your emails, then you're never going to be able to write a decent email yourself. Yeah. And at the same time, I think for college students, you know, if you believe that everyone else is getting A's from using ChatGPT, you know, you're in sort of this impossible arms race.
30:12But you're right. Like you're not going to learn if you only ever use these tools. And so I think having any line of defense against this stuff for professors or really any professional in a workplace who needs genuine human work is really, really very important. I'm curious whether your job gets harder as time goes on because these models change all the time. And I'm curious if you think they will evolve in ways that get harder to detect. I'm also curious, you know, you have to train on human writing. How do you guarantee that you're getting human writing and fresh human writing? Yeah. Okay. A lot to unpack here.
31:03So I do think models have gotten much more capable in the last six to nine months. I think they've gotten much better at using their context and using a lot of context. So we've gone from like a really short question to ChatGPT to instead people who are writing essays with Claude code. Like let's do a deep dive on the like Strait of Hormuz and like the Iran situation. and then Claude will go do a bunch of research, make a bunch of web calls, write down a bunch of context and then compile a big essay. And so even though all of this is autonomous, it's using a lot more of its context and so its output is much more detailed than you would have previously expected.
31:48And so part of that, part of our job is trying to understand what AI writing looks like in this new paradigm where these models are much more capable. And then the other side is, yeah, the human side. I think a lot of our human data today comes from the pre-2022 internet, pre-ChatGPT, where we know for sure that most text was written by AI. And I think there's like, what we're looking at today is like, if we're going to use a data set from 2026, we need to be incredibly certain that it's not contaminated and there's not AI or chat GPT outputs in there. Is there a future where you are paying people to sit in a room at a notepad and handwrite essays in order to get like true human output?
32:45We've literally discussed this. We've talked about having an essay contest or something, but we have to supervise you to make sure that what people are writing are legit. I think we've landed on some simpler, lower-tech solutions, like just looking for trusted writers and sources that we believe are either not using AI or using AI in a way which we don't think we need to detect. For example, I'm just going to ask Claude, this idiom's on the tip of my tongue. And so I'm going to ask Claude for some wording suggestions. I think that's totally fine. And we don't want that to be flagged as AI. That's interesting.
33:26So when you say trusted sources, you mean like the New York Times? Or what does that look like to you? Yeah, I think that means professionals who write and who have a history of writing before AI. So I think if we're looking at books, for example, there's a lot of authors who have been prolific and published a lot of novels before AI. And then there's some that suspiciously started putting out like three or four books a year starting in 2024. And they just self-publish on Amazon. And those are the ones that we want to avoid. Are you working directly with any of those authors who you trust? Not today.
34:10I think this is a really ongoing question for us of making sure Pangram still works as well on 2026 human-written content as it does on 2020 human-written content. And so I think we're going to have some interesting projects around that later in the year. But it's mostly just if we do our job right, then it's invisible. Then Pangram continues to work and nobody really notices. So what does come next for Pangram? We were chatting before we started recording, and you mentioned that there might be some advancements coming soon. Yeah, so we have some interesting models that are in the oven. I think that I'm most excited for this, for our next release, for our text model, which is going to do much better on humanizers.
35:02There's this whole crop of tools on the internet called humanizers, which people can use to basically cheat, to have an AI model paraphrase your AI text in a way that it doesn't trigger an AI detector anymore. And so this has been kind of an adversarial battle, but our new model is going to be much, much better here. And it's also going to be better at understanding the degree of AI assistance in a piece of text. That's fascinating. It takes a remarkable dedication to laziness to make an AI essay and then run it through another AI program to hide that you ran it through AI in the first place. It's a big business, actually.
35:46Like, there's so many of them out there. And they astroturf Reddit. And I think they're just kind of like some of the scum of the earth. They're the worst of the worst. So do you have to train on their outputs specifically? Yes, we do. So we ran a big data collection campaign to collect a lot of data from these humanizers. And then what we did is we built our own internal humanizers that mimic what these external humanizers do. So we could go from a small amount of data to a large amount of data, and then we train our model on it. Long term, I think tools like this are going to be essential for the world to be able to tell what is human-made and what is not.
36:30How do you think about turning Pangram into like a real sustainable business? And are you worried that, you know, one day Google goes, this is important. We're just going to add a little button to Chrome and, you know, poof, that's that. You know, if Google does this, I'm happy. Then, you know, my work here is done and the problem is solved. I think there's a lot of competing factors here. Sort of how I think about Pangram in general is we're building this core technology, this core infrastructure for a future where we have these powerful generative AI models. If ChatGPD stopped existing and Claude and all the others, then we wouldn't have a business.
37:15But I think they're going to stick around and they're going to continue to have societal effects that we need to solve. So how do you make sure that Pangram lasts? You think that this is just going to be essential enough of a business that people will keep coming to you? I think so. I genuinely think we're really well positioned because it's such a hard problem. And because these AI models are always improving, we also have to always improve. And so I think it's a bit of a winner-takes-all scenario where if we have the best technology and we're building technology that improves faster than the others, then we could keep up when these other AI detectors that already kind of don't really work that well are just going to stop working completely.
37:57I wanted to ask one last question, which is on the website, formerly known as Twitter, your bio says that you are a slop janitor. You know, your whole startup is about spotting AI. You've spoken throughout this interview, I think, you know, somewhat derisively of AI-generated text. I'm curious, you know, are you opposed to AI writing? Do you think there's a place for it? And maybe most importantly, what to you qualifies as slop? I think there's a place for AI-generated text. I think, especially if it's properly disclosed, I think there's not a problem. It's a great tool for synthesizing information and providing it in a clean and readable format.
38:44But what I really don't like is when people are dishonest about AI content. When people say, I wrote this myself and it wasn't written by them, it was written by ChatGPT, that's where I feel like it's dishonest. I think there's the other side of it where AI content today is significantly worse than human content in a lot of ways. It's lower information density. It's optimized to be easy to read and pleasing rather than to actually effectively get ideas across. but I think as we like no matter how far into the future we look this core concept of writing being proof of thought is like not going to disappear so like if I if I want to really like think about a concept then I could write about it I can like workshop and edit my writing and then ultimately I'm going to have something that's like really concise and talks about my opinion in a way that I'm happy with.
39:47Whereas if I'm instead asking ChatGPT to generate a piece on my writing, then I will just simply not have thought about it as deeply. So I think the problem of using AI to kind of be like this cognitive offloading tool where AI will do the thinking instead of you, I think this is just going to be a problem in perpetuity. That's interesting. So your distinction is less about the text and more about whether there was real human thought behind what goes into it. Definitely. Yeah. I think that's a useful explanation. Cool. Well, Max, I'm glad somebody is diving into this very hard problem. I appreciate you joining us.
40:27Cool. Thanks so much, Jake. That's it for The Verge Cast. Remember to subscribe to The Verge for ad-free episodes, exclusive newsletters, and a whole lot more at theverge.com. We'd love to hear from you. You can email us at vergecast at theverge.com or call the hotline 866-VERGE-11. Shout out to everyone who's still demanding the thunder round in the comments. I see you. I hear you. One day we'll return. The Verge Cast is a production of The Verge and Vox Media Podcast Network. Today's show is produced by Josh Cajas, Eric Gomez, Brandon Kiefer, Travis Larchuk, and Aaron Locasio. We'll see you tomorrow.
41:03Support for this show comes from St. Germain Elderflower Liquor. Nothing says summer like a sunny backyard hang. Think light-colored linens everywhere and an easy conversation on your lips, your favorite hors d 'oeuvre in one hand, and the bright, refreshing taste of a Saint-Germain spritz in the other. Those are the details that stick with us. And the Saint-Germain spritz, an elevated take on the Hugo spritz, is the simplest part. Just mix sparkling water, wine, and Saint-Germain for your go-to summer cocktail. Visitst-chermain-liquor.com to shop and explore recipes. Enjoy responsibly.
From the publisher
AI text detectors have been notoriously unreliable, but that's starting to change. This year, Pangram keeps coming up as the trusted source in identifying AI-written text. We sit down with Pangram CEO Max Spero to find out how the system was made, how much we should trust it, and where the line is between useful AI and AI slop.
Further reading:
OnePlus officially gives up on the US and Europe
Google ordered to open Android and Search to rivals in Europe
Brendan Carr plans to let broadcast giants dominate the airwaves
The literary world isn’t prepared for AI
Subscribe to The Verge for unlimited access to theverge.com, subscriber-exclusive newsletters, and our ad-free podcast feed.
We love hearing from you! Email your questions and thoughts to vergecast@theverge.com or call us at 866-VERGE11.
Learn more about your ad choices. Visit podcastchoices.com/adchoices
