In short
Eye On A.I. Podcast Episode Notes
Episode Title 191 Jonathan Gillham: AI or Human? Exploring Originality AI's Detection Technology
Episode Description This episode features Jonathan Gillham, founder of Originality AI, a company focused on developing AI detection technology. The discussion centers on distinguishing between human and AI-generated content, the technology behind detection, its applications, ethical considerations, and the future of AI in content publishing.
Episode Sponsors
- NetSuite by Oracle: Offers cloud financial systems for efficient business management.
- Vanta: Security and compliance platform for businesses.
Key Concepts and Discussions
Background of Originality AI
- Founding Motivation: Originated from the need to differentiate human-written content from AI-generated text, especially in the content publishing industry.
- Development Process: The technology employs supervised learning to create models that predict the presence of AI in written content.
AI Detection Technology
- Methodology: Originality AI utilizes supervised learning on millions of data sets to predict content origins (human vs. AI).
- Challenges: The detection of AI-generated content is complicated by adversarial prompts and the evolving nature of AI writing models (e.g., GPT-3, GPT-4).
Accuracy and Limitations
- Originality AI boasts over 95% accuracy in detecting AI-generated text.
- There’s a 3% false positive rate, which can complicate its use in sensitive areas like academia.
Use Cases
- Web Publishers: The primary users who need to ensure content authenticity to mitigate risks associated with AI-generated spam.
- Academia: While detection is useful, it is recommended to use it for informational purposes rather than disciplinary actions due to potential false positives.
Ethical Considerations
- Transparency: There is a call for clear labeling of AI-generated content to maintain integrity in publishing and inform readers of content origins.
- Impact on Writers: Discusses the ethics of using AI tools in content creation, emphasizing that humans should benefit from AI’s efficiency.
Future Developments
- Originality AI aims to enhance its detection capabilities, including automated fact-checking and better understanding mixed-use content (AI-assisted and human-generated).
- The company is self-funded and focuses on remaining profitable while investing in technology improvements.
Major Takeaways
- AI Detection is Imperfect: While Originality AI offers high accuracy, the technology is not foolproof, and users must understand the limitations of detectors.
- Importance of Context: The effectiveness of AI detection can vary significantly based on context, such as the history of a writer’s work.
- Societal Implications: The rise of AI-generated content raises questions about the future of writing, content quality, and the role of human creativity.
Conclusion Jonathan Gillham discusses the complexities of AI detection technology and its implications for the future of content publishing. The episode emphasizes the balance between leveraging AI's capabilities and maintaining ethical standards in content creation.
---
Listen to the episode for a deeper understanding of AI detection technology and its impact on the content landscape. Follow Craig Smith and Eye on A.I. on Twitter for more insights into AI developments.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00People wish the world was clean sometimes and it was yes or no. and I want to rely on that with 100 % certainty. The world is messier than that with AI now and people need to understand that detectors are not perfect. And the flip side of that is we can't just dismiss detectors because they do work depending on the level of accuracy that you require. I think there's a few things that the world hasn't fully understood yet and I think that's on average and I think that's one of the key ones as it relates to AI detection. Hi, I wanted to jump in and give a shout out to our sponsor, NetSuite by Oracle.
0:38A little quick math. The less your business spends on operations, on multiple systems, on delivering your product or service, the more margin you have and the more money you keep. But with higher expenses on materials, employees, distribution, and borrowing, everything costs more. So to reduce costs and headaches, smart businesses are graduating to NetSuite by Oracle. NetSuite is the number one cloud financial system, bringing accounting, financial management, inventory, HR into one platform and one source of truth. With NetSuite, you reduce IT costs because NetSuite lives in the cloud with no hardware required, accessed from anywhere.
1:26You cut the cost of maintaining multiple systems. You improve efficiency by bringing all of your business processes into one platform, slashing manual tasks and errors. Over 37 ,000 companies have already made the move. So do the math. See how you'll profit with NetSuite. Now through April 15th, NetSuite is offering a one-of-a-kind flexible financing program. Head to netsuite.com slash IonAI. IonAI all run together, E-Y-E-O-N-A-I. That's netsuite.com slash IonAI. Again, netsuite.com slash IonAI. Hi, I'm Craig Smith, and this is IonAI. In this episode, I speak with Jonathan Gillum, founder and CEO of Originality.ai, a company that has developed an AI-powered detector for identifying AI-generated content.
2:31Jonathan shares his journey from working in content marketing to recognizing the need for a reliable tool to distinguish between human-written and AI-generated text. We discuss the challenges of detecting AI content, the architecture of originality.ai's detection models, and the importance of balancing accuracy with minimizing false positives. Jonathan also provides insights into the potential risks for web publishers using AI-generated content and the future of AI detection in an ever-evolving landscape. I hope you find the conversation as intriguing as I did. Background as a mechanical engineer by training, worked in that field for a period of time, started building different online businesses, generally in the content publishing space.
3:25Off the back of that, I built a content marketing agency where we were having writers create content. We were very heavy users of AI for a portion of that business, where we were transparently communicating to our clients that the writers were using AI to create content. And then we had a all human written division. And then it came clear, this is predating ChatGPT, but sort of post GPT-3 and some of the AI wrappers that were produced around GPT-3, that a lot of uncertainty existed between what was human and what was AI generated. a lot of people are happy to pay writers a hundred dollars or a thousand dollars an article they're not very happy to find out that it was copied and pasted from chat gpt in five seconds and that's where originality was was born out of yeah um and what was the content marketing agency or the content creation agency that you uh were working was that your own small shop or was that a Yeah, so it was our own shop.
4:29It was called Content Refined. It was then purchased by a company called Crowd Content and then folded into Crowd Content. But it was a content marketing agency where publishers of websites that wanted to rank in Google would come to us for specialized writing that was focused on ranking in Google. Yeah. I just asked because I have a nephew that works in that industry. And I keep telling him that industry is going to disappear. But maybe I'm wrong. I don't know. So that's the problem is how do you tell what text is human generated and what is AI generated? And as I said, I was very skeptical when I first heard about you guys because there are many ways to massage AI-generated text to make it have enough variation from what came out of the AI for people not to notice it.
5:41I mean, I can spot straight AI-generated text often just because of the format that a lot of these models use. You know, they always say, in conclusion, or at the beginning they say, in the realm of such and such. But those are easy things to change. change uh so i tried originality ai and i was impressed yeah it it picked up when something was 100 uh ai generated and it picked up when something was 10 ai generated so i wanted to hear how you guys are doing this yeah so i think first i mean that the do a detectors work question is i think one that is should be asked and i think it's such a new um such a new problem that people are facing that they're trying to attach um existing tools and the understanding on how existing tools are used so it's like oh it's like plagiarism checking and it's and it's really not the same because with plagiarism checking you can produce this sort of um clear proof on hey here's where these are here are the 10 words that were copied that's that's a very different use case for the detector plagiarism detector than it is for an AI detector.
7:04And so I think there's a lot of question mark around, do they work? Because they aren't perfect. Originality is not perfect. It's a prediction. So to answer the question, how does it work? It's our own AI that is built to predict the supervised, using supervised learning to predict the difference between, is this AI or is this human. So Fed, at this point, millions of data sets, millions of records of text, known human, known AI, some synthetic content in there with AI light editing to try and make sure that we are able to identify what is human and what is AI. And I think to your point around, there's some telltale signs on, especially with like GPT-3, 3.5, GPT-4 can go a few more directions or some of the latest LLMs from other companies.
8:01But if there's any adversarial prompt that is introduced and said, write like X or write in this style, then those telltale signs for a human to be able to tell are gone and it can become a flip of a coin for a human's capability of detecting. But with an adversarial prompt, your system can still pick it up? Correct, yeah. So with any adversarial prompter or AI-enabled adversarial techniques, so paraphrasing is a very common technique that gets used to try and beat detectors. With our detector, we're trained on, we have sort of a red team and a blue team. Our red team is always trying to beat our detector, use the latest strategies and tactics to beat our detector.
8:49and then our blue team is building the most advanced detector and then we try and sort of push out models that are suited for specific use cases. So we have our 3.0, we have our turbo model that is near impossible to beat, but it ends up with getting the trade-off to that is it ends up with a few higher false positives. And then we have our sort of standard model where it's like a more balanced approach to a little bit of AI and things that are allowed If somebody used an AI SEO tool to provide some keywords, you don't necessarily want to tag them as 100 % AI. But there's some that want to say just 0 % chance of AI, and our turbo model is built for them.
9:34That's fascinating. And how large are these models or how large is the training data that you train them on? Yeah, so we went through a few different models, a few different data sets. So we're up to over a million, well over a million text records that have been used in the process of the creation. So sort of exact numbers I won't get into in terms of what we're at. I'm not 100 % sure on current model and the number of text records that went in there. I know we have well over millions of data records. And then in terms of number of models that we're at, so we're at sort of model about 20, fourth deployed model, but we've tested and built up well over 20 models that have been created with many of them not producing a better result than what we currently are using.
10:32yeah um i had on the podcast a while ago uh the chairman of uh newsguard are you familiar with them yeah yeah are they using your tech or at the time that i spoke to him he said they were identifying ai uh really uh with just human judgment uh which certainly you can catch a lot As he said, a lot of times in very sloppy AI-generated texts, the copy will refer to, as an LLM, I can do this or I can't do that. But nonetheless, there's a lot of copy out there that I would guess a human editor or evaluator wouldn't be able to pick up. Yeah, so familiar with NewsGuard, familiar with that problem that they're aiming to address in terms of the depth around fact checking.
11:39I'd say there's some studies that have been done around humans' capability, especially this has been more done in the academic setting where teachers kind of care about what technology they can rely on and can they detect AI-generated content themselves. pre-adversarial prompt, they end with historical understanding of their students' writing. They could detect in that sort of neighborhood of 75 % of the time that it was AI generated. As soon as the student did a little bit of sort of adversarial prompting in terms of making sure that the content was written similar to their style, that it turned into a full flip of a coin on the study.
12:19So the capability, and I think that that capability of humans to detect is going to continue to diminish. GPT-2, you know, go back three years, capability of humans to be able to tell the difference between AI and human, pretty high. You know, as things have evolved, that has really turned into a problem that other AIs seem to be the most capable of addressing. yeah who's your biggest use case you mentioned uh content generation studios yeah so copy editors so basically any any user that is receiving writing from a from a writer receiving text from a writer need to make sure that it meets a certain standard and then pushing pushing that out publishing it so web publishers are our largest our largest customers that's who we're built for that's who we're focusing on we have a lot of use cases specific to that.
13:16We have a lot of users from academia. We don't love AI detection in academia. Although this problem exists in academia, the detection capabilities are not perfect. You know, depending on our tests, we're 99.9 % accurate, a non-adversarial prompted GPT-4 dropped down to like 98 % with adversarial prompts. And then, but we'll have like a 3 % false positive rate. And although that's small and definitely a useful setting for web publishers who are managing writers and have some history with them, it becomes a lot more problematic when it needs to be applied with a sort of binary classification of pass or fail in an academic setting.
14:00If you produce false positives, very hard to use that tool. So although the problem's there, we have tons of teachers that are using it. We like it only to be used from an informational standpoint and not from an academic disciplinary standpoint. Yeah, there was that famous case of a professor. I don't know if he actually failed his class, but because there were a lot of international students whose English is already challenged. And I don't remember the details of that case. But yeah, it was a tech, it was a professor that used ChatGPT. So especially in the early days of ChatGPT, they have some controls around this now.
14:48But if you were to ask an LLM, it always wants to sort of answer the question and please, it's trying to do what's asked. And so it says, was this written by an LLM? And then ChatGPT would produce an answer that, well, based on ABCDE, it could have been produced by an LLM, which is kind of true, but definitely not. And the prof thought that asking ChatGPT, did you write this, was essentially enough to get a positive answer. And then he failed, ran all his students' paper through ChatGPT asking that question. ChatGPT made up an answer as we know it can. And, yeah, unfortunately, he failed his entire class.
15:27Yeah. Although this is something I asked my nephew because he's concerned about his writers using these large language models. And I said, well, if if you can't tell if the human reader can't tell what difference does it make? I mean, if I were a writer for a content generation studio, I would just, you know, crank out, you know, you increase your output, which, you know, depending on your arrangement could either increase your income or increase your, you know, your profile within the company. Why should it matter? Yeah, so two reasons. The one sort of obvious one is if you're okay with using AI, which I do on some of my sites, I think there's a great use case for the use of AI with generating text to publish.
16:34But the person that is accepting that risk of publishing the content and receiving that should be the publisher and they should be the ones that receive the efficiency benefit. it. And so again, probably a lot of people happy to pay you simple math, a lot of people happy to pay a writer$100 for an article, not very happy to pay them$100 for a minute of work to copy and paste it out of chat GPT. So that's sort of the fairness and efficiency. And if AI is going to produce all this sort of excess capacity for humanity, who gets to reap the benefits. And I think that's sort of the one of the sort of the fairness component to it.
17:11And then why do they care at all component is if there is a risk of publishing AI-generated content in the eyes of Google. If Google is filled with nothing but AI-generated content, then why would you go to Google? Why wouldn't you just go to the AI? So I think there's an existential threat for search engines, aka Google. And I think publishers, and we have seen this, that Google is anti-spam. They've sort of left it vague about AI content versus human content. But if you're publishing hundreds or thousands of articles a day using AI with limited or no human oversight, that is bad in the eyes of Google.
17:52That's an existential threat to Google. And I think my view and a lot of other publishers' view is that publishing AI-generated content is riskier than human-generated content. Yeah, and certainly with a hallucination problem, some bad information can slip into text if you're not careful. So it's a big supervised learning system. Is the generated text evolving? How do you keep up with the market? Yeah. So at first we were, it's been an interesting process. So at first we were, we were sort of thinking this is going to be very labor intensive effort to always stay on top of the next new model, the next new model.
18:47and especially when sort of we were first building we built pre-chat gpt um for the gpt3 wrappers of writers of the day and then chat gpt launched and so we always felt like we were in this game that we're going to play catch up that okay grok mixtro llama um as as each of these new models launch would we now need to to train up on on that new model and we do there's some component of that But what we found is that all of these large language models are based on some similar sort of technology around transformers trained on similar machines using the same data sets to build their model on. And the output, although different from a user standpoint, from our AI's capability of detecting standpoint, we see very little drop off in detection efficacy with each launch of each new model.
19:40So continue to build a dataset based on every new model being launched. But the sort of performance drop off that we see has actually been diminishing and not increasing as we've had new models being released. Yeah. How do you protect your market? because I would guess, I mean, I'm getting a lot of pitches now for detectors of AI-generated video or AI-generated images, you know, deepfake detectors and that sort of thing. And it just seems as the tech advances, there are going to be more and more companies doing what originality does. and some of the big ones. I mean, Amazon or OpenAI or Google could build an AI detection model into their products.
20:46So how do you protect your market? Yeah, so I think first, if the societal problems associated with AI-generated text, So the fairness component, the societal harm due to mass propaganda, if that all went away with watermarking that was enforced, then originality would die. And I'd be a little bit unhappy, but also happy in that sort of the societal problems that we're aiming to solve would have been addressed. I'd say the second piece of that is it's very easy to spin up a very bad AI detector. And so I think there's open source models that are available. There's been papers that have been created.
21:27Very easy to create a low budget, very poor AI detector. And that's what we saw. I think there's 40 of them that have launched that I'm aware of, and maybe there's more. As there's been more models, the investment that we've put into our model significantly outstrips what I am certain is at least 90 plus percent of those other detectors that have launched. I think what we're actually seeing is a consolidation around sort of understanding of what detectors have invested in their technology and have a research team and are pushing the detection capabilities forward and the responsible use of AI versus those that are sort of just spun up an open source model and said, here's a detector, even though it's terrible.
22:20yeah and and how do you guys say is this a subscription service or is that uh pay as you go so people can use it on one piece of text uh on the fly yes we have two we have a subscription model and then we have a a credit base so subscription plus credits if you use more than the the monthly allotment of credits um but yeah that that's the the payment structure and then we have some free um parts so we have a we have a free chrome extension that lets um you watch a google document get created so basically takes all the metadata behind a google document and then recreates that document's creation so that you can sort of essentially watch a writer write so if there is ever a false positive which sort of falls in that three percent camp um there they have a the writers have a free tool that they can rely on to show that no look i really did produce this document.
23:16Yeah. And in that Chrome extension, does that look at a website text or is that something that you could do in the future where you open up the New York Times, a New York Times article and you click on the extension and it says this article was high probability that 20 % of this article is AI generated. Yeah. So, yes, it does that. So it will sort of direct connect into the originality's model and say, was this AI generated or not? It only works in the Google document for that recreation of the writing experience. And so that doesn't work on a website. Yeah. And who, I was talking about use cases before.
24:11Or who are your biggest customers? You don't have to name names. What kind of companies are using this? So end users, web publishers is the biggest segment. Biggest customers are the biggest web publishers in the world. So some of the groups that have published, that have consolidated a lot of websites, a lot of media companies that have consolidated multiple websites. and then marketing technology companies that they're in customer is our web publishers. So the biggest customers for us are the biggest web publishers and then marketing technology companies that use our AI detection within their solution for web publishers.
24:57To ensure that they're writers, that they're not passing along AI-generated content to web publishers. Correct. Or giving web publishers the capability of understanding their content's risk of their content's chance of being AI generated. Right. I'm not really familiar with the term web publishers, but who would be a big web publisher? What kind of companies? Are you talking about a WebMD website? Like any New York Times would be a big. So anyone that is publishing content on the web, just trying to clarify the difference between sort of non-web publishers will be more reliant on the human. So like a print publication will be more reliant on humans in the loop for ensuring that the content meets a certain quality.
25:53Whereas web publishers seem to end up on a more sort of systematic, wanting a more systematic approach and mitigating that Google, the Google risk. Yeah. And so the Google risk is that are they actively punishing AI generated content in their in their what they show to users or how is that affecting people? Yeah, so it is Google's position on this is unclear. Their position is spam is bad. We don't want spam. The nuanced answer becomes far more nuanced when you dig into it, where, yes, they have clearly actively punished websites that they have identified as AI-generated spam. So if you're publishing thousands of articles on your site with AI, no human in the loop, that is clearly going to be viewed by Google as spam.
26:56And they are actively working on punishing those sites. And you will see many case studies of sites that have had this sort of skyrocket number of articles. And then their traffic will, some period of time later, crater. So clearly Google is acting on it. Are they, has it been worked into their algorithm? And are they suppressing AI-generated content ahead of or behind, punishing it behind human-written content? That becomes a far more unclear answer. I think most publishers would agree that people that are publishing content on Google or on the web in order to try and rank in Google understand that there is a risk associated with AI-generated content.
27:43And then they make their risk-adjusted decision on whether or not they publish AI-generated or only human-generated content. Yeah. I did a podcast episode a few months ago with a guy. There have been a number of papers like this on the collapse of model output toward the mean as the model is increasingly trained on AI-generated content. In other words, it's generating content, and then if it's being trained on generated content, you lose the long tail on either end of the distribution, and the content becomes increasingly uninteresting. Yeah. Is that something, are you familiar with that problem?
28:42Is that something that has motivated you guys or that clients are interested in? Yes. So I think definitely it is a problem that we're highly interested in. It is a – I think it's a fascinating question, and we've done some of our own research around this, looking at the – as we're trying to identify other indicators that would be helpful in our model. model we were looking at the readability score and the and the when we get human content rewritten the range on the readability score for that human content will be far more significant and then it'll be very tightly a very tight normal distribution on the readability score for ai produced content and so i think that does really produce that sort of standardized a standardized output and as more synthetic data gets consumed those models are going to be producing even more average content.
29:40So it is a use case that our detector has been used for in the assistance of creating known human, known AI data sets. We need to be careful when we do that because we can create our own bias because we know we have false positives, we know we have false negatives. So it is a problem that our technology has been used for in the help of creating data sets for AI models, ensuring that they are helping to ensure that they're human. Yeah. Who are, can you name, are they the big model providers that are doing this? Not open AI, but names that would be very recognizable. Right. Yeah. To weed out AI generated content in their training data sets.
30:33Right. Yeah. And I wouldn't know. So big customers using it for assistance in creating data sets, whether those are the data sets that are going into the model, we don't have any insight into that. I wanted to jump in and give a shout out to our sponsor this week. When it comes to ensuring your company has top notch security practices, things can get complicated fast. Vanta automates compliance for SOC 2, ISO 27001, HIPAA, and more, saving you time and money. With Vanta, you can unify your security program management with a built-in risk register and reporting, and proactively manage security reviews with AI-powered security questionnaires.
31:21Over 7 ,000 global companies like Atlassian, Flow Health, and Quora use Vanta to build trust and prove security in real time. Listeners get$1 ,000 off Vanta at vanta.com slash ionai. That's I-O-N-A-I, E-Y-E-O-N-A-I, all run together. And that's Vanta, V-A-N-T-A dot com slash ionai, E-Y-E-O-N-A-I. Give them a try and get$1 ,000 off Vanta. your security program management with built-in risk register and reporting. Right, right. And you were saying that your model has been working across new models that are being introduced. What about something like Gemini 1.5? Have you guys had a chance to test that model?
32:33I mean, the sort of most powerful models that are coming out. Yeah, so every time a new model comes out, the red team jumps on it, builds a data set, and then we test and sort of check where our efficacy lies. And so, yeah, on every model that has come out, we haven't seen, it's in the last six months, we've not seen the efficacy of a new model dip below, efficacy of our current model dip below a 95 % accuracy rate on new models for a standard data set. And Gemini Pro falls into that camp as well. But it's been pretty interesting to see. And sort of my theory around, I think the models are going to really start to become far more commoditized than they are now, the foundational models, because the output does seem to be all starting to become very similar between them with differences.
33:39But the similarities amongst them from our AI detector's view has been converging, not diverging. And is that because, as you said earlier, they're all training on the same data? I'd say I'm not enough of an expert to know, you know, to why. But I think that's, you know, transformer technology, same data, same hardware, similar process, range of output, but not a wildly different output in terms of how these models are trained and the data set these models are used to be trained on. Yeah. So the other, the sort of subtext in what you're doing is that there is and will continue to be real value in human generated text.
34:32I think that's our view, that I think humans will continue wanting to know with content was generated by a human or generated by an AI. If you're going online to read a review about a product that you're wanting to buy, you would like to not have to complete a Turing test while reading every review, trying to guess if it was AI generated or human generated. And so I think there's a lot of times where humans want to know that the content that they are reading was written by a human and not just generated by AI. Yeah. Where do you stand on, you know, we were talking about international students and one of the great things about these language models, I mean, it's a pretty marginal use case, is that people who have trouble writing well in English, you know, they're not necessarily stupid.
35:32They just are not very good writers, maybe because English is a second language or they don't have a very strong linguistic capability. and these things are a wonderful tool for them because you can write what you want to write and the meaning is clear, even though the English is not well constructed, you put it through an LLM and you've got a clean. Yeah, so how do you guys feel about that use? Yeah, so personally, I love it and I think we use it within our team. We publish content that was generated by assisted with AI. So, you know, personally, as an engineer, I'd rather communicate in spreadsheets than words.
Read the full transcript
36:26And I often use AI to dump my thoughts into and then get it to be like, oh, that's that's what words are supposed to communicate. That's how they're supposed to look great. And so I'll copy and paste paste at my my own writing. And I have my own trained GPTs to try and sound like me. And then so personally, I use it for that exact use case. English is not second language, but I'm sure there's people with English as second language that are significantly better at English than I am. And then internally in our business, we use our AI research team are all English as second language individuals. They produce a lot of content for us to help document our training and where we're at in our efficacy tests.
37:09And AI is used heavily in the content that they produce. And so I think it's a great use case. I think it becomes a question mark on, is it a great use case? So would I have that same answer if I was a English literature 101 teacher and assigning an English writing assignment and AI could do 99 % of the work? That would be that would I would probably have a different answer in that use case. I think it's use case specific, but in sort of our own internal use cases, we love it. Yeah. And do you feel that people publishing AI-generated tech should label it as such? I think it's a – so I think bulk AI-generated with no human in the loop, absolutely, should have a label.
38:01I think the line between AI edited, human thoughts, AI edited, human involvement on ensuring that content communicates what the human wanted to, I think that is more questionable. I think the author should always be, and I think that the world is moving to this, but I think the author, the connection between the output and a name is important. And I think that that needs to meet a certain threshold in terms of your own standards for if content is attached to your name in the web anywhere. And I think that's ultimately the sort of final mitigating step of every piece of content should have a name that is understandable on who it is.
38:50Yeah. And so originality AI has, you know, like 95 plus percent accuracy in detecting 100 % AI generated text. I mean, if I run something on Cloud 2 and I put it into originality, bang, it says 100 % AI generated. But as you said, people are using these tools. You know, I do a lot of research with perplexity right now, and I'm kind of back and forth between something I'm writing and perplexity. And if some of that language gets copied into something I'm writing, and I haven't gone through and done the math, but I've tried that with originality. And yeah, it'll say, you know, 10 % AI generated.
39:51How accurate is that? Yeah, so it's a prediction score. So it's a prediction machine. And so it's saying the confidence that AI was involved in the creation of that document. So if it says 60 % AI, 40 % human, it's saying we're 60 % machine, 60 % confident that AI was used in the creation of this document. Not that 60 % is AI, 40 % is human. And so our classifier and so our accuracy is reporting is based on that, the binary classification of over 50, under 50, and whether we knew it was AI, knew it was human. In terms of mixed use cases, AI outlined, human written, AI edited, we're really, really interested in that problem with multiple sort of research veins that are pursuing our capability of people want to understand.
40:48Our tool is being used for people to understand the creation of that document, where AI was used, how it was used, and does that meet our criteria? For your example, Craig, there's times where you use it, and that's great. And there's times where I use it, and I think it's great. But there's other times where other people are wanting to enforce a certain level of absolutely no AI can be assisting you in the creation of this text. If that's their view, we want to have a model for that. if it's a mixed use case, we won't have a model for that to be able to communicate the efficacy and be able to communicate the efficacy well.
41:22So the answer is
41:27all human, all AI, easy. We're moving into a world where it's going to be a mixed use case for most people. And how do we best serve that mixed use case? And that's an ongoing sort of line of research for us. Right. So if I load a document into originality and it says 20%, that's a 20 % probability that AI was used in the development of the content, not the 20 % of this content is AI generated. Exactly, yeah. Yeah, in that case, I guess somebody like my nephews who's working with writers and wants to ensure that he's not being given AI generated, 100 % AI generated content. He would, if something came back 20 % or 50%, then it's a judgment.
42:34if this is, you know, then he could go back and ask, you know, we need to know how much of this content was generated. Is that how you would use that in that case? Yeah, so the way we recommend it being used is ideally not on an individual article basis, but across a writer's body of work. And if that writer has a long history of sort of 100 % human, 100 % human, 100 % human, and then has one that maybe triggers like 40 % chance of AI, 60 % chance of being human, that still we would want that sort of company to still have that article pass through as human because we're predicting to be human and you have enough of a history of that writer's work showing up as human.
43:27And so that's the ideal use case is to have a sort of a body of knowledge of that writer's work and then making judgment call based on your company's risk profile of publishing AI generated content on what those thresholds are for you. And ideally, you have your own data to sort of get a feel for where that should sit. So we're hesitant to say above X threshold is the target because that target is dependent on everyone's personal risk tolerance with publishing AI content. yeah yeah uh and and i guess uh if somebody uh was is a hundred percent human generated for for a couple of years worth of uh of content and then suddenly they start hitting 60 percent uh starting at a certain day can you kind of assume oh they've started using ai I think that that would be my, that would be the assumption that I would run with for sure, is that if you have this long history of 100 % human and then a new level that is consistently at something different, they have changed their writing process.
44:46Unless they're writing a totally new type of article, like potentially if they switched from personal stories to scientific paper writing, where it becomes a little bit, the accuracy can be a little bit more uncertain. then I think that would be a very safe assumption that they have changed the writing process. And if it still fits within the specs that you're comfortable with, then great. If that doesn't fit within your – you now at least are the one that's capable of making that decision. Yeah. And how – you know, you're a private company, so you don't have to answer. But how quickly are you growing and how far are you from profitability?
45:37Because I would presume it costs a lot of money to build and train these models. Yeah, so we're self-funded. So we haven't had some previous exits. Self-funded. we are marginally profitable at this point and and will continue continue to be as the the aim yeah we have hundreds of thousands of users that are using it some some you know the biggest companies in the world to to individuals across the board but we're really focusing on building for the digital marketing, web publishing world. Yeah. And the profitability, is it the sort of thing once you've amortized your initial costs and the profitability will go up?
46:44I mean, right now, it's in a very exciting space and we're not cash constrained on our growth efforts. And so I'd answer the question that, I mean, I think at this point we are targeting a profitability amount that allows us to reinvest and not need to secure outside funding and our growth efforts are not cash constrained. And so that's sort of where we're sitting. Yeah. And do you think, again, as you were saying, if suddenly somebody, Google or OpenAI, came out with a stronger model and the problem went away, you'd be out of business. But how do you see the future for this particular niche? Yeah, so I think there's a chance that that would happen.
47:46I don't think that's the most likely scenario. I think the most likely scenario, so OpenAI has come out with the model. They were viewed as the experts, and their model, although they had it built so that their own detector, own classifier, although they had it built so that the false positives were extremely low, they still weren't zero. And so anytime you're building a classifier, you're balancing between detection and false positives. In their case, they were very biased towards reducing false positives, understandably so. And that would be the same sort of criteria that I would have enforced if that was at OpenAI.
48:28But the result was they had a detector that was very inaccurate and then a false positive rate that was low but still not zero. And so the result was given their constraints, they needed to make it free. They needed to make it very low on false positives. The result was a detector that was pretty useless and they ended up needing to shut down. I think watermarking is sort of talked about as the other sort of silver bullet that would solve this problem and kill detectors. I don't think that's going to be an effective means of detection because there will always be these open models that will not have watermarking.
49:08And there will also be methods of spinning that content to avoid the watermark and potentially reverse engineering the watermark to be able to verify if you were capable of passing it or not. So I think watermarking is going to be pretty low chance. I think whatever technology is in the hopper not yet known and will produce a new model that our detector will then be incapable of telling the difference. I think that's probably the, like, where does this industry end in 10 years? I don't know, two years, one year, six months, who knows? I think if, you know, Qstar was a thing and there was a totally new method of AIs being capable of producing text and our detector then became incapable of telling the difference between the two, that's probably, you know, GPT-10 comes out.
50:03Will we still be effective? I don't know. Yeah. Yeah. Well, it's fascinating. and as I said, I've tried it a couple of times. I'm not a subscriber because I don't really have a demand for it, but I was impressed by the results. So you're building new models. I mean, what's next in your roadmap? Yeah, continuing to build models that will communicate to users the source of that content. And so whether that was AI edited, human written AI edited, continuing to build models that will continue to help people really understand where AI was involved in the text. That's where we're pursuing on the on the AI side.
50:55continue to have our red team and blue team so that we're always at the forefront of detection accuracy and then continue to build additional features that leverage what we're doing and help out any editor that receives a piece of text and needs to publish it. So making sure that that text meets the criteria and the sort of quality specifications that the company that they work for want to see all their content meet. So we have automated fact checking. It's not great, It's not, it's not, we're still in beta, but we've automated fact checking to help ensure that hallucinations don't get into the, kind of the facts, making sure content is at a certain reading level, grammar level, and just building out, continue, continue to build out tools that help editors, copy editors do their jobs.
51:44Yeah, that fact checking is interesting. That would be a model that somebody could run on the end product. I would imagine it's, I mean, you know, perplexity does that. You can load text into that model and it'll tell you whether there are any inaccuracies. but you're building if that's an LLM then that you're building and you're doing it on top of an open source? Correct. Done. We built it on top of an open source, our own LLM. Accuracy rates, it still has the challenges of AIs with hallucinations and its accuracy rate in terms of answering questions correctly exceeds existing models, but doesn't come close to the point of being able to replace editors.
52:52Yeah. Well, coming from a journalism background, I'm all for protecting the uniqueness of human writers and editors. So that's great. Is there anything I didn't talk about that you think you want to mention?
53:13I think we've talked about it a bit, but it's sort of an interesting, and I touched on this story, but it's an interesting sort of, the way the world views detectors right now is evolving, but there's sort of this incorrect binary assumption around AI detectors don't work or AI detectors are perfect. and it's people wish the world was clean sometimes and it was yes just give me a yes or a no and i want to rely on that with 100 certainty the world is messier than that with with ai now and people need to understand that detectors are not perfect and teachers can't rely on them with absolute certainty and fail students just because they got a detector said it was written with ai because there are false positives.
54:03And the flip side of that is we can't just dismiss detectors because they do work depending on the level of accuracy that you require. And I think that's something that, you know, it's funny that the world has sort of gone on this learning journey together with generative AI. And there's a few things that I think the general consensus has gotten right in terms of how to use them. I think there's a few things that the world hasn't fully understood yet. And I think that's on average. And I think that's one of the key ones as it relates to AI detection. Hi, I wanted to jump in and give a shout out to our sponsor, NetSuite by Oracle.
54:47A little quick math. The less your business spends on operations, on multiple systems, on delivering your product or service, the more margin you have and the more money you keep. But with higher expenses on materials, employees, distribution, and borrowing, everything costs more. So to reduce costs and headaches, smart businesses are graduating to NetSuite by Oracle. NetSuite is the number one cloud financial system, bringing accounting, financial management, inventory, HR into one platform and one source of truth. With NetSuite, you reduce IT costs because NetSuite lives in the cloud with no hardware required, accessed from anywhere.
55:35You cut the cost of maintaining multiple systems. You improve efficiency by bringing all of your business processes into one platform, slashing manual tasks and errors. Over 37 ,000 companies have already made the move. So do the math. See how you'll profit with NetSuite. Now through April 15th, NetSuite is offering a one-of-a-kind flexible financing program. Head to netsuite.com slash IonAI. IonAI all run together, E-Y-E-O-N-A-I. That's netsuite.com slash IonAI. Again, netsuite.com slash IonAI. That's it for this episode. I want to thank Jonathan for his time. If you want to read a transcript of today's conversation, you can find one on our website, IonAI.
56:32That's E-Y-E hyphen O-N dot A-I. In the meantime, remember, the singularity may not be near, but A-I is changing our world, so pay attention.
From the publisher
This episode is sponsored by Netsuite by Oracle, the number one cloud financial system, streamlining accounting, financial management, inventory, HR, and more. NetSuite is offering a one-of-a-kind flexible financing program.
Head to https://netsuite.com/EYEONAI to know more.
In this episode of the Eye on AI podcast, join us as we sit down with Jonathan Gillham, founder of Originality AI, a company specializing in AI detection technology.
Jonathan delves into the origins of Originality AI, born out of the need to distinguish between human and AI-generated content in the content publishing industry.
Discover how Originality AI uses supervised learning to build sophisticated models that accurately predict the presence of AI in written content, addressing the challenges of AI detection and adversarial prompts.Learn about the diverse use cases of Originality AI, from web publishers to academia, and how the technology helps mitigate the risk of AI-generated spam in the eyes of Google.
Understand the ethical considerations and societal impacts of AI-generated content, and why transparency and labeling are crucial in this evolving landscape. Jonathan also shares insights into the company's business model, growth trajectory, and future developments, including automated fact-checking and enhancing mixed-use case detection.
Tune in to gain valuable insights into how AI detection technology is shaping the future of content publishing and ensuring content authenticity.
Don't forget to like, subscribe, and hit the notification bell for more on groundbreaking AI technologies.
This episode is sponsored by Vanta, The security and compliance platform trusted by more than 7,000 customers.With Vanta, you can unify your security program management with a built-in risk register and reporting, and proactively manage security reviews with AI-powered security questionnaires.
Listeners get $1,000 off Vanta at vanta.com/eyeonai
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI




