Perplexity Is Bypassing AI Blockers!?

7 Aug 2025 · 13 min · 6 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Debate over whether Perplexity bypasses website AI-blocking rules. Cloudflare published research claiming Perplexity ignores “no-crawl” directives using stealth/undeclared crawlers and hides bot activity. Perplexity counters that user-driven fetching via its tool is different from automated scraping, and that Cloudflare’s detection can’t reliably distinguish legitimate AI assistance from threats.

Guest backgrounds

No named guests in the transcript. Hosts discuss with “Jaden” and “Jamie” (AI Hustle School community).

Key claims

Cloudflare’s “name and shame” and that bots are over 50% of traffic (Imperva). Defenders argue GET requests to public HTML aren’t illegal and that blocking all AI would harm UX.

Notable examples

Wikipedia bot-traffic complaints; Cloudflare’s marketplace/token compensation idea; cars.com-style “research on my behalf” scenario.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

AI Community and Subscription Offer

1:28 to 2:41

Learn about exclusive content available in the AI Hustle School community.

“You know, they're one of our favorite AI companies to talk about around here.”

Cloudflare's Accusations Against Perplexity

2:41 to 4:16

Discussing the claims made by Cloudflare regarding Perplexity's practices.

“Okay, let's get into all the drama with perplexity because everyone's like a bunch of people have accused them, but a lot of people are coming to their defense, which is just a drama I love to see with perplexity.”

The Debate: AI Tools vs. Scraping

4:16 to 6:52

Examining the ethical implications of AI tools scraping data from websites.

“A lot of people are like, oh my gosh, this is terrible.”

Legal Perspectives on AI Crawling

6:52 to 8:27

Analyzing the legality of AI tools accessing public data and the implications.

“But I think the argument actually gets a lot better than that.”

Perplexity's Response to Cloudflare

8:27 to 10:33

Reviewing Perplexity's defense against Cloudflare's claims and the ongoing debate.

“It's a fine line, I feel like, because I could see the frustration.”

Future of AI Regulations

10:33 to 13:27

Discussing potential future regulations affecting AI companies and copyright issues.

“training off of, you know, copyrighted information, particularly if you, if you look at like YouTube or some of the music industry stuff.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Moto Casino, American Social Casino. Welcome to Moto Casino, where the excitement never ends. With thousands of the hottest free-to-play social casino games, fastest payouts, and the best promotions in the industry. No tricks or gimmicks. Owned and operated in the USA. Moto Casino is a free-to-play social casino. No purchase necessary. 21 plus to play. Void word prohibited. Sign up today for a generous welcome bonus. Moto Casino, American Social Casino. Download the Moto Casino app today. You're listening to a podcast right now. Driving, working out, walking the dog. If you're into podcasts, chances are you have something to say too.

0:39With RSS.com, starting your own is free and easy. Upload an episode and we distribute it to Apple Podcasts, Spotify, Amazon Music, and hundreds more. Track your listeners, see where they're from, and start earning from ads like this, even with just 10 listeners a month. If you've been thinking about starting a podcast, this is your sign. Start free at rss.com. Today in the podcast, we're talking about perplexity and they are actually catching some heat being accused of scraping websites that have explicitly blocked AI scraping. So kind of breaking the rules a little bit, impersonating real people with, you know, specific browsers and things, trying to hide the fact that they're bots, all kinds of, all kinds of accusations.

1:26So we're going to talk about them today. You know, they're one of our favorite AI companies to talk about around here. So we're going to get deep into it. Before we do, Jaden, why don't you tell me about our school community? Yeah. So every week, Jamie and I record a bonus piece of content and we post it all over on our school community. We don't post anywhere else and you get access to it for$20 a month. And basically this content is us explaining different tools, strategies, and techniques that we're actively doing in our own businesses. So it's stuff we don't share publicly, but it's all inside of this community.

1:56There's over 300 members. It's amazing. 19 bucks a month. This week, we recorded a whole video. Basically, I got interviewed by the Wall Street Journal yesterday on ChatGPT's deep research tool and the best tips, how to use it effectively, what I'm actively using it for, some of those pros and cons. And so we did a deep dive, basically demoing some of my prompts and tools and showing how it works and some interesting things. So if that or over 60 other videos that we posted over the last year are interesting to you basically in growing and scaling your business with AI tools, go check out the AI Hustle School community.

2:29It's discounted right now at$19 a month. We'll raise the price in the future. But if you want to lock in that price, you can go check it out and the price won't ever be raised on you. Amazing community. We'd love to see you there. Okay, let's get into all the drama with perplexity because everyone's like a bunch of people have accused them, but a lot of people are coming to their defense, which is just a drama I love to see with perplexity. Basically, what's going on was on Monday, Cloudflare calling them out, naming and shaming. They published some research. And over on the research, they said that they saw that basically perplexity ignores blocks and it would hide its crawling and scraping activities.

3:13So Cloudflare has kind of become famous recently because they set up this thing where if you are a company, you can block AI models from scraping your website. And there's all this drama because companies like Wikipedia say they spend so much money, like a third of their traffic is bots and they spend so much money serving data to bots. And they're like, it's not even real people. And so Wikipedia was mad. And Cloudflare basically came up with a solution that I thought was cool where you can, if you're using Cloudflare as an intermediary, they will, basically your server kind of runs through them, your web traffic runs through them and they detect that something's a bot, they'll block it.

3:48You have the ability as a website owner to block it. They also have another option that they're rolling out right now, creating a marketplace so you can say bots can scrape your website, but they have to compensate. They have to use tokens and they basically compensate you for scraping it so you actually get paid for it so it's not such a drag on you. And so this is kind of in the solution Cloudflare is using. Now, in their recent research, they say that Perplexity is using stealth undeclared crawlers to evade website no-crawl directives. Now, off the bat, I feel like this sounds bad. A lot of people are like, oh my gosh, this is terrible.

4:22Not everything is as it seems, but I guess, Jamie, what did you think when you first saw the story? I mean, I think it's interesting. I think, you know, how are we, and I think what the article is going to get at and what we're talking about here today is what's the difference between a robot text crawler that's training an AI model compared to someone who's using an AI agent to research and find information for them? Is there a difference between those two and should we, you know, allow one and not the other? So I think it's a really interesting debate. And I think I can see why some people are, you know, defending perplexity, because if someone is doing some deep research on something, and they're just trying to save time with perplexity, why wouldn't they be able to access, you know, a certain website?

5:06You know, I don't know. I don't know. What are your thoughts on it? Because I feel I could see both sides of it. Yeah, I know what you're saying. I feel like, first of all, I would say that Cloudflare was pretty aggressive in the call out here. they said they said some supposable reputable ai companies act more like north korean hackers time to name shame and hard block them this is a cloud flare really really ripping a new one on to perplexing they don't like this what's interesting though um there's a bunch of people over on x kind of posting more what's happening lies which is choker underscore zi shout out for your great analysis.

5:45You just got to love some Twitter users. There'll be like this massive controversy and some really technical person with like really deep understanding will give his breakdown on the situation. Anyways, he said, they are not accessing private data. Doing a GET request to a publicly available HTML endpoint is not illegal and there is legal precedent for this. If you don't want scrapers, put the data behind off walls and there you will have your terms and service to standby. Otherwise you got nothing kind of has a good point, which is, um, if you don't make it, like if you don't make it, so you have to have an account to, to see the data, it's technically publicly available.

6:25There's no way you really can stop all scrapers. Um, I know people don't like that. Cause like, well, we told them not to scrape, but there's no legal precedent and you can't force anyone to. So that's his argument. I understand that argument. Personally, I'm on the side of just like, I don't know, I'm on the side of the bots because I think it's great that perplexity exists. I think it's a cool tool and all those kind of tools. I know a lot of people on the content side are mad. That's just my personal opinion. So that's the argument. But I think the argument actually gets a lot better than that.

6:55This is what was written over on Hacker News. Someone said, if I as a human request a website, then I should be shown the content. Why would the LLM accessing the website on my behalf be in a different legal category as my Firefox web browser. This is a really good point because this is one of the situations where basically Cloudflare was all up in arms about is that if I'm using perplexity and I just say like, hey, go over to cars.com and find me the top seven most affordable Honda Acuras in my city. I don't know if cars.com is even a website, by the way, so don't quote me on this. Basically, right?

7:36Like you say, go to this specific website, look for these things, bringing back this information. And basically using perplexity like a search engine, that's just way more useful. Perplexity can do this kind of thing. Now, let's say cars.com is like, no, we have a no AI, you know, no AI bot thing. So we like, we don't want anyone to scrape us. Well, sure. But like, do you want me to just go and manually have to do this and parse through the data myself? Like, no, that's horrible. That's bad user experience. I want to use perplexity to do it. So like, literally, that's no different than me doing it myself.

8:05It's just I'm using a tool to save me time. And so the problem is that a lot of these AI companies get swept up in basically being called crawler bots, but like a human is actively directing them to do that, to retrieve information that's useful. So I think that's basically where the argument gets pretty flawed. And I don't think perplexity is doing anything wrong in that regard. Yeah, I don't know. It's a fine line, I feel like, because I could see the frustration. Well, and I think they could be upset, Cloudflare, about because basically it sounds like perplexity is finding a way around their software that they just created to charge people for data?

8:43Do you think? I mean, or at least maybe they're realizing there's some holes in their plan. I don't know. What do you think? Yeah, so basically what I think was actually happening because Perplexity came out and published a whole article on this. It's kind of funny. The drama was pretty good. They called out Cloudflare's blog post. They called it a sales pitch for Cloudflare. They made their own blog post defending themselves. They said that the behavior was from a third-party service that occasionally used. I think they said, quote, the difference between automated crawling and user-driven fetching isn't just technical.

9:18It's about who gets to access the information on the open web. This controversy reveals that Cloudflare's systems are fundamentally inadequate for distinguishing between legitimate AI assistance and actual threats. So shots fired because basically Cloudflare came out and said, Perplexity is supposedly reputable company, but they're not listening to anyone's do not crawl things and breaking all the rules and they're a bad company. We're going to name and shame them. And then Perplexity came out and said, actually, our users are just using our tool to go to websites and retrieve information. And if you can't tell the difference, then your services are inadequate for detecting real threats.

9:56Basically, your whole company is useless. So I kind of love this drama of the beef between probably the real truth is somewhere between the middle perplexity probably had a couple tools in there they shouldn't have it seems like a majority was good there's probably a few things they were doing that weren't good and cloudflare honestly does need i think to have some sort of system to uh differentiate between the two of those because like how could that possibly be like how could you block any ai tool that a user is actively directing so i'm i see i see both sides of the argument here. I think perplexity is going to win this probably, but yeah, I do love all the drama.

10:32Yeah. Cause I mean, you know, there really is a big problem with a lot of these AI companies training off of, you know, copyrighted information, particularly if you, if you look at like YouTube or some of the music industry stuff. And so I could see there definitely is a need there for, you know, if an AI model is training off of something, they should be compensated somehow. But then, you know, kind of, you know, if someone's doing research on a particular topic and it finds comes across this information does it have it's you know getting sucked up into the the system anyways i don't know it's it's it's a tricky situation to figure out but it is a tricky situation especially considering the fact that um according to a report which is done by it's the it's imperva they have a report called the bad bot report and uh apparently ai traffic accounts were over 50 % for the first time ever in internet history, AI bot traffic is over 50 % of internet traffic.

11:29So if you have a web, 50 % of the visitors to your website are going to be bots. Now, I do think we need a way to differentiate this because I know everything's getting sucked up into this people that are self directing versus bots that are scraping data that should be treated differently, probably or at least classified differently in case we make rules where you need to pay for training data. But yeah, if I'm using a tool like OpenAI or ChatGPT and I tell it to go to a website and do a thing, I want to just do the thing, especially with agents, right? Because once we get into age of agents, you might have seven agents running and I might be saying, go to YouTube and watch every video about this specific topic and write a detailed report on what we should be doing differently in how we're treating this topic versus what you're seeing on popular YouTube tutorials.

12:14And so all of a sudden, if every single bot was getting blocked as like a bot, but like I'm actively trying to, you know, do something, it just feels like a human was directing that. So I think it'd be good to differentiate. That's my opinion. We'll see where this falls. It feels like, to be honest, based off of the regulatory, what's happening. Trump gave this whole like AI speech recently where he said he didn't, there wasn't, I don't think any regulations specifically on this, but in his speech, he seems sympathetic with AI, more sympathetic with AI companies and said something along the lines of, you know, it's impossible for these AI companies to stay competitive if they have to pay for all copyrighted material that they take in and train off of.

12:53And he put forth the argument that you've heard both sides of this, but the argument that an AI model taking in data, learning from it, and incorporating into his model is the same as a human doing that. And so humans don't have to go pay for every article they read online. Therefore, an AI model shouldn't, yada, yada. So I think because that's kind of the line of reasoning that he has been putting forth in that, I think we'll probably expect policy to come out of the lighthouse along those lines. And we'll probably see AI companies not have to pay for a lot of this copyrighted data. That's in my opinion where this is probably going to go, but we'll definitely watch that further because that'll play into all of this as well.

13:27Hey, if you enjoyed this episode, please be sure to leave us a review wherever you are listening from. We really appreciate those. They help us out and we read every single one. and again be sure to check out the ai hustle school community if you're interested in growing your business using ai or also making money on the side with ai thanks for listening and we'll see you next time hi uncle laser here america's built on fast cars fast food and even faster women photo casinos built in america you know what they're fast out fast cash rise redemption you see only moto is u.s owned and operated moto it's american made that's who we are and that's who we care about.

14:04Play for free at moto.us. Moto Casino, the social casino void were prohibited. No purchase necessary. Visit moto.us for more details. Must be 21 plus. Moto Casino, America's social casino.

14:33Amazon Music, and more. Track your listeners, see where they're from, and start earning from ads just like this. If you've been thinking about starting a podcast, this is your sign. Start your new podcast for free today at rss.com.

From the publisher

In this episode, Jamie and Jaeden discuss the controversy surrounding Perplexity, an AI company accused by Cloudflare of unethical scraping practices. They explore the ethical implications of AI scraping, the legal distinctions between human and AI access to information, and the future of AI traffic regulation. The conversation highlights the ongoing debate about the responsibilities of AI companies and the need for clearer guidelines in the rapidly evolving digital landscape.



Chapters


00:00 The Controversy Surrounding Perplexity AI

01:55 Cloudflare's Accusations and Perplexity's Defense

05:51 The Legal and Ethical Debate on AI Scraping

10:01 The Future of AI Traffic and Regulation

See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

More from AI Hustle: Make Money with AI

All 178 episodes
Perplexity Is Bypassing AI Blockers!?AI Hustle: Make Money with AI · 13 min
Listen in VO