The short version

Within a year of each other, Cloudflare and AWS both shipped a way for site owners to charge AI bots for access, using the same HTTP 402 "Payment Required" response. Most of the coverage is written for large publishers who can afford to turn AI traffic away. If you run a small or mid-sized business, your situation is the opposite. You want to be crawled and cited, because AI answers are becoming how people find you. So the real risk here isn't the toll. It's getting swept into someone else's default setting and quietly dropping out of the AI results your buyers now use, without anyone deciding to.

Two of the biggest infrastructure companies on the internet just did the same thing. Cloudflare, which sits in front of a large share of the web, launched a feature that lets you charge AI crawlers to read your pages. About a year later, AWS added the same capability to its firewall. When two companies at that scale move in the same direction, it stops being a curiosity and starts being the way things are going to work.

The headlines frame this as good news: publishers can finally get paid when AI companies take their content. And for a handful of very large publishers, that's true. But for most of the businesses I work with, the framing is off. The money question is a distraction. The question that actually matters is whether you can still be seen.

What actually changed

Here's the mechanic, in plain terms. When an AI bot asks for one of your pages, your site can now answer in one of three ways: let it in, block it, or respond with a "Payment Required" message that says access costs money. That "Payment Required" response is an old, rarely used part of the web (HTTP status code 402) that suddenly has a job.

Cloudflare introduced its version, called pay per crawl, on July 1, 2025. Site owners set one rule per AI crawler: allow, charge, or block. Around the same time, Cloudflare started blocking AI crawlers by default for new customers, which is the part most people skipped over. Big names signed on, including TIME, The Atlantic, Fortune, Stack Overflow, and Quora. These are exactly the sites with deep archives worth licensing.

AWS followed on June 15, 2026, adding AI traffic monetization to its Web Application Firewall for CloudFront customers at no extra charge. Same 402 response, but built on an open payment standard called x402, with pricing set per page and settled in stablecoin through Coinbase. AWS says its bot controls can now classify more than 650 different AI bots and agents. The reason they give for building it: AWS states that AI bot traffic has grown more than 300% year over year, and that bots now account for more than half of web traffic for many content providers.

One more detail matters, because it tells you where this is heading. On July 1, 2026, exactly a year after launch, Cloudflare announced it's already moving away from charging per crawl toward charging per use, paying publishers when their content actually gets used in an answer rather than every time a bot fetches a page. Their reason: more than 50% of crawl traffic from good bots goes to re-fetching pages that haven't changed. As Cloudflare put it, a single page might be crawled once and then cited in thousands of answers, or crawled over and over and never used at all. The toll booth is being rebuilt around the thing that actually has value, which is the answer, not the crawl. Do not treat today's settings as the final version. This is going to keep moving, which is exactly why we plan to keep this post updated.

Timeline: Cloudflare launched pay per crawl on July 1, 2025 using HTTP 402; AWS added AI monetization to its Web Application Firewall on June 15, 2026 via the x402 protocol; Cloudflare shifted to pay per use on July 1, 2026. Sources: Cloudflare, AWS.
The toll model shifted from pay per crawl to pay per use in a year. Sources: Cloudflare, AWS.

Why this is a visibility decision, not a revenue one

Cloudflare published the number that makes the whole thing click. For every visitor an AI company sent back to a site, here's roughly how many times it crawled, measured across the first week of August 2025: Anthropic crawled about 50,000 times per referral, OpenAI about 887 times, and Perplexity about 118 times. On news and publication sites the gap narrows but the shape holds: Anthropic around 2,500 to 1, OpenAI around 152 to 1, Perplexity around 33 to 1. Nearly 80% of that crawling is for training, not for answering a live question.

Bar chart: for every visitor referred in the first week of August 2025, Anthropic's bot crawled about 50,000 times, OpenAI about 887 times, and Perplexity about 118 times. Source: Cloudflare.
For every visitor referred, how many times each AI bot crawled the web. All industries, Aug 1 to 7, 2025. Source: Cloudflare.

Read those numbers as a big publisher and you get angry: they're taking a lot and sending back almost nothing. Read them as a small business and you notice something else. The crawl and the citation come from the same companies. Block the crawler to stop the taking, and you can block yourself out of the answer at the same time. Search Engine Journal put it well in its piece on this shift: don't put a toll on traffic you've never looked at.

Donut chart: nearly 80% of AI crawler activity is for training the model, which sends no visitor and rarely cites you. The remaining roughly 20% is search indexing plus live answer fetches, each under 5%. Source: Cloudflare.
Nearly 80% of AI crawling trains the model and sends nothing back. Source: Cloudflare.

And being in the answer is starting to matter in ways you can measure. Adobe reported AI-referred traffic to US retailers up 393% year over year in its 2026 data. Meanwhile, as AI answers take over the top of the search page, the clicks that ranking used to earn are shrinking. The job is shifting from "rank and get the click" to "be the source the AI answer is built on." If your content is what the answer quotes, you win even when nobody clicks. If you've been walled off from the crawl, you're not in the running.

There's a slower danger too. A model that was blocked from learning your content during training can technically still reach you through live search later, but it won't really know who you are or why you matter. Cutting off the crawl today can cost you understanding tomorrow, not just a visit.

The part almost no one is writing for small businesses

When I bring this up with small business owners, I hear two reactions. The first is "wait, that's a setting?" The second is "is anyone actually buying from AI search yet?" Both are fair. The answer to the first is yes, and it might already be switched on. The answer to the second is that the traffic is showing up faster than most people expect, which is what that 393% number above is really about.

Most of the advice out there assumes you have leverage. Charge the bots, license your archive, make the AI companies pay. That advice is for publishers sitting on decades of content someone wants to buy. A small or mid-sized business is not that. You don't have a licensable archive. Your leverage to charge is close to zero. And you need the visibility far more than you need the toll.

So for an SMB, the danger isn't the toll booth. It's the default. Cloudflare and AWS sit in front of an enormous number of small business websites, and plenty of those sites were set up by a host, a past developer, or a template that flipped these controls on without anyone choosing them on purpose. A setting that blocks AI crawlers by default doesn't feel like a decision. It feels like nothing. But it can be the reason your business stops showing up in AI answers while your competitor down the street keeps showing up.

This isn't hypothetical for me. When I audit small and mid-sized business sites, I regularly find AI search bots blocked, and every so often Google's own crawler blocked right next to them. Nobody chose that. It arrived with a setting, a plugin, or a template, and it sat there quietly. That's the whole risk in one picture. Not a toll someone decided to pay, a door that was never opened.

It's common enough now that the major SEO platforms have started scoring whether a site blocks AI bots, right alongside the usual site-health checks. When the tools build a metric for something, it's because they're seeing it everywhere.

That's the whole story for most SMBs. Not "how much should we charge." Instead: do we even know what our own setup is doing to AI crawlers right now, and did we choose it?

Do I need to do something, or is my team already on this?

When a name as big as Cloudflare or AWS is in the headline, three people in a small company tend to have the same reaction at the same time, from three different angles. Here's the question each of them should actually be asking.

If you build or run the website

You're the one who'd actually see these controls, because they live in the CDN or firewall settings, not in the CMS. The reaction to skip is "we're fine, I didn't turn anything on." The reaction to have is to go look, because the defaults changed underneath you. The question: what are our current rules doing to AI crawlers right now, and did we choose them or inherit them?

If you lead marketing

You own whether the brand shows up where buyers look, and buyers increasingly look inside AI answers first. The trap is assuming this is an IT setting that has nothing to do with you. It has everything to do with you, because a server rule you never see can undo months of content work. The question: is anyone measuring whether we appear in AI answers, and does the person who controls the server know that visibility is on the line?

If you're the founder, owner, or on the board

You don't need to know what a 402 response is. You need to know that a decision affecting whether your company is findable is currently being made, and it might be getting made by default, by a setting nobody owns. That's the risk to name out loud. The question: who owns this decision for us, or is it being made for us by a host and a default we've never looked at?

Notice that none of these questions is "which vendor do we buy." They're all versions of the same thing: is this being decided on purpose, or by accident? For most small businesses, moving it from accident to on purpose is the entire win.

What we're watching

This is early, and it's moving fast, so a few threads are worth keeping an eye on. The model itself is still forming, with Cloudflare already shifting from paying per crawl to paying per use. The payment plumbing is being built on stablecoin rails and open standards like x402, which will decide how quickly smaller sites could ever participate. And the fight over training data is getting louder: in June 2026, a group of US publishers sent a cease and desist to Common Crawl, the free web archive whose data helped train early versions of GPT, arguing that copyright law is not an opt-out regime. How that lands will shape what "access" even means. We'll update this post as the picture changes.

I'll be honest about the part I can't tell you. I don't know exactly where this settles. Whether tolls ever reach businesses your size, whether the pay-per-use model wins, whether the fights over training data rewrite the rules again. It's early, and anyone who tells you they know how it ends is guessing. What I do know is narrower and more useful. The businesses paying attention now will make these calls on purpose. The ones that aren't will have the calls made for them, by default.

Where this fits with the work we do

The pattern underneath all of this is one we see constantly: a small business's visibility gets shaped by settings and defaults nobody chose, sitting in places the marketing team never looks. Figuring out what your current setup is actually doing to AI crawlers, what you'd lose if the wrong bot got blocked, and which access decisions are worth making on purpose is exactly the kind of thing our engagements are built to diagnose. If the questions above made you realize nobody at your company can answer them, that's a good reason to talk.

Want to know where your business actually stands in AI search? Let's talk.

Common questions

Can you charge AI bots to crawl your website?

Yes. Cloudflare introduced pay per crawl on July 1, 2025, and AWS added the same capability to its Web Application Firewall on June 15, 2026. Both use the HTTP 402 "Payment Required" response, and both let you allow, charge, or block each AI crawler.

What is HTTP 402 pay per crawl?

HTTP 402 is a rarely used "Payment Required" web response. When an AI bot requests a page, your site can answer in one of three ways: let it in, block it, or return a 402 that says access costs money. Cloudflare and AWS both use it to meter AI crawler access.

Should a small business charge or block AI crawlers?

Usually not. For a small or mid-sized business it's mostly downside, because the crawler that takes your content is often the same one that cites you in AI answers. Blocking it can remove you from the answers your buyers use. The bigger risk for most SMBs is a default host or CDN setting that blocks AI crawlers without anyone choosing it.

How do I know if my site is blocking AI crawlers?

Check the bot and firewall settings in your CDN, not just your CMS, since these controls live at the network layer. The major SEO platforms have also started scoring whether a site blocks AI search bots, right alongside classic site-health checks.

What did Cloudflare change in 2026?

On July 1, 2026, one year after launch, Cloudflare announced it is shifting from charging per crawl to charging per use, paying publishers when their content is actually used in an answer rather than every time a bot fetches a page. The reason: more than half of good-bot crawl traffic was re-fetching pages that hadn't changed.

Related reading
Generative Engine Optimization: How to Show Up in AI Search Is Your Site Invisible to AI Search? How to Tell and Fix It AI Search Sends Ready Buyers. Can Your Site Close? Do Google AI Overviews Reduce Clicks to Your Site? What Is llms.txt, and Do You Actually Need It?

Sources