Charge AI Crawlers

Block, Allow or Charge AI Crawlers? A Publisher Decision Framework for 2027

TL;DR. Allow, charge or block is presented as a menu of three choices. It behaves more like a menu of three requests. Each one is honoured at the discretion of the operator receiving it — and the operators that honour it are largely the ones that would have paid you, while the operators that pay nothing are largely the ones your policy never reaches. There is also no neutral setting to retreat to: allowing writes copies you cannot recall, and blocking writes an absence you cannot backfill. The framework below replaces the three-option menu with three prior questions — which options do you actually hold, which of them are reversible, and who in your organisation is entitled to take them.

1. The menu that is not a menu

The clearest statement of the standard advice comes from the company that built the control surface. Cloudflare’s pay-per-crawl announcement gives domain owners three options for any given crawler: allow it free access, charge it at a configured per-request price, or block it outright with no option to pay. The trade press compressed that into a workflow — keep the crawlers that send visitors, block the ones that only take, price the rest. By July 2026 the same company had scheduled a default flip: from 15 September 2026, mixed-use crawlers are blocked by default on ad-carrying pages for new customers, new sites of existing customers, and every existing free-tier account (TechCrunch, 1 July 2026).

Much of this is right, and worth saying before disagreeing with any of it. Machine traffic is now the majority of the open web — Cloudflare Radar put automated requests at 57.5% of HTML traffic against 42.5% human on 3 June 2026 — and where the product is the page view, an uncompensated fetch is a real cost against a real revenue loss. Publishers with distinctive archives spent two years unable to say “yes, for a fee”. Now they can, and nothing below argues for going back.

The assumption hiding inside the three options

The menu assumes a permission is something you set. It is not. A robots directive is a request written on your own property and honoured, or not, by somebody else’s software. A 402 Payment Required is an offer that becomes revenue only when a counterparty accepts. Even a hard block at the edge works only as well as your ability to recognise the traffic you meant to stop. All three options are proposals filed with a third party, and the third party decides.

Two consequences follow, neither priced into the menu. Your policy applies unevenly across the operators it targets, and the unevenness correlates with exactly the wrong thing. And unlike almost every other setting a team touches, this one writes durable artefacts in both directions, so there is no safe position to sit in while deciding.

What does it mean to block AI crawlers?

Blocking means refusing an automated client access to your pages — either through a robots.txt directive the operator chooses to honour, or a rule at your CDN (content delivery network) or WAF (web application firewall) that refuses the request outright. The first is a request; only the second is enforcement, and only enforcement works on an operator that has decided not to co-operate.

2. Compliance is the variable nobody puts in the model

Between October 2025 and July 2026, Goodie recorded 31 million AI citations across eleven answer surfaces and audited the access policies of 105 US and UK publishers against 25 AI user agents. Its central finding is the most important input to this decision and is almost never stated: blocking works, but only against the labs that choose to honour a block. Grok, AI Overviews and DeepSeek together produced roughly half the news citations in that sample, and none offers a functioning opt-out.

The illustration is stark. The New York Times, which blocks aggressively, was cited zero times by ChatGPT, zero by Gemini and once by Claude — and 37,642 times by Grok, 8,007 times in AI Overviews and 4,303 times by Perplexity, a company it is suing. The Associated Press, meanwhile, blocks every OpenAI crawler and still draws around 80% of its citations from ChatGPT, because licensed content travels through the contract rather than the crawl.

The pattern the numbers make

Read those cases together and a rule falls out. The operators you can block are, roughly, the ones that publish crawler names, respect the exclusion protocol, sign licences and answer letters. The ones that pay nothing are, roughly, those whose crawlers were never publicly named — the Tow Center noted in March 2025 that the crawlers behind DeepSeek and the Grok models were not publicly known, while ChatGPT, Perplexity, Copilot and Gemini had disclosed theirs. Your access policy therefore selects against your best counterparties and is invisible to your worst.

This is not an argument that blocking never bites. It is an argument that the bite lands opposite to the intent. Evidence sits either side of the Goodie work: BuzzStream’s March 2026 analysis of four million citations across 3,600 prompts and ten industries found the link between crawler access and citation far weaker than publishers assume, and the Pebblous report put post-blocking retention between 70.6% and 92.3% depending on the bot. TollBit data in The Register showed 13.26% of AI bot requests ignoring robots.txt in Q2 2025, up from 3.3% two quarters earlier. The surface with no opt-out at all is the one most brands care about most, which is why how AI Overviews actually treat backlinks is a better use of an afternoon than a bot policy review.

The Policy Reach Table

Before arguing about which option to take, establish which operators your decision reaches at all. Sort them by what they have already told you about themselves.

Operator classDoes your policy reach it?What blocking costs youWhat you can charge it
Declared, compliant, commercially engaged — publishes crawler identities, honours directives, has a licensing programmeYes, fullyCitation eligibility inside the products that also link back to youA real price, if what you hold is hard to substitute
Declared and compliant, not buying — publishes identities and honours directives, but has no purchasing route at your tierYesThe same eligibility, with no offsetting paymentNothing yet: a 402 is an offer, not a sale
Declared but selectively compliant — honours some directives, disputed on others, verification patchyPartly, and you cannot tell which partUncertain, and unmeasurable from your sideNothing enforceable short of litigation
Undeclared operators — no published crawler identity, no documented opt-out, no contact of recordNoNothing. You were never able to withhold anything from themNothing
User-triggered agents and AI browsers — fetches carrying an ordinary browser signatureNot at the HTTP layerYou block people, not machinesNothing

The instruction that makes the table useful: add up only the rows your policy reaches and treat that total as your policy. The rest is a preference. A publisher whose citation volume concentrates in the two red rows has no access policy — it has a statement of disapproval, which is a legitimate thing to publish but should not be scored as protection.

Key takeaway. A crawler policy is not a control surface applied evenly to the AI ecosystem. It is a filter with a known bias: it binds the operators that behave and passes through the ones that do not. Decide with that bias in view, or you spend your leverage on the counterparties least in need of discipline.

3. Both directions write history

The second flawed assumption is that a permission can be held open while you decide. Most settings behave that way: change a bid, watch a fortnight, change it back, and the world is where you left it. A crawl permission does not, because each direction creates something that outlives the setting.

What each fetch actually creates

A training fetch creates a frozen parameter. Whatever generation trains during your open window carries you; whatever trains during your closed window does not, and cannot be amended afterwards. You can be included in the next model. You cannot be added to this one. A search fetch creates an index entry that persists on the operator’s refresh schedule, not yours. An agent fetch creates nothing durable — composed into one answer, then discarded. A paid fetch creates a contract, the only one of the four with an agreed termination clause.

Blocking does not clear the record either. The Pebblous analysis makes the mechanism plain: shutting the pipe does not empty the tank. Copies already learned, index entries already written and third-party quotations of your material keep feeding answers. That is why retention stays high after a block — and why the reverse move is slow. Corpora refresh on their own cadence with no submit button; practitioners tracking citation-loss recovery report first movement three to four months out, with the substantial part around month six.

The best available estimate of the cost

Hangcheng Zhao of Rutgers Business School and Ron Berman of Wharton looked at exactly this in Strategic Response of News Publishers to Generative AI, using historical robots.txt snapshots from the HTTP Archive to date each publisher’s block and a staggered difference-in-differences design across three independently constructed traffic panels. The April 2026 version estimates roughly a 7% reduction in weekly visits within six weeks of a publisher beginning to block, a pattern that also shows up in household-level browsing data. An earlier version of the same work reported 23% of total traffic and 14% of human traffic for large publishers; the authors’ revision brought that down, and widening the sample changes the picture again for mid-sized sites. Around 75% of the publishers in the sample blocked an OpenAI-related crawler at some point.

Handle that carefully. It is an association measured on news publishers, revised downward by its own authors, not a law of nature for a firm selling industrial testing. What it establishes is direction: the commonest defensive move was followed by less traffic, not more, across three panels. For the same evidential standard applied to citation data, see our guide to measuring entity authority. The closest familiar analogue is a manual action and the recovery that follows it — quick to incur, slow and unglamorous to undo.

Why irreversibility deserves its own line in the decision

Economics has a name for the extra cost carried by a decision you cannot take back. Arrow and Fisher, and Henry, called it quasi-option value in 1974; Dixit and Pindyck built the general treatment. The claim applies here without modification: when the future is uncertain and one branch permanently destroys information or capability, that branch carries a premium over and above its measured cost. Under uncertainty, reversibility is worth paying for.

THE RESTORATION TEST

Write this down before the block goes live, not after. (1) What exactly do you expect to be restored when you reverse this? (2) By when? (3) Who has to act for that to happen — name the party, and note whether it is you. (4) What will not come back at all?

If you cannot name the actor in question three, the decision is not reversible, whatever the dashboard implies. If your answer to question four is “nothing”, you have not thought about the model generation that trains while the door is shut. Keep the sheet; it is the only control arm you will get.

If I block AI crawlers, can I undo it later?

The setting reverses in seconds; the effects do not. Index entries and cached copies expire on the operator’s schedule, model generations trained during the closed window cannot be amended, and competitors that occupied your citation slots do not vacate them on request. Treat reversal as a recovery project with a timeline, and read our notes on diagnosing and recovering lost AI citations before you assume the switch is symmetric.

4. “Charge” is only an option if you have a counterparty

The third option sits alongside the other two as though it were the same kind of thing. It is not. Allow and block are unilateral: you act, something happens. Charge is bilateral: you act, and nothing happens until somebody agrees. Three preconditions must hold before it is genuinely on your menu.

The three preconditions

First, a willing buyer for your corpus — content whose substitute is expensive. If an equivalent fact sits on forty other domains, your price is capped near zero whatever you type into the pricing field; the 2026 link building statistics give a fair sense of how rarely a mid-market corpus clears that bar. Second, metering both sides accept, which is why the model moved from charging per fetch to paying on use: Cloudflare’s July 2026 argument was that a crawl is a poor proxy for value, since one fetch may feed thousands of answers or none. Third, an enforcement path costing less than the licence is worth — which alone disqualifies most mid-market businesses.

The market has already run this experiment at scale. Cloudflare’s AI Crawl Control disclosures put the network above one billion HTTP 402 Payment Required responses served to AI crawlers per day — a billion offers a day, with no visible bidding war among the labs to meet publisher prices. That volume has produced a price floor and a negotiating posture, which is worth having. But an offer at scale is not a market, and a dashboard field is not revenue.

What you are actually choosing when you choose to charge

For the few businesses with a genuinely scarce corpus, direct licensing is the right conversation. For everyone else, the practical route is to borrow somebody else’s counterparty: an intermediary that already holds identity, permissioning, metering and settlement, and fronts a substantial share of the web. Pay Per Use launched on that logic, with Ceramic.ai and You.com as its first named partners. The trade is worth naming plainly: choosing “charge” usually means choosing a landlord. That can be correct. It is simply a different category of decision from the other two, and it belongs to someone who understands they are accepting a dependency rather than flipping a toggle.

Key takeaway. Allow and block are things you do. Charge is a thing you propose. If you cannot name the buyer, the meter and the remedy, the third option is not on your menu — and a pricing field that nobody accepts is indistinguishable, in outcome, from a block.

5. The decision has an owner, and it is usually the wrong one

Here is the part that costs real money, and it has nothing to do with which option is correct. In most organisations the block-allow-charge decision is not taken by anyone who can see its consequences.

ParseAI’s analysis of roughly 3,000 sites, weighted towards US and UK B2B software and ecommerce, found 27% blocking at least one major LLM crawler — and most of that blocking happening at the CDN or WAF layer rather than in robots.txt. Not in a file the content team reads. In a console the content team cannot open. Playwire’s write-up is blunter: security teams filter AI crawlers at the firewall, often silently, and often remove precisely the bots that decide whether a brand is eligible to be cited. Several popular SEO plugins ship block-AI-bots toggles switched on by default. The AI Visibility quarterly tracker has the share of sites blocking at least one AI crawler moving 11.6%, 13.5%, 13.7% across three quarters — a slow ratchet, largely unremarked by the people it affects. It is the same blind spot that makes negative SEO defence hard: the damage is done to a surface nobody on the team is watching.

Why the misallocation is structural, not careless

Jensen and Meckling’s 1992 treatment of specific knowledge and organisational structure states the problem cleanly: knowledge that is expensive to transfer should be co-located with the right to decide, and the control system should reward the holder for using it well. Both halves fail here. The knowledge — which questions we must be present for, what a citation is worth, which page classes carry the linkable assets — sits with marketing. The right to decide sits with infrastructure, a function measured on bandwidth, request volume, attack surface and cost, every one of which improves when a crawler is refused.

That is the mechanism, stated without blame: a decision routed to a team whose metrics all improve in one direction has effectively already been taken. One error arrives as a line item; the other arrives as nothing at all — no notification, no invoice, no control arm. Organisations choose the error that does not argue back. Whoever carries the link building specialist remit is usually the only person in the building who would have noticed.

The default flip compounds it, because inheritance is silent. Any property spun up after 15 September 2026 carries the new defaults: campaign microsites, a new regional subdomain, the standalone build for a research asset your digital PR and guest contribution programme is about to point sixty links at. Nobody chose that. Somebody deployed it.

Who should own the decision to block or allow AI crawlers?

A named accountable owner on the demand side of the business — the person answerable for organic and AI visibility — with infrastructure holding implementation and veto on genuine security grounds. The test is not seniority; it is whether the owner’s scorecard contains a number that gets worse when the site becomes less retrievable. If no such number exists on anyone’s scorecard, create it before you touch the policy.

6. Running the framework: the permission decision record

The framework that replaces the three-option menu is deliberately small, because heavy process here would be worse than none. One page per decision, and one rule about which decisions need the page.

The rule first. Anything that reverses on your own clock is an experiment: run it, measure it, stop discussing it. Anything that reverses only on somebody else’s clock, or not at all, is a decision, and gets a record and a named owner. Live agent access is an experiment; index access is a decision with a lag; training access is a decision with a permanent component. Most estates have this backwards — they leave the permanent settings on whatever shipped and endlessly relitigate the reversible ones, in the way teams once relitigated PageRank sculpting.

FieldWhat it must stateFailure mode when left blank
ScopeWhich page classes, named — not which bots. The technical library, the pricing pages, the research hub, the blogA site-wide policy set by whoever wrote the first rule, applied to an estate of wildly unequal value
PurposeWhich access purpose is being changed and why, in one sentence a non-specialist can checkCategory confusion: a training exclusion that quietly removes search eligibility too
Expected effectThe number you expect to move, its current value, and by how muchNothing to falsify later, so the decision becomes permanent by default
Reversal planWho acts, how long restoration takes, and what will not returnThe reversibility illusion — the assumption that undo is symmetric with do
Owner and review dateOne name, one date, both visible outside the team that implemented itInheritance: new properties adopt the setting silently and nobody re-reads it

The evidence the record runs on

Three inputs, none exotic. Read logs at the edge, not the origin: a managed edge refuses a request before robots.txt is consulted, so origin logs give a clean bill of health for traffic that never arrived. Verify operators by published signature or IP range rather than user-agent string, since forgery is cheap and the fastest-growing machine visitors carry ordinary browser signatures anyway — the behaviour covered in our work on AI browsers as a discovery surface. And keep a fixed panel of buyer questions across the engines, alongside whatever structured feeds and AI-readable endpoints you publish, so any change you make has a before.

The link building consequences, specifically

Link building’s output is a page other people point at. A permission decision decides whether the pointing still resolves for a machine. Three consequences follow. Check the permission status of a page class before commissioning the asset that will live on it: a data study behind an agent block keeps earning links and stops earning citations, the most expensive way to be half-right. Audit the properties that inherit defaults rather than the ones you open daily — subdomains, microsites, documentation hosts. And the same decision is being taken about you on every domain where you hold a placement, which is why niche edits and inserted links on unfamiliar hosts deserve the same infrastructure check as your own estate, and why local citations and directory surfaces are worth more than their link equity suggests: they are third-party, machine-readable and not subject to your policy. If your content only assembles after JavaScript runs, none of this matters yet — fix client-side rendering that hides links from crawlers first.

7. Worked example: Marchwood Laboratories

Marchwood Laboratories is invented and its figures are illustrative. A Southampton materials and product-compliance testing group, roughly £46m revenue, 340 staff. Its estate splits into four page classes: about 2,900 technical library pages (method notes, standards explainers, tolerance tables), 400 commercial pages, a 600-post blog, and a gated client portal. Its 2027 visibility budget is £165k. In August 2026 its infrastructure lead flags the September default and proposes tightening bot policy across the group.

Two paths from the same starting point

Path A treats it as one decision about crawlers: block training and agent network-wide, keep search, apply uniformly, reuse the existing bot-management preset. Implementation: an afternoon. Path B treats it as four decisions about page classes. The technical library stays fully open on all three purposes and is explicitly excluded from the preset, on the argument that it is the firm’s linkable asset and its job is to be quoted. Commercial pages allow search and agent, block training. The blog allows everything. The portal is already gated. Each non-obvious decision gets a one-page record with a named owner in marketing and a reversal plan. Implementation: two weeks, mostly arguing.

At month zero both are identical: 14% citation share on a fixed panel of 55 buyer questions across four engines, 96 referring domains earned in the prior year, and a technical library that accounts for 71% of the group’s organic entrances.

Months one to twelve

Months one to three: Path A looks like a clean win — verified bot requests fall 38% and a bandwidth line goes down. Path B has nothing to show, and its owner is asked twice why the library was exempted from a group security decision.

Months four to six: Path A’s preset turns out to have been written against a broad verified-AI-crawler category, catching two mixed-purpose fetchers that were feeding the firm’s presence on two engines. Panel citation share drifts from 14% to 9.2%. Nobody attributes it: the drift is slow and the dashboard that would have shown it was never built. Path B’s library, untouched, starts appearing in agent-fetched answers on standards questions, against a baseline that predates the change.

Months seven to nine: a trade association republishes six of Marchwood’s method notes with attribution, and the citations follow the notes. Path A spots the gap and reverses in month eight. Restoration is not immediate: index entries expire on their own schedule, one engine has settled on a competitor for the four questions that matter most, and the model generation trained during the closed window carries no Marchwood material at all. That last part cannot be fixed by unblocking.

Month twelve: Path A sits at 11% citation share, having recovered part of the drift, with roughly £7k of annual bandwidth saved and 88 new referring domains. Path B is at 22% with 103 new referring domains, no bandwidth saving worth reporting, and two decision records that let an incoming head of marketing understand in ten minutes why the estate is configured as it is.

The line the board understood: Path A made one decision about crawlers and applied it to an estate where the pages were worth wildly different amounts. Path B made four decisions about pages and was able to say, afterwards, who had made each one and what they expected.

What this does not claim: Path A’s saving was real, its reasoning was defensible on the information available, and for an image-heavy estate under genuine scraper load the same policy could be correct. The failure was not the direction of the decision. It was deciding at the granularity of the crawler when the value was distributed at the granularity of the page — and leaving no record, so reversal took four months to become thinkable.

8. The strongest objection to all of this

The hardest counter is not that a specific claim above is wrong. It is that the whole apparatus is disproportionate. It runs like this: for well above 95% of businesses reading this, the answer is allow everything, it takes ten minutes to verify, and wrapping decision records and named owners around a robots.txt file is the sort of governance theatre that makes marketing departments slow and unpopular. Publishers with real rights exposure already have lawyers. Everyone else should open the gates and get back to work.

That objection is substantially correct and I am not going to soften it. For most readers, the output of this framework is “allow”, and it should be reached quickly. Four things bound it rather than refute it.

One: arriving at the right answer and holding the right default are different things. Most sites are not running “allow” because they decided to; they run whatever their platform, plugin or CDN shipped. When the shipped value changes — and it changed on 15 September 2026 — a policy you never chose becomes a policy you never noticed changing. Two: the population that gets hurt is not the publisher weighing licence terms; it is the mid-market firm whose infrastructure team enabled a preset during an incident, which is what the ParseAI CDN-layer finding describes. Three: the prescription is small. One page, for the two or three decisions that do not reverse on your own clock, and nothing for the ones that do. If your version has a committee in it, you built the wrong thing. Four: the asymmetry applies only to the irreversible rows. Everything green in the table above is a ten-second setting and deserves the contempt a ten-second setting earns.

What would falsify this?

Two things. If the operators that currently ignore exclusion protocols adopted and honoured them — or were compelled to — then the compliance bias in section two disappears and a crawler policy becomes an even-handed control surface. And if any operator shipped genuine re-inclusion, an interface that back-fills a previously withheld domain into the artefact currently being served, the irreversibility argument collapses into an ordinary reversible setting. Watch for the second one in the licensing and payment programmes, which is where the incentive to build it lives.

9. What survives the next default change

Strip the dates out and a durable principle remains. When access to your content is mediated by an intermediary and honoured voluntarily by the party you are addressing, your access policy is not a control device. It is a selection device: it decides which half of an ecosystem you are legible to. Selection devices should be judged by who they select, never by what they forbid — so the first question about any new permission control, whatever it is called in 2029, is which counterparties are bound by it and how they differ from the ones that are not.

The second rule is narrower and more useful day to day: never let a reversible experiment and a one-way door share a control panel. Interfaces flatten decisions into settings, and settings invite the same casual treatment whatever sits behind them — something worth remembering when you next evaluate the tools that promise to manage all of this for you. Governance here is almost entirely the work of re-separating those two things after a dashboard has merged them.

And the thing none of it touches: whether a machine has a reason to include you once it can read you. Access is a gate, not a case. Clearing the second bar takes the same corroboration that has always decided which sources get quoted — independent, off-domain, re-established every time somebody asks. Our overview of what link building is and how it works and the fifteen strategies that still earn links in 2026 remain the load-bearing part, precisely because no permission setting can grant it to you or take it away.

The Monday checklist

1. Fetch your own key pages as a verified AI crawler from outside your network and record what comes back — before reading any dashboard.

2. List every property you control, including microsites and subdomains, and note which were created after your platform changed its defaults. Those inherited a policy nobody chose.

3. Ask in writing who holds the block-allow-charge right today, and whether any number on that person’s scorecard gets worse when the estate becomes less retrievable.

4. Classify your estate into no more than four page classes by what each is for, and stop treating the site as one asset.

5. Sort the operators you care about into the five rows of the Policy Reach Table, then recalculate what your current policy actually covers.

6. Write one permission decision record for every setting that does not reverse on your own clock — scope, purpose, expected effect, reversal plan, owner, review date.

7. Set a fixed panel of buyer questions across the engines and take a citation baseline this week, so any future change has a before.

None of this argues for keeping every door open, or for shutting them. It argues for knowing which of your doors swing both ways — and for making sure the person who closes one is the person who will notice what stops arriving.

Leave a Reply

Your email address will not be published. Required fields are marked *

Pay-Per-Crawl Previous post Pay-Per-Crawl vs Pay-Per-Use: Choosing a 2027 AI-Access Model
RSL 2027 Really Simple Licensing Next post RSL in 2027: Really Simple Licensing as a Standard for Link Publishers