TL;DR — Cloudflare’s three AI crawler permissions sort bots by why they came. Publishers care about what comes back. Those are different variables, and for most businesses they run in opposite directions.
Sort the three classes by the distance between the fetch and a human being and the picture inverts: Agent is distance zero (a person is already waiting), Search is deferred (someone may arrive later, through a link, on a surface you don’t own), Training is infinite (nobody ever arrives from that fetch). The 15 September default — Search allowed, Agent and Training blocked on ad-carrying pages — is precisely correct for an advertising publisher and close to backwards for a lead-generation, SaaS or transactional business that inherits it.
And none of the three permissions mentions attribution. The field that does — the optional fourth signal, use, shipped 1 July 2026 and defaulting to reference — is the only place in the entire regime where anyone asks to be linked back to.
1. The consensus that formed in six weeks
On 1 July 2026 Cloudflare replaced its single “Block AI Bots” toggle with three separate controls, available to every customer including the free tier. The three classifications are Search (crawling that builds an index so the system can answer questions about your site later), Agent (automation acting in real time on a person’s behalf — chat fetch bots and browser-driving agents), and Training (crawling that absorbs your content into a model’s weights). On 15 September 2026, new domains, new sites belonging to existing customers, and existing free-tier customers get new defaults: Training and Agent blocked on pages that display ads, Search left allowed.
Within six weeks the advice had converged. Allow Search, because Search is the class that historically sent visitors back. Block Training, because Training is the class that takes without returning. Decide on Agent case by case, and if in doubt, block it. That recommendation is now in every configuration guide written since July, and it is a genuine improvement on the blunt instrument it replaced. A single all-or-nothing switch forced small sites into what Cloudflare itself called a Faustian bargain — show up in search and accept training, or protect the content and disappear. Three dials beat one.
What the taxonomy is actually built on
The classification is behavioural, not branding. Cloudflare’s stated reasoning is that “AI” has stopped being a useful category boundary now that a search result page is itself an answer engine, so the sorting questions became: what is this bot doing on my site, what is it storing, and how will it reshare what it took. That is the right set of questions. It is also worth noticing how much of the taxonomy the three dials leave untouched. The underlying bot database classifies eleven behaviours — alongside Search, Agent and Training sit Transact, Data Collection, Security Testing, SEO, Ads Verification, Social and Link Preview, Feed Fetching, and Monitoring and Operations. Three of the eleven were promoted to universal controls.
Which three got promoted is not arbitrary. They map onto the harms an ad-funded publisher experiences: my content trained a competitor, my content answered the question so the pageview never happened, my content still earns its impression through search. The other eight are operational noise to a publisher — and not to everyone.
Why the advice converged so fast
Speed of convergence is usually a sign that a recommendation was inherited rather than derived. When a vendor ships a taxonomy, the taxonomy quietly ships an implied objective function with it — and the one embedded here is the protection of a monetisable pageview.
This piece takes the enforcement question as settled and asks a different one. Assume the dials work exactly as advertised on the operators that honour them. Are they the right dials, pointed at the right things, for the site setting them? For most of the businesses now inheriting these defaults, the answer is no — and the reason has nothing to do with compliance.
What are the three AI crawler permissions?
Search, Agent and Training — three behaviour classes a site owner can allow or block independently. Search builds an index to answer questions later; Agent fetches in real time because a person asked; Training absorbs content into model weights. They describe why a bot arrived, not what the bot leaves behind.
2. Purpose is not consequence
Each of the three permissions is a statement about the requester’s intent. None is a statement about your outcome. The market has assumed the two move together — that a benign purpose returns value and an extractive one does not. That assumption is doing enormous work and has never been examined.
Sort the same three classes by a different variable and the ordering changes. The variable that matters to a business is not why the fetch happened. It is how far that fetch sits from a human being.
THE ARRIVAL TEST
For each crawler class, ask one question: does a human arrive because of this fetch, and when?
Agent — distance zero. A person is already waiting on the other end. The fetch is not a precursor to the visit; the fetch is the visit, conducted by proxy.
Search — deferred distance. A person may arrive later, through a link, on a surface you neither own nor control, after an intermediary decides whether to show you.
Training — infinite distance. Nobody arrives from this fetch. Ever. The return, if there is one, is a recommendation inside a future conversation you will never observe.
The rule: permission decisions should be ordered by arrival distance, not by extraction. A class you can only be paid for indirectly and a class with a buyer attached should never share a setting because they share a vendor category.
Re-sorted, the recommended default inverts
Line the consensus up against arrival distance and the mismatch is stark. The class the consensus protects most carefully — Search — sits at deferred distance and is, in 2026, the class most likely to end in an answer that replaces the click entirely. The class the consensus is most relaxed about blocking — Agent — sits at distance zero and is the closest thing left to a visitor. The class everyone blocks first — Training — is the only one that cannot be re-run at query time, and the only one that determines whether a model knows you exist before any retrieval happens at all.
For an advertising publisher this ordering is correct as it stands, because the harm being defended against really is the substituted pageview, and an impression that never renders is revenue that never exists. The default is well designed for its intended customer. The problem is who else is now receiving it. Cloudflare sits in front of roughly 21.3% of all websites as of January 2026, and the 15 September change applies automatically to new domains and to existing free-tier customers — which is to say, to the long tail of small businesses, agencies’ client sites, and side projects that never made a permission decision in their lives.
The coupling nobody set deliberately
A second-order effect turns the three-dial mixing desk back into something cruder. Multi-purpose crawlers are governed by the most restrictive applicable rule, so from 15 September a site that selects “block Training” also blocks Googlebot, Applebot and Bingbot, which crawl for both purposes with one agent. Your Training setting is a floor under everything a mixed-purpose operator may do — and the mixed-purpose operators are the incumbents whose indexes still feed the link equity and crawl paths the rest of your site depends on.
Key takeaway
The three permissions describe extraction. Your business cares about arrival. Because those variables are not correlated — and for lead-generation and transactional models they are close to inverted — a permission set copied from a publisher’s harm model will reliably block your best traffic and protect your least valuable.
3. Agent is a person, not a crawler
The Agent classification is misfiled. Cloudflare’s own definition concedes the point in passing: the bot “visits your web application in order to complete a job, and often there’s a human waiting on the other end.” That sentence describes a visitor, not a scraper. It has a user-agent string instead of a browser fingerprint, and that is the entire difference.
The volume is not marginal either. In July 2026, Claude-User — the agent that fetches a page because a person in a conversation asked about it — accounted for 8.31% of individual bot traffic on Cloudflare’s network, the second-busiest bot on the internet behind Googlebot at 12.88%. Every one of those requests had a person attached to it.
What the arriving human is worth
The value of the sessions that do produce a click has moved sharply and in one direction. Adobe Digital Insights’ Q1 2026 analysis found that in March 2026, visitors arriving from AI assistants converted 42% better than non-AI traffic, generated 37% more revenue per visit, and spent 48% longer on site. Twelve months earlier the same channel converted 38% worse — an 80-point swing in a year. Seer Interactive’s breakdown of a B2B client put ChatGPT referral sessions at a 15.9% conversion rate and Perplexity at 10.5%. Conductor’s 2026 benchmarks put the whole channel at roughly 1.08% of all website visits, which is exactly the point: it is small, dense, and late-stage.
The explanation is obvious: the comparison already happened inside the assistant. By the time a person follows a link out, they have described their problem, seen a synthesis of the options, and chosen one. This is the same mechanism that makes agentic browsing worth more per click than the raw session counts suggest, and it is what a block on the Agent class actually turns away.
The honest caveat about those multiples
The 2026 referral figures are measured on sessions that produced a click — a survivorship-heavy sample — and headline growth multiples are inflated by platform growth that would have happened anyway. A 2026 arXiv preprint disentangling answer-engine optimisation from platform growth found ChatGPT referrals to a live site growing 5.7× while untouched control pages grew 3.5×, leaving a net effect of about 1.82× (95% CI 1.31–2.54). Treat any single-source multiple as directional. This argument does not need it to be large, only non-zero.
Should you block AI agent traffic?
Only if a rendered ad impression is how the page makes money. An agent fetch is a person’s visit conducted by proxy, so blocking it forfeits your densest late-stage traffic. If the page monetises through a lead, a quote, a signup or a sale, the agent is the customer’s hands, not a competitor’s crawler.
4. The permission that decides whether you get a link
Read the three permissions again and notice what is missing. Search, Agent and Training all govern whether a fetch may happen. Not one of them says anything about attribution. You can allow all three and be summarised without credit; you can block two of three and still be cited. Nothing in the classification system asks for a link.
The signal that does was added on 1 July 2026 as an optional fourth field, and almost nobody has looked at it. Cloudflare is testing use, an extension to Content Signals that expresses what a bot may keep and reshare after it has fetched. It takes three values, from least to most permissive.
THE ATTRIBUTION DIAL
use=immediate — interact, but store and reuse nothing. The fetch happens, the content leaves no trace, and no citation is possible because nothing was retained.
use=reference — index, excerpt, and link back. This is the default now being prepended to managed robots.txt files, and it is the only value anywhere in the regime that names a link.
use=full — summarise and reproduce. Maximum reach, no obligation to attribute, and — on the current rules — bots that reproduce in full cannot hold Verified status at all.
The rule: purpose permissions govern the fetch; use governs the credit. If your business is built on being cited rather than on being read, this is the field you should have an opinion about, and the three you have been arguing over are secondary.
Why a link builder should care more about use than about ai-train
The default matters more than the mechanism. When millions of managed robots.txt files quietly acquire use=reference, the web’s stated default preference becomes, in plain language: you may excerpt me if you link back. That is the first time the open web has expressed a collective position on attribution in a machine-readable field, and it is the same bargain that has underwritten what a backlink is since 1998 — restated for a surface where the citation, not the click, is the unit being traded.
The mechanism has teeth in one place and none in another. As a robots.txt line it is a preference, and preferences bind nobody. But the same three content-use levels are becoming a classification inside Cloudflare’s bot database, combinable with behaviour classes into rules like “allow everything used for Search, SEO and Ads Verification, but only up to reference.” There, it is enforced at the edge. There is also a reputational lever: a verified bot caught abusing the signal loses Verified status, and losing trusted standing across a fifth of the web’s domains is a deterrent with real cost attached.
For anyone who earns citations rather than buying them, the formats most likely to be excerpted — comparisons, definitions, ranked lists, tables of figures — are precisely where the difference between reference and full decides whether the excerpt carries your name. That is a bigger lever on citation volume than any decision about training crawlers, and it costs one line in a text file.
5. Three taxonomies, no agreed referent
There is a further problem with speaking confidently about “the three permissions”: the three parties who would have to agree on them do not. What looks like a settled vocabulary is one vendor’s product taxonomy, running ahead of a standards process that has not adopted it and a search engine that does not parse it.
What the standards track actually contains
The IETF’s AI Preferences working group — chartered specifically to standardise how content owners express these preferences — has a standards-track vocabulary draft, draft-ietf-aipref-vocab-06, published 28 April 2026 and running to October 2026. It defines two categories: train-ai and search, each taking allow, disallow, or unstated. There is no Agent category. The draft states in its own text that it does not yet have working-group consensus. Minutes from the April 2026 interim record editors being asked to draft a concrete proposal for a “use/input” category covering retrieval-augmented generation before the next interim — meaning the real-time retrieval case, the single largest AI use of the open web, did not yet have an agreed name in the standard as of spring 2026.
The companion attachment draft, which defines a Content-Usage header and a matching robots.txt rule and is co-authored by Google’s Gary Illyes with Mozilla’s Martin Thomson, carries an August 2026 milestone to reach the IESG; its published revision lapsed under the six-month rule with no newer version posted. Normal working-group churn, not failure — but not a settled standard either.
What Google does with the vocabulary
Nothing, by its own account. In April 2026 Google added content-signal and content-usage to the list of robots.txt tags its parser does not support, alongside crawl-delay, noindex, nofollow and noarchive — the tags having been identified as frequently used via HTTP Archive data. In July 2026, asked directly on Reddit, Google’s John Mueller said no crawler or LLM he was aware of uses the content-signal directives, that the directive was made up by a CDN, and that adding it simply creates maintenance overhead. He said the same of llms.txt.
The disagreement is not evenly spread, and where it concentrates is the tell. All three parties recognise training. All three recognise search. The category with the least agreement — present in the vendor taxonomy, absent from the standard, unparsed by the search engine — is the one with a human attached, and it is the one scheduled to be blocked by default from 15 September. A category that has not yet been named in a standard is being switched off before the argument about what it means has finished.
None of which makes the declaration worthless. A dated, published statement of terms is the document any future licence negotiation or enforcement action will reference, which is the same logic behind treating published AI content terms as a legal artefact rather than a technical control. Just do not confuse a record of intent with a mechanism, and do not report the setting to a board as though it changes what models do.
6. Your permission set is a function of how you make money
If purpose does not determine consequence, something else has to. The variable that does the work is your revenue model — which event on the page you are paid for. An advertising publisher is paid for a rendered impression, so a fetch that answers without rendering the page is a loss. A lead-generation business is paid for a submission by a qualified buyer, so a fetch that ends with one arriving is the best thing that can happen. Identical infrastructure, opposite settings.
The table below gives a defensible starting position for four common models. The reasoning column matters more than the cells.
| Revenue model | Search | Agent | Training | Why |
| Ad-funded publisher | Allow | Block on ad pages | Block or licence | The rendered impression is the product, so a fetch that answers without rendering destroys the unit being sold. The shipped default is built for you. |
| Subscription or data publisher | Allow | Allow on free tier only | Block behind auth | Free editorial is a shop window and wants agents; the proprietary archive is the asset and wants a contract. Split the estate, not the site. |
| Lead-gen, B2B and SaaS | Allow | Allow everywhere | Allow selectively | Nothing on the page is sold by the pageview. Agent fetches are late-stage buyers; blocking them removes the densest traffic you have. |
| Transactional retail and local services | Allow | Allow everywhere | Allow | Availability, price and specification need to be current inside the answer. Being absent from an agent’s fetch means being absent from the shortlist. |
THE PERMISSION MATCH — the setting is downstream of the revenue model, never of the crawler’s intent. Shading marks the cell where the shipped default is most likely to be wrong for that model.
The inherited-default problem
The most common permission configuration in 2027 will not be one anybody chose. It will be whatever arrived with the CDN account, the security plugin, or the managed ruleset an agency enabled years ago. The 15 September change applies automatically to new domains and existing free-tier customers, and the opt-out has to be exercised before the date. Sites that make their money from qualified enquiries rather than impressions will inherit an advertising publisher’s defence posture by default, and most will never know it happened, because nothing visibly breaks. Citations simply stop appearing and the cause looks like content.
7. Worked example: Calderwell Instruments, Halifax
Calderwell Instruments is an invented but deliberately ordinary case: a West Yorkshire supplier of industrial test and measurement equipment, roughly £28M revenue, 190 staff, selling calibration hardware and service contracts across the UK and Germany. Its site runs to about 4,100 pages — a product catalogue plus a large applications library of technical guides inherited when it bought a trade publisher’s content assets in 2021. Those guides still carry programmatic display advertising worth about £61,000 a year.
The trigger and the reflex
On 3 August 2026 their agency circulates a two-page memo about 15 September. The recommendation is the consensus one: block Training, block Agent, allow Search. The marketing director is inclined to sign it. The finance director asks a question that turns out to matter — which pages actually carry the ads?
The answer is that the ads sit on the applications library, because that is where the inherited inventory was. And the applications library is where every technical enquiry starts. A four-week audit of edge logs against the CRM produces three findings. First, agent-class fetches to guide pages had risen from 4,200 to 19,700 a month between January and July 2026. Second, 31% of quote requests in the previous two quarters came from sessions whose first touch was a guide page and whose referrer was missing or an AI assistant. Third, the average value of a quote request from that path was £4,900 against £3,100 from paid search.
The decision and the arithmetic
Applying the shipped default would have blocked Agent on exactly the £61,000 of pages that fed a lead channel worth roughly £390,000 a year in gross quote value. The board takes a different route in four parts. Ad slots are stripped from the 340 most-fetched guides, sacrificing about £38,000 of the £61,000 display revenue and removing those pages from the ad-page condition entirely. Agent is allowed sitewide. Training stays blocked, but only on the 260 pages carrying proprietary calibration reference data, negotiated separately. And use=reference is set across the estate, with use=full set on nothing.
Twelve months later: agent fetches to guide pages at 44,000 a month; assistant-attributed quote requests up 61%; measured citation share in the category from 9% to 21%. The guides also earned 47 new referring domains, because rewriting them as clean, current, first-party technical references made them the thing other suppliers and trade bodies linked to — the same dynamic that makes a maintained reference asset earn links long after the campaign that launched it has ended.
The honest negatives: the £38,000 of display revenue never came back, and the network’s CPM on the surviving inventory fell about 12% because the highest-traffic pages had left the pool. The German catalogue pages, governed by a separate reseller’s CDN account, kept the inherited block for a further five months before anyone noticed, which cost an estimated 900 agent fetches a month in a market where European visibility was the year’s growth target.
The finance director’s line, recorded in the minutes:
“We were about to protect sixty-one thousand pounds of advertising by switching off the pages that generate our leads. The setting was described to us as caution. It was a decision to sell the wrong asset.”
8. Where this argument is weakest
The strongest objection is not that the reasoning is wrong, but that it is aimed at a target that does not exist. It runs roughly like this.
The ad-page condition already does the work. Cloudflare blocks Agent only where ads are served — precisely where the pageview is the product. On a page with no ads, nothing changes on 15 September. So the “inversion” is a strawman: nobody is telling a SaaS company to block agents, and the scoping is exactly the discrimination being demanded. Worse, the Arrival Test smuggles in an assumption that an agent fetch is worth something, when the plainer reading is that agent fetches replace visits rather than being them — and the conversion figures come from the small minority of sessions that produced a click, which is survivorship of the loudest kind. And the use field being elevated as the real lever is, on Google’s own account, parsed by nobody.
All four points are fair and the last two are largely correct. Four bounds, in order of strength.
- The scoping is careful; the coupling is not. The narrow ad-page condition is a good piece of design, but it sits next to a rule that is not narrow at all: selecting “block Training” blocks Googlebot, Applebot and Bingbot under the most-restrictive-rule mechanic, everywhere on the site, ads or no ads. The setting most people will select first is the one with the widest blast radius, and it is not scoped to ad pages.
- The survivorship point is right, and the argument survives it. Conversion multiples measured on clicked sessions are inflated, and the disentangling preprint quantifies how much. But the Arrival Test is ordinal, not cardinal. It claims only that the probability of a human arriving is greater than zero for Agent and exactly zero for Training. No dataset disputes the sign, and the ordering is all the instrument needs.
- The use field really is unenforced as a robots.txt line. Set it and expect a link and you will be disappointed. Its value is that it is enforced at one edge as a bot classification, it is where the negotiating vocabulary for attribution is forming, and it costs one line. Report it as a declaration, never as a control.
- The framework does not say allow everything. It says do not let a taxonomy built around one harm model choose settings for a business with a different one. For an ad-funded publisher the consensus advice is correct and should be followed as written.
What would falsify this
If agent fetches turn out to substitute for human visits at close to one-to-one — if answers get good enough that the person never needs the page — then the distance-zero class collapses into the same bucket as Search and the advertising publisher’s default becomes the correct default for everyone. The observable that would show it: a sustained fall in click-through per agent fetch across multiple sectors while agent fetch volume keeps rising. Anyone running the measurement stack for this should be tracking that ratio monthly, because it is the single number that decides whether this argument holds.
A second falsifier is standardisation. If the IETF vocabulary adopts an agent or retrieval category and the major engines implement it, the vendor-taxonomy critique dissolves and the three permissions become three real ones. Watch the working group’s milestones, not the vendor announcements.
9. What to do before 15 September
Seven steps, none needing a licensing conversation or a legal opinion.
- Open your live robots.txt and read it. Not the one in the repository — the one being served, including anything a CDN or plugin prepends. Note whether a Content-Signal line is present and what it says.
- Find out who owns the toggle. In most organisations bot policy sits with security or infrastructure, in the same console as rate limiting, set by people measured on cost and risk rather than on demand. Name the owner in writing before the date.
- Answer one question about every page template: is it paid for by an impression, or by an action? That answer, not the crawler’s purpose, sets your Agent policy for that template.
- Pull thirty days of edge logs and split bot traffic by class. You are looking for the ratio of agent-class fetches to sessions with a missing referrer. That is your proxy for arrivals you are currently invisible to in analytics.
- Check whether “block Training” is selected anywhere. If it is, confirm you understand that from 15 September it also blocks the mixed-purpose search crawlers, and decide whether you meant that.
- Set use=reference explicitly. It is one line, it is the only place attribution is named, and it is the declaration you will want dated if a licensing conversation ever happens.
- Set a review date. The vocabulary is a draft, the defaults are a year old, and both will move. A permission set with no review date becomes an inherited default the moment the person who set it changes job.
The durable principle
Every access-control vocabulary the web has produced encodes the business model of whoever wrote it. The 1994 robots exclusion standard encoded a research crawler’s etiquette, which is why it has no concept of payment. The 2026 purpose taxonomy encodes an advertising publisher’s harm model, which is why it can express “do not answer without rendering my page” and cannot express “index me but say my name.” The next vocabulary will encode whoever is loudest in 2028 — most likely the agent platforms, which will produce categories about transactions and none about attribution.
The defence is a habit, not a setting: when a new permission vocabulary appears, work out whose revenue it was designed to protect before adopting its categories as your own. Then ask the only question that has ever mattered about an automated fetch — what does it leave behind? Historically there have been exactly two acceptable answers: a human, or a citation. Everything else is bandwidth. Permissions are worth setting to the extent that they raise the share of fetches ending in one of those two, and worth ignoring to the extent that they simply describe why somebody wanted your page.
That is also why the permission argument has a ceiling. Access policy governs what happens on your own domain, and the corroboration that decides whether a model reaches for you at all is assembled from what other people publish — the trade press, the launch and community surfaces where your category gets discussed, the institutional pages with no advertising on them that stay fetchable when the ad-funded web has finished pricing itself. Getting permissions right stops you being your own obstacle. It does not make you findable, and the strategies that earn it are the same ones that worked before any of this vocabulary existed.
If you are setting these dials for the first time, start with the fundamentals of how link equity and citation actually accrue, record your current visibility benchmarks before and after so you can attribute any movement, and watch acquisition pacing across the transition — a permission change that coincides with a content push is one you will never be able to evaluate. If your pages compete for extracted answer formats, the attribution dial is the setting that decides whether the extraction carries your name — which, on a surface with no blue links left to click, is the whole of the return. The training question, meanwhile, is a separate strategic decision about corpus membership and deserves to be argued on its own terms rather than folded into a security default.
