TL;DR
- The detection question the industry asks — was this written by a machine? — has no decision attached to it. Google’s scaled content abuse policy is deliberately method-agnostic, and no ranking or retrieval system takes production method as an input.
- The answerable question is what a rival’s estate costs to keep. That is readable from public data, without a detector, because it is a property of the publishing schedule rather than the prose.
- Google defines crawl budget as the URLs Googlebot can and wants to crawl. The profession spent a decade on can. The wants side runs on perceived inventory, popularity and staleness — and popularity is the one term you buy with links.
- An index is a cache. Indexing Insight’s 1.4-million-page study puts a 99% chance of deindexing at 130 days without a recrawl, and a 90% chance of being forgotten outright at 190.
- Ahrefs found 88% of ChatGPT’s citations arrive through its general search channel, while only 6.82% of its retrieved results sit in Google’s top 10. Index membership is the gate; position is not.
- So you do not out-compete a scaled estate by out-publishing it. You hold a smaller estate where every page is financed by outside demand, and you spend where eviction will not reach.
Two verbs, and only one of them has an answer
The genre that has grown up around this problem is short and consistent. Run your competitors’ pages through a detector. Find the ones that come back synthetic. Then out-produce them, usually with a better-supervised version of the same pipeline. Every part of that sequence is sold as a paid tool. The first part does not work.
It does not fail because detectors are unreliable, although they are. It fails for a duller reason: a perfect detector would output a variable that nothing downstream consumes. When Google formalised scaled content abuse — its policy term for mass-producing pages with little value for readers — in March 2024, it wrote the definition around intent and outcome rather than production method, and it has said so repeatedly since. The March 2026 spam update, which made scaled content abuse its primary enforcement target, changed the severity and not the criterion.
What is scaled content abuse?
It is Google’s policy against generating many pages primarily to manipulate rankings, with little or no value added for users. The policy is method-agnostic by design: hand-written thin content and machine-written thin content are treated identically, and machine-written useful content is not covered at all. This is why a competitor audit built on authorship gives you a finding you cannot act on. You end up holding a true fact about a rival that no system in the chain — ranking, retrieval, or the AI Overviews layer that sits on top of both — is scoring.
The second half of the standard advice is worse, because it is expensive. Out-producing a scaled estate means entering that estate’s cost structure. It assumes the estate’s page count is its advantage. It is not. Page count is the thing the estate has to pay for, every period, forever, and the bill is not denominated in production cost.
The reframe
You cannot detect a machine in a sentence. You can read a cost structure off a calendar — and the cost structure is the part that tells you where to attack.
Crawl budget was always two numbers. The profession only argued about one
Google’s own definition is precise and almost nobody reads it closely: a site’s crawl budget is the set of URLs Googlebot can and wants to crawl. Two terms, multiplied.
Can is capacity. It is the term every crawl budget article for the last decade has been about, and it is genuinely not most people’s problem. Gary Illyes has confirmed the rough one-million-page threshold is unchanged since 2020, and in January 2026 Google cut the HTML crawl limit from 15MB to 2MB while Illyes pointed out that expensive database calls cost a server far more than page count does. Capacity is a server question, and the technical side of link building has always been comfortable there.
Wants is demand, and Google documents exactly three drivers of it: perceived inventory (how many URLs it believes you have, and how many are junk), popularity (in Google’s words, URLs that are more popular on the Internet tend to be crawled more often to keep them fresher), and staleness (recrawling often enough to catch changes).
Read that list as a link builder rather than as a technical SEO and something obvious falls out. Perceived inventory is a self-declaration: you tell Google what you have, through sitemaps and internal links. Staleness is also a self-declaration: you tell Google you changed something. Popularity is the only one of the three that is a statement other people make about you. It is the only term in the crawl demand equation that cannot be manufactured from inside your own domain, and it is the term the profession has spent ten years filing under somebody else’s job.
Why the capacity side got tighter anyway
Google rewrote its crawl budget documentation on 22 July 2026. Officially for clarity; substantively, it now says every site starts on a conservative default limit, that capacity is shared across all Google crawlers, and that not every crawled page will necessarily be indexed after crawling. That last clause is the one to sit with.
Meanwhile the pipe filled up. Cloudflare’s Matthew Prince reported on 3 June 2026 that automated requests had passed half of all HTML web traffic at 57.5%. In Cloudflare’s 23 June to 21 July 2026 window, the three AI bot categories combined reached 35.4% of bot traffic and overtook search-engine crawlers at 26.2% for the first time. Illyes has stated publicly that his goal is to work out how to crawl the web less, with smarter scheduling aimed at URLs that deserve the visit.
Key takeaway
Capacity was never the interesting number. Demand was. And demand is allocated, not capped — which means it is taken from somewhere.
An index is a cache, and everything in it is provisional
Adam Gent’s Indexing Insight research is the most useful public dataset on this and it is barely discussed outside technical circles. Working from 1.4 million pages across 18 sites, pulled through Google’s URL Inspection API, it produced two thresholds.
- 130 days. An indexed page not recrawled within 130 days has roughly a 99% chance of flipping to a not-indexed state.
- 190 days. At 190 days without a recrawl there is a 90% chance the page is being forgotten entirely — its status reverses to URL is unknown to Google, which carries zero crawl priority.
This is not passive decay. In late May 2025 Gent recorded an active purge in which more than 25% of monitored URLs moved into removal states, with some sites losing between 15% and 75% of their indexed pages. On one site monitoring a million URLs, 16% sat in the forgotten condition, many with historic performance data proving they had once been indexed and earning.
What does it mean when Google forgets a page?
It means the URL is no longer a candidate for anything. Not ranked low — absent from the set the systems draw from, with no scheduled reason to return. Recovery requires rediscovery, and rediscovery requires a signal from outside the page.
Which gives you the sentence the rest of this argument hangs on. A published page is not an owned asset. It is a rented position in someone else’s cache, and the rent is paid in demand. A 6,000-page estate has not built a fortress. It has taken out 6,000 tenancies on the income of a few hundred.
Why this became a weapon in 2026 and not in 2018
Crawl economics was a footnote for most of its life, and rightly. Three things moved it.
Pages became free to make. The index did not become free to hold
Ahrefs found across 600,000 pages that 86.5% now contain some AI-assisted content while only 4.6% are wholly machine-written, which tells you the interesting variable was never a binary. What changed is that the marginal cost of an additional page collapsed toward zero while the cost of keeping a page eligible did not move at all. Maintenance cost is a function of how many checkable, dated claims a page carries. It is identical whoever wrote it. Page count is not.
Capacity is now shared with parties that do not finance it
The old bargain was legible. A crawler took pages and sent readers, and the exchange rate was roughly visible. Cloudflare Radar has been publishing that exchange rate as a crawl-to-refer ratio, and through 2026 it ran near 5 to 1 for Google and in the hundreds or thousands to one for the dedicated AI crawlers — 2,237 to 1 for Anthropic and 217 to 1 for OpenAI on the July 2026 window, DataDome logged 17.7 billion AI agent requests in the second quarter of 2026, up 45% on the first, which is the supply side of the collapsing value of an agentic fetch.
The consequence for this argument is narrow and specific. Google’s July 2026 documentation says host capacity is shared across its crawlers, and every one of those requests lands on the same server budget as everything else arriving from everyone else. So the capacity term, the one that was safely ignorable for a decade, is being consumed by parties with no interest in your refresh schedule. That does not make capacity the binding constraint for most sites. It does mean the headroom that made it ignorable is thinner than it was, and thinnest on exactly the large estates this article is about.
One fetch now feeds two consumers
Google’s own guidance through 2026 has been consistent that AI Overviews and AI Mode are rooted in core Search ranking and quality systems. An index discard therefore cascades: the page is not merely unranked, it is unavailable to the layer that writes the answer — including the deep research modes that compile from dozens of sources.
No amount of on-page work reaches a document nothing re-reads, which is the failure mode most presence audits never think to test for.
The reconciliation nobody has published
Two Ahrefs findings from spring 2026 look contradictory. Louise Linehan and data scientist Xibeijia Guan analysed 1.4 million ChatGPT prompts and found the model pulls roughly 16.6 cited and 16.6 non-cited URLs per prompt, that just 49.98% of retrieved URLs end up cited, and that the general search retrieval channel is cited at 88.46% — supplying 88% of all citations. Reddit, by contrast, had 16.18 million retrievals in the same dataset and a 1.93% citation rate.
The same team then measured overlap with Google and found only 6.82% of ChatGPT’s retrieved results appear in Google’s top 10, and 16.61% anywhere in Google’s organic results. A separate 15,000-query study put 12% of AI citations in Google’s top 10.
So which is it — does ranking matter or not? Both numbers are correct because they measure different objects. Index membership is the gate. Position is not. What the answer layer buys from a search index is a candidate pool, not an ordering. That reframes the whole citation-set question: the asset being competed for is continued membership of a general web index, and membership is precisely what the 130-day and 190-day thresholds take away.
Google is explicit that crawling is not a ranking signal. Correct, and it does not weaken any of this. Eligibility is not position. You are not trying to be crawled into a better rank; you are trying not to be forgotten out of the pool.
Instrument one: the Industrialisation Signature
Here is the detection method that replaces the detector. Six observables, all public, none of them a property of the writing. A rival can rewrite every sentence on their site and not move a single row.
How do you detect AI competitor content without a detector?
You stop looking at the content and start looking at the publishing schedule. Scale leaves a signature in the calendar and in the balance between page count and outside demand, and neither can be edited away by improving the prose. The six rows below are readable from a sitemap, a backlink index and a few weeks of patience — and unlike a detector score, each one tells you something you can act on.
| What you read | Where it comes from | Maintained estate | Industrial estate |
| Indexed share | Indexed URL count against declared sitemap count | Stable, high, moves with publishing | Falls as page count rises — the eviction front, made visible |
| Lastmod spread | The XML sitemap file itself | Update dates diverge from publish dates | Every lastmod equals its publish date; nothing has been revisited |
| Time to index | Track new URLs from publication to appearance | Flat across months | Lengthens quarter on quarter — rationing already biting |
| Pages per named author per week | Bylines plus publish dates | Low single digits | A byline shipping ten or more a week is a masthead, not a person |
| Referring domains per indexed page | Any backlink index, divided by index count | Rises or holds as the estate grows | Collapses — demand per page is what buys the recrawl |
| Topic entry rate | New top-level sections per quarter, from nav and sitemap diffs | Adjacent expansion | Unrelated categories entering together, on one schedule |
The two rows that matter most are the fifth and the first, and they are related. Referring domains per page is the estate’s income statement; indexed share is its solvency. An estate that has multiplied pages while its referring domain profile stayed flat has diluted the only crawl demand driver it does not control, and the indexed share will follow it down on a lag of one to three quarters.
Instrument two: the Refresh Budget
Run this on your own estate before you run anything on anyone else’s. It takes an hour and produces one number.
THE REFRESH BUDGET
1. In Search Console, open Crawl Stats and filter by file type to HTML. Record average requests per day. Call it R.
2. From the Pages report, record indexed page count. Call it N. Mean refresh interval is N divided by R, in days.
3. Ignore the mean. Crawl is allocated by demand, so it is concentrated — a handful of URLs take most of it. Pull Days Since Last Crawl for every revenue-relevant URL through the URL Inspection API and sort descending.
4. Eviction exposure is the share of those URLs sitting above 130 days. Bands: under 5% healthy, 5 to 20% thinning, above 20% you are losing pages you have not noticed losing.
5. Divide total referring domains by N. That is demand per page, and it is the lever that pulls step four down.
The reason to compute it in this order is that steps one and two produce a reassuring number and steps three and four produce the real one. A 400-page site with a two-day mean interval can still have forty pages nobody has fetched since February. Averages hide eviction because eviction happens at the margin, and the margin is where the expansion happened.
Where a scaled estate is structurally weak — and where it is not
An industrial estate is optimised hard for one variable, cost per published page. Every other variable is unmanaged, and unmanaged variables are attack surfaces.
Refresh: production scales, maintenance does not
An operation that built 6,000 pages at £18 each cannot revisit them at £18 each, because revisiting requires reading, checking and deciding, and none of those got cheaper at the same rate. So the pages age. Staleness is a documented crawl demand driver, so the neglect compounds directly into the thing that removes them. A 300-page estate can hold every page genuinely current. That is not a virtue claim. It is an arithmetic one.
For example: a comparison page carrying twelve prices, four version numbers and two regulatory references has roughly eighteen ways to become wrong in a year. At 300 pages that is a manageable editorial calendar. At 6,000 it is an impossible one, and the operation that made those pages priced none of it. The estates that avoid the trap do so by publishing pages with almost nothing checkable on them, which solves the maintenance problem by removing the reason to fetch the page at all.
Concentration: the inventory they cannot stop declaring
Google calls perceived inventory the factor a site owner can most positively control. An industrial estate maximises it in the wrong direction, and the same expansion thins internal authority across thousands of URLs — which is why questions about sculpting internal equity come back at exactly this size of site, and why pruning is a link decision rather than a content one. Cutting inventory raises effective demand for every page that survives.
Consequence: a bimodal outcome distribution
The March 2026 spam update completed its rollout in roughly 72 hours. Reported drops for sites running scaled content operations clustered between 40% and 90%, and manual actions in this category act at the domain level rather than the page level. Lily Ray had predicted the crackdown in December 2025 and compared its likely severity to Panda and Penguin.
Practically, this changes three things about how you treat that competitor. Do not buy placements from them or from anything that shares their footprint. Do not co-locate on the same aggregators if you can help it. And when you evaluate a recovery scenario for a client whose rival looks unassailable, price in that you are competing with an entity whose outcome distribution has a fat left tail — and note that the same domain-level logic now governs how automated link patterns are detected.
The eviction dividend
Now the part that runs against instinct. The obvious move against a scaled rival is to attack their thinnest pages, because they look like easy pickings. They are — which is exactly why you should not spend there. Positions held by pages the index is already forgetting return to the market at no cost to you. Spending budget to win them means paying full price for something about to be vacated, on the lowest-value queries in the set.
Put numbers on it. If a rival holds 200 ranking pages against your target set and their indexed share has fallen to 46%, a material fraction of those positions is being held by pages on a clock. Say a quarter of them clear the 130-day mark within two quarters. That is fifty positions returning to the market whether you publish against them or not. A programme that spends evenly across all 200 has committed a quarter of its budget to winning ground that was arriving free, and has under-funded the fifty or so head pages that will still be there in a year.
THE EVICTION DIVIDEND
Estimate the share of a rival’s ranking pages that sit above 130 days since last crawl or outside their indexed set. That share is your dividend — you get it without spending. Put your marginal budget on the pages eviction will not reach: their high-demand, well-linked, frequently-refreshed head pages. Those are the only ones you actually have to buy.
What this changes about link building
Four changes, none of which is about domain authority scores.
A link is a recrawl instruction
Popularity is one of three documented crawl demand drivers and the only one that is not your own testimony. That is a different reason to want links from the one the discipline usually gives, and it survives in a world where the click is optional. Even if a placement sends no traffic and moves no ranking, it changes how often a fetcher returns to the page it points at. Anyone still framing what a backlink is purely as an equity transfer is describing half of it.
Point links at the starving pages, not the winning ones
Standard practice sends placements to money pages that already have demand. That is where marginal return on crawl demand is lowest. Reverse it against your Days Since Last Crawl list from instrument two: the pages at the top of that list are the ones where an external citation buys the largest change in behaviour. It also gives link velocity a purpose beyond looking natural — a steady trickle across the estate keeps more of it warm than a burst at one URL.
Screen the placement for how often it is itself refetched
This is the genuinely new prospecting column. Discovery propagates: your page gets refetched partly because the page linking to it gets refetched. A placement on a document that is crawled daily is worth more than the same placement on a dormant page, independent of any authority metric, and no backlink tool reports it. You can approximate it in ten minutes from the linking domain’s own sitemap lastmod spread and publishing cadence.
This reorders the target list. Frequently-updated trade and news properties — the kind you reach through journalist request platforms or through timing a story to a live news cycle — score higher on this dimension than an aged resource page that has not been touched since 2021, even where the resource page wins on every metric your tools display.
Deletion is a link-building lever
Perceived inventory and popularity are multiplied together, not added, which means removing junk raises the effective demand reaching everything that remains. A 900-page estate where 500 pages have never earned a referring domain or a fetch in six months is not a 900-page estate; it is a 400-page estate carrying 500 declarations that dilute its own crawl demand. Consolidating those 500 into 60 genuinely maintained pages is the cheapest link-equivalent move available, because it needs no outreach and no budget. It is also, for most teams, psychologically the hardest, which is why it is usually the last thing tried and should be among the first.
Build assets that are re-read rather than re-published
A page that changes on a schedule earns staleness-driven demand permanently, which is a structural argument for calculators and interactive assets and against the annual set-and-forget guide. It is also why rendering-dependent content is expensive twice over: it costs more per fetch, and it competes for a shared capacity limit that Google now says is conservative by default.
Worked example: Pennard Metrology
Pennard Metrology is a UKAS-accredited calibration and dimensional inspection lab in Swindon: £8.4M revenue, 62 staff, 310 published pages built up over nine years. In September 2025 a private-equity-backed national competitor began publishing at scale. By August 2026 that rival’s sitemap declared 6,240 URLs.
Pennard’s first instinct was the standard one. They priced a detector-led audit and a 900-page programme of their own at £137k over eighteen months. Instead, in June 2026, they ran the signature.
- Indexed share. The rival’s indexed count against declared sitemap fell from 91% to 46% over eleven months. Their absolute indexed count rose from 364 to 2,870, so the dashboard read as growth.
- Cannibalisation. 41 of the rival’s original 400 pages left the index in the same period — including two that had ranked in the top three for Pennard’s most valuable term.
- Demand per page. Rival: 1,410 referring domains across 6,240 pages, 0.23 per page. Pennard: 890 across 310, 2.87 per page — twelve times the density on a fifth of the footprint.
- Time to index. The rival’s new URLs went from four days to appearance in September 2025 to thirty-nine days by June 2026.
- Their own exposure. Pennard’s mean refresh interval was a comfortable 1.9 days. The bottom quartile sat at 47 days, and 22 pages had not been fetched in 210 days. Nineteen of those were already gone.
They did three things across March to August 2026. They froze page count. They moved £31k of the £137k into fourteen external placements, chosen against the Days Since Last Crawl list and screened for the linking page’s own refetch frequency. And they put dated, substantive revisions on 41 commercial pages, driven by real changes to accreditation schedules and tolerance tables rather than a timestamp edit.
By August, Pennard held 310 of 310 pages in the index, the bottom-quartile interval had fallen from 47 days to 9, and eleven of the 34 head terms they tracked had lost the rival’s page from page one entirely — with no competing content published against any of them. The rival’s indexed share, meanwhile, had drifted to 31%.
Four things that went wrong
- The dividend went somewhere else. Of the eleven vacated positions, seven were filled by a national marketplace rather than by Pennard. Eviction creates a vacancy, not an inheritance — the selection rule fills it with whoever it prefers next, and that is rarely the second-best specialist.
- The instrument diagnoses supply, not demand. Four of the fourteen placements pointed at starving pages that turned out not to be worth feeding. Days Since Last Crawl tells you which pages are undernourished. It says nothing about which ones earn.
- Dating the pages created a liability. Once 41 pages carried visible revision notes, a stale one became conspicuously stale. Two went eight months without a substantive update and read worse than if they had never been dated at all.
- The rival stopped being a scaled estate. In July 2026 the parent acquired a fourteen-year-old regional competitor and began migrating the estate onto that domain. Capital bought the one input this framework treats as unbuyable, accumulated demand, and the crawl argument stopped applying almost overnight. What Pennard had was a timing advantage, not a structural one.
Where this argument is weakest
The strongest objection is not that the data is wrong. It is that the mechanism is being oversold.
Google says crawl budget is not a concern below roughly a million pages, and Illyes has confirmed that threshold has not moved. Most rivals are nowhere near it. So this is thin-content-does-not-rank in accounting clothes. That objection lands on the capacity term, and on capacity it is right — I would concede it without hedging. Almost nobody reading this is capacity-constrained.
It misses on four counts. First, the claim is about demand, and demand is not capped, it is allocated — a return threshold, not a ration, which punishes low-demand URLs harder than an even division would. Second, it is a claim about the margin, not the mean: an estate can be healthy in aggregate while its bottom decile is being evicted, and the bottom decile is exactly where the scaled expansion happened. Third, the instrument reads observed intervals from the URL Inspection API, so it does not require you to believe any particular mechanism. Fourth — the honest bound — for a 200-page site acquiring links at a normal rate, none of this binds. The argument is about your rival, and about not becoming one.
The second objection is harder and I have no complete answer to it. The fastest-growing version of this problem is not a spam site. It is an established brand with an aged domain bolting a scaled programme onto real accumulated demand. Against that opponent the crawl argument gives you very little for twelve to eighteen months, because they have demand to spend down. What it predicts is that they spend it — that refresh gets pulled away from the pages that were earning and toward thousands that generate none. That is a slower and less satisfying prediction, and Pennard’s fourth negative is what it looks like when the timeline runs against you.
The Monday checklist
- Pull Days Since Last Crawl for every revenue-relevant URL through the URL Inspection API. Sort descending. Look at the top twenty.
- Compute eviction exposure: share of those URLs above 130 days. Write the number down; you will want the baseline.
- Divide your referring domains by your indexed page count. Do the same for your two closest rivals.
- Download each rival’s XML sitemap. Compare declared URLs to indexed count, and check whether any lastmod value differs from its publish date.
- Publish one new URL per rival domain and time how long it takes to appear. Repeat in ninety days and compare.
- Take your next four placements off the Days Since Last Crawl list rather than off the money-page list.
- Add one column to your prospecting sheet: how often is the linking page itself refetched? Approximate it from the domain’s sitemap.
- Identify rival ranking pages above 130 days since last crawl. Remove them from your target list — that is the dividend, and you do not pay for it.
- Pick the ten pages you would genuinely defend and commit to a real revision schedule. Delete or consolidate anything you would not.
- Stop paying for AI detection on competitor content. Move the budget to the row that moves: referring domains per page.
The rival’s page count was never the thing to be afraid of. It was always the thing they had to pay for. Every article they publish enlarges a liability they are servicing out of demand they did not earn, and the only counter-position worth holding is a small estate where every page is individually financed. That is not a content strategy. It is what link building is for, stated in the units the index actually settles in — and the wider strategy set, the tooling and the benchmark data all read differently once you price a page as rented rather than owned.
