Proving Human Authorship

Proving Human Authorship in an AI-Saturated Web

TL;DR

You cannot prove human authorship. There is no test for it and there never was one. What you can do is satisfy a standard, and six institutions currently run six different standards. Five of them are satisfied by content a model drafted. The sixth, the stylometric score sold by detector vendors, is the only definition under which machine drafting is disqualifying, and it is also the only one with no forum, no appeal and no standing.

The flagship human-authorship certification in the world does not examine a single manuscript. It verifies an identity, takes a signed promise and charges a fee. That is not a flaw in the scheme. That is what proof of authorship has always been, in every field that has ever had to rule on it.

The operational consequence for earned media: the only authorship evidence you cannot manufacture on demand is a dated record of a human act held by somebody who is not you. Commission the record of the act, not the article.

The word doing all the work

Every demand to prove human authorship carries a hidden assumption: that human-authored names a property the document holds, the way a file holds a size or a modification date. It does not. Text does not record its own production. A paragraph you rebuilt nine times over an afternoon and a paragraph a model produced in 800 milliseconds arrive at the reader as the same object, and there is nothing inside the object that separates them.

That would be a philosophical curiosity if the demand were rare. It is not. In 2026 it arrives from clients, professional bodies, procurement teams, marketplaces, publishers and regulators, and it almost always arrives phrased as though the answer were a document you could go and fetch. The first useful question is not how to prove it. It is what exactly is being claimed.

What does “human-authored” actually mean?

It means whatever the party asking has defined it to mean, and no two of them have defined it the same way. There is no shared referent. Every institution that has had to rule on authorship has set a threshold somewhere on a continuum, and the thresholds sit so far apart that one article can be human-authored to five of them and machine-generated to the sixth, at the same instant, without a word changing.

You can watch the instability inside the most-quoted number in the field. Graphite’s study of the AI-saturated web, the source of the finding that roughly half of new articles are machine-written, classifies an article as AI-generated when a detector attributes at least half the content to a model. The 50% line is a choice. So is the instrument: Graphite’s May 2026 update ran 55,400 Common Crawl URLs through Pangram, GPTZero and Copyleaks rather than a single detector, and the headline share came in an average of 3.3 points lower than the earlier version. Axios, reporting the update, added the caveat that matters more than the figure: most articles are no longer written purely by a person or purely by a machine.

So the headline statistic about machine-written content is a threshold applied to a continuum by an instrument with a known error rate, and it moves when you change the instrument. That is not a criticism of Graphite, whose methodology is unusually transparent. It is the shape of the whole problem. The category is manufactured at the point of measurement, which means the sentence this was written by a human has no truth value until somebody names the standard.

Six institutions, six definitions

Six bodies currently issue rulings on authorship that carry consequences. It is worth setting them side by side, because the pattern is not subtle once you do.

The US Copyright Office applies a test of expressive control. Works generated wholly by a machine are not registrable; works where a human determined the expressive elements are. Prompting alone does not qualify, however detailed the prompt. The Office had registered more than 6,000 human-AI collaborative works by April 2026, which is the point: AI drafting does not defeat the standard, it triggers a disclosure obligation and a limitation of claim. Nobody runs a detector. The applicant describes the human contribution, under penalty of perjury, and an examiner reads the description.

UK statute runs in the opposite direction. Section 9(3) of the Copyright, Designs and Patents Act 1988 grants 50 years of protection to a computer-generated work with no human author, and names as its author the person by whom the arrangements necessary for creation were undertaken. On the Government’s own reading in the March 2026 report on copyright and AI, for a general-purpose model responding to a prompt that person is usually whoever typed the prompt. British law currently makes the prompter the statutory author of a purely machine-generated text, and the Government’s stated preference is to repeal the provision. The legal definition of authorship in this jurisdiction is mid-flight.

Scholarly publishing has the most developed position, and it is a position about responsibility. The ICMJE revised its Recommendations in January 2026 with a formal AI section: a model cannot be listed as an author because it cannot be accountable for the accuracy and integrity of the work. Authors must disclose if, where and how AI tools were used, down to language and grammar correction, and undisclosed use may constitute misconduct. COPE, Elsevier, Springer Nature, PLOS and Wiley converge on the same rule for the same reason. AI assistance is a disclosure matter, not an authorship matter.

UK press regulation does the same thing without writing a new rule. IPSO’s position, published in January 2026, is that the Editors’ Code applies to all published content whatever tools were used in its creation, and that Clause 1 already prohibits inaccurate, misleading or distorted information including through misleading attribution. IMPRESS went further and amended its Standards Code to require human editorial oversight of AI-generated content, clear labelling and verification steps. Two regulators, two postures, and neither of them adjudicates on whether a machine drafted a sentence. They adjudicate on whether somebody stands behind it.

Google has held the same line since February 2023 and restated it through 2026: the focus is the quality of content rather than how it was produced, and there is no AI-detection ranking rule. What the quality rater guidelines do flag is fabricated attribution, the article that says in my experience as a dentist when no dentist was involved. That is a rule against lying about a person, not against using a tool. Oversight, not authorship, is what is being measured, which is worth holding onto if you are still budgeting against AI Overviews and the backlinks that feed them.

Then there is the detector vendor, whose definition is stylometric: distance between your text and a reference distribution. This is the only one of the six under which a model in the drafting loop is disqualifying by construction. It is also the only one with no hearing, no appeal, no named adjudicator and no remedy.

THE SIX AUTHORSHIPS

Who asksTheir definitionWhat discharges itForum, and does AI drafting defeat it?
Copyright registrar (USCO)Human control of the expressive elementsDisclose the AI material, describe and claim the human contributionExaminer and the courts. No: 6,000+ human-AI works registered by Apr 2026
UK statute (CDPA s.9(3))The person who made the arrangements for creationNothing. Authorship is assigned automaticallyThe courts. No: the prompter is the statutory author (repeal proposed Mar 2026)
Scholarly publisher (ICMJE, COPE)Accountability for all aspects of the workNamed human authors plus disclosure of if, where and how AI was usedEditors, retraction, misconduct findings. No, if disclosed
Press regulator (IPSO, IMPRESS)Accurate and not misleadingly attributedHuman editorial oversight, a named responsible party, prompt correctionsComplaints and adjudication. No: IMPRESS requires labelling, not abstention
Certification mark (Authors Guild)The expression emanated from human intellect, de minimis tool use allowedIdentity verification, a signed licence, a fee, a public registration numberLicence revocation. In principle yes, but nothing is examined
Detector vendorStylometric distance from a reference distributionNothing you can produce. The score is computed from the text aloneNone. Yes, by construction

Five of the six are definitions of responsibility. One is a definition of production. The field has spent two years optimising against the one that no law recognises, no tribunal convenes and no engine consumes.

The certification that certifies nothing

If human authorship could be proved, the organisation with the strongest incentive to prove it would have built the proof. The Authors Guild launched Human Authored in beta for members in January 2025 and took it public on 2 March 2026, opening it to all authors of US-published titles and, a week later, to publishers buying in bulk. More than 3,000 authors had certified around 5,000 titles by then. Each title gets a numbered registration in a public database and a trademarked mark for the cover.

The process is: register, submit to third-party identity verification, sign a licence agreement, pay ten dollars. Certification is free for members because members are pre-verified. Authors are capped at ten titles a year without special permission.

What the process does not include is any examination of the manuscript. Guild chief executive Mary Rasenberger said so directly to Publishers Weekly in March 2026: the programme does not vet manuscripts for AI content before issuing certification, because no reliable detection method exists. The multi-step process and the fee are the deterrent, on the reasoning that scammers do not put much effort into things.

Read that again with the thesis in hand. The most serious human-authorship certification in the world verifies who you are, binds you to a promise with a revocable trademark licence, and imposes friction. It verifies nothing about the text. And it is not a broken scheme: it is a well-designed one, because it correctly identified that the only tractable object in the transaction is an accountable person.

The carve-out is the other tell. The mark permits AI for spell-checking and research, and the qualifying language is that the literary expression emanated from human intellect. That is the Copyright Office’s expressive-control test in different words. Even the purist scheme could not write a bright line, because there isn’t one to write.

Key takeaway

Every certification scheme that has actually shipped resolves to the same three components: an identity, a promise and a consequence for breaking it. None of them contains a test. When somebody offers you proof of human authorship, the useful question is which of those three they are actually selling.

Why the sixth definition has no forum

Detector vendors are the only party whose definition would make your drafting method dispositive, so it is worth being precise about what their instruments do and do not measure.

The accuracy gap is not the main problem

The published gap between vendor claims and independent testing is wide and well documented. Copyleaks advertises 99.12% accuracy and Scribbr’s twelve-tool benchmark measured it at 66%. GPTZero claims 99% with a zero false-positive rate; the same benchmark put it at 52% overall. Liang and colleagues, publishing in Patterns in 2023, found false-positive rates as high as 61% for non-native English writers, a finding nobody has overturned. Turnitin’s roughly one percent flags hundreds of genuine essays at institutional scale. Vanderbilt, Georgetown, UC Berkeley and Curtin have all switched their detectors off.

But accuracy on clean samples is the wrong thing to argue about. A 2026 evaluation by Hadra, Cambridge and Mesbah at Sultan Qaboos University found leading tools at 69% and 61% overall accuracy, with performance on hybrid human-AI texts falling to nearly zero, and accuracy on scientific prose 28 to 38 points below humanities prose.

Hybrid is the modal case. Graphite’s own caveat says most articles are no longer purely one or the other. So the instrument collapses precisely on the population it is being asked to classify, and it collapses hardest on technical writing, which is what most B2B content is. A tool that cannot separate the majority category is not a weak tool. It is not a measurement at all.

An accusation with no hearing attached

The structural point is worse than the statistical one. Take the five institutional standards: each names an adjudicator, a route to be heard, and a remedy. An examiner can be argued with. A journal has a corrections and retraction process. IPSO publishes adjudications and requires them to run with due prominence. A trademark licence can be revoked and disputed.

A detector score has none of that. It produces a number, the number travels, and there is no body that convenes, no evidence that is admissible, and no finding that can be appealed. The freelance market has been living this since 2025: work rejected and invoices unpaid on the strength of a percentage, with no process behind it. Upwork has never documented platform-level scanning of deliverables; its guidance turns on client agreement rather than automated checking. The scanning that hurts people is happening at the client’s desk, informally, with a free tool.

This matters for anyone treating detector scores as a link building tool category. A metric with no forum cannot be optimised against, only feared.

What the proof stack actually proves

The incumbent playbook for proving human authorship is now well formed, and it is worth taking seriously before taking it apart. It has four components: retain version history, deploy an authorship-tracking layer such as Grammarly Authorship or Turnitin Clarity, publish author schema and contributor pages, and publish an AI-use policy.

Each of these is a good idea. None of them proves what it is sold as proving.

What does version history prove?

It proves the shape of an insertion, not its source. Google Docs records that 900 words appeared at 14:32. It does not and cannot record where they came from: a model, a notes app, an earlier draft, an email, a colleague. A slow accretion of small edits across three sessions is consistent with human composition and it is also consistent with a person retyping generated text, which is why extensions now exist for the sole purpose of replaying a paste as simulated typing. The evidence is one-directional. A messy history supports your account; a clean one does not refute it, and cannot be made to.

For example: a technical writer who drafts in a plain editor, pastes the finished piece into the CMS once and never revises it has produced a perfect human-authored article with the worst possible version history. A content farm that runs generated text through a typing simulator has the best. The instrument sorts working habits, not authorship.

Schema, policies and the self-attestation ceiling

Author markup and a published AI policy do real work, but the work is legibility, not credibility. Both are self-emitted. They tell a reader and a parser what you say about yourself, and a party that has decided to doubt you is precisely the party for whom self-assertion carries no weight. This is the same ceiling that governs signed and unsigned content and the same one that limits what C2PA can contribute to a link: a manifest is a statement by the party with an interest in the answer.

There is also a labelling regime forming around all of this that operates on generation rather than authorship. The EU AI Act’s transparency duties came into force on 2 August 2026 and fall on providers and deployers of the systems, not on the person who published the page. Google Merchant Center already requires AI-generated product imagery to carry IPTC DigitalSourceType metadata and AI-generated titles and descriptions to be labelled separately. YouTube has run mandatory disclosure since May 2025. None of these is an authorship test either, for the reasons set out in the SynthID and AI-content labelling analysis.

Key takeaway

Every element of the proof stack is retro-fittable. You can add author schema, a policy page, a contributor bio and a version-retention habit this afternoon and backdate none of them. That is precisely why they carry so little evidential weight, and it points at what does.

The one thing you cannot manufacture

Set the detectors aside and ask the question a tribunal would ask. If you had to establish, to a sceptical party, that a human being did the work behind a piece of content, what would you produce?

Not the text. Not metadata about the text. You would produce evidence that a person did something on a date: attended a site, ran an inspection, gave evidence, sat on a panel, answered a journalist’s question, was named in a case report. And the evidence would be worth having in proportion to who was holding it. A record you keep about yourself is an assertion. A record somebody else made about you, at a time you did not control, is a different class of object.

That is a description of earned media. Not of its ranking value, which is the usual argument and is covered adequately in the link building strategies guide, but of its evidential value, which nobody prices. When a trade title quotes a named person at your firm, three facts become externally recorded: that the person exists, that they said the thing, and that they said it before the date on the page. You cannot generate any of them.

The Act-or-Text Split

Run your referring domains through two questions and the estate sorts into four cells with very different evidential value. The first question: does this artefact record a human act or only text? The second: who holds the record, you or a third party?

 Third party holds the recordYou hold the record
Records a human actTESTIMONY A named quote in a news story, a case note, a panel listing, a witness statement, a published inspection. Dated, unfakeable, not yours.ASSERTION Your own case study, methodology page or field diary. True, useful, and worth exactly what an interested party’s account is worth.
Records text onlyPUBLICATION A guest post, a contributed column, a directory entry. Evidence that an editor accepted words. Not evidence that a person did anything.INERT The owned estate. No amount of schema, provenance metadata or policy language moves a cell in this quadrant.

The distinction cuts across everything the field normally screens for. A quote in a regional trade title and a guest byline on the same domain have identical authority metrics and sit in different cells. Most competitor backlink analysis will tell you nothing about which is which, because the split is a property of what the page records, not of the domain that hosts it. It also cuts in unexpected directions: a Product Hunt launch or a Hacker News thread records a dated act, whereas a listicle placement records only text. And even an original interactive asset sits in the Inert cell until somebody who is not you cites it.

The operating rule follows: commission the record of the act, not the article. A digital PR brief that asks for an explainer placement produces a Publication. The same budget spent putting a named specialist on the record commenting on a tribunal decision, a regulatory change or a published dataset produces Testimony, and it usually costs less because you are supplying a comment rather than an asset.

The Backfill Rule

No external record can carry a date earlier than the day you asked for it.

Everything else in the authorship stack is retro-fittable: schema, bylines, policy pages, review workflows, version retention. A third-party dated record is the only component that accrues strictly forward. Which means the evidential capacity of your estate in 2028 is a function of what you commissioned in 2026, and no budget applied later can close the gap. It is the same asymmetry that makes link velocity a history rather than a lever.

Worked example: Ledbury and Nash

Ledbury and Nash is a Bristol employment-law firm with about 5.4 million pounds in annual fee income and a 240-article guidance library that had been its main acquisition channel for six years.

In February 2026 a professional body running an approved-resource list for HR practitioners introduced an AI-content policy and ran the firm’s library through a commercial detector. Sixty-one pages scored above the threshold. The firm was removed from the list. Referrals from it had been running at 14% of new enquiries.

What the proof stack cost, and what it bought

The first response was the incumbent playbook, and it took ten weeks and about 19,000 pounds: author schema across the library, contributor pages with practising-certificate details, a published AI-use policy, Grammarly Authorship rolled out to nine fee-earners, and a version-retention rule. On re-scoring, 44 of the 61 pages still flagged. Nothing in the stack was being read by the instrument, because the instrument only ever reads the text.

What resolved it was a different move. The firm stopped trying to rebut the score and asked the professional body what its actual admission standard was. The answer, once written down, was the ICMJE shape: a named responsible reviewer holding a current practising certificate, a documented review step, and a corrections route with a published response time. That cost 6,200 pounds a year, mostly the supervising partner’s time. The firm was readmitted in nine weeks.

The sixty-one flags were never withdrawn. They sit on file. The firm was readmitted because it satisfied a different standard, not because it defeated the accusation, which is the entire argument of this article compressed into one administrative outcome.

Applying the split to the estate

Running the Act-or-Text Split across 180 referring domains produced an uncomfortable number: 23 recorded a human act. Nine were named quotes in HR and legal trade press, six were tribunal case notes where a partner was on the record, five were conference and panel listings, three were consultation responses published by a government department. The other 157 were Publications and Inert: guest posts, resource-page entries, directory listings and the firm’s own library.

The brief changed accordingly. Over the following twelve months the firm stopped pitching explainer articles and put a named partner on retainer-style availability for comment on tribunal decisions, which two trade titles took up within a month. By June 2026, answers naming the firm across 40 monitored employment-law prompts had gone from 7 to 24, and approved-list referrals had recovered to 12% of enquiries. The pattern is consistent with what the SERP-less audit framework predicts when the missing locus is corroboration rather than retrieval, and it is worth noting that the recovery is concentrated in exactly the query space that recruitment and HR-tech link building competes in.

What it cost that nobody puts in the case study

The named-editor model concentrated roughly six hours a week of unbillable review on one partner, and that partner is now a single point of failure the firm has not solved. Two of the flagged pages were deleted rather than defended, because defending them cost more than their traffic was worth; both had been genuine human work. And one trade title walked away from the commentary relationship because the firm would not comment on live cases, which is the correct professional answer and an unrecoverable loss of a Testimony source. Buying Testimony means accepting that the counterparty sets terms you sometimes cannot meet.

Where this argument is weakest

The strongest objection is not that detectors will get better. It is that enforcement follows what is cheap to measure, not what is correct. Bad proxies become binding all the time, precisely because they are administrable: breathalysers, standardised tests, credit scores. If enough clients, marketplaces and procurement teams start running classifiers, the stylometric definition acquires a forum by fiat, and the claim that it has no standing becomes false. In education and in freelance contracting this has already happened. Ledbury and Nash lost a channel to it.

That objection is correct and it is not adequately answered by pointing at accuracy studies. But it is bounded in four ways.

One: a proxy survives where it improves a decision. Education has a binary decision to make and no better instrument. A search or answer engine’s decision is which source to cite, and a stylometric flag does not predict accuracy, usefulness or corroboration. Google has stated for three years that it does not target AI detection and its rater guidelines flag fabricated attribution rather than machine production. The proxy has nothing to improve.

Two: the false-positive cost has to fall on someone who cannot leave. A student cannot exit the institution. A publisher can exit a platform, and a platform that torches its supply loses its corpus. This is why Vanderbilt, Georgetown, UC Berkeley and Curtin turned their detectors off rather than defend them, and it is a real constraint on how far the practice spreads into commercial content.

Three: the instrument degrades as the base rate rises. Near-zero accuracy on hybrid text is survivable when hybrid is rare. When roughly half of published articles sit somewhere on the continuum and most professional work is assisted somewhere, the flag stops separating anything and its administrative usefulness decays on its own.

Four, and decisively: even where a flag binds, the remedy is not proof. Ledbury and Nash did not rebut their scores and could not have. They supplied a different standard. Wherever a stylometric flag has real consequences, the exit is always through one of the other five doors, which means the playbook this article recommends is the correct response to the objection rather than a casualty of it.

There is a second, quieter objection: that disclosure regimes will simply resolve the ambiguity. The CHI 2026 study of AI disclosure in freelance work, by contrast, found that policies were so inconsistent in defining AI use that workers systematically misread them, and that the most common strategy was passive disclosure, revealing use only when explicitly asked. Where AI was prohibited outright, workers read the prohibition as a reason to stay silent. Disclosure regimes do not fix an undefined predicate. They relocate it.

What to do on Monday

A literal sequence. It takes about a fortnight of part-time work and most of it is not content production.

  • Name the standard in force. For every party that could ask you to prove authorship, write down who adjudicates, what would satisfy them and what happens if it does not. Most of the list will have no adjudicator, which tells you the exposure is reputational rather than procedural.
  • Install the accountability furniture. A named responsible reviewer with real credentials, a documented review step, a corrections policy with a stated response time, and a published AI-use note. This is the shape every institutional standard converges on, so it clears five of the six doors at once.
  • Run the Act-or-Text Split across your top 150 referring domains. Count the Testimony cell. If it is under 15% you have an estate with no external evidential capacity, whatever it is worth for ranking. Do this before your next backlink review rather than after.
  • Rewrite one digital PR brief. Take a live brief asking for a placement and change the ask to a record of an act: a named comment on a decision, a data submission, a consultation response, an on-the-record inspection. Same budget, different cell.
  • Identify three recurring events you already attend or produce and arrange for a third party to record them. Panels, industry consultations, published test results, regulator submissions. These are cheap because you are already doing the underlying work and only the record is missing.
  • Screen prospects for a forum. An outlet with a corrections policy and a regulator behind it produces a placement with an appeal route attached. An unregulated contributor network does not. Add the question to your prospecting sheet alongside whatever your link building specialist already screens for.
  • Stop paying for detector scores as a QA gate. If you must know, know it, but do not build a workflow around a number with a 61% false-positive rate on non-native writers and near-zero accuracy on the hybrid text your team actually produces.
  • Retain the process record anyway. Not because it proves anything, but because in the one setting where it matters, a client dispute, an account of your process supported by contemporaneous artefacts is what a reasonable counterparty will accept. Cheap insurance, honestly labelled.

The reframe underneath all of it is small and load-bearing. You are not in the business of proving that a machine was absent. You are in the business of making it unambiguous who answers for the work, which is the only question any institution has ever actually asked, and the only one a model cannot answer on your behalf. That is also, incidentally, why the authenticity premium accrues to settlement rather than to sincerity, and why training-source strategy and technical SEO for link building cannot substitute for it. The 2026 link building statistics and the beginners guide set the baseline.

Leave a Reply

Your email address will not be published. Required fields are marked *

AI-Content Labelling Previous post SynthID, Watermarking and AI-Content Labelling: Implications for Earned Media
Verifiable Author Identity Next post Verifiable Author Identity: sameAs, Credentials and Trust Chains