TL;DR: Agent marketplaces are shipping identity, reputation and validation registries, and the standard advice is to get in early and start accumulating ratings. The first field studies of a purpose-built agent reputation registry found that most identities were placeholders and most reviewers were coordinated Sybils. That is not an implementation bug. A reputation carries information only in proportion to what the identity holding it would cost to abandon — and an agent identity costs almost nothing to mint. Every workable system therefore binds the cheap identity to an expensive one and scores the expensive one. Which means you cannot build a reputation inside an agent marketplace. You can only import one.
1. The first purpose-built agent reputation system, and what the field data showed
The pitch is clean, and you have probably heard some version of it. Autonomous agents are about to transact across organisational boundaries at volume. They need a way to answer one question about a counterparty they have never met: can this one be trusted? So the ecosystem builds registries — identity, reputation, validation — and the advice that follows is immediate. Register early. Start accumulating ratings. Build your agent reputation before your competitors build theirs, because the agents doing the buying will read those scores and the longest, cleanest histories will win the work.
That advice is wrong, and the reason it is wrong is not a matter of opinion. We now have field data.
ERC-8004, the “Trustless Agents” standard, is the first permissionless trust layer built specifically for agent economies — three on-chain registries covering identity, reputation and validation, designed to let agents discover and evaluate each other without any pre-existing relationship. It is exactly the thing the advice assumes. In mid-2026, a team from Imperial College London, Ohio State, Bristol, CSIRO and Manchester published the first empirical study of it, crawling registration events, off-chain files and payment transactions across Ethereum, BNB Smart Chain and Base through 13 May 2026.
Two findings matter here. On identity, most registrations were placeholders rather than working agents: only 3%, 4% and 15% of registrations across the three chains exposed a valid registration file with at least one live service endpoint. On reputation, the registry could not function as a trust signal at all — values were not commensurable across raters, feedback records were rarely grounded in verifiable interactions, and the score could be manipulated at minimal cost. Consistent with that, the authors found that 73.5%, 59.2% and 90.6% of reviewers across the three chains exhibited coordinated Sybil behaviour.
A separate June 2026 measurement study of the same standard found the ecosystem “registration-heavy but operationally shallow”: of 10,000 registered agents, 67 exposed service records, 628 had received any reputation feedback at all, and 19 combined metadata, services, feedback and cross-chain registration. Ownership was concentrated — the top ten owner wallets held 51.4% of agents — and so was the feedback: a single client account contributed 65.82% of all feedback records in the dataset.
Read those numbers as a market, not as a technology story. The cheapest thing to manufacture in an agent marketplace is a five-star history. The second cheapest is the identity that holds it.
What is an agent reputation system?
An agent reputation system is any mechanism that records the outcomes of an agent’s past interactions and exposes a summary of them so other agents can decide whether to transact. In 2026 these take three forms: on-chain registries such as ERC-8004, venue-native feedback attached to a marketplace account, and validation layers where an independent party or a trusted execution environment attests to what an agent actually did. All three answer “what happened before?” None of them, on their own, answer “and who exactly did it happen to, and to whom?”
That second question is the subject of this article. It is also the oldest question in the fundamentals of link building — a citation is only worth something because somebody with something to lose put their name on it.
2. Why a reputation system is only as strong as the name it hangs on
In 2001, Eric Friedman and Paul Resnick published a paper in the Journal of Economics & Management Strategy called “The Social Cost of Cheap Pseudonyms.” Their question was simple: what happens to cooperation when participants can costlessly change identity? Their answer was formal and bleak. When names are free and replaceable, misbehaviour carries no reputational consequence, because the misbehaving party simply becomes someone else. Cooperation can still emerge, but only through a convention in which newcomers “pay their dues” by accepting poor treatment until they have built a history worth protecting. And they proved that no equilibrium sustains substantially more cooperation than that dues-paying one. The mistreatment of newcomers is not a design flaw in such systems. It is the price of allowing free name changes.
That result, twenty-five years old, is the most useful thing anyone has written about agent marketplaces.
Three attacks, one underlying property
The classic failure modes of reputation systems look distinct and are usually treated as separate engineering problems. They are not. Each is priced by the same variable.
- Whitewashing: shedding a bad history by re-registering. Priced by the cost of a new identity.
- Sybil attacks: forging many identities to outvote or out-rate honest participants. Priced by the cost of a new identity, multiplied.
- Ballot-stuffing and reciprocity rings: manufacturing feedback between accounts you control or trade with. Priced by the cost of the accounts doing the rating.
When the cost of a name approaches zero, all three approach free simultaneously. This is why the ERC-8004 numbers are not surprising and why a better scoring algorithm would not have saved them. You cannot compute your way out of an identity that costs nothing.
The formulation worth carrying into every conversation about this: reputation is credit, and identity is the collateral. A lender extends credit against something the borrower would forfeit by walking away. A market extends trust against something the counterparty would forfeit by becoming someone else. A key pair posts no collateral, so it can borrow nothing — no matter how good its payment record looks, because that record was written by parties who also posted nothing.
For example, freight has been running this experiment in the physical world for years. A US motor carrier authority costs a few hundred dollars. The standard fraud pattern is a freshly registered authority that books loads for two to four weeks, collects payment and vanishes; double-brokering complaints to the Federal Motor Carrier Safety Administration ran to roughly 8,000 in 2025, against around 2,000 in 2021. The remedy now deployed is instructive: the FMCSA’s Motus system adds government ID verification and a facial scan at registration. Nobody built a cleverer rating algorithm. They made the name expensive.
Key takeaway: A reputation score is a claim about the past attached to a name. Its information content is bounded above by what that name costs to abandon. If the name is free, the score is free, and everyone in the market eventually prices it that way — which is why newcomers get punished by default in every cheap-identity venue ever studied.
3. The Identity Collateral Table
Several identities are in play in a single agent transaction, and they are routinely discussed as though they were one thing. They are not, and they differ on the only dimension that matters. The table below sorts them by collateral — by what the holder would forfeit if the identity were abandoned and re-minted tomorrow.
Read the third column downwards; it predicts the other three completely.
| Identity in play | Collateral posted | Cost to abandon and re-mint | What it can therefore underwrite |
| Agent key or instance (key pair, registry row, Agent Card) | None | Seconds, at near-zero cost | Nothing. It names a running process, not an accountable party. |
| Marketplace or registry account (listing, feedback history) | Its own accumulated history | Low — a fresh account, sometimes a fresh wallet | Continuity inside one venue, until that venue changes its trust model. |
| Posted stake, bond or escrow | Cash at risk | The value of the stake — real, but purchasable | A bounded promise on a single transaction, sized to the bond. |
| Payment-rail identity (KYC’d merchant, Visa TAP credential, KYA binding to an operator) | Banking relationship, chargeback exposure, underwriting | High — re-onboarding, and the history restarts | That a real, accountable party stands behind the request. |
| Domain and published estate | The record others have built against it | Ten pounds for the name; years for the record | Continuity of identity across venues and across engines. |
| Legal entity and its external record (registrations, audited credentials, trade-body membership, press, independent testing, citations) | Everything the entity has ever done in public | Cannot be re-minted — you do not hold the pen | Standing with a counterparty who has never met you and cannot ask around. |
Two things fall out of the table that are not obvious before you build it.
The first is that the domain row is where most teams misread their own position. The name costs ten pounds. What is expensive is the record other parties have built against it over years — and that record is precisely the part you do not control, which is exactly why it is worth something. The same logic underwrites measuring entity authority as a discipline: an entity is credible in proportion to the corroboration it did not write.
The second is the direction of the last row. Every identity above it can be replaced by spending money. The bottom row cannot be replaced at any price, because its collateral is held by other people — a liability if your public record is bad, an unassailable asset if it is good, and the only row with that property.
4. What real marketplaces are actually building (and it is not reputation)
If the argument above is right, you would expect the serious infrastructure players to be doing something other than building better scoreboards. They are.
By April 2026 every major payment network had shipped or announced a primitive under the loose label KYA, Know Your Agent, an identity framework that cryptographically binds an agent’s request to a registered operator and an authorised user so a merchant can verify both before settling. Visa’s Trusted Agent Protocol issues an identity credential to the agent and verifies three signatures at transaction time: legitimacy, delegator, and payment method. Mastercard’s framework bundles developer provenance, user binding, permission scopes, telemetry and risk scoring into a Digital Agent Passport. Google’s AP2, the Agent Payments Protocol, carries signed user mandates so the authority behind a purchase is provable after the fact. A2A version 1.2 added signed Agent Cards tied to domain verification. Web Bot Auth, built on HTTP Message Signatures, lets an agent prove it is the crawler it claims to be rather than a spoofer wearing its name. Validation registries and trusted execution environments attest to what code actually ran.
Line those up and the pattern is unmistakable. Not one of them is a reputation system. Every single one is a mechanism for binding a cheap identity to an expensive one — an operator with a company behind it, a merchant with an underwritten bank relationship, a human with a verified mandate, a domain with a public record, a chip that cannot lie about what it executed. The industry is not trying to make agent scores more accurate. It is trying to make agent names expensive, because that is the only known way to make any score mean anything.
Regulation pushes the same way. The EU AI Act requires operator identity in the behaviour logs of high-risk systems, and NIST has flagged agent identity management as a priority standards area. Compliance obligations are, in economic terms, a mechanism for making an identity costly to hold and costlier to discard — which is why they strengthen reputation systems rather than burden them, a dynamic worth watching if you operate across European markets.
Do I need to register my agent to be trusted?
Register it to be reachable, addressable and verifiable — those are real, and skipping them means agents cannot transact with you at all. But registration is a gate, not a reputation. You clear it once and stop. The mistake is treating the registry listing as an asset that appreciates, when the listing is worth exactly what the identity behind it would cost to replace. Handling the signing keys, the manifest and the verification chain correctly is technical SEO work for link building by another name: necessary, table-stakes, and never the thing that wins the decision.
Key takeaway: The entire 2026 identity stack — KYA, TAP, Agent Passports, signed mandates, signed Agent Cards, bot authentication, attestation — is a collateral-import layer, not a trust layer. Collateral is the input. Reputation is the output. Teams that budget for the output and skip the input end up with a score nobody weights.
5. The two very different things people call “agent reviews”
“Reviews” in an agentic context refers to two objects with almost nothing in common. Conflating them produces most of the bad advice in this category.
What the venue keeps
The first is settlement telemetry: dispute rate, fulfilment rate, cancellation rate, whether the delivered specification matched the advertised one, whether the price charged matched the price shown. It is generated automatically, largely accurate because it is a by-product rather than an opinion, and the most predictive record of whether your next transaction goes well. It is also private, venue-owned and non-portable by default.
What the world keeps
The second is published assessment by parties outside the transaction: review platforms, trade press, independent testing, benchmark participation, professional bodies. It is noisier, slower and less precise about your operations. It has one property the first does not: it exists before your first transaction in any given venue, and it survives your exit from all of them.
The rule that unifies both, and that tells you what any given review is worth: a review is worth what its publisher would lose by publishing it falsely. That is why a testimonials page on your own site carries approximately nothing, why a five-star average on a registry where ratings cost nothing carries nothing, and why an assessment on a platform with regulatory exposure and a brand to defend carries a great deal.
The United Kingdom has made this unusually literal. Under the Digital Markets, Competition and Consumers Act 2024, fake and misleading reviews became banned practices from 6 April 2025, and platforms hosting reviews must take reasonable and proportionate steps to prevent, detect and remove them. The Competition and Markets Authority can now decide a business has broken consumer law and fine it directly, without going to court, up to 10% of global annual turnover. In March 2026 it opened five investigations across the funerals, food delivery and car sales sectors, and in the regime’s first year imposed £4.7 million in fines. The CMA estimates as much as £23 billion of UK consumer spending is influenced by online reviews annually.
Notice what that legislation actually did in the terms of this article. It did not improve any rating algorithm. It attached a large, enforceable cost to publishing a false assessment — it posted collateral on behalf of an entire layer of the market. The regulated human-facing review layer became more expensive to fake in the same period that the machine-facing agent registry layer, which has no regulator at all, was running at up to 90.6% coordinated Sybil reviewers.
Do AI agents read reviews?
They lean on them heavily, but on the accountable ones. SE Ranking’s study of 129,000 domains found that domains listed across multiple review platforms earned an average of 4.6 to 6.3 citations, against 1.8 for domains absent from them. Trustmary’s 2026 analysis reports that ChatGPT references reviews in 58% of responses, Perplexity in effectively all of them, and that 34.5% of Google AI Overviews cite at least one review platform — a pattern consistent with what we already know about how AI Overviews use backlinks as corroboration rather than decoration. The striking part is that those same platforms have lost between 76% and 92% of their organic search traffic. The humans left; the machine reliance grew.
The signal being read is not the star rating. It is the accountability of the publisher. An engine leaning on a regulated platform is leaning on collateral, and this is the mechanism behind most of the factors driving AI product recommendations — which is why a generic five-star profile moves less than a detailed, specific, mid-range assessment on a platform that could be fined for faking it.
6. Reputation you cannot take with you
There is a second problem with venue-native reputation, and it is strategic rather than technical.
A reputation held inside a marketplace is a switching cost the marketplace owns. It is building an audience on rented land, with one twist that makes it worse: the more predictive that record is, the stronger the venue’s incentive to keep it inside. A venue whose settlement history genuinely forecasts supplier reliability holds an asset it would be irrational to export. Accuracy and portability are pulled apart by the same incentive.
Meanwhile the venues are multiplying. Google’s agent marketplace, Anthropic’s enterprise marketplace for Claude-powered tools, AWS and cloud catalogues, on-chain registries, and a growing set of vertical procurement networks — with A2A itself now at 150 organisations running production traffic under Linux Foundation governance. Every new venue resets you to zero, and the Friedman–Resnick result tells you what zero means: newcomers get treated badly, and that treatment is the equilibrium rather than a bug someone will fix.
Which produces the strategic fact of this whole category. If venues keep proliferating, you are permanently a newcomer somewhere. Cold start stops being a phase you get through and becomes a steady state you operate in. And the only thing that arrives with you on day one in a venue you joined this morning is the record that lives outside all of them.
For example: a supplier with four years of spotless settlement history on one network and no external footprint enters a new one with literally nothing. A supplier with identical operations, an audited credential and three years of independent coverage enters already legible to any agent that consults the open web — the same asymmetry visible in AI browsers and what they change for SEO, where what an agent can find about you off your own property decides what it does on it.
There is a diagnostic, and it takes twenty minutes.
THE CLEAN-SLATE TEST
Suppose that tonight every identifier you control is destroyed: agent keys and signing certificates, every marketplace and registry account with its feedback history, API credentials, and your domain. Tomorrow you re-enter the market with the same people, the same warehouse, the same processes, the same prices — under a new name.
Score three lines. (1) What you lose: list every trust asset that disappears with the identifiers. (2) What survives: list every asset held by a third party that would still describe your organisation — audited credentials tied to the legal entity, trade-body membership, independent test results, press coverage, published datasets others have cited, named individuals with their own public record. (3) The ratio: survivors as a percentage of your total trust surface.
Verdicts. Under 20%: your reputation is a tenancy, and your landlords are your counterparties. 20–50%: mixed, and the venue-held half will not travel to the next marketplace. Over 50%: you own it, and every new venue you enter starts warm.
The test is uncomfortable on purpose. Most organisations discover that the assets they spent the most on last year are in line one.
Key takeaway: Venue-held reputation and portable reputation are not two grades of the same asset. One is a tenancy that ends when the venue changes its rules or a new venue opens; the other is a holding that follows you everywhere and cannot be reset by anyone, including you. Budget them separately, because they behave nothing alike.
7. The strongest objection: the settlement ledger beats your press cuttings
Here is the serious counter-argument, at full strength rather than in a form convenient to knock down.
Behavioural settlement data is enormously more predictive of transaction outcomes than any amount of editorial coverage. On-time rate, dispute rate, refund rate, specification accuracy, price-integrity at checkout: these are hard, continuous, machine-verified measurements of the thing an agent cares about. Trade press coverage is soft, sparse, lagging and frequently bought in all but name. As venues accumulate volume, their internal records will dominate selection inside those venues — and deserve to. On this view, “import external corroboration” is a marketer’s consolation prize, offered because marketers cannot influence the ledger.
Concede it. All of it. Where a venue has genuine settlement history on you, that history should and will outrank anything a journalist wrote, and any framework claiming otherwise is selling something. But four boundaries hold, and each one is load-bearing.
First, a settlement ledger can only price the second transaction. It is silent on the first, by construction, because there is nothing in it yet. In a market with one dominant venue that is a brief inconvenience. In a market where venues proliferate — which is the market we have — first transactions never stop occurring, and the ledger is structurally absent exactly when the decision is hardest.
Second, predictiveness and portability trade against each other. The better a venue’s record forecasts outcomes, the more valuable it is as a moat and the less likely it is ever to be exported to a rival. So the most accurate reputation in the market is also the least transferable, and betting your standing on it is a bet that one venue wins permanently.
Third, the ledger measures execution reliability, not suitability. A spotless dispute rate says you deliver what you sold. It cannot tell an agent whether you are the right counterparty for an unusual requirement, a regulated edge case, or a specification nobody has ordered before. Those judgements are made from published, external, human-authored material — the material an agent can actually read about you.
Fourth, and most damaging to the objection: the ledger is only as good as the identity binding underneath it. A behavioural record attached to a cheap name is a behavioural record you can walk away from on a Tuesday and rebuild on a Wednesday. The 73.5% to 90.6% Sybil rates in the ERC-8004 data are what an unbound ledger looks like in the field. Every venue that wants its record to mean something must therefore import binding from outside itself — from a payment network, a company registry, a regulator, an audited credential. The external layer is load-bearing even inside the venue that is trying to replace it.
What would falsify this? A genuinely portable behavioural reputation: settlement history that travels across venues in a standard format, honoured by competing marketplaces, with real adoption rather than a specification and a press release. If an agent passport carrying verified dispute rates becomes something rivals accept from each other, the venue record becomes durable and portable at once, and the ranking in section three inverts. Watch the KYA credential schemes and the payment networks for exactly this — they are the only parties with both the data and the cross-venue standing to attempt it. So far, every deployed scheme exports identity and keeps performance.
8. Worked example: two ways to spend £180,000
Ryburn & Coe is invented, and the figures are illustrative, but the shape is one you will recognise. It is a Leeds-based UK customs brokerage and freight forwarder: roughly £38 million in revenue, 130 staff, about 1,100 SME importer clients, and a growing share of enquiries arriving through agent-mediated procurement rather than a human phoning for a quote. The board approves £180,000 for 2027. Two plans compete.
| Budget line | Plan A — farm the registries | Plan B — import the collateral |
| Venue integration and listings | £70,000 — four venues, deep integration | £25,000 — same four venues, minimum viable |
| Ratings programme | £60,000 — incentivised post-shipment feedback plus a reciprocal review arrangement with two partner firms | £0 |
| Credentials and identity binding | £0 | £45,000 — AEO scope extension, BIFA membership upkeep, signed Agent Cards, entity records reconciled across every registry an agent might consult |
| Published performance dataset | £0 | £50,000 — quarterly clearance-performance data, third-party verified |
| Earned corroboration | £30,000 — paid placement in venue promotion slots | £40,000 — trade press, independent benchmarks, profiles on the two platforms buyers’ agents actually cite |
| Measurement | £20,000 — rating dashboards | £20,000 — shortlist rate, first-booking win rate, portability audit |
Baseline at month zero was identical: Ryburn appeared in 6% of observable agent-generated shortlists, and took 41 first-time agent-mediated bookings in the previous quarter.
Months one to three. Plan A moves faster and looks better in every board pack. Roughly 300 feedback records accumulate across two venues at a 4.9 average. Plan B has almost nothing to show: an AEO scope extension in progress with HMRC, the first quarterly clearance dataset published — entries cleared without amendment, median clearance time, detention rate — and two trade titles that picked it up because nobody else in the sector publishes the numbers.
Months four to six. One venue runs a Sybil sweep and deduplicates feedback by payer identity. Plan A’s 300 records resolve to 34, because most originated from three related accounts inside the reciprocity arrangement. The 4.9 average survives; the volume signal, which was doing the actual work, does not. Worse, that arrangement is an incentivised-review practice without disclosure, which under the DMCC Act is not a fragile tactic but a banned one, carrying CMA exposure up to 10% of global turnover. Meanwhile Plan B’s dataset starts being cited back to it — an agent asked for UK brokers publishing verified clearance performance returns Ryburn, because there are three such datasets in the country and one of them is theirs.
Months seven to nine. A new shipper-side procurement network launches and both plans join. Plan A starts at zero — four years of nothing transferred, because none of it ever could. Plan B arrives with the AEO credential, the trade-body membership, the published dataset and the coverage, all of it legible to any agent that consults the open web, and is shortlisted in its second week.
Month twelve. Plan A: four venues, three with usable ratings, an 11% shortlist rate, and one suspended account following the reciprocity flag. Plan B: five venues, a 19% shortlist rate, and procurement teams quoting its published clearance figures back at it inside RFQs — what it looks like when a dataset becomes a category’s reference point rather than a marketing asset, the same dynamic behind interactive calculators that earned over 100 links.
The single sentence for the board: Plan A built a reputation it could not take anywhere, attached to a name anyone could copy. Plan B built one it could not lose.
9. What to do on Monday
Seven things, in order, none requiring an agent strategy document.
- Run the Clean-Slate Test. Twenty minutes, a whiteboard, three columns. Get the survival ratio before you argue about anything else, because it determines the entire budget.
- Clear the identity gates once, then stop. Signed Agent Cards, verifiable operator binding, a KYA-compatible credential path, consistent entity records across every registry an agent might consult. This is a gate, not an asset. Do not fund it like an asset.
- Audit every credential you hold that a third party audits. Industry accreditations, regulated statuses, trade-body memberships, standards certifications. These are the cheapest real collateral most firms already own and almost nobody publishes in machine-readable form.
- Stop paying for ratings you cannot take with you. Any spend whose entire output lives inside one venue is rent, not investment. Price it that way in the plan, and check the incentivised-review rules before you run any feedback campaign in the UK.
- Publish one number nobody else in your category publishes. Verified, dated, quarterly, on your own domain. Operational truth that is checkable is the most under-supplied citable asset in most sectors, and it is the fastest route to being the source an answer is built from — the logic that underpins the strongest of the link building strategies that still work in 2026.
- Fund corroboration from parties with something to lose. Trade press, independent benchmarks, accountable review platforms, expert-request channels such as Connectively, Featured and Qwoted, and technical communities where a claim gets argued with rather than published — earning attention on Hacker News is a harsh but honest test of whether a claim survives contact with people who know the subject.
- Replace the blended “AI reputation score” in your reporting with three separately owned numbers: shortlist rate in venues where you have no history, first-transaction win rate, and the portability ratio from step one. The first is marketing’s, the second is sales and operations’, the third is the board’s.
And a note on timing. Manufactured reputation is cheap right now in every venue with a cheap identity layer, so the market is pricing agent scores at roughly what they are worth — very little. That changes as binding tightens. Firms that spend the interval building collateral rather than scores arrive already holding the thing that gets counted; the ones farming ratings arrive holding a number that just got deflated, exactly as manipulated link velocity patterns stopped working the moment the signal got audited.
The through-line is unglamorous and old. Trust has always been extended against something a counterparty would hate to lose, and no protocol has repealed that. Agent marketplaces changed only the speed at which the cheap version can be manufactured — which, if anything, makes the expensive version worth more. There is more context in the current link building statistics, in the tooling that tracks this properly, and in the mechanics of getting a genuine finding picked up quickly through newsjacking.
