NLWeb and Conversational Endpoints

NLWeb and Conversational Endpoints: What an /ask Interface Actually Wins You

TL;DR

A conversational endpoint — NLWeb (natural-language web), Microsoft’s open project that turns your existing structured data into an /ask interface that is itself an MCP server — is being sold as a new agentic discovery channel. Stand one up, the pitch goes, and the agents will come. They won’t, and that is not what it is for.

An endpoint does not create demand or win selection. It governs the fidelity and control of an exchange that happens only after an agent decides — for reasons upstream of the endpoint — to consult you.

Its real, under-priced payoff is grounding control: becoming the canonical first-party narrator of your own data, so the answer an agent gives about you is current, correct, and attributable — not a months-stale paraphrase of scraped HTML. Build it as an answer-integrity asset, measured on fidelity, not traffic.

And because models draw the large majority of what they say about you from third-party sources — one analysis put it at 85% — an endpoint perfects the slice you control and multiplies the earned trust that routes agents to you. It never substitutes for it.

The pitch, and the category error

The claim worth dismantling

The line has spread through 2026 marketing decks with the confidence of something obvious: the agentic web needs machine-readable front doors, NLWeb gives your site one, so publish an endpoint and you have opened a new place for agents to find you. It sounds like the early days of search — expose your content in a format the machines read, and the machines will bring you traffic. The analogy is where the error hides. Search indexed pages and then ranked them, so being in the index was a step toward being chosen. A conversational endpoint is not an index entry and it does not rank anything. It is a door — and a door does not decide who walks through it.

The confusion matters because it drives money to the wrong place. Treated as a discovery channel, an endpoint gets an acquisition budget and an acquisition expectation: build it richly, “promote” it, and wait for agent demand to arrive. When the demand does not materialise — and it will not, because the endpoint generates none — the reasonable-sounding conclusion is that the technology underdelivered. It did not. It was filed under the wrong job.

What the wrong model costs

Picture the failure in miniature. A shopper’s agent is asked to compare three suppliers and, for one of them, grounds its answer on a product page a crawler cached fourteen months ago. It confidently quotes a price that has since risen, a lead time that has since doubled, and a feature that was quietly discontinued. The supplier is not absent from the answer — it is present and wrong, which is worse. No endpoint that was pitched as a traffic channel would have been budgeted to fix this, because misrepresentation is invisible on a traffic dashboard. This is the tell: the value a conversational endpoint actually delivers does not show up in the metric its sellers point at.

What NLWeb actually is

What is NLWeb?

NLWeb (natural-language web) is Microsoft’s open project, introduced at Build 2025 and led by R.V. Guha — the creator of RSS, RDF and Schema.org — that turns a website’s existing structured data into a natural-language /ask endpoint. Every NLWeb instance is also an MCP server (Model Context Protocol, the open standard that lets an agent call a tool or data source), so agents can query your content directly instead of scraping the page.

The moving parts

NLWeb is deliberately unglamorous. It sits on the web plumbing most sites already have — Schema.org (the shared vocabulary for tagging web data), RSS, sitemaps and product feeds — and wraps a small service around it that answers natural-language questions from that data, returning structured results. It exposes two endpoints: /ask for conversational queries and /mcp for agents. Microsoft shipped a Python reference implementation on GitHub, added an enterprise path through Microsoft Foundry at Ignite in November 2025, and by early 2026 Cloudflare offered a managed deployment via its AutoRAG (automated retrieval-augmented generation) integration, so a team no longer has to run the vector database itself. There is even a Yoast plugin for WordPress. The through-line is that the technical bar keeps dropping.

This lineage is the point Guha keeps making: NLWeb is the syndication idea — RSS for a world of agents — rebuilt so a model can hold a conversation with your data rather than parse your markup. The structured-data and technical foundations it depends on are the same ones good technical SEO has always rewarded; NLWeb simply makes them queryable.

Cheap to add is not the same as a channel

Because standing an endpoint up is now a weekend’s work for a competent team — or a plugin for everyone else — it is tempting to read the low cost as evidence of high opportunity: surely something this easy to adopt must be a land-grab. It is the opposite. Low adoption cost means everyone will have one, which is precisely what stops an endpoint from being a differentiator. A capability the whole market can add in an afternoon converges to table stakes; it does not confer advantage. The ease is real. The channel is imagined.

Three things happen in an agentic exchange — an endpoint owns one

Is a conversational endpoint a discovery channel?

No. It controls the accuracy and framing of an exchange after an agent decides to consult you; it does not decide whether the agent consults you. The deciding — discovery and selection — happens upstream, on trust you earn elsewhere, and an endpoint changes none of it.

Know, get, choose

Strip an agentic interaction to its bones and three separate things have to happen, in order, for you to win it. First, the agent has to know to consult you at all — you have to be in the set of sources it considers. Second, once it consults you, it has to get your answer rather than a stale paraphrase assembled from whatever it cached or retrieved. Third, what it gets has to make it choose you over the alternatives. These are not three views of one event; they are three different jobs with three different owners.

A conversational endpoint owns the middle job and only the middle job. It has no influence on whether an agent knows to come — that is set by what the model already carries about you and what its retrieval surfaces, registries and directories put in front of it. It has no influence on whether you are chosen — that is a judgement the agent makes on trust and relevance. What it changes, and changes completely, is the get: whether the agent that has already decided to consult you receives your live, first-party answer or a lossy reconstruction of it. That single job is worth owning. It is just not the job the word “discovery” names.

Make it concrete. An independent review names you, so a shopper’s agent resolves to consult you — the know is done, and the review, not the endpoint, did it. Now the agent either scrapes a pricing page a crawler last cached in 2024 or queries your /ask endpoint. Same visit, same intent, opposite outcome: one path hands the agent a phantom discount and a discontinued model; the other hands it today’s truth. The endpoint changed nothing about why the agent came and everything about what it left with. That is the whole job — smaller, and more valuable, than “discovery” makes it sound.

The Grounding Gap

To see exactly what an endpoint fixes and what it leaves untouched, take the facts an agent might report about you and, for each, compare what it gets today with what your endpoint would give it — then ask the only two questions that matter. Does the endpoint close the fidelity gap (is the answer now accurate, current, attributable)? And does it change whether you are chosen? The map below runs six kinds of fact through both columns.

The Grounding Gap

Fact about youWhat an agent gets todayWhat your endpoint gives itCloses fidelity gap?Changes selection?
PriceA cached figure, often months oldThe live price, as of this queryYesNo
Availability / lead timeA guess, or a stale “in stock”Current stock and real lead timeYesNo
Spec / compatibilityInferred, sometimes inventedThe authoritative spec, queryableYesNo
Policy / termsAn outdated or hallucinated versionThe policy as it stands todayYesNo
Provenance (“who you are”)A generic or wrong descriptionYour own first-party framingYesNo
Competitive positioningWhatever third parties say — or silenceNothing extra: the endpoint can’t vouch for itselfPartlyNo

Read the last two columns together and the whole argument is in the contrast. The fidelity column is a wall of green: on every kind of fact an agent might get wrong, the endpoint makes it right. The selection column is a wall of red: on not one of them does being right, by itself, make the agent choose you. The one row that breaks the pattern is the most revealing — an endpoint cannot improve your competitive positioning, because a source narrating its own superiority is exactly the input a wary agent discounts. Fidelity and selection are different gaps, and a conversational endpoint closes only one of them.

The “partly” is worth being precise about, because it is the whole temptation in one cell. Your first-party framing does make a model’s description of you more accurate once it has decided to include you — the endpoint can ensure you are characterised in your own terms rather than a competitor’s. That is a genuine, if faint, benefit adjacent to selection. But it is strictly downstream of being included, and inclusion is the earned decision the endpoint never reaches. Sharpening how you are described in a race you were already entered into is not the same as being entered — and mistaking the first for the second is precisely how the budget goes wrong.

Key takeaway

An endpoint turns a wrong answer about you into a right one. It does not turn a right answer into a chosen one. Budget it against misrepresentation, not against traffic — and keep the selection budget where it belongs, in earned corroboration.

Why selection stays earned

Will adding NLWeb get my brand cited more by AI?

Not on its own. Models draw most of what they say about a brand from third-party sources, so citations still track earned corroboration, not the presence of an endpoint. What the endpoint changes is the accuracy of the citations you already earn — a real gain, but a different one.

The 85% you don’t own

The clearest evidence that an endpoint cannot own selection comes from where AI answers actually source their claims. Lantern’s February 2026 analysis found that 85% of brand mentions in AI answers originate from third-party pages, not the brand’s own domain; Muck Rack’s May 2026 study of 25 million citations put earned sources at 84% of what AI reads, against 0.3% paid. Whatever you publish on your own endpoint sits inside the small minority of grounding a model treats as self-interested. The large majority — the part that decides whether you are named at all — is written by other people about you. You influence that only by earning the third-party corroboration the model actually weights.

This is also why an endpoint cannot fix the most expensive failure of all: being left out of the comparison entirely. When a model builds a shortlist, it names the entities it can clearly resolve and confidently differentiate. A brand whose distinctiveness is asserted only on its own endpoint, and nowhere in the independent record, gets described in generic terms or omitted from the “best of” set — a pattern documented repeatedly across 2026 brand-visibility audits. Being resolvable as a distinct entity, and being differentiated by sources other than yourself, is upstream of anything an endpoint can do.

Why a self-narrated source is discounted

There is a structural reason models lean on third parties, and it is the same reason this publication has returned to all year: an answer is only as trustworthy as its most manipulable input. A source describing itself has every incentive to flatter, so a system trying to be reliable weights independent corroboration above self-description by design. That is not a temporary quirk of today’s models; it is the property that makes them worth consulting. It means the harder you lean on your own endpoint to do your selling, the more a careful agent discounts it. The endpoint earns its keep by being accurate, not persuasive — and the things that actually drive an agent’s recommendation sit almost entirely outside it.

The two also behave differently over time, which settles how to fund them. Fidelity is a step-change that then holds: fix the data and stand up the endpoint, and the answers are right — but rightness does not compound, it plateaus. You cannot get more than accurate. Earned selection is the opposite: each independent citation makes the next one easier and the entity a little more resolvable, so the returns build on themselves without ceiling. That is why the endpoint is a fixed cost you pay once and maintain, while corroboration is the investment you keep compounding — and why a budget that pours into the plateau and starves the compounding curve gets the shape of the opportunity exactly backwards.

Key takeaway

Roughly 15% of what a model says about you is yours to state directly; the other 85% is earned or it is absent. An endpoint makes your 15% perfect. Only corroboration moves the 85%.

The real payoff: grounding control

What does an NLWeb endpoint actually win you?

Grounding control. When an agent does consult your site, it answers from your live first-party data instead of a stale, scraped paraphrase — so the answer about you is current, correct, and attributable to you. That is an answer-integrity asset, and it is worth building for its own sake — just not as a way to be found.

The misrepresentation tax

Once you stop expecting traffic, the endpoint’s value comes into focus, and it is not small. The AllAboutAI 2025 hallucination report put the average error rate on general-knowledge questions — the category that most overlaps with brand facts, pricing and product detail — at 9.2%, roughly one wrong answer in eleven, delivered with no hedge and no way for the reader to tell. Wrong prices, phantom features, outdated policies and confused competitor comparisons are not edge cases; brand-monitoring vendors report them as the norm for any company that has not actively managed its AI accuracy. And the audience is primed to be burned: only about a quarter of US adults say they trust AI to give them accurate information, so every confident error lands on someone already half-expecting one. A single false “fact” surfaced to a prospect mid-evaluation can cost a deal outright.

This is the tax a conversational endpoint is built to cut. When the agent grounds its answer on your live /ask data rather than a cached page, the wrong price becomes the right price and the discontinued feature disappears from the pitch. The gain is defensive and unglamorous, which is exactly why it is under-priced — and why it is real.

And the cost is not only lost deals; a confident error can be expensive even when it wins the sale. Suppose an agent, grounding on a stale page, tells a buyer you offer a 60-day return you retired last year, and closes the order on that promise. You have now booked revenue against a commitment you did not make — a support dispute, a refund, or a chargeback with your name on it, seeded by a fact you never stated. Misrepresentation leaks in both directions: it loses business you should have won, and wins business on terms you cannot honour. A first-party endpoint is the only thing that lets the agent transact against what is actually true today.

Recency is the half of fidelity nobody budgets for

Fidelity is not only about being correct; it is about being current, and models punish staleness harder than most teams realise. Muck Rack’s May 2026 data shows AI systems strongly favour content from the past twelve months; Lantern found that pages not refreshed in over three months are three times more likely to lose their citations than recently updated ones. A static owned file cannot keep up with that clock. An llms.txt file — the emerging convention for pointing crawlers at your key documentation — is a signpost frozen at publish time; it cannot answer “what is the lead time today.” A conversational endpoint composes the answer at query time from live data, which is the difference between a fact and a fossil. The same freshness pressure that governs earned coverage governs your first-party answers too.

In-place answering

The publishers who piloted NLWeb early — O’Reilly Media, Tripadvisor, Allrecipes, Eventbrite — did not do it to be discovered; they are already known. They did it so that an agent asking “which of your books covers this topic” or “does this property allow dogs” gets the answer from their catalogue rather than a model’s guess about their catalogue. That is the honest use of the technology: not a new front door for strangers, but a truthful mouth for the visitors — human or agentic browser — who were already coming. Everything an endpoint does well, it does for an exchange that was going to happen anyway.

A worked example: Tenby Instruments builds an endpoint

Consider Tenby Instruments, a UK distributor of laboratory and scientific instruments with about £25M in revenue and thousands of SKUs whose specifications, compatibilities, lead times and compliance certifications change constantly. This is precisely the catalogue an agent will get wrong from a cached page, so an endpoint genuinely makes sense. Before spending, Tenby runs the audit every team should: fifteen questions a procurement agent would actually ask — current price, lead time, protocol compatibility, cert status — put to five engines. Nine of the fifteen come back wrong, stale or misattributed: a discontinued analyser still described as available, an eight-week lead time quoted as two, a superseded certification cited as current. That 60% misrepresentation rate is the number the endpoint addresses. It has a £90,000 “agent-readiness” budget for 2027 and one real decision: what the endpoint is for.

The traffic framing — the one to avoid — treats the endpoint as a discovery channel: roughly £45,000 to build and integrate an elaborate conversational layer, £30,000 to “promote the channel,” and £15,000 of contingency, on the expectation that agentic demand will arrive. The fidelity framing spends almost inversely: about £20,000 for a lean endpoint plus the Schema.org and feed hygiene the audit exposed; £15,000 for measurement — re-running the audit monthly and tracking a first-party-grounding rate; £45,000 into the earned corroboration — trade coverage, reviews, practitioner citation — that gets Tenby into agents’ consideration sets and trusted enough to be consulted; and £10,000 to make its entity and attribution machine-resolvable.

Tenby’s £90k, two ways

Budget lineTraffic framingFidelity framingWhy
The endpoint itself (build + integrate)£45k£20kPlumbing. Lean and first-party beats elaborate.
Data / feed hygiene (Schema.org, feeds)incl.What actually makes the answers right.
Measurement (paraphrase audit + grounding rate)£15kTracks the metric that proves the point.
Earned corroboration (coverage, reviews, citation)£45kThe tie-breaker that gets you consulted at all.
Owned / machine-resolvable entity£10kSo a model can resolve who you are.
“Promote the channel”£30kSpends on demand an endpoint cannot generate.
Contingency£15kPadding for a plan aimed at the wrong job.

The traffic column fails on its own terms before any agent arrives. The £30,000 to “promote the channel” assumes demand to capture, but an endpoint captures demand some earlier, earned signal produced — it does not create it. The £45,000 elaborate build buys conversational depth no one has been given a reason to reach for. The plan spends most on the job an endpoint cannot do — generate demand — and nothing on the job it can: being right when consulted. The fidelity column is not merely cheaper; it is aimed at jobs that exist.

The timeline separates them cleanly. In month one both endpoints are live; the traffic build looks busier, but inbound “discovery” stays flat, because the endpoint summons no one. By month three the fidelity build’s misrepresentation rate on the queries agents do run has fallen from nine-in-fifteen to two, and its earned coverage is starting to route procurement agents to it in the first place; the traffic build is still waiting for demand it never generated, its misrepresentation rate untouched because nobody funded the data hygiene. By month six the fidelity build is the trusted first-party ground truth for Tenby’s catalogue: when an agent consults it, the answer is current, correct and attributable, and Tenby wins the specs it genuinely qualifies for instead of losing them to a stale error. The traffic team concludes that “NLWeb doesn’t drive discovery.” They are right, for the wrong reason — they built a fidelity asset and judged it as an acquisition one. The endpoint never brought more agents; it stopped the ones already arriving from leaving with the wrong answer.

There is a quieter compounding effect the traffic framing misses entirely. Because Tenby’s earned coverage now routes agents to an endpoint that answers correctly, every citation it earns is worth more than it would have been — a review that lands an agent on accurate, current data converts better than one that lands it on a fossilised page. Fidelity and earned corroboration are not competing line items; they are complements, each raising the other’s yield.

Where this could be wrong

The strongest case against everything above is that a conversational endpoint really is becoming a discovery channel, and the evidence is not trivial. Microsoft is wiring NLWeb into its own surfaces — Copilot, Foundry, the broader agent tooling it keeps shipping — and an NLWeb instance is natively an MCP server, so it can be listed, indexed and invoked inside the agent ecosystems Microsoft and others are building. Cloudflare’s AutoRAG turned deployment into a managed click; Shopify and major publishers expose endpoints their platforms then surface. Being NLWeb-enabled, the argument runs, does put you inside agent workflows and directories — so it drives selection, not merely fidelity. Concede all of it: the distribution is real and growing, and dismissing it would be the strawman this column is meant to avoid.

Three bounds

Grant the distribution in full, and it still does not become a discovery channel for you, for three reasons. First, distribution of the protocol is not distribution of your endpoint. Microsoft shipping the rails, and Cloudflare paving them, makes you callable inside those surfaces — but each surface still selects which endpoint to call by relevance and trust. The plumbing being everywhere is exactly why having it wins nothing: it is table stakes, and the choice among the qualified is made on grounds the endpoint does not touch.

Second, where an endpoint is auto-invoked — on your own Shopify storefront, for a user already on your domain, by an agent you are already connected to — discovery has already happened off the endpoint. That is warm fulfilment, the conversational-endpoint twin of a direct install: the agent arrived because of something upstream, and the endpoint served it well. Serving a visitor well is retention; it is not the same as summoning one. Third, universal adoption is subtractive, not additive. Once every site exposes an /ask, having one is a floor and selection reverts to trust — and models already draw 85% of brand mentions from third parties precisely because a source narrating itself is the input they most distrust. The more endpoints proliferate, the more selection leans on the earned layer they cannot occupy.

The falsifier

What would prove this wrong? One clean condition: a dominant assistant routing agent queries to endpoints by a mechanism the site controls through publishing or paying — a pay-to-be-the-grounded-source slot, or presence-only routing that consults whoever has an endpoint with no trust weighting at all. If that shipped and held meaningful share, the endpoint would become a genuine discovery lever, and this argument would fold. Watch for it. But it runs against the grain of what assistants are built to do: grounding on an untrusted, self-interested first-party source is the exact hallucination-and-manipulation risk they are racing to reduce. A platform that routed to self-narration over independent corroboration would be degrading the reliability it competes on — the same answer-independence logic that fences paid influence out of the answer fences self-interested grounding out of the ground truth.

And NLWeb itself might not win

There is a fair objection from the other direction, too: NLWeb is early and might not become the standard at all. Seasoned observers call it a visionary specification that still needs ecosystem validation, tooling and reference integrations before mainstream pilots, and put substantial enterprise adoption two to three years out. That caution is warranted — and it does not touch the argument, because the argument is about the category, not the brand. Whether the winning conversational endpoint turns out to be NLWeb, a bare MCP answer server, or something not yet named, it will still own the get and not the know or the choose. Bet on the category’s shape, not on which logo wins it.

Key takeaway

The endpoint becomes a discovery channel only if a dominant assistant sells or presence-routes grounding. Until then, protocol distribution is table stakes and selection stays earned — so plan for fidelity and let the falsifier, not the hype, tell you when to change.

What to do Monday

The work splits into two clean tracks: build the endpoint for fidelity, and measure it with the audit that proves the point. Start with the audit, because it sizes the prize and doubles as a citable asset.

The Paraphrase Audit

Pick 15 questions a buyer’s agent would actually ask about you — branded, pricing, spec, compatibility, policy, comparison — and put each to the five engines your buyers use (ChatGPT, Gemini, Perplexity, Copilot, Claude).

Score every answer as Correct, Stale, Wrong, Misattributed or Absent. The share that is not Correct is your misrepresentation rate — the exact value a conversational endpoint is there to cut.

Re-run it monthly and track the rate as your first-party-grounding KPI. Publishing the anonymised method and findings is a legitimate first-party study — data only you hold — and a citable asset that earns corroboration the way a static page never will.

  • Run the Paraphrase Audit first. Do not build anything until you know your misrepresentation rate. If engines already answer accurately about you, a lean endpoint or better feeds may be all you need.
  • Fix the data the audit exposes. Most wrong answers trace to stale or missing structured data. Schema.org markup, current feeds and clean sitemaps do more for fidelity than the conversational layer on top of them.
  • Stand up a lean endpoint, and wire the MCP server. Use Cloudflare AutoRAG or the reference implementation; keep it first-party and accurate rather than elaborate. Its job is to be right when consulted, not to be impressive.
  • Instrument fidelity, not traffic. Report first-party-grounding rate and the monthly misrepresentation rate. Counting endpoint “hits” as discovery is the miscategorisation this whole piece is about. The right monitoring tools watch what AI says about you, not who visited your /ask.
  • Keep funding the layer that gets you consulted. An endpoint perfects the 15% you own; the earned corroboration that moves the other 85% is still the core link-building and digital-PR work — and it is what routes agents to the endpoint in the first place.
  • Budget the endpoint as answer-integrity infrastructure. It belongs in the retention and reliability column, not the acquisition one. Filing it correctly is what stops you overspending on a channel it was never going to be.

One line to leave with: a conversational endpoint does not get you found — it gets you quoted correctly. NLWeb and its successors own the moment an agent already at your door asks a question; they make the answer yours, current and attributable, and they do nothing to decide whether the agent shows up or picks you. That is a real and underrated asset, worth building on its own terms. Build it for fidelity, earn the discovery the ordinary way, and — whether you sell in Manchester or across the European markets Tenby ships into — measure the endpoint by whether the answer about you is right, not by a traffic line it will never move. For the wider signal picture, our guide to how AI Overviews assemble their sources and our running statistics on AI citation, alongside the freshness discipline that newsjacking keeps sharp, map the earned terrain the endpoint sits on top of.

Leave a Reply

Your email address will not be published. Required fields are marked *

MCP Server Previous post Publishing an MCP Server as a Discovery Channel: What It Actually Wins You
Verified Agents Next post Web Bot Auth and Verified Agents: The New Access-Control Layer