Share of Model

Share of Model: Measuring Category Dominance Across Five Engines

TL;DR

Share of model borrows its credibility from share of voice and share of search. Both of those had a denominator closed by somebody other than the person doing the measuring. Share of model has two candidate denominators — the competitor list you supply, or the set of brands the engine happens to name — and an interested party supplies each one.

The consequence is arithmetic, not opinion: your reported share moves when your configuration moves, and it can move in the opposite direction to your actual presence. What survives is the pairwise contrast between you and one named rival, because the denominator cancels. And the shape of placement you should buy flips depending on which side of parity you sit.

1. A metric with an author, a date and a family tree

Share of model is not folk terminology that drifted into the industry. It has a paper trail. Jack Smyth and Tom Roach, then at Jellyfish, part of The Brandtech Group, introduced the concept during 2024, and the Share of Model platform launched on 4 December 2024. Researchers at INSEAD formalised the definition in mid-2025, and David Dubois and colleagues carried it to a Harvard Business Review audience that June. By 25 June 2026, Jellyfish was reporting findings drawn from more than 27 million model response data points across ten industries, seven models and fifteen-plus markets.

What is share of model? It is the proportion of AI-generated answers in a category where your brand appears, sometimes weighted by how prominently and how favourably it appears. It is sold as the answer-engine successor to share of voice, and it is now the headline number on most AI visibility dashboards.

The family tree is not decoration. Roach’s own essay placed share of market as the patient zero of every share-of metric and argued that the new one stands on the shoulders of the old. That argument is doing real work. It is what licenses a marketing director to put the number in a board pack without first establishing that it means anything, because the ancestors earned that licence decades ago. Which makes the lineage claim testable. If share of model inherits the property that made its ancestors trustworthy, the inheritance holds. If it inherits only the word share, it does not.

What the ancestors actually had

Two properties, and both of them sit outside the measurer.

  • An exogenous denominator. Somebody with no stake in your reading closes the total. Share of voice, as Nielsen has defined it for decades, is a brand’s media expenditure set against total category spend — a figure assembled from invoices, which exist because money moved. Share of search divides your brand queries by all brand queries in the category, counted by the search engine’s own logs.
  • An external criterion. Each metric was validated against something it did not itself produce. Les Binet presented share of search at the IPA’s EffWorks Global conference in 2020, testing it across automotive, energy and mobile handsets, and found it led market share by six to twelve months, up to a year in cars. James Hankins later put it at around 83% of a brand’s market share across 30 cases in 12 categories and seven countries. The metric earned its authority by predicting sales.

Share of model has neither. That is the argument of this piece, and the rest of it is the arithmetic.

2. The four shares, side by side

Set the family out and the break is visible in one column.

MetricWhat the denominator isWho closes itValidated against
Share of marketTotal category units or revenueAuditors, trade bodies, retail panelsIt is the outcome — nothing to validate against
Share of voiceTotal category advertising spendMedia owners, through invoicesSales, via decades of excess-share-of-voice work
Share of searchAll brand queries in the categoryThe search engine’s own query logsMarket share, six to twelve months ahead
Share of modelYour competitor list, or whatever the engine namesYou, or the engine being measuredNothing published

Three rows have a denominator produced by a party indifferent to the result. One does not. That single column is the whole difference between a measurement and an index, and it survives every improvement to sampling, cadence and engine coverage that the last two years of methodology work has produced.

3. Two denominators, and you own both of them

Open any AI visibility tool and it will ask you to name your competitors. Three to five is the standard ask, and the standard advice matches it: run monthly share-of-voice reviews against three key competitors, compare share trends against your three closest rivals over a trailing ninety days. The dashboard then divides your mentions by the mentions of everyone on that list.

This is the closed denominator, and it has a defect that no amount of sampling discipline touches. The competitive set is an input. In organic search the set was an output: whoever ranked on the page was, by definition, competing with you for that query, and you discovered rivals you had never heard of by reading the SERP. A share of model reverses the direction of that information. It cannot report a competitor you have not already thought of, which means it is structurally blind to the entrant — precisely the event that a category dominance metric is supposed to warn you about.

The field has noticed the smaller half of this problem. The sophisticated correction now circulating is the open denominator: count every brand the model names, not only your pre-selected list, on the grounds that a closed list inflates your score by shrinking the pool. Several vendors advertise open-denominator counting as a default. The reasoning is sound as far as it goes.

Why the open denominator is not the fix

It does not remove the endogeneity. It relocates it. Under an open denominator the competitive set is defined by the engine’s own output, so you are now drawing the boundary of the market using the very behaviour whose contestation you are trying to measure. If the engine is currently naming six adjacent brands that no buyer would ever shortlist against you, they are in your market. If it stops naming them next month, they are not, and your share rises with nothing having happened.

Competition economists have had a name for this error since 1956. In the du Pont cellophane case, a market defined by the substitution actually observed at prevailing prices came out absurdly wide, because those prices were already the product of the market power under investigation. Observed substitution at prevailing conditions is not evidence of the boundary when the conditions are the thing in dispute. An engine’s current slate of named brands is prevailing conditions.

Key takeaway

There is no third denominator available. Share of market has audited sales, share of voice has invoices, share of search has query logs. There is no register of category membership for AI answers, no body that certifies who is in a category and who is adjacent to it. You will either declare the set or let the measured system declare it, and both parties have a stake in the answer.

4. The arithmetic of parts, and why editing a list moves your score

What is a compositional metric? Any set of numbers that are parts of a whole and therefore sum to a fixed total — shares, percentages, budget splits. Compositional numbers obey their own algebra, and applying ordinary statistics to them produces results that look fine and are not.

Karl Pearson set out the hazard in 1897, warning against interpreting correlations between ratios whose numerators and denominators contain common parts. If three quantities are independent, two ratios formed by dividing them by the same third quantity will not be. John Aitchison built the modern framework in 1982 and 1986 and named the two properties any honest compositional statement must satisfy: scale invariance, and subcompositional coherence — an inference drawn about a subset of the parts must not depend on what else is in the set.

A share of model fails the second property by construction, and the failure is large enough to see in a single month of data. Suppose a tool tracks you and five rivals. You are named 38 times across the sample; the six tracked brands are named 240 times between them. Your share is 15.8%. Nothing about your presence changes over the next paragraph.

  • Switch to an open denominator and the engine turns out to have named eleven distinct brands, adding 74 mentions to the pool. Your 38 mentions are now 12.1%.
  • Drop the weakest tracked rival, who accounted for 12 mentions, because a colleague argues they are not a real competitor. Your 38 mentions are now 16.7%.

Same month, same answers, same brand, same 38 mentions: 12.1%, 15.8%, 16.7%. A 4.6-point spread generated entirely by list editing. Set that against the size of the movements teams actually report to a board — a two-point quarterly gain is a good quarter — and the artefact is twice the signal.

Instrument 1 — The Basket Swap

Purpose: establish how much of your reported share is a property of the category and how much is a property of your configuration. Run it once, on one month, before you present a share number to anyone who allocates money.

1. Pull the raw answers for one category on one engine for one month. Do not use the dashboard total; you need the underlying responses.

2. Count every distinct brand named in those answers, not only your tracked set, and record mentions per brand.

3. Compute your share three ways: against the tracked set only; against every brand named; and against the tracked set minus your weakest-scoring tracked rival.

4. Record the spread between the highest and lowest of the three figures.

5. Compare that spread with the change you intend to report this quarter.

Bands. Under 2 points, the share is stable enough to quote as a level. Between 2 and 6 points, report pairwise contrasts only and stop quoting the percentage. Over 6 points, the metric is measuring your configuration and should not appear in a board pack at all.

5. The hidden variable is slate size

There is a second problem hiding inside the same word, and it is arithmetically nastier because the two versions of the metric move in opposite directions.

Two formulas circulate under the name share of model. The first divides the number of answers in which you appear by the number of answers tested. The second divides your mentions by all brand mentions in those same answers. One widely circulated industry definition manages to contain both halves in a single sentence, describing the metric as the percentage of answers that mention your brand, measured relative to all brand mentions in those answers. Those are different quantities. The first does not sum to 100% across brands and is a prevalence rate; the second does and is a share.

What is slate size? The number of brands an engine names in one answer. It is a presentation decision made by the engine, not a fact about your category, and reporting suggests ChatGPT typically names three to four brands per recommendation answer.

Now watch what happens when the engine becomes more generous and the typical slate widens from four names to six. Every brand’s prevalence rises, because every brand appears in more answers. No brand’s share need move at all, because the pool of mentions grew in proportion. If the engine tightens instead, prevalences fall across the board and shares again sit still. The two metrics wearing one name respond to a purely presentational change with opposite signs, and neither movement contains information about your category.

This is why a rise in your reported visibility can coincide with a fall in your reported share in the same dataset, a pattern teams usually attribute to a competitor’s campaign. It is worth checking the average number of brands named per answer before you attribute anything to anybody. That number is free to compute and almost nobody reports it.

6. Five engines is not one market

The cross-engine composite is where the errors compound, because pooling shares requires weights and the weights do not exist.

Start with the arithmetic, which is worse than it looks. A dashboard can produce an overall figure two ways: average the per-engine shares, or pool the raw counts. These are not the same number and they can rank you differently. Suppose on a sparse engine you take 60% of ten mentions and a rival takes 40%, and on a dense engine you take 45% of 190 mentions and the rival takes 55%. Average the two shares and you lead, 52.5% to 47.5%. Pool the counts and you trail, 46% to 54%. Neither calculation is wrong; they answer different questions, and the dashboard rarely says which one it ran.

Pooled counts are verbosity weights in disguise

When you pool, each engine’s influence on the composite is proportional to how many mentions it produced, which is its citation density multiplied by how many prompts you sent it. Neither of those is its share of your category’s demand. Perplexity has been measured at 5.9 citations per query; measured citation rates across engines have ranged from 27% down to under 1% and zero in the same study. Pool across that spread and the engine that talks most decides your number.

The honest correction is to weight each engine by its share of category usage. Try to source those weights and you find the recursion. Across broadly the same period in 2026, published estimates of ChatGPT’s share of AI activity landed at 62.6% of measurable B2B AI referrals in one panel, 77.9% of chatbot market share on StatCounter, and 92.4% of trackable AI referral traffic across 166 analytics properties. To weight a share, you need another share, and the other share is contested by thirty points. None of them is broken down by category anyway, which is the level at which you actually compete.

And the underlying dispersion is severe. The INSEAD work found the detergent brand Ariel holding close to 24% share of model on Meta’s Llama and under 1% on Gemini — same brand, same category, same moment. A composite over five engines is dominated by which engines you chose to include, and it describes a buyer who consults all five and weights them the way your tool does. No such buyer exists. Every real one is somewhere on a single surface at the moment of decision, which is why per-surface behaviour is worth more to a plan than any blend of it.

Key takeaway

Report engines separately or do not report them. A blended share of model is an average over populations that do not mix, weighted by how talkative each engine is, expressed as a fraction of a set you nominated. Each of those three steps is defensible in isolation and their composition is not a quantity about the world.

7. What survives is the contrast, not the share

Aitchison’s framework does not conclude that compositional data is useless. It concludes that only statements about ratios of components are meaningful, because a ratio is invariant to the closure. That result transfers directly and it is the practical payload of this article.

Take your mentions divided by one named rival’s mentions on the same answer set. Add ten brands to the tracked list, remove three, switch from closed to open denominator: the ratio does not move, because both numerator and denominator sit inside every version of the set. The quantity that destroys the share leaves the contrast untouched.

The same cancellation handles a deeper problem. Whatever bias your prompt basket carries — over-weighting the ground you already hold, under-representing the questions buyers actually type — that bias applies to you and your rival equally, because you are both being measured on the same prompts. It is common-mode, and it subtracts out of a ratio in a way it never subtracts out of a level. This is the one place in AI visibility measurement where an unfixable sampling problem stops mattering, and it costs nothing to exploit: you already hold the raw counts, and you simply stop dividing by a set you invented.

Instrument 2 — The Pairwise Ladder

1. Name three rivals, not five and not ten — the ones a buyer would genuinely shortlist against you. Fewer pairs, each of which means something.

2. For each engine separately, compute your mentions divided by that rival’s mentions on the identical answer set. Never blend engines at this stage.

3. Report it in words: named 1.4 times for every one time they are. No percentages appear anywhere in the output.

4. If you must summarise across engines, use the geometric mean — multiply the ratios and take the nth root. Averaging ratios arithmetically is wrong in a way that flatters you: a 2.0 on one engine and a 0.5 on another average to 1.25 arithmetically and to exactly 1.0 geometrically, which is the truth.

5. If either brand records zero mentions in a stratum, report the raw counts and mark that pair undefined. Do not impute a floor; a zero and a small number are different objects.

6. Rank the pairs by distance from 1.0. The pair nearest parity is the only one a single quarter of work can realistically move.

A worked case: fourteen weeks in field service software

Merrow Field Service Software is a Guildford firm with 54 staff and £7.4m of annual recurring revenue, selling scheduling and job-management software to UK maintenance contractors. It pays £1,450 a month for an AI visibility platform tracking itself and five named rivals across four engines, and its April 2026 board pack reported a share of model of 15.8%.

The basket swap run on 6 April produced a spread of 4.6 points, which put the metric in the middle band: contrasts only. So the team switched reporting. Against its nearest rival, Calder Systems, Merrow stood at 38 mentions to 44 on ChatGPT, a ratio of 0.86, and led on Perplexity at 22 to 15. It was a challenger on the surface that mattered most to its buyers and a leader on one that barely fed its pipeline.

Between mid-April and late June the team earned nine placements, chosen by a rule set out in the next section. It re-measured on 13 July. Merrow’s mentions had risen from 38 to 45. Calder’s had fallen from 44 to 40. The pairwise ratio had crossed parity, 0.86 to 1.12. And the headline share of model had fallen, 15.8% to 14.9%, because the engine had widened its typical slate from four names to six and the mention pool had grown 26% while Merrow grew 18%.

One dataset, three readings: absolute presence up, share down, competitive position up and across the line. The share was the only one of the three that pointed the wrong way, and it was the only one that had been in the board pack.

8. What the evidence says about dominance itself

There is one substantial study of category dominance in AI answers, and it deserves reading closely rather than quoting. Semrush and Kevin Indig tracked 1,094 US categories monthly in ChatGPT from January to June 2026, five prompts per category covering definition, comparison, alternatives, use case and buying question, across more than 50,000 brands and 600,000 citations. They defined a category owner as the brand with the highest share of mentions, appearing in at least four of five prompts, with at least a five-point lead over the runner-up.

  • 15.2% of categories had a clear owner. 31.2% had an emerging leader. 53.7% were unsettled, with no brand appearing in even three of the five prompts.
  • Among the top half of categories by demand — which carried 98% of the AI search volume in the sample — only 11.3% had a clear owner. Dominance is rarest exactly where the volume is.
  • Where a clear owner existed, it held first place in 90.4% of month-over-month comparisons. Where leadership flipped, the median lead had been 1.3 percentage points; where it held, 2.9.
  • Domain-level SEO strength predicted ownership at close to a coin flip: owners had higher branded search volume in 55.7% of pairs, higher organic traffic in 48.4%, a higher Authority Score in 52.5%.

Reading the margin honestly

Put those two facts together. The margins that separate an owner from a challenger are on the order of one to three percentage points. The basket swap in section 4 moved a share by 4.6 points without anything happening. The ownership classification is therefore finer than the instrument’s sensitivity to its own configuration — which does not make the study wrong, and this distinction matters. Their comparisons are within a single fixed instrument, the same brands on the same prompts month after month. Under those conditions a share behaves as an ordinal contrast and the finding holds. What does not survive is portability: their 15.2% is a fact about their basket, and your 15.8% is a fact about yours, and the two numbers cannot be laid beside each other even though they share a name and a unit.

One more result from the same study belongs here, because it separates two things dashboards routinely blend. Only 21% of the most-cited domains in a category were also the most-mentioned brand, and the two correlated slightly negatively at −0.229. The page an engine reads and the brand an engine names are close to independent, which is a caution against reading citation counts as a proxy for category standing, or the reverse.

9. Where this argument has to survive its strongest objection

The best counter is not that the denominator is fine. It is that nobody claimed otherwise. Vendors themselves describe the metric as directional rather than absolute. The argument runs: a board needs one number, everyone knows it is an index, the comparison is against ourselves over time on a frozen basket, and the Semrush study just demonstrated that within-instrument comparison works. Precision is not the job; direction is.

That is a serious position and the first half of it is correct. A frozen basket does support ordinal comparison. Three things bound it.

  • The direction is not preserved. The defence assumes the error is a level shift — crude but pointing the right way. Merrow’s fourteen weeks show the sign itself inverting: presence up, share down. A metric that can reverse the sign of a real change is not a rough guide to that change.
  • A communication device becomes an allocation device. Every number that appears on a slide for two quarters starts governing budget in the third, and the compositional artefacts are largest when a category is consolidating, which is exactly when the allocation decision is live.
  • The honest crude number is already available and cheaper. Named twice as often as our nearest rival on ChatGPT is at least as communicable as 15.8%, carries no closure, and requires no extra data collection.

The institution with the most reason to disagree

If shares carried decisions well, the strongest evidence would come from the body that computes them for a living. The UK’s Competition and Markets Authority published revised Merger Assessment Guidelines on 18 March 2021, its first update in over a decade, and the direction of travel was away from shares. Market definition was explicitly diminished, the guidelines noting that in most mergers the evidence gathered in the competitive assessment captures the dynamics more fully than formal market definition. Shares of supply were retained as useful evidence when assessing closeness of competition — evidence in service of a pairwise question, not the analysis itself. The old market-share safe harbour was dropped. Fifty years of practice pushed the regulator toward asking which two firms constrain each other and away from asking what the market is and who has what percentage of it.

There is an honest complication, and it cuts the right way. The CMA’s pairwise tools have gone quiet at the first phase of review — diversion ratios generally require consumer surveys, which do not fit a Phase 1 timetable, so shares persist partly because they are cheap. That excuse does not transfer. In AI visibility the pairwise number is not the expensive option; it is the same raw counts with one division removed. There is a regulatory reporting context in which measurement conventions get settled slowly and expensively, and this is not one of them.

10. What this changes about the placements you buy

Adopt the pairwise contrast and one practical rule falls straight out of the arithmetic, and it contradicts the advice everybody gives everybody.

Placements come in two shapes. A joint document names you and your rival together: the comparison page, the category roundup, the shortlist, the buyer’s guide. A solo document names you and not them: the case study, the interview, the sponsored research, the trade write-up of your work. Their effect on a ratio is not the same, and it is not the same in the two directions.

Suppose you sit at 6 mentions to a rival’s 4, a ratio of 1.50. Earn one joint document and you are at 7 to 5, a ratio of 1.40 — you added a mention and your position got worse. Earn one solo document and you are at 7 to 4, or 1.75. Now run it as the challenger at 4 to 6. The joint document takes you from 0.67 to 0.71; the solo takes you to 0.83. Both help, but the direction of the asymmetry has flipped: joint documents pull every pair toward parity, which is what a challenger wants and what a leader is paying to prevent.

The joint-mention rule

If you trail: buy joint. Comparison pages, category roundups and listicles, shortlists, alternatives pages — anything where the incumbent is already named and you can stand beside them. You are buying co-membership, and co-membership is worth more to you than to them.

If you lead: buy solo. Sponsorships, commissioned research, original interviews, case coverage where no rival is named. Every roundup you win is also a roundup that certifies your challenger as a peer.

The screen: before commissioning anything, ask whether the target document already names your rival. That is now a first-order question about the placement, not a detail. Sort your prospect list by it.

The sizing: dominance in the one study that measures it turns on margins of one to three points. You are buying margin, not volume. Nine well-shaped documents moved Merrow across parity; ninety poorly shaped ones would not have.

This is a genuine reversal of standard practice. Get into the comparison roundups is challenger advice, and it is sold identically to leaders and challengers by everyone from tooling vendors to agencies and consultants. For a category leader, systematically buying joint placements is paying to compress your own lead — and the compression shows up first in exactly the pairwise contrast that the share of model number was hiding.

One caveat worth stating plainly, because the rule is sharper than the evidence beneath it. The asymmetry is arithmetic and certain; the assumption that a mention in a joint document is worth roughly the same as a mention in a solo one is not. Comparison pages plausibly carry more weight per mention on shortlist-shaped prompts. Treat the rule as the default and let your own re-measurement overturn it if it does.

11. The Monday checklist

  • Run the Basket Swap on one category, one engine, one month of raw answers. Write down the three-way spread before you look at any dashboard number.
  • Compute the average number of brands named per answer for your category and record it every month. If your share moves and this moved, the engine moved, not you.
  • Ask your vendor two questions in writing: which denominator the score uses, and whether the cross-engine figure averages shares or pools counts. If either answer is unavailable, the score is not reportable as a level.
  • Replace the headline percentage in your next board pack with three pairwise ratios, one per named rival, reported per engine and in words.
  • Sort your current prospect list by whether the document already names your closest rival, then buy joint or solo according to which side of 1.0 that pair sits on.
  • Set the review cadence against margin, not against the calendar: a pair inside 1.1 of parity is worth watching monthly, a pair beyond 2.0 is not worth watching at all.

The underlying discipline has not changed as much as the vocabulary suggests. Being named alongside the right rivals, in the right documents, on the surfaces your buyers actually use, is the same earned corroboration problem it has always been, and the tactics that build it are the ones that were already working. What changed is the reporting layer on top of it. A share is a claim about a whole, and nobody in this market can tell you what the whole is — so make claims about pairs instead, where the arithmetic is honest, the tooling you already pay for holds the inputs, and the benchmark data you cite is doing work it can actually support.

Leave a Reply

Your email address will not be published. Required fields are marked *

AI Visibility Variance Previous post Variance, Drift and Noise: Why Your AI Visibility Score Moves
Benchmarking Citation Share Next post Benchmarking Citation Share vs Competitors: A Reproducible Method