Measurement & Research
AI Share of Voice: Define What Your Percentage Means
This guide is part of the King of AEO learning library.
The short answer
AI share of voice describes a brand’s share of a defined set of observed mentions, citations or other visibility units. Its meaning depends on the denominator, competitor set and sample. State those rules before calculating it, show absolute counts alongside percentages, and avoid presenting a selected prompt panel as a measurement of the whole market.
In this guide
Choose whether you mean presence or competitive shareDefine the counted event preciselyFix the competitor set and name its limitsKeep the prompt mix stable across periodsReport absolute movement beside relative movementHandle weighting and missing observations openlyUse share to locate a competitive questionWork through two different denominatorsSourcesDefine the counted event precisely
You can count individual name occurrences, answers containing a brand, citation instances or answers containing an owned-domain citation. Each choice rewards different behaviour. A verbose answer that repeats one company name can dominate a raw occurrence count. Counting a brand once per answer reduces that effect but ignores how extensively it was discussed. Citation counts measure references, which may involve a company's own page or third-party coverage. Choose the unit according to the question and retain the raw data needed to explain the result.
The citation tracking guide separates instances, pages and publishers so that source-based shares can be built consistently. If the measure concerns mentions, establish entity matching rules. A common word that is also a company name can generate false positives. Product names, subsidiaries and abbreviations need a deliberate mapping. Do not add an alias halfway through a series without checking whether historical data should be reclassified. Otherwise an apparent increase in voice may simply reflect better detection of appearances that were previously missed.
Eligible answers: Fixed panel and period
Counted unit: Mention or citation rule
Competitive pool: Defined brands and other class
Brand numerator: Eligible brand units
Share percentage: Numerator divided by pool
Share of voice depends on an explicit counting unit and competitive pool. Brand presence uses a different denominator.
Fix the competitor set and name its limits
A tracked competitive pool might contain direct alternatives, a broader category or every identifiable brand in the answers. State which one applies. If only three competitors are monitored, the resulting share is within that tracked group. Excluding a major alternative can make everyone else's percentage look larger. If new brands appear, retain them in an other or unclassified category where appropriate rather than silently dropping them. A denominator should reflect the claimed competitive universe, not just the names that were convenient to enter into a tool.
Consider an illustrative report where Brand A receives eight appearances and two tracked rivals receive six each. Brand A has forty per cent of the twenty tracked appearances. If another untracked rival receives twenty appearances, that earlier percentage does not describe the full observed category. It remains a share within a selected pool. The remedy is not to hide the earlier calculation but to label it accurately and improve the competitor coverage where the decision requires it. Competitive percentages are only as comprehensive as the counting rules that create their denominator.
Keep the prompt mix stable across periods
A shift towards questions that explicitly name the brand can increase visibility without any change in neutral discovery. A shift towards topics where the brand has strong coverage can do the same. Use a stable core panel and report major strata separately, such as branded versus non-branded or comparison versus informational prompts. Prompt tracking explains how to preserve the panel and collection conditions. Share calculations cannot rescue a comparison whose underlying questions changed materially between the two periods being presented.
NIST's sampling guidance highlights systematic error as well as precision. Applied here, a larger number of favourable prompts does not correct a biased prompt selection. Repeated runs can reveal variability, but repeated versions of the same narrow task still represent that task. Keep the reporting language tied to the sample: share among observed answers in the defined panel. The sampling bias guide covers why that qualified statement is more defensible than a claim about all AI answers or the entire buying market.
Report absolute movement beside relative movement
Share can rise because a brand gains appearances, because competitors lose them or because the observed pool shrinks. Those situations may require different actions. Always show the numerator and denominator alongside the percentage. An illustrative brand with ten appearances out of fifty has a twenty per cent share. If it retains ten while the pool falls to twenty-five, its share doubles without any additional appearance. A chart showing only the percentage can make that change look like growth in exposure when the brand's own count is unchanged.
The reverse can also happen. A brand can gain appearances while its share falls because the category expands faster. That may still be a useful outcome, depending on the objective. Avoid treating relative share as the only performance measure. Pair it with presence or owned-citation counts and relevant downstream evidence. AI referral traffic may supply a separate view of observed visits, although it cannot be inferred from mention share. The measurements answer complementary questions and should be allowed to diverge without forcing a single narrative.
Handle weighting and missing observations openly
Some teams weight prompts by commercial importance or estimated frequency. That can be useful for an internal planning index, but it changes the interpretation from a simple observed share. Document the weights, their source and whether they remain fixed over time. A subjective importance weight is not measured user demand. If the same answer is counted under several overlapping prompt groups, ensure the aggregate does not accidentally double-count it. Preserve unweighted results where useful so readers can see how much the weighting changes the conclusion.
Missing runs also affect the pool. If one platform fails more often on certain tasks, dropping those runs can change the apparent competitive balance. Report collection coverage and use a predefined comparable subset when necessary. The primary paper on valid generative AI measurement emphasises clear contexts and metrics. For share of voice, that means specifying the platform mix and eligible answer set as part of the result, rather than treating them as hidden technical details that cannot influence the headline percentage.
Work through two different denominators
Consider the illustrative observations below. Each row reports whether a brand appears at least once in an answer, so repeating one name five times in a single response does not add five observations. Twenty valid answers were collected. Several answers mention more than one brand, which is why the brand-answer observations total thirty rather than twenty.
For Brand A, answer coverage is 12 divided by 20, or 60%. Its share of the observed brand-answer mentions is 12 divided by 30, or 40%. Both calculations are correct, but they answer different questions. Keep the brand set fixed when comparing mention share over time: adding more tracked competitors can change that denominator without changing Brand A coverage. Neither result estimates population market share, and neither establishes that a mention was favourable.
| Illustrative brand | Answers mentioning brand | Coverage of 20 answers | Share of 30 brand-answer observations |
|---|---|---|---|
| Brand A | 12 | 60% | 40% |
| Brand B | 10 | 50% | 33.3% |
| Brand C | 8 | 40% | 26.7% |
Sources and further reading
- NIST sampling schemesSampling design affects precision and systematic error.
- Valid measurement of generative AIEvaluation validity requires explicit concepts, contexts and metrics.