What Is AI Share of Voice?
AI share of voice is the percentage of AI-generated answers across a defined prompt set in which your brand appears as a named vendor, a cited source, or a recommended option. It is measured per assistant, per prompt set, and per time window, and it functions as the answer-engine equivalent of search impression share. The metric matters because a buyer who asks an assistant to shortlist vendors sees three to five names, not ten blue links. If you are absent from that shortlist, you are absent from the evaluation entirely, regardless of how healthy your organic rankings look on a rank tracker.
The metric is deliberately narrow. It does not measure sentiment, it does not measure traffic, and it does not measure conversion. It measures presence in the answer layer, which is the only surface most AI-assisted buyers will actually read. Marketing leaders who conflate presence with performance end up optimizing the wrong variable. Treat AI share of voice as a leading indicator that sits upstream of referral sessions, branded search lift, and eventually sourced pipeline, and instrument each of those outcomes separately rather than folding them into one number.
Most enterprise teams calculate the metric for the first time and find it lower than expected. Categories with three to five entrenched incumbents typically show those incumbents capturing 60 to 80 percent of all vendor mentions, leaving challengers to compete over a thin remainder. That concentration is the strategic reason the metric earns a permanent place on the marketing dashboard. It exposes a winner-take-most dynamic that traditional keyword rank tracking smooths over and that pipeline reporting only reveals two quarters too late.
Why Are Boards Suddenly Asking About AI Share of Voice?
Boards are asking because assistant-mediated research has moved from novelty to default in the earliest stages of B2B buying. In most enterprise programs, 20 to 40 percent of buyers now open an assistant before they open a search engine when scoping a category, and that proportion runs higher among technical evaluators and younger buying committees. The consequence is a measurement gap. A company can hold perfectly stable organic rankings while quietly losing the discovery conversation happening one layer above the search results page.
The second driver is that classic top-of-funnel metrics have lost their explanatory power. Organic sessions decline while pipeline holds steady, or sessions hold while pipeline softens, and nobody in the room can explain the divergence. AI share of voice restores a legible input variable. When a CMO can report that the company appeared in 34 percent of category answers this quarter against 21 percent last quarter, the executive conversation moves from anecdote and screenshots to a trend line with a defensible methodology behind it.
The third driver is competitive urgency. Once one vendor in a category starts publishing structured comparison content, original benchmark data, and clearly attributed claims, assistants begin favoring that vendor as a reliable citation source. The advantage compounds, because retrieval systems reinforce sources they have already surfaced successfully. Teams that wait two or three quarters before they start measuring typically find the gap materially harder and more expensive to close than the same gap would have been at the outset.
How Do You Calculate AI Share of Voice?
The base calculation divides the number of answers in which your brand appears by the total number of answers generated across your prompt set, multiplied by 100. Run the identical prompt set across each assistant you care about, repeat every prompt at least three times to account for response variance, and record the modal result rather than a single sample. Variance between identical prompts run minutes apart typically ranges from 10 to 25 percent, so single-run measurement will not survive scrutiny from a finance-literate audience.
A raw presence percentage is only the starting point. Serious programs separate three sub-metrics. Mention share counts any appearance of the brand name. Citation share counts answers that link to your owned domain or attribute a claim to your material. Recommendation share counts answers where the assistant explicitly proposes you as a fit for the buyer's stated requirement. Recommendation share is the hardest to move and correlates most closely with downstream pipeline, so weight it heaviest when you build the reporting view.
Normalize across assistants before you aggregate anything. ChatGPT, Perplexity, Gemini, Claude and Copilot retrieve differently, cite at different rates, and refresh their underlying indexes on different cycles. Reporting one blended number hides exactly where you are weak. Most mature teams publish a per-assistant grid alongside a weighted composite, with the weights set by the assistant mix their own referral data shows rather than by public market share estimates that rarely reflect a specific buyer population.
The Weighted Presence Ladder Turns Messy Answers Into One Number
The Weighted Presence Ladder is a four-rung scoring model that converts qualitative answer data into a single defensible figure. Rung one is absence, scored zero, where the brand does not appear at all. Rung two is passive mention, scored one, where the brand is named in a list with no elaboration. Rung three is cited presence, scored two, where the answer links to owned content or attributes a claim to your material. Rung four is recommended fit, scored three, where the assistant actively proposes the brand against the buyer's stated constraint.
To score a prompt set, sum the rung values earned across every answer and divide by the maximum possible score, which is three times the number of prompts. A brand named in every answer but never cited or recommended scores 33 percent. A brand appearing in only half the answers but recommended in most of those will score higher, which is the correct outcome. The ladder deliberately rewards depth of presence over breadth of name-drops, because depth is what changes a shortlist.
Run the ladder monthly against a fixed prompt set and quarterly against an expanded one. Fixed sets protect trend integrity; expanded sets catch category drift as buyers adopt new terminology. Teams that maintain both typically need 8 to 12 weeks before the first meaningful movement appears, because retrieval indexes and model refresh cycles lag content publication. Set that expectation with the executive team before the first report ships, or the program will be judged as failing while it is still working.
Which Prompts Belong in Your Measurement Set?
Your prompt set should mirror the questions a real buyer asks across an evaluation, not the keywords your rank tracker happens to monitor. In practice that means 40 to 80 prompts spanning five intents: category definition, vendor shortlisting, direct comparison, pricing and commercial structure, and implementation or integration risk. Below 40 prompts the numbers are too noisy to trend reliably. Above roughly 100 the maintenance burden usually exceeds the analytical value the extra coverage delivers.
Weight the set toward the bottom of the funnel. Category definition prompts are comparatively easy to win and rarely influence a deal. Comparison prompts of the form vendor X versus vendor Y for a specific use case, and shortlist prompts of the form best tools for a given job in a given industry, are where revenue is actually decided. A workable split is 20 percent definitional, 40 percent shortlist, 25 percent comparison, and 15 percent commercial and risk questions.
Version-control the prompt set and freeze it for at least two consecutive quarters. The most common measurement failure is a team quietly editing prompts between reporting cycles, then celebrating an improvement that came entirely from asking an easier question. Store the set, the run dates, the assistant versions where they are exposed, and the raw answer text. Auditable inputs are what let this metric survive its first skeptical review from a CFO or a board member.
What Counts as a Good AI Share of Voice Score?
There is no universal benchmark, because the denominator is your own prompt set. What matters is relative position against a named competitive set and the direction of travel over time. In most categories, the leading vendor sits between 45 and 70 percent weighted presence, a credible challenger sits between 15 and 30 percent, and any vendor below 10 percent is functionally invisible in assistant-mediated research no matter how strong its traditional search footprint appears.
Set targets as gap closure rather than absolute scores. A realistic first-year objective is moving from single digits into the 15 to 25 percent band, or narrowing the distance to the category leader by roughly a third. Movement of 3 to 8 percentage points per quarter is typical for programs publishing consistently against a focused prompt set. Programs that promise substantially more than that are usually relying on prompt sets narrow enough to be gamed by a handful of pages.
Pair the headline score with two guardrails. The first is accuracy: track what proportion of your mentions describe the product correctly, because assistants confidently repeat outdated positioning pulled from stale third-party pages. The second is framing quality: a mention inside a list of budget alternatives can actively damage an enterprise vendor. Presence without accuracy and appropriate framing is not a win, and reporting it as one erodes the credibility of the whole measurement program.
How Do You Turn the Metric Into Budget and Pipeline Decisions?
The metric earns its place when it changes allocation. Map every weak prompt cluster to the specific content, third-party source, or partner asset that would plausibly fix it, then cost that fix. If comparison prompts are your weakest cluster, the remedy is usually structured comparison pages, analyst and review-site presence, and clearly attributed original data, not additional blog volume. Most teams find that three to five well-built assets move a cluster further than thirty generic posts ever will.
Then close the loop to revenue. Tag the accounts arriving through assistant referrals, watch branded search volume in the weeks after presence improves within a cluster, and check whether opportunity creation follows in the same accounts. In most enterprise programs the sequence runs 8 to 12 weeks from publication to measurable presence, and another full quarter before sourced pipeline becomes legible. Reporting the metric without explaining that lag invites the wrong conclusions from impatient stakeholders.
This is the discipline Lemniscate Growth builds into the AI intelligence pillar of its five-pillar model: measurement wired to pipeline rather than to vanity dashboards, with prompt sets owned by the same team accountable for sourced revenue. The free AEO Checkers, AI Citation Checkers and GEO Scorers inside The GrowthGPT give teams a practical way to establish a baseline before committing budget to a full measurement program, which is usually the right first move.
Ready to build measurable pipeline?
30-minute strategy session. No pitch. Just pipeline advice.
Get Your Free Strategy Session