Best GEO Software
All posts
By Best GEO Software Teamshare of voicemeasurementbrand mentionsgeobuying guide

AI Search Share of Voice: Measure Brand Mentions in Generative AI Answers for Visibility

Write the share-of-voice spec before you buy a tool: prompt set, denominator, mentions vs citations, position, sentiment, and how often to sample each engine.

Share of voice is the number most AI visibility tools put on their first screen. It's also the number buyers accept with the least scrutiny. Two tools can report different share-of-voice figures for the same brand in the same week and both be correct, because they measured different things. If you compare vendors before you've decided what the metric means, you end up buying whichever definition flatters you during the demo.

So write the spec first. This post is that spec, laid out the way a buyer would hand it to a vendor. Our general explainer on AI share of voice covers the idea. Here the focus is on turning it into requirements you can check during a trial.

Pick the denominator and write it down

There are two common ways to calculate share of voice in AI answers, and vendors rarely say which one they use.

Answer share is the fraction of answers in which your brand appears at least once. Mention share is your mentions as a fraction of all mentions of the brands you track, across the same answers.

An illustration with made-up round numbers, just to show the arithmetic: you run 40 prompts once each and your brand is named in 12 answers. Your answer share is 30%. Across those 40 answers, the brands you track were mentioned 80 times in total, and 14 of those were you. Your mention share is 17.5%. Same data, same week, two honest numbers that differ by more than ten points.

Neither one is wrong. Answer share tells you how often you're in the conversation at all. Mention share tells you how much of the conversation you hold against a named competitor set. Your spec should say which one goes in the board report, and it should fix the competitor list, because adding one rival to the set changes mention share overnight without any change in what the engines say.

Freeze the prompt set

The prompt set defines the market you're measuring. Change it and you've changed the market. A team that adds 20 easy branded prompts in March will report a share-of-voice jump in March that has nothing to do with visibility.

Build the list from questions buyers actually ask: sales call notes, support tickets, on-site search, Search Console queries that read like questions. Tag each prompt by topic and buying stage so you can report share of voice for "comparison" prompts separately from "how do I" prompts. Then version the list. When you add or retire prompts, log the date and report the old and new sets side by side for one cycle.

Prompts aren't equal, either. A prompt thousands of people ask matters more than one nobody asks. If your tool offers prompt volumes, weight by them or at least sort by them, so a win on a dead prompt doesn't look like a win on a busy one.

Report each engine separately

ChatGPT reports 820M+ weekly active users. Perplexity reports 22M+ monthly. Gemini reports 650M+ monthly. A share-of-voice figure that averages those engines equally treats a small audience as if it were the size of the largest one. Report per engine first. If leadership wants one number, weight it on purpose and footnote how you did it.

Per-engine reporting also catches a common failure. A brand can hold steady in ChatGPT while sliding in Perplexity, and a blended number hides that until the slide is large.

Mention, citation, position, and sentiment are four columns

A mention means the engine named you in the answer text. A citation means it linked one of your pages as a source. You can have either one without the other, and a competitor's review site can be cited in an answer that names you. Our post on citations vs mentions goes deeper on why they diverge.

Position matters as well. Being the first recommendation and being the fourth name in a list with a caveat both count as one mention in raw share of voice. Sentiment is the fourth column: a mention that calls you "expensive and hard to set up" isn't the same as one that calls you the default choice. Your spec should require the tool to store all four per answer, so you can filter share of voice to "named first, positive framing" when that's the number that matters.

Sample often enough to see a trend

AI answers vary between runs of the same prompt. One check per prompt per month gives you a noisy number that swings for reasons unrelated to anything you did. Daily checks, read as a trend over several weeks, are far more stable. Ask the vendor how often each prompt runs on each engine and whether that cadence holds on the plan you're actually buying.

Personas and location change answers as well. "Best CRM for a small agency" asked as a founder in Texas can return a different list than the same prompt asked as an operations lead in Germany. If you sell in several markets, your spec needs country, and sometimes state or city, as a dimension.

The spec as a checklist

RequirementWhy it mattersAsk the vendor
Stated denominatorAnswer share and mention share differWhich formula does your share-of-voice chart use?
Fixed competitor setAdding a rival moves mention shareCan I lock the set and see history when it changes?
Versioned prompt listNew prompts shift the marketIs there a change log for prompts?
Per-engine breakdownAudiences differ in sizeCan I report each engine separately?
Mention vs citationThey diverge oftenDo you store both per answer?
Position and sentimentNot all mentions are equalCan I filter share of voice by them?
CadenceAnswers vary run to runHow often does each prompt run on my plan?
Personas and locationAnswers vary by askerWhich dimensions can I set per prompt?

How Promptwatch maps to the spec

We recommend running this spec in Promptwatch. Its share-of-voice and competitive benchmarking views sit on top of prompt tracking that supports topics and tags, personas, and country, state, or city targeting, so the freeze-and-segment discipline above can live inside the tool instead of a spreadsheet. Prompt volumes and difficulty scores give you the weighting. Prompt trends show how a prompt's visibility moved over time and what changed between checks, which is what you need when share of voice drops and someone asks why. Citation analytics keep the mention column and the citation column apart, down to the page and domain. Sentiment analysis covers the framing column.

Data comes from the real interfaces of ChatGPT, Perplexity, Gemini, Claude, and the other engines it covers, plus Google AI Overviews, rather than from API calls alone. Essential is $95/mo for 50 prompts on one project. Professional is $245/mo for 150 prompts across two projects, which is enough to split a brand's prompt set by market. Business is $579/mo for 350 prompts. There's a free Explore tier, but it only covers ChatGPT and 10 prompts, which is too small a sample for a share-of-voice number you'd show anyone.

If you're at the shortlist stage, our comparison of share of voice tracking tools for ChatGPT and Perplexity runs the same spec against the other vendors. The Promptwatch trial is at promptwatch.com.

Reporting it without misleading anyone

Keep the first report plain. Same prompt set as last period, or a note saying what changed. One row per engine. Answer share and mention share both shown, with the formula in a footnote. Position and sentiment as a second table for the prompts where you were named. When the number moves, link the prompt trend that moved it. A share-of-voice chart that can't be traced back to individual answers will be questioned the first time it drops, and you'll want the trail ready.