Best GEO Software
All posts
By Best GEO Software Teamvisibilitymetrics

What Is an AI Visibility Score, and Why Do Vendors Calculate It Differently?

AI visibility scores compress brand presence and prominence into one number, but collection methods, denominators, and weighting rules change the result.

An AI visibility score is a summary of how often and how prominently a brand appears across a tracked set of AI answers. It is a calculated metric, not a value reported directly by ChatGPT, Gemini, Perplexity, or any other answer engine.

That second point explains why two GEO tools can monitor similar prompts and display different scores. Each vendor decides which answers enter the calculation, what counts as brand presence, how position affects the result, and whether an absent brand contributes a zero. The collection route can differ too. A score is only comparable when those choices remain stable.

Start with the unit being scored

Before looking at a dashboard average, ask what happens to one response. A basic system might assign one point when the brand appears and zero when it does not. That is effectively a mention rate, even if the interface calls it visibility.

A prominence model asks a harder question: how much of the answer belongs to the brand? Being the sole recommendation is different from appearing fourth in a list. A detailed comparison is different from a name in a closing sentence. The model turns those distinctions into a per-response value, then combines the values across the selected prompt set.

Promptwatch publishes its method in its visibility score documentation. It scores each analyzed response from 0 to 100. A substantive mention receives a score based on placement, share of attention, depth, repetition, structural emphasis, and relevance. An answer with no substantive mention receives zero.

Placement supplies the starting range in that method. A sole recommendation scores 100. A first entry in a ranked list starts in the 90 to 94 range, second place in the 65 to 80 range, and third in the 50 to 64 range. Fourth or later starts between 25 and 49. Passing or minimal mentions fall below those ranges. Attention and depth can adjust the result because list order alone does not describe the whole answer.

The denominator changes the meaning

Promptwatch calculates the displayed score by adding all response scores and dividing by all analyzed responses in the selection. Zeros remain in the denominator. The result therefore reflects coverage and prominence at the same time.

Consider a clearly labeled hypothetical panel. A brand could dominate the few answers where it appears, yet be absent from most of the panel. Averaging only the mentioning answers would make it look strong. Averaging across the full panel would produce a lower value that reflects limited coverage. Neither calculation is mathematically invalid, but they answer different questions.

This is the first item to check when scores disagree:

  • Does the denominator include every tracked response?
  • Does it include only responses where the brand appears?
  • Are failed runs, answers without web search, or unsupported engines excluded?
  • Does a date filter change the prompt set as well as the time period?

A vendor should be able to answer those questions without treating its formula as a trade secret. Buyers do not need every internal implementation detail, but they do need enough information to interpret movement.

Visibility is not sentiment

A negative answer can give a brand high visibility. If most of the response discusses problems with one product, that product dominates the reader's attention even though the tone is unfavorable.

Promptwatch keeps sentiment separate from prominence. This is a sound reporting distinction because blending the two can hide a reputation problem. A highly visible brand with poor sentiment needs different work from an absent brand. One may need product clarification or source correction. The other needs broader category association and inclusion.

Mentions and citations are separate as well. The citations versus mentions guide defines a mention as a brand named in the answer text and a citation as a URL used as a source. A brand can score for visibility without its own site being cited. Its page can also be cited without the answer naming the company. A score that quietly treats citations as mentions changes the question again.

Weighting choices produce different rankings

Even when two products use the same denominator, their weighting can differ. One may rely mostly on position. Another may reward answer share, headings, comparison detail, or repeated references. A system may give every prompt equal weight, while another weights prompts by estimated importance. A blended score may combine several engines, locales, or personas.

These decisions are not automatically flaws. Commercial prompts may deserve a separate view from broad research prompts. A local company may care more about responses in its service area. Trouble starts when the weighting is hidden or changes silently. Then a chart can move even though the underlying brand appearances did not.

For purchasing and reporting, retain the raw measures beside the composite:

  • Total analyzed responses
  • Responses with a substantive mention
  • Mention rate
  • Average answer position
  • Citation rate for the brand's domain
  • Visibility by engine and prompt group

The composite helps a reader scan. The raw counts explain it.

Collection affects the score before the formula runs

A perfect formula cannot repair an unrepresentative input. The product must first collect the type of answer the audience sees. Interface monitoring and API monitoring can return different sources and response structures. Locale, logged-in state, model selection, search activation, and run timing can also affect the sample.

Repeated sampling matters because AI answers vary between identical runs. Promptwatch's AI answer volatility definition, updated August 1, 2026, recommends around three runs per prompt and platform within a rolling seven-day window as a floor. More runs are needed for confidence in one prompt. A large aggregate can tolerate more variation than a claim about a single question.

This means a weekly score should not be read like a fixed search ranking. Small movement may reflect ordinary response variation. A sustained change across a stable prompt panel deserves investigation.

How to compare vendors without comparing labels

During a software trial, give each product the same prompt set and ask for the response-level evidence behind its score. Pick one answer where your brand leads, one where it appears weakly, and one where it is absent. Check how each state enters the average.

Then test five details:

  1. Whether zeroes remain in the denominator
  2. Whether prominence and sentiment are independent
  3. Whether citations are counted separately from mentions
  4. Whether the prompt panel stays fixed across periods
  5. Whether every aggregate opens into stored responses

Our Promptwatch review is the local starting point for its product scope. For teams that want a documented prominence formula plus the responses, citations, crawler data, and traffic around it, Promptwatch is our recommendation.

Keep the score on the dashboard, but put its denominator in the report. That one habit prevents a polished 0 to 100 number from being mistaken for an objective grade issued by the answer engine itself.