How Prompt Volumes and Difficulty Scores Are Calculated for AI Search
What the volume bars and difficulty scores in an AI visibility tool actually measure, from the published methodology, so you know what you're trusting.
Every AI visibility dashboard shows you a volume number and a difficulty score next to each prompt, and almost nobody buying these tools asks where the numbers come from. That is a mistake we can help with, because at least one vendor publishes its methodology, and reading it changes how much weight you should put on the bars.
The vendor is Promptwatch, and the following comes from its own published documentation, fetched this week. We are summarizing, not endorsing the numbers as ground truth; no one has ground truth for what people type into ChatGPT.
What prompt volume is, and is not
Keyword volume measures how often a term gets typed into a search engine. Prompt volume tries to measure how often a question is actually asked of AI models, which is a different and usually smaller pool of longer, conversational language. Nobody outside the model providers sees that traffic directly, so any prompt volume figure is an estimate, and the interesting question is how the estimate is built. Google's AI features page is the official description of Overviews and query fan-out. It does not publish ChatGPT prompt volume. These bars are a vendor estimate, not a Google metric.
Promptwatch builds it from keywords, not from the prompt sentence itself. Each prompt carries attached keywords, and each keyword has a monthly search volume figure for Google, Bing, and AI search, refreshed at most once every three months. The prompt's volume is a weighted average across those keywords, and the weights are the revealing part: AI search counts five times, Bing one and a half times, and Google half. The output is deliberately not a precise integer. You get a band from 1 to 10k+ shown as one to five bars, and the field reads "No data" until at least one attached keyword has volume.
The weighting tells you what the number is for. By counting AI search five times and Google half, the vendor is signalling that the bar is meant to reflect assistant demand, not classic web demand. That is the right instinct for a prompt tracker, but it also means the bar will not match the number you remember from Keyword Planner. The two tools measure different things on purpose.
Two practical consequences. First, a prompt with no attached keywords shows nothing, which is honest behavior; be suspicious of tools that always have a number. A tool that returns a volume for every prompt, including ones with no keyword anchor, is filling gaps with a guess. Second, because the refresh cycle is quarterly at fastest, do not read week-to-week volume moves into these bands. They are for prioritization, not trend analysis. A prompt that looks like it dropped this week probably did not move at all; the underlying keyword data simply has not refreshed.
What difficulty actually blends
The difficulty score estimates how much effort visibility for a prompt would take, based on the past 30 days. It blends two things: keyword competition across major search engines in the topic area, and the authority of the sources AI models already cite for that prompt. High score, tough fight. Low score, opening.
The second ingredient is the one traditional SEO tools do not have. If the models currently cite three high-authority domains for a prompt, displacing them is a different project than displacing two thin affiliate posts, even at identical keyword competition. Difficulty that looks at the actual cited sources reflects the fight you are really in.
This also explains a finding that surprises new buyers: your Google ranking does not predict your AI visibility. Models do not rank pages; they assemble answers from whichever sources they judge fit for the specific prompt. Page one in Google with zero citations is common, and so is the reverse. A site that ranks nowhere in classic search can still be the source a model picks, because the model cares about fit for the prompt, not about classic rank. Difficulty is the closest thing to an honest read on how hard that specific prompt will be.
How to use the numbers
Treat volume bands as a sorting key and difficulty as a budget estimate. Sort candidate prompts by volume, then attack the low-difficulty ones first, and re-check quarterly as the underlying keyword data refreshes. Pricing context: prompt tracking with volumes, difficulty, and query fan-outs is on paid plans from Essential at $95/mo; the free Explore tier covers 10 ChatGPT prompts. Comparisons with the rest of the field are in our directory, and the ranking method is here: how we rank.
The pairing of the two numbers is where the work is. A high-volume, low-difficulty prompt is the rare opening you take immediately. A high-volume, high-difficulty prompt is a longer project you budget for, not one you skip. A low-volume, low-difficulty prompt is cheap to win but may not move revenue. The point of having both numbers is to stop sorting on one and over-investing in prompts that look big but are unwinnable, or under-investing in prompts that look small but are the actual buyer question.
FAQ
Why does a prompt show "No data" for volume?
The field reads "No data" until at least one attached keyword has volume. A prompt with no attached keywords shows nothing.
How often does Promptwatch refresh keyword volume?
At most once every three months. Do not read week-to-week volume moves into these bands. They are for prioritization, not trend analysis.
What does difficulty blend?
Keyword competition across major search engines in the topic area, plus the authority of the sources AI models already cite for that prompt, based on the past 30 days.