Best GEO Software
All posts
By Best GEO Software Teamtools

Brand Sentiment Analysis in ChatGPT, Claude, and Gemini with Sentiment Actions

How sentiment scores, trends, and Sentiment Actions work when ChatGPT, Claude, or Gemini talk about your brand.

A mention can still be a warning. "Acme is cheap" and "avoid Acme" both count as visibility. Sentiment analysis is the function that scores how the model talked about you, shows the trend, and highlights the phrase in the answer. Sentiment Actions turn a shift into a task, and a task is what stops a trend from being a chart nobody reads.

The chart-nobody-reads problem is the one Sentiment Actions exists to solve. A trend chart is a picture. A task is a commitment. A team that looks at a chart can decide the trend is not urgent. A team that receives a task has to decide what to do with it, and deciding what to do with it is the work. The difference between a chart and a task is the difference between information and obligation, and sentiment analysis without the action layer is information that never becomes obligation.

Promptwatch is the tool that has this on the same prompt ledger as visibility and citations. Checks run daily on paid plans against the real UI, including ChatGPT, Claude, and Gemini, plus Perplexity, Grok, Llama, DeepSeek, Mistral, Copilot, AI Overviews, and AI Mode. This is not a social listening suite. It will not ingest Twitter, Trustpilot, or G2. Site: promptwatch.com. Knowing what it is not is half the setup, because a client who expects Twitter ingest will be disappointed by a tool that never promised it.

The same-ledger point is what makes the sentiment score useful instead of orphaned. A sentiment score on its own is a number. A sentiment score next to the visibility score and the citation for the same prompt is a story. The story says: the model mentioned you, cited this URL, and said this about you. The three numbers together tell you whether the mention was good, bad, and whether it came with a link. A sentiment score without the other two is a mood reading without context, and a mood reading without context is a number that produces more questions than it answers.

Score, trend, and the highlighted sentence

Sentiment Score is the tone on the answers you already track. A 30-day trend tells you whether a model started hedging after a pricing change or an outage story. Highlighting in the answer is what makes the score usable. You need the clause, not a color tile, and a color tile is what you ignore while a clause is what you act on.

The clause is the unit of action because the clause is the unit of language. A color tile says the answer was negative. A clause says what the answer said. The first tells you there is a problem. The second tells you what the problem is, and what the problem is, is the only thing you can fix. A team that acts on a color tile acts on a feeling. A team that acts on a clause acts on a sentence, and a sentence is a thing you can rewrite, refute, or confirm.

Claude, ChatGPT, and Gemini do not share a personality. One engine can stay neutral while another repeats a forum complaint. Keep the series per model. A blended "AI sentiment" average will hide the engine your buyers use, and hiding the engine your buyers use is the exact failure mode sentiment analysis is supposed to prevent. The split also tells you where to spend: if Claude is the one going negative and your RFPs name Claude, the Claude series is the one that earns the work, and the ChatGPT series is the one you watch.

The per-model split is the difference between a sentiment report and a sentiment average. An average is one number built from three engines that do not agree. A report is three numbers that each mean one thing. The average hides which engine moved. The report shows it, and showing it is what lets you spend on the engine that matters. A blended average is the failure mode the per-model split exists to prevent, and the prevention is the whole point of doing sentiment per engine instead of across engines.

Paid Promptwatch plans refresh daily. That is a check cadence, not a ping the second a token changes. Ask any vendor what they are actually polling, and how often, because the answer is usually less impressive than the marketing implies. A daily check is enough to catch a pricing-change hedging pattern inside a week, and a week is the unit a content or PR team actually works in.

The cadence question is the question that separates the marketing from the product. A vendor that says "real-time sentiment" is a vendor that owes you a definition of real-time. Real-time could mean every second, every minute, or every day, and the three mean very different things at very different costs. A daily check is a cadence a team can keep up with. A second-level check is a cadence that produces more alerts than a team can read, and more alerts than a team can read is a cadence that gets muted.

Sentiment Actions

When tone shifts, Unified Actions can add a sentiment card: update a page, answer a claim, or brief a PR line. You still decide whether the model is echoing a true issue. Fixing the product beats rewriting the H1, and rewriting the H1 is what teams do when they do not want to fix the product.

The H1-rewrite reflex is the reflex that produces content without progress. A model that says "Acme is slow" is not complaining about an H1. It is complaining about a product. Rewriting the H1 to say "Acme is fast" does not change the product, and the model will keep saying "Acme is slow" because the product is still slow. The sentiment card that says "fix the product" is the card that fixes the sentiment. The sentiment card that says "rewrite the page" is the card that hides the sentiment, and hiding it is what makes it come back next week.

The weekly action digest can carry those cards. Agent Chat can pull the same series if you ask which prompts went negative. None of that replaces a human reading the highlighted sentence, because the sentence is where the truth is and the score is only a pointer to it.

What this is not, and who else scores mentions

Brandwatch and similar suites listen across the open web. Promptwatch scores AI answers on prompts you configured. Use both if you need both. Do not cancel social listening because a GEO tool has a sentiment column, because the two instruments watch different surfaces and a sentiment column on AI answers is not a social listening replacement.

The two surfaces are the two things the two instruments watch. Social listening watches what people say about you on the open web. A GEO sentiment column watches what models say about you in answer to a prompt. The first is a crowd. The second is a machine. The crowd is large and noisy. The machine is small and consequential, because the machine is the one a buyer asks before a purchase. Cancelling the crowd to keep the machine is fine. Cancelling the machine to keep the crowd is fine. Cancelling one because the other exists is the mistake, because the two are not substitutes.

Otterly.AI from $29 per month includes sentiment on its mention checks, four base engines with Gemini and Claude add-ons, up to 7-day lag. We have seen Otterly mark a neutral mention as positive, so treat that column as directional. Peec AI from $95 per month is three models and a score. Profound Starter is $99 per month annual, ChatGPT-only. Method: how we rank. Tools list.

The directional caveat on Otterly is the caveat that keeps the column honest. A sentiment column that marks a neutral mention as positive is a column that overcounts good news. Treating it as directional means treating it as a hint, not a verdict, and a hint is what a column with a known overcounting problem should be. A verdict from a column that overcounts is a verdict that overcounts, and the overcount is what produces the false comfort.

Explore is free, ChatGPT with 10 prompts. Essential is $95 per month with a 7-day trial and 50 prompts. Professional is $245 per month if you need more engines' worth of history in one project pair. Promptwatch is 4.7 on 5 on G2, 1,840 or more brands.

FAQ

Is a negative score a PR crisis?

Not by itself. Open the highlighted sentence. If the model recited a true limitation, change the product or the docs. If it recited a stale rumor, write the correction on a page crawlers can fetch, and a page crawlers can fetch is the operative phrase. A correction on a page a crawler cannot fetch is a correction the model will never see, and a correction the model never sees is a correction that does not correct anything.

Will we get a ping the moment sentiment flips?

No. Paid plans check daily. Plan a daily or weekly review. Do not staff an on-call for token-level mood, because token-level mood is not a thing the product measures. An on-call for a daily check is an on-call that waits a day anyway, and the waiting is the part that makes the on-call pointless.

What to do this week

  1. Add 10 prompts that invite an opinion, "is X worth it" or "X vs Y," not only navigational ones.
  2. Load them into Promptwatch and read the highlighted clauses on Claude, ChatGPT, and Gemini.
  3. Note any prompt where one model is negative and the others are not.
  4. Accept one Sentiment Action that points at a real claim you can fix.
  5. Schedule a weekly read of the trend chart. Skip building an alert theater.