Measurement

Answer Sentiment: It's Not Whether AI Names You, It's How

A mention is not a win. When an AI assistant names your business, it wraps your name in a sentence — and that sentence carries adjectives, caveats, comparisons, and hedges that do more to move a customer than the mention itself. The metric almost nobody tracks is how AI talks about you when it talks about you. I call it Answer Sentiment, and it's the difference between "they're the established leader" and "they're an option, though some reviewers mention slow support."

For two years the entire AEO conversation has fixated on presence: did the model name us or not? That was the right first question. But presence is now table stakes, and it's a blunt instrument. Two businesses can both be "mentioned by ChatGPT" and be in completely different competitive positions — one described as the obvious choice, the other buried in a list with a quiet asterisk. If your dashboard scores both as a win, your dashboard is lying to you.

Why "did AI mention us?" is the wrong scoreboard

Think about what actually happens in the buyer's head. They ask an assistant for help choosing. The assistant returns a paragraph. The buyer doesn't count mentions — they absorb a vibe. They come away with an impression of who's the safe pick, who's the cheap pick, who's the risky pick, and who to avoid. That impression is built almost entirely out of the framing language around each name, not the presence of the name.

So a business can be technically winning the visibility game and losing the actual decision. You can be named in 80% of relevant answers and still lose because in most of them you're the hedged option — "worth considering, but they're newer and have fewer reviews than the alternatives." Meanwhile a competitor named in half as many answers is winning because every time they appear, they appear as the default. Presence tells you that you're in the room. Answer Sentiment tells you what the room is saying about you.

This is the natural next layer on top of the metrics I've written about before. Share of model measures how often you show up. Answer position measures where in the answer you land. Answer Sentiment measures the thing those two miss entirely: the emotional and evaluative charge of the language attached to your name.

What Answer Sentiment actually measures

I break it into three layers, because they move independently and you fix them in different ways.

Layer one: framing

The adjectives and role the model assigns you. Are you "the leading" / "the established" / "the go-to," or are you "an option" / "a smaller player" / "a newer entrant"? Framing is the model's compressed verdict on your standing, and it comes straight out of how the web collectively describes you. If the corpus talks about you as the authority, the model repeats it. If the corpus is neutral, you get neutral, forgettable framing — which in a competitive answer is nearly as bad as a negative.

Layer two: caveats

The "buts." This is where mentions quietly bleed value. "Great work, though they're on the pricier side." "Highly rated, but availability can be limited." "Solid, but some customers mention slow response times." Every caveat is the model surfacing a pattern it found in your evidence base — usually your reviews — and attaching it to your name at the exact moment of decision. One recurring caveat can neutralize an otherwise glowing mention.

Layer three: comparison

How you're positioned relative to whoever else is in the answer. Are you the anchor everyone else is measured against, or are you the contrast case that makes someone else look better? Comparative framing is the most decision-relevant of the three, because AI answers to "help me choose" are inherently comparative. Being named alongside stronger-framed competitors can actively hurt you.

How to measure it without a lab

You don't need special tooling to start. You need discipline and a rubric.

Build a real prompt set

Start from the questions your customers actually type — the ones with buying intent, phrased the way a human phrases them, not your keyword list. This is the same prompt-market-fit discipline I preach everywhere: if you measure sentiment on prompts nobody asks, you're grading a test that doesn't count. Twenty to forty prompts across your real categories and locations is plenty to see the pattern.

Capture the whole answer, not just the mention

Run each prompt across the assistants your buyers use and save the entire answer, not a "yes/no, were we mentioned" flag. The sentence around your name is the data. Repeat on a fixed cadence, because AI answers drift — I've written about answer volatility and prompt drift at length, and sentiment drifts right along with presence. A single snapshot is an anecdote; a weekly series is a signal.

Score the framing on a simple scale

For each mention, score three things: framing (negative / neutral / positive), caveats (list every "but"), and comparative standing (disadvantaged / equal / advantaged). Don't overthink the scale — a consistent three-point rubric applied every week beats a precise one applied once. What you're looking for is the pattern across the set: which caveats recur, which competitors consistently out-frame you, which prompts flip you from positive to hedged. That pattern is your work list.

One tip that saves a lot of arguing: score the answer as a buyer would read it, not as you wish it read. It's tempting to count "highly rated, but pricey" as a positive because it's mostly complimentary. To a buyer comparing three options, that "but" is the whole point — it's the reason they click the option without the caveat. Grade on the effect the sentence has on the decision, not on how flattering it feels. And keep the raw answers. Six months of saved answers is the single most persuasive artifact you can put in front of a skeptical owner, because the drift from "an option" to "the established choice" is something they can read with their own eyes.

What actually moves Answer Sentiment

Here's the part people find uncomfortable: you can't prompt-engineer your way to better framing. The model's language about you is downstream of the evidence about you. Move the evidence, move the sentiment.

Framing improves when the independent web starts describing you as the authority — earned coverage, being cited by others, a presence in the reference layer. That's slow, compounding entity authority work, and it's exactly why authority can't be faked at answer time.

Caveats improve when you fix the real pattern behind them. If the model keeps saying "some customers mention slow support," the durable fix is to make support faster and let the review corpus catch up — not to argue with the model. AI is reading your reviews as structured evidence; the recurring complaint in your reviews becomes the recurring caveat in your answers. Sentiment measurement is, in this sense, the best voice-of-customer report you'll ever get, because it tells you which complaint the AI has decided is representative.

Comparative standing improves when you close the specific gap the model is using to rank you — the missing specifics, the thinner proof, the fact that a competitor has a clear answer to a sub-question and you don't. Read the comparison sentence literally; it usually names the exact thing to fix.

An example, so this isn't abstract

Picture two home-services companies in the same city, both asking, "are we winning in AI?" Both are named by the major assistants in most relevant answers. On a presence dashboard they look identical — green across the board. Then you read the actual sentences.

Company A shows up as "one of the established, highly-rated options in the area," full stop. Clean framing, no caveat, and it's usually the first name in the answer. Company B shows up as "also worth a look — they're newer, and a few reviews mention scheduling delays, though pricing is competitive." Same category, same "we got mentioned" checkmark, completely different outcome in the buyer's head. Company B is paying for its visibility in undermining caveats it doesn't even know are there, because its dashboard only ever told it "mentioned: yes."

When Company B finally reads the whole answer, the fix writes itself. The scheduling caveat is coming straight out of its reviews, so the durable move is operational — tighten scheduling, let the review corpus update — not a plea to the model. The "newer" framing is an entity-authority gap that earned coverage and corroboration close over time. None of that is visible until you stop counting names and start reading sentences. That's the entire argument for tracking sentiment: the work list is hiding in the language, and the language is invisible to a presence metric.

The mistakes I see

The first is celebrating presence and ignoring framing — declaring victory because you're "showing up in AI" while every appearance quietly undersells you. The second is treating a bad caveat as a PR problem to spin rather than an evidence problem to fix; you cannot talk the model out of a pattern that's true in your review base. The third is measuring once. Sentiment moves, and the businesses that win treat it like a vital sign they check on a cadence, not a trophy they won once.

My read

My read — and I'll flag this as a forecast, not a fact — is that Answer Sentiment becomes the metric that actually matters within a year or two, and presence becomes the thing everyone already has. Once every serious competitor is "mentioned by AI," the fight moves to how you're mentioned, and the businesses that were only counting mentions will wonder why all that visibility isn't converting. It's not converting because the framing is doing the selling, and they were never watching the framing.

There's a quieter reason I care about this metric: it keeps AEO honest. Presence can be chased with volume and technical tricks, and some of that will always feel a little hollow. Framing can't — the only durable way to get AI to describe you as the leader is to become the kind of business the evidence describes that way. Answer Sentiment drags the whole discipline back toward something real: be genuinely better, make it verifiable, and let the models repeat what's true. Start reading the whole sentence now, while your competitors are still counting names.

Key takeaways

  • Answer Sentiment is how an AI talks about you when it names you — the framing, caveats, and comparisons around your name, not just the mention itself.
  • Presence is now table stakes. Two businesses can both be 'mentioned by AI' and be in opposite competitive positions depending on the language attached to each.
  • Measure three layers: framing (are you 'the leading' or 'an option'?), caveats (every 'but'), and comparative standing (anchor or contrast case?).
  • You can't prompt-engineer better framing — the model's language is downstream of the evidence. Move the evidence, move the sentiment.
  • A recurring caveat is usually your reviews talking. Sentiment tracking is the sharpest voice-of-customer report you'll get: it names the complaint AI decided is representative.
  • Measure on a cadence, not once — sentiment drifts with the answers. A weekly series is signal; a single snapshot is an anecdote.

Frequently asked questions

How is Answer Sentiment different from share of model or answer position?
Share of model measures how often AI names you; answer position measures where in the answer you land. Answer Sentiment measures something both miss: the evaluative charge of the language wrapped around your name — whether you're framed as the leader or the hedged option. You can win on presence and still lose the buyer if the framing consistently undersells you.
Can I improve how AI describes me by changing my website copy?
Only at the margins. The model's framing is downstream of the broader evidence about you — earned coverage, third-party corroboration, and especially your reviews. Better site copy helps you get quoted, but to change 'they're the pricier, newer option' into 'they're the established leader' you have to change what the independent web and your review corpus say, which is slower, compounding authority work.
How often should I measure Answer Sentiment?
On a fixed cadence — weekly is a good default for most businesses — because AI answers drift and sentiment drifts with them. Run a stable set of real, buyer-intent prompts, capture the full answer each time, and score framing, caveats, and comparative standing. What matters is the trend across the series: which caveats recur and which competitors keep out-framing you.
Scott Tischler

About the author

Scott Tischler is the Founder & Chairman of AIrecommend.ai and a practitioner-authority on AI search and Answer Engine Optimization. With 20+ years in marketing technology — including American Express, MetLife, and UBS — and executive and professional study at Wharton, Harvard, and Oxford, he helps businesses become the ones AI recommends.

Want to be the business AI recommends?

See how AIrecommend.ai builds the entity authority answer engines reward.

Explore AIrecommend.ai