Skip to content
All writing
6 min read

Sentiment scores are a distraction

Your sentiment score moved from 6.2 to 5.8. Now what? The number is real, and it's still useless — because it compressed away the only part you could have acted on.

Almost every feedback tool will give you a sentiment score. It's the easiest thing to compute and the easiest thing to put on a dashboard, which is roughly why it's everywhere. It's also, for the purpose of deciding what to build next, close to worthless.

Here's the test. Your score drops from 6.2 to 5.8 this month. What do you do on Monday?

You can't answer that, and neither can the score. It has already thrown away the only part of the feedback that was actionable: which specific thing broke, for whom, and how many of them said so.

A score is a compression, and it compresses the wrong axis

Sentiment analysis takes thousands of specific statements and collapses them onto a single dimension: positive to negative. That dimension is real. It just isn't the one you make decisions on.

Consider two months with an identical sentiment score. In the first, a hundred people are mildly annoyed that your empty states are ugly. In the second, twelve people cannot complete checkout. Same score. Wildly different Mondays.

The score can't distinguish them because severity and frequency trade off against each other inside the average. That's not a tuning problem — it's what an average does.

It's also easy to move for the wrong reasons

Sentiment is confounded by who happened to be talking. Ship a feature a vocal segment loves and your score climbs — while the checkout bug that's quietly costing you signups goes on costing you signups. The number went the right way. Nothing got better.

This is the failure mode that matters most, because it's invisible. A metric that moves for reasons unrelated to what you care about doesn't just fail to help; it actively reassures you at the wrong moment.

What to measure instead

The useful unit isn't a score. It's a named problem with a count attached.

  • "Checkout fails at the CVV step" — 17 mentions, 8 of them naming the CVV field specifically
  • "Can't find the export button" — 14 mentions
  • "App crashes on files over ~20MB" — 12 mentions

Each of these is a decision you can actually make. You can assign it, estimate it, argue about it with real stakes, and — critically — check next month whether it shrank.

That last part is the one teams skip. A ranked list is a snapshot, and a snapshot can't tell you whether you're winning.

The measurement that actually matters is the delta

The question worth answering every month isn't "how do customers feel?" It's four narrower ones:

  • What's newly broken that wasn't before?
  • What's getting worse?
  • What's improving?
  • What did we fix, that has now actually stopped being mentioned?

That last one is the proof a fix landed — the thing that's almost impossible to demonstrate with a score, and straightforward with named, counted problems compared across two periods.

A score tells you the mood. A ranked, counted problem list tells you the job.

The honest caveat

Sentiment isn't useless everywhere. If you're tracking brand perception across a market, or watching for a PR incident, an aggregate mood signal is exactly right — you genuinely want the average, and you're not trying to file a ticket at the end of it.

The mistake is importing that metric into product prioritisation, where the question is categorically different. "How do they feel" and "what should we fix" are not the same question, and the tool that answers the first cannot answer the second.

If your feedback tool's main output is a number between 1 and 10, ask what you'd do differently if it moved. If there's no answer, it isn't measuring your problem.

Start turning feedback into decisions.

Upload your first batch and see the problems worth fixing in minutes.

14 days · no credit card required