Blog

A 4.2 Rating Feels Like a Red Flag. Even Though That's 84%.

Why does a 4.2 rating feel mediocre when it's technically 84%? Turns out it's not just math — culture, psychological bias, and how different platforms calculate ratings all shape what that number actually means.

Article contents
Reading progress

Be honest for a sec — ever seen a 4.2 rating on Google Maps and instantly thought, "eh, probably mid"? Then seen a 4.8 and immediately trusted it, "okay this has to be good"?

Same, all the time actually. Open Maps, search for food, scroll the rating, skim two or three reviews, decision made in under a minute. Classic behavior. Until one day my brain randomly asked itself: What is a rating number actually telling us?

Why a 3 out of 5 feels like failing

Mathematically, a 1-5 scale has its midpoint at 3, so a 3 should mean "average," not "bad." In real life though, it plays out completely differently — a 3.0 restaurant is an instant red flag in our heads, a 3.5 still feels kind of meh, and it's only around 4.5+ that we go "okay, this is safe."

That got me curious. If almost every place people consider "worth visiting" clusters at 4.0 and above, are we actually using the full 1-5 scale, or is the scale we use in practice way narrower than that? Something more like:

- 4.0 = not great - 4.3 = decent - 4.5 = good - 4.8 = must-try Not an official rule, but that's exactly what made it worth digging into.

Then came the more annoying question

If people's rating standards are all over the place, can ratings even be compared directly? Say Restaurant A has a 4.7 and Restaurant B has a 4.4 — we'd automatically pick A. But what if A only has 30 reviews while B has 15,000? What if A is in an area where every restaurant averages around 4.6, while B is in an area where the average is closer to 4.1? Suddenly that 4.4 for B looks a lot more impressive than it did a second ago.

That's when it hit me: an average rating alone doesn't tell the full story. It's not that existing rating systems are wrong — it's that we're used to reading the number without the context behind it.

Turns out it's not just a feeling — different cultures, different standards

This isn't just a feeling, it turns out. Look at how people rate on Tabelog, Japan's biggest restaurant review platform — the standard is wildly different from the Google Maps most of us default to. On Tabelog, a 3.5 already counts as excellent, and only about 0.05% of roughly 800,000 listed restaurants (as of 2022) ever cross 4.0 — that's Michelin territory. The exact same restaurant can sit at 4.5-4.7 on Google Maps, largely because tourist reviewers tend to rate on a far more generous curve.

So a 3.5 that our brains automatically read as "meh" actually means "worth planning a trip for" in a different rating culture. A rating isn't an absolute number. It's shaped by the habits of whoever is doing the rating.

Then there's the question of who even bothers to leave a rating in the first place. Research on online reviews describes something called a J-shaped distribution: most people only leave a rating when their experience was extreme, either really great or really bad. People who felt "it was fine" tend to skip writing a review entirely. The result is that the ratings we see publicly are already skewed toward both extremes, lots of 5-stars, a noticeable cluster of 1-stars, and surprisingly few in the middle. Even though the people who genuinely felt "it was fine" are, in real life, probably the largest group of all.

The more I dig, the more questions pile up

So when you really sit with it: a rating looks like just a number, but it's actually carrying a lot, the math behind it, cultural habits that differ from place to place, and bias around who bothers to leave a review versus who doesn't.

And the more I dig into this, the more questions show up, not fewer.

If everyone's standards are different, who actually gets to decide what "good" means? If the ratings we see are already skewed because people who felt "it was fine" don't bother reviewing, how much can we really trust that number? And if the problem is genuinely this complicated, is there maybe a more honest way to read a rating than what we're doing right now?

I don't have the answer yet. But maybe that's exactly what makes this worth thinking about further.

It all started with something as ordinary as staring at a star rating on a screen, and one small question I couldn't unsee: Does the rating we see every day actually mean what we think it means?

I'm still trying to find out.