A single aggregate number now travels further than any individual review of a game. That number is produced by a conversion process with rules most readers never see.
Outlets do not use one scale
Some reviewers publish out of ten, others out of a hundred, others use letter grades, and a growing number publish only a recommendation with no score at all.
An aggregator has to translate all of that into a shared scale before averaging anything. Letter grades and recommendation verdicts are mapped to numeric equivalents chosen by the aggregator, not by the outlet.
That mapping is a judgment call. A review intended as mildly positive can land noticeably higher or lower depending on which number the site assigns to it.
Weighting decides which outlets matter
Not every aggregator treats reviews equally. Some weight by perceived reach or track record, meaning a handful of large outlets pull the average more than many smaller ones.
Because the weights are usually not published, two aggregate scores for the same game can differ while both are technically accurate averages of the same underlying reviews.
Inclusion is its own filter. A site that only accepts approved outlets is measuring a narrower slice of criticism than one accepting anything indexed.
Timing changes the number
Aggregates move as reviews arrive. Early scores often come from outlets given advance access, and those tend to be the ones with the resources to finish a long game quickly.
Later reviews frequently come from writers who bought the game and played it under normal conditions, including whatever launch problems exist. That group can pull an average down.
The result is that the number quoted on launch day and the number a month later can describe meaningfully different reception.
Averages hide the shape of the distribution
A game that everyone finds decent and a game that half of critics love and half dislike can produce the same average. The average tells you nothing about which case you are looking at.
Divisive games are exactly the ones where a reader most needs the spread, because the split usually tracks a specific design choice some players will accept and others will not.
Reading two or three full reviews from opposite ends of the range gives a better prediction of personal fit than any midpoint can.
The number acquires uses reviews never intended
Aggregate scores get used commercially in ways that have nothing to do with criticism, from store page placement to internal performance reporting inside publishers.
Once a number carries that weight, pressure builds around it, and reviewers become aware that a half-point difference in a personal verdict may be read as a business outcome.
None of that changes what a review says. It does change how much interpretation a reader should hang on a two-digit summary of dozens of separate opinions.