Ratings are compressed judgements, not measurements
The number is a summary of a conversation, and the conversation is the content.
Assigning likelihood and consequence produces a rating that sorts a list, which is genuinely useful. It does not produce a measurement, and the difference matters as soon as the rating travels beyond the people who set it. A board paper showing a risk at twelve conveys precision that the underlying judgement does not contain.
The specific loss is the reasoning. Two people can arrive at the same rating for entirely different reasons, one worried about frequency and the other about severity, and the rating conceals that they disagree about the nature of the exposure. A short note recording what the group actually worried about carries more information than any refinement of the scale.
There is also a well-known distortion at the extremes. An event with severe consequence and low likelihood scores modestly and is deprioritised, which is why serious frameworks handle catastrophic exposures through a separate route rather than through the general matrix. Anything whose occurrence would end the organisation should not be competing on a grid with a recurring nuisance.
Used properly the matrix is a structure for a discussion among people who know the operation, and its output is a record of who was in the room and what they concluded. Used as arithmetic it converts that discussion into a colour so it does not have to happen again.
There is one further problem with the matrix worth naming, which is that the scales are usually ordinal and are treated as though they were interval. Moving from a three to a four does not represent a consistent increase, and multiplying two such scales produces a value with no defensible meaning. This is why the same underlying exposure can be rated very differently by two competent groups using the same tool, and why the resulting number should never be the basis for a comparison between unrelated risks.