The Metric Can Be Right and the Decision Can Still Be Wrong

Suppose we are evaluating two models for a task.

Model A is substantially cheaper than Model B. We run an experiment and Model A also has slightly better measured accuracy.

Two three-dimensional forms share the same measured height and width, while one extends much farther in an unmeasured depth.
Two alternatives can look equivalent when the dimension that separates them is not being measured.

That seems like a fairly easy decision.

Except there is something hidden in that conclusion.

We are assuming that cost and accuracy capture what matters.

Maybe they do. But the experiment hasn’t actually established that. That may not matter for the experiment itself. It matters when we use the experiment to make a product decision.

A metric can be completely correct about what it measures and still give us an incomplete picture of what we are trying to optimize.

1. The metric is not the objective

Suppose two models both have 95% accuracy.

That is useful information.

But what does the other 5% look like?

One model might mostly fail by abstaining: it gives no answer.

The other might almost never abstain. It always provides an answer, but sometimes that answer is wrong.

If the metric treats both cases simply as:

then both types of failure count the same.

But from the perspective of someone using the product, they may be very different.

A blank answer and a plausible but incorrect answer create different problems.

So both systems can honestly have:

while creating very different experiences for the person using them.

The metric is not wrong. It is answering a narrower question.

Metrics compress information.

A large set of experimental results becomes:

94.7% accuracy.

A complicated process becomes:

18 minutes average completion time.

Thousands of customer interactions become:

52 seconds average wait time.

That compression is the reason metrics are useful. I don’t want to inspect thousands of individual cases every time I need to make a decision.

But something is necessarily lost in the compression.

  • An average hides the distribution.
  • Accuracy hides the types of errors.
  • Model cost may hide the work created somewhere else.

So there is another question that I’ve found useful:

What information did we throw away when we created this metric?

Most of the time, that information probably doesn’t matter.

The point is not to keep adding metrics forever.

The more interesting question is:

Could any of the information we removed change the decision?

If the answer is no, great.

If the answer is yes, then perhaps we are measuring the wrong boundary of the problem.

2. The model may not be the system we care about

This becomes particularly important for AI products where a person handles the output.

It is easy to think of the experiment as:

We compare accuracy, cost, latency, maybe other quality metrics, and decide which model is better.

But that may not be the actual product.

The real workflow might be:

The model is only one component of that system.

An improvement in the model-level metrics does not necessarily mean that everything downstream also improved.

3. Review effort can be part of model quality

So, what does “better” mean in a human-in-the-loop system?

One approach is:

But there may be other dimensions that matter.

A model can be cheaper to run while being more expensive to verify.

That verification cost can be time. It can be the number of times someone needs to go back to the source material. It can be the attention or effort spent deciding whether an output can actually be trusted.

It can be the difference between immediately seeing that something failed and spending thirty seconds deciding whether an answer is correct.

So maybe what we care about is closer to:

The exact function isn’t important.

“Better” was multidimensional from the beginning.

We just compressed it into the dimensions that were easiest to measure first.

4. Choosing the metric is part of the experiment

Metrics often look extremely objective.

There is nothing ambiguous about the arithmetic.

But before those numbers existed, someone decided:

  • what counts as correct;
  • which cases belong in the evaluation;
  • whether different types of errors count the same;
  • whether human review matters; and
  • where the system being evaluated begins and ends.

Once those decisions are made, the calculation can be perfectly reproducible.

But choosing what to calculate is part of the experiment.

I think this is why I’ve become more careful when a metric and what I observe seem to disagree.

Sometimes the answer is that the metric is wrong.

Sometimes the observation is misleading.

And sometimes both are right.

The metric is accurately measuring one part of the system, while the observation is pointing to something outside the boundary we originally chose.

The metric can still be right

None of this means that an evaluation needs to measure everything.

That would make experimentation almost impossible.

If accuracy and cost are enough to make the decision, then they are enough.

The question is whether we know that.

A useful sequence for me is becoming:

Is the metric calculated correctly?

Then:

What information did we throw away when we created it?

And finally:

Could any of that information change the decision?

If not, move on.

If yes, maybe the experiment needs to expand.

The dangerous metric is not necessarily one that was calculated incorrectly.

It can be completely right.

We may simply have asked it a smaller question than the decision required.


Pablo

Sep 25, 2026