A Result Is Not the Decision
The Shape of the Evidence
Years ago, I worked on machine learning for quantitative trading.
My manager was a trader. He knew probability and statistics extremely well, but he was not particularly interested in sophisticated ML for its own sake. He was open to trying it, and we used different models, features, and approaches.
But his questions about experiments were usually very simple.
Suppose we ran a set of experiments over some parameter and got something like this:
4 → 515 → 526 → 737 → 508 → 52
The obvious result is 6.
But that was not necessarily the result he trusted. He cared about what happened around it.
Compare it with:
4 → 515 → 646 → 697 → 678 → 63
The maximum is lower.
But the second set of results tells a different story.
There appears to be a region where something consistently works.
At the time, I mostly interpreted this as caution about overfitting.
That is part of it.
But the parameter example is only one case. The broader question is what structure around a result makes the conclusion believable enough to act on.
The result is only the beginning
An experiment gives us some numbers.
But eventually someone has to make a decision: use the model, ship the feature, change the system, allocate more capital, move something into production.
So the question is, how do we get from:
This experiment produced a better result.
to:
We should act on it.
Of course, the magnitude of the improvement matters, but is that enough?
What if the experiment was repeated 100 times and this was the single best run?
What if the improvement disappears under a very small parameter change?
What if the new component performs well independently but behaves badly with something already in the system?
What if we cannot explain the underlying behavior that produced the new result?
What if the person running the experiment has a long history of reliable results and we trust their judgment?
None of these questions changes the reported number.
But all of them inform the decision.
The surrounding results are also evidence
Take the two parameter examples again.
In the first:
In the second:
If I only ask which experiment performed best, the answer is obvious.
But if I ask which result gives me more confidence that I am seeing something that may survive outside the experiment, the answer might be different.
The neighboring results matter.
Not because a smooth region is automatically correct.
Real phenomena can have sharp thresholds. Complex systems can behave discontinuously.
But if I claim that parameter 6, the one with the highest value, is capturing a robust underlying phenomenon, then the fact that nearby parameters 5 and 7 behave completely differently is information I need to explain.
The shape of the surrounding experimental results is additional evidence that itself needs an explanation.
This happens in more than one dimension
The same idea shows up in different forms.
If I claim a parameter captures a robust phenomenon:
What happens around it?
If I claim two strategies diversify risk:
How do they behave relative to one another?
If both make money independently but fail under the same event, I may not have two strategies in the sense that matters. I may have the same underlying exposure twice.
If I claim an accuracy number represents quality:
What kinds of cases constitute that accuracy?
A single aggregate can hide very different kinds of errors.
If I claim something generalizes:
How did seeing earlier results change what I tried next?
The final experiment may be clean while the path used to arrive at it was heavily influenced by previous results.
In each case, the number is real.
The problem is deciding what the number allows us to conclude.
This is not just statistical significance
There are formal ways to reason about repeated experiments, multiple comparisons, uncertainty, variance, and sensitivity. Those still matter, but I think the practical question is broader.
A result should have structure consistent with the explanation you are giving for it.
If I say there is a stable effect, I should expect some form of stability.
If I say two systems are independent, I should see evidence that they fail differently.
If I say a metric captures quality, I should understand what cases make up that metric and what important behavior it excludes.
If I say a result should generalize, I should understand how much of my research process has already adapted to the historical data.
This does not give a mechanical rule for making the decision.
It gives more dimensions along which to understand the uncertainty.
I have noticed myself asking the same questions later
Years later, while reviewing ML experiments, I noticed myself asking questions that sounded very similar.
The headline metric looked good.
But something about the observed behavior did not quite fit the conclusion.
The useful questions became:
- What cases make up this number?
- What does the metric fail to capture?
- What assumptions are embedded in the evaluation?
- What would I expect to see if our explanation were actually right?
I realized I had started asking some of the same annoyingly simple questions that had once been asked of me.
The point was not that the number was wrong.
The point was that the number was not the decision.
Decisions are made under uncertainty
Maybe the simplest version of this is that moving an experiment into production is a choice under uncertainty.
We never know everything.
Repeating the experiment 1,000 times does not eliminate every unknown.
Understanding the mechanism does not guarantee the future will behave like the past.
A stable region can still fail.
An expert can still be wrong.
So the objective cannot be certainty.
The objective is to understand the uncertainty well enough to distinguish:
what we know,
what we think we know,
what we do not know,
and which of those unknowns actually matter to the decision.
The result matters, but it is just the starting point.