Picture a forecast-error chart running from 28% down to 19% over three years. Honest work, and a real improvement. Now ask what the business started doing differently as a result. The answer is very often nothing at all.

That gap is the thing this essay is about. Organisations invest in forecast accuracy on the assumption that accuracy converts into value on its own. It does not. A forecast creates value only through a decision it changes. If the rule that runs downstream of the model is the same on Monday as it was before the model existed, the accuracy gain has no route into the profit and loss account.

The accuracy trap

Accuracy has a threshold, not a slope. Below the threshold, the forecast is too noisy for anyone to act on it, so the organisation falls back on judgement, buffers and standing rules. Above the threshold, the decision has already absorbed everything the model can tell it, and further precision falls into a gap between what the model knows and what the process is capable of expressing.

Consider replenishment. If a store orders in cases of twelve, a forecast accurate to within half a case and one accurate to within a fifth of a case produce the same order, on the same day, for the same store. The second model is genuinely better. It is also worth nothing. The decision has a granularity, a cadence and a tolerance, and none of the three moved.

A forecast is worth the difference between the decision you make with it and the decision you would have made without it. Frequently that difference is zero.

This is uncomfortable because accuracy is the one thing an analytics team can improve without asking anyone's permission. Changing an ordering policy requires operations, finance and four hundred store managers. Retraining a model requires a sprint. So the effort flows to the place with the least friction, and the number on the wall keeps improving while the business result stays flat.

Work backwards from the decision

The fix is not more modelling discipline. It is starting at the other end. Before anyone specifies features or a backtest window, we write down the decision the forecast is supposed to serve and interrogate it with four questions. If the answers are not available, the forecasting work is premature, whatever the business case says.

  1. What action does this forecast trigger, and who takes it? Name a role and a moment, not a department. "The regional planner raises or holds the safety stock parameter at the Thursday review" is an answer. "Improves supply chain visibility" is not.
  2. What threshold flips the action from one choice to another? Every decision is a partition of the number line. Find the cut points. If nobody can tell you where they are, the decision is being made on feel and the model is decoration.
  3. What does it cost to be wrong in each direction? Not the average cost of error — the cost of over, and separately the cost of under, in currency, for one unit, in this business. The two are almost never equal.
  4. How late can the answer arrive and still change what happens? A forecast delivered after the purchase order is cut has an accuracy of zero, whatever the backtest says. Latency is not an engineering detail. It is part of the specification.

These four questions are the spine of the two-week Diagnose phase we run at the start of every engagement. They are also the fastest way we know to cancel a project that should never have been funded. A brief that cannot answer question one has saved you a year, not cost you one.

Being wrong is not symmetric

The default loss function is a lie about your business. Mean absolute percentage error and root mean squared error both treat an over-forecast and an under-forecast of the same size as equally bad. Almost no real decision works that way. Over-forecast fresh produce and you write off stock at the end of the day. Under-forecast it and you lose the margin, the basket around it, and some fraction of the customer's belief that you are worth the trip.

Once you accept the asymmetry, the modelling changes. You stop optimising the central estimate and start optimising the quantile that matches the cost ratio. For an electricity distributor, being short of load is typically several times more expensive than being long, because the shortfall settles at the balancing price. Forecasting to the mean was statistically defensible and commercially expensive. Re-cutting the objective around the true cost of each direction is usually where the value actually sits.

A worked example: the threshold, not the model

Take a published one. A response model on 41,188 contacts, scored on a held-out set, under a stated cost structure. The obvious lever is the model: try a better family, tune it, buy more features.

Moving the decision threshold is worth EUR 2,828 per thousand contacts. Upgrading to a better model is worth EUR 48.42. The threshold is worth roughly fifty-eight times the model, and the default of 0.5 sits about five times too high for the economics.

Nothing about that is a modelling result. It is arithmetic on the cost of a wrong call in each direction, and it was available before anyone fitted anything. The model was never the constraint; the rule downstream of it was.

The project, with the method and the cost assumptions stated in full.

The so-what test

We now run one question in every model review, and we run it before anyone opens the accuracy report. If this number came out ten per cent different, what would we do differently? If the answer is specific — a different order quantity, a different staffing roster, a different price, a different escalation — the model is decision support and the accuracy work is worth funding. If the answer is "we would know more", the model is reporting, and reporting has a much lower ceiling.

It is a blunt test and it fails a lot of respectable work. That is the point. A model that cannot name the decision it changes will not survive contact with the operating rhythm anyway; it is cheaper to find out in week two than in month nine.

Where this leaves you

None of this is an argument against accuracy. Accuracy below the threshold is the thing standing between you and a decision worth automating, and getting there is hard, technical work that we do for a living. It is an argument against treating accuracy as the objective rather than the input. The objective is a decision that is faster, finer-grained or braver than the one you can make today.

So before the next model, spend an afternoon on the decision. Write down the action, the threshold, the cost of being wrong in each direction and the latency it can tolerate. If that page is easy to fill in, your forecast has somewhere to land. If it is hard, you have just found the more valuable problem — and it was never a modelling problem. Read how that plays out across twenty-two published projects, or bring us a decision you think is stuck.

Key takeaways

Four things to take into your next model review

Accuracy has a ceiling of usefulness

Past the point where the decision changes, additional precision is invisible to the business. Find that point before you fund the next percentage point.

Specify the action before the model

Name the role, the moment and the threshold that flips the choice. A forecast without a named decision has no mechanism for creating value.

Price both directions of error

Symmetric loss functions describe almost no real business. Optimise the quantile that matches the cost ratio, not the central estimate.

Latency is part of accuracy

A number that lands after the commitment is made scores zero, however good the backtest was. Design the delivery window with the decision.

Nothing to subscribe to yet

There is no mailing list. When there is something worth sending, it will be announced here first. Until then the published projects are the newsletter — they change when the analysis does.

See the published work
Keep reading

Related insights

Three more pieces on the same argument from different angles — pipelines, model risk and what demand data can tell you between seasons.

Bring us a decision, not a data set.

Book a 20-minute scoping call. We will take one recurring decision apart with you — action, threshold, cost of error, latency — and tell you honestly whether a forecast would change it.