Sectors we take work in
Published projects across twelve datasets
Public datasets, every figure reproducible
0Numbers here you have to take on trust

Weekly SKU demand · UCI Online Retail II, 1,344 SKUs

0.685WMAPE, 13-week holdout
0.724Best baseline
02 · Financial services & insurance

Claims leakage never announces itself. It arrives as a slightly worse loss ratio.

Banks and insurers rarely have a modelling shortage. They have a drift problem, an ownership problem and a capacity problem. Scorecards are recalibrated on an annual cycle regardless of whether the population moved in March; challenger models sit in a notebook because nobody agreed what would trigger a switch. We put monitoring, a documented recalibration trigger and a named owner around models that already exist before we propose building anything new.

On the claims side the binding constraint is almost never recall — it is investigator hours. A fraud model tuned to theoretical performance floods a team of forty with three hundred referrals a week and gets switched off within a quarter. We size triage to the capacity that actually exists, agree a false-positive budget with the business in writing, and route the marginal cases to a cheaper check rather than a full investigation.

Typical engagements

  • Credit scorecards with population-stability monitoring and a written recalibration trigger
  • Claims triage sized to investigator capacity, with an agreed false-positive budget
  • Fraud detection combining network features with behavioural signals, scored at first notification of loss
  • Reserving and capital analytics documented to survive a regulator's second question
03 · Healthcare & life sciences

Capacity is a forecasting problem that arrives disguised as a staffing problem.

Emergency arrivals are more predictable than most hospital boards believe. Day of week, hour of day, public holidays, school terms, weather and a handful of local events explain the majority of the variance, and a 72-hour forecast by unit is achievable with data the trust already holds. The genuinely hard part is the other side of the ward: discharge timing, which depends on a pharmacy round, a consultant signature and a transport booking that nobody has ever modelled together.

In life sciences the same pattern repeats at a different tempo. Trial recruitment is forecast at programme level and managed at site level, so under-performing sites are carried for months because there is no agreed stop rule. We build site-level enrolment curves with an explicit decision attached: keep, support or close by a stated date. Adherence work follows the same logic — predict the drop-off, then target the intervention at the patients whose behaviour the programme can actually change.

Typical engagements

  • 72-hour patient flow forecasts by unit, refreshed hourly and visible on the bed meeting screen
  • Discharge-readiness models that surface tomorrow's blockers on today's ward round
  • Site-level trial recruitment curves with a stop rule agreed before first patient in
  • Adherence and persistence modelling for therapy support programmes
04 · Manufacturing & supply chain

Unplanned downtime is paid for long before anyone measures it.

Most plants have ten years of historian data and no failure history worth training on, because failures are rare and the maintenance log is free text. So we stop pretending it is a classification problem. We frame it as survival — time to the next event, conditioned on load, temperature, vibration and the age of the last intervention — and we rank assets by the cost of the stoppage rather than by the probability of it. A cheap pump on the critical path outranks an expensive one with a spare.

Inventory is the same argument at a different altitude. Safety stock is usually set from a service-level table that assumes lead times are stable, and lead times have not been stable since 2020. Modelling the actual lead-time distribution per supplier, per lane, usually releases more working capital than any forecast improvement — and it makes supplier risk measurable from delivery variance instead of from an annual questionnaire nobody reads.

Typical engagements

  • Time-to-failure models on historian data, ranked by the cost of the stoppage they prevent
  • Multi-echelon inventory optimisation using observed lead-time distributions per supplier and lane
  • Supplier risk scoring built from delivery variance rather than questionnaire responses
  • Yield analytics that separate operator effect, batch effect and slow process drift
05 · Energy & utilities

One point of load forecast error has a published price.

Energy is the sector where our work is easiest to value and hardest to fudge. Imbalance is settled, invoiced and visible, so a percentage point of half-hourly error converts directly into a number the finance director already recognises. That also means the modelling has to be honest about weather: a single deterministic forecast is not enough, and ensembles need to be carried through to the load model rather than collapsed to a mean at the first opportunity.

The load shape itself is moving under everyone's feet. Rooftop solar has hollowed out the middle of the day, heat pumps have sharpened the winter morning, and EV charging is quietly rebuilding the evening peak in a pattern that has no historical analogue. Models trained on five years of settled data will under-forecast the new peak and over-forecast the old one. On the asset side, we rank transformers and feeders by customer-minutes at risk, which is the only measure the regulator and the control room both accept.

Typical engagements

  • Half-hourly load forecasts using weather ensembles, holiday effects and settlement-class detail
  • Embedded generation and EV charging profiles built explicitly into the baseline
  • Transformer and feeder health models ranked by customer-minutes at risk
  • Balancing cost attribution — which forecast error cost what, by half hour
06 · Technology & SaaS

Churn is decided months before the renewal call.

The signal is almost always in the product, not in the CRM. Seat utilisation, depth of feature adoption, the departure of the admin who ran the original rollout, a support ticket pattern that changes shape — these move first. Sentiment fields filled in by an account manager who is compensated on the renewal move last, and move in the wrong direction. Building the health model on telemetry is what buys you the two months of notice that make an intervention possible.

The second problem is aggregation. Net revenue retention is a single number covering three different businesses — expansion, contraction and logo churn — each with its own drivers and its own owner. Modelled together, they cancel out and the executive team argues about the average. Modelled separately, the conversation becomes specific: pricing owns contraction, product owns expansion, customer success owns logo churn, and each has a forecast they can be held to.

Typical engagements

  • Account health models built on product telemetry rather than CRM sentiment fields
  • Net revenue retention decomposed into expansion, contraction and logo churn, forecast separately
  • Cohort-level unit economics with payback measured in cash, not in bookings
  • Infrastructure capacity and cost forecasting tied to the product roadmap and the pricing model
Cross-sector

Most of the method transfers. The loss function never does.

Sector experience is worth having, and it is routinely oversold. Here is our honest split of what a consultant carries from a utility into a hospital, and what has to be learned again from scratch.

What transfers

Much of the work is the same problem wearing different vocabulary. A consultant who has done it in one sector does it faster in the next.

  • Hierarchical forecasting and reconciliation — the store rolls up to the region in retail exactly as the feeder rolls up to the substation
  • Backtesting discipline: rolling origin, honest holdouts, and a baseline that a naive model has to be beaten by
  • The engineering underneath — calendar, weather and event features, versioned data, reproducible training runs
  • Model operations: monitoring, drift alerts, retraining cadence and a rollback path agreed before go-live
  • The habit of tying every forecast to a named decision with an owner and a deadline

What does not

The rest is where transplanted playbooks fail. We spend the first two weeks of every engagement learning it rather than assuming it.

  • The cost of being wrong. An overstock is a markdown, an under-forecast ward is overtime, a missed load is an imbalance charge — the asymmetry differs every time
  • Data latency. Half-hourly settlement data and a monthly claims close demand completely different architectures
  • Explainability thresholds. A regulator and a category buyer will accept very different answers to "why did it say that?"
  • Seasonality shape — weather-driven, promotion-driven, clinically driven or contract-driven, and rarely more than one of those
  • The unit of decision, and who is allowed to make it without asking anyone

Not sure your sector belongs on this list?

Book a 20-minute scoping call. We will tell you plainly whether we have done this work before, what we would go after first, and what it would be worth if it went well. If the answer is that someone else is better placed, we will say that too.