Error: forecast distance
Bias: persistent direction
Level: category, SKU, location
Validity: stock-out and launch flags
Impact: service, inventory, cash
A small total error can hide unusable allocation
Over-forecasting one item by 100 and under-forecasting another by 100 can cancel in the total while the second item goes out of stock.
Evaluate both magnitude and direction at the grain where replenishment decisions happen.
Use more than one summary metric
MAPE becomes unstable near zero; WAPE is easier at portfolio level but can be dominated by large items.
Report coverage, zero-demand share, median error, and value segments rather than selecting a model from one score.
Observed sales may be censored demand
A stock-out records what could be sold, not everything customers wanted. Promotions, launches, price moves, and channel restrictions also break comparability.
Flag these periods and document any exclusion or lost-demand estimate.
Translate error into asymmetric cost
Under-forecasting can create lost service and expedite cost; over-forecasting can tie up cash or expire.
The same percentage error has different consequences across products, lead times, and margins.
Accept with rolling-origin backtests
Recreate forecasts at historical decision dates using only information available then, across horizons, levels, and simple baselines.
BI0 can be evaluated for demand planning; algorithms, censored-demand treatment, backtesting, and replenishment integration need project confirmation.
Separate magnitude from direction with explicit formulas
Define error sign once, for example actual minus forecast. Use MAE for average scale, RMSE for heavier tail penalty, WAPE for portfolio magnitude, and signed error for bias. Always report coverage, zero share, quantiles, and value segments because each aggregate metric has blind spots.
A good company-month total can hide severe SKU-location allocation errors. Evaluate at the replenishment grain and check whether bottom-level forecasts reconcile with category and company totals.
Recognize censored demand and event contamination
Sales during stock-out are what could be sold, not latent demand. Mark availability and use additional evidence only with documented assumptions. Promotions, launches, closures, price changes, and one-off orders also change comparability.
Rolling backtests must use only information visible at each origin. Later refunds, corrections, and campaign outcomes are future leakage when evaluating an earlier decision.
Use rolling-origin and horizon-specific evaluation
Repeat evaluation over historical decision dates and forecast horizons that match lead time. Compare against seasonal naive and moving-average baselines, not only sophisticated models.
Translate results into stock-out days, service, average stock, expiry, and expedite cost. A slightly lower WAPE may be worse when bias is concentrated in strategic items.
Monitor drift without blindly retraining
Track error distribution, directional bias, input completeness, event coverage, and planner overrides. Persistent bias or assortment change triggers diagnosis; first separate real demand change from pipeline failure and stock-out censoring.
Version every forecast and model so the organization can reconstruct the decision. Automatic retraining without diagnosis can learn incidents as normal behavior.
State the BI0 boundary
BI0 may be evaluated for demand and inventory analysis. Algorithms, hierarchy reconciliation, censored-demand treatment, backtesting, monitoring, and replenishment writeback require project confirmation.
Forecasting does not eliminate stock-outs. Sparse launches and structural breaks need intervals and human judgment, and no single metric is universally best.
Public references
Take the next step with BI0.AI
Talk through a real business scenario and see how governed AI BI can fit your team.
Explore BI0.AI