Supply Chain

Your Forecast Accuracy Number Is Measuring the Wrong Thing

Demand planning process measurement for the supply chain: The main process measurement for demand planning – in terms of numbers on a monthly basis – is the so-called forecast accuracy. Unfortunately, the numbers that are communicated most of the time are wrong for several reasons – not necessarily because of incorrect arithmetic – but because of two wrong decisions taken prior to the actual arithmetic.

Demand Planning measures the forecast accuracy for the product families on a monthly basis and reports an 85% average accuracy number in S&OP meetings with senior management. As a result, the team seems to do a good job. However, the plant is expediting, the DC is holding 6 weeks’ stock of the wrong SKU’s and service to the top 10 items is deteriorating. The reason for all these problems is that the team is planning at SKU-location level and the aggregation up to family level hides the inaccuracy of the individual SKU-location plans.

Why Aggregate Forecast Accuracy Flatters You, Bias Is Better Than Error for Predicting Inventory Problems, and Most Importantly Why You Should Measure the Value that Your Demand Planning Process Adds.

 

Aggregation Is Where Accuracy Goes to Look Good

The errors for over- and under-forecasts of individual SKUs will cancel out as you aggregate up to the product family level. This means that an over-forecast of one product in a family and an under-forecast of another will result in high accuracy for the family, although the individual forecasts were terrible. The errors will cancel out almost mathematically as you aggregate up to the highest level of aggregation (e.g. the company as a whole).

 

Aggregation Is Where Accuracy Goes to Look Good

The error of individual forecasts cancels out as you go to a higher level of aggregation (e.g. product family). Over-forecasting one SKU and under-forecasting another SKU in the same family results in a very good family-level forecast accuracy. The individual forecasts were wrong, but the aggregation looks excellent. As a result, the typical accuracy report of demand planning looks very different from the operational reality of your inventory, production, and service. All the key decisions regarding inventory, production, and service are made at the lowest level of aggregation (i.e. at SKU-location).

 

This means that in addition to reporting the accuracy metric(s) that you use to measure your process, you should also report out the numbers aggregated up to higher levels of planning (e.g. Product Family, Department, etc.). A good forecast accuracy report will indicate the specific metric(s) used as well as the level(s) at which the numbers were aggregated. A simple number without any indication of the specific metric(s) used and the level(s) at which the numbers were aggregated is not useful for measuring a demand planning process.

Add a simple heatmap to review error by SKU and location. Also, create a simple Pareto to review cumulative absolute error by SKU. Be surprised at how few items contribute to total error and how a focus on those few items can have significant value.

 

Pick a Metric That Matches Your Demand

MAPE (Mean Absolute Percent Error) also explodes on low-volume, intermittent items since a small actual demand value results in a huge percentage error from a very small unit error (3 units versus 30,000 units, for example). Thus, MAPE must be avoided in supply chain metrics for the scorecard.

A secondary problem with using MAPE as a measure is that it explodes on low volume items with intermittent demand. This is because the error is expressed as a percentage of the actual demand, and therefore a small error in units on a small actual can result in a large percentage error. The average error is therefore dominated by a few infrequently ordered SKUs, with the result that the process evaluation may incorrectly identify a value-adding planning process as being value-destroying.

While WAPE or WMAPE can be used to report error for a portfolio of SKUs (weighting errors by volume), for most people MAE (Mean Absolute Error) is simply the error in units. For S&OP reporting, it is best to use MAE in addition to the above metrics. A further useful measure is MASE (Mean Absolute Scaled Error) which measures the performance of a forecast against a simple naive forecast.

A scorecard can include several different measures for measuring various aspects of a demand plan. A weighted average error, such as WAPE (Weighted Average Percentage Error) or WMAPE (Weighted Mean Absolute Percentage Error), is typically used for portfolio reporting. These types of error measures weigh the errors of individual SKUs by the volume of each SKU. In order to find the absolute value of errors in dollars, a planner could use a measure such as MAE (Mean Absolute Error). In order to see if the errors are positive or negative, a planner could use a measure such as bias. In order to see how well a particular method of forecasting stacks up against a naive method, a measure such as MASE (Mean Absolute Scaled Error) can be used.

 

Bias Is the Metric That Predicts Your Inventory Problem

Measuring the error on the number of errors (i.e. the size of the error) is very different from measuring whether the error was positive or negative (i.e. the direction of the error). Errors that, on average, are 15% too high are a completely different operational problem than those that, on average, are 15% too low. Errors that are randomly too high and too low create variability that can be mitigated by safety stock. But errors that are consistently too high create excess inventory that will eventually decay to write-offs. And errors that are consistently too low create a series of problems including expedites, premium freight and service failures.

Over-forecasting leads to decayable inventory that can become waste. Under-forecasting leads to expedites, premium freight and potential service failures – all costs that are not revealed by an absolute error measurement, as they all have a negative sign to them.

Bias as a trend is a much more useful metric. It will surface up the underlying commercial influence in the process that is difficult to surface. It will not magically go away. So as a rough rule of thumb, I would look for the bias to be within plus or minus 5% on the planning family averages over time.

For demand bias, it is practical to average by segment (e.g. product category). Bias for average can hide two large biases of opposite sign for individual items. Bias for individual items can vary strongly between new products, promoted items and regular business.

 

The Question Neither Metric Answers

None of these metrics address the one question of real importance to a CFO trying to justify the existence of a planning team to the rest of the senior management team: the Question Neither Metric Answers.

Accuracy does not tell you whether the accuracy is good. A score of 78% is good compared to what? It is only the forecasting process manager and the planning team leader who care about this. The CFO of a company that needs a planning team cares about whether or not it is worth the money that he or she is paying for this.

FVA (or Value Add) is the change in performance measures caused by each step in your process. In other words, each step in your process must cause your forecast to improve versus your naive forecast. If not, then you are destroying value and should be laying off the people whose salaries you are paying to do this work. FVA is the only metric that matters to a CFO looking at the cost of a planning team.

 

The only sensible way to evaluate a process step is to see if it adds value, i.e. improves the forecast quality over some naive alternative (e.g. a simple historical average). If it does not add value, then it destroys value and you are paying for it.

 

What FVA Typically Reveals

Typically, FVA results are displayed in a simple graph or chart such as a stairstep diagram. The Naive forecast is at the base of the step, followed by the time series or ML model output, then the manual adjustment by the individual planner, the average of all contributions by all planners and finally the output of the consensus process.

Studies on the FVA of individual time-series-forecasts and their corresponding actuals for over 147,000 forecast observations across 10 organizations and 22 business units have shown that judgemental refinement steps on average added very little value above and beyond the base model. Where value was added, negative adjustments (i.e. reducing a forecast) were found to be more effective than positive ones. Illustrative worked examples of FVA for practical planning scenarios are provided by practitioners in the reference.

For example, in the worked examples a process which had been established to add value to model output decreased its accuracy by several percentage points of WMAPE. The cost of this was three days of meetings per month of a planning team.

Of course, this does not mean that you should eliminate the human from the forecasting process. There are many things that a human planner will know about demand that a model will not. For example, a customer’s promotional calendar, a competitor leaving the market, a plant’s qualification status etc. But what the results of this type of analysis does mean is that instead of adding no value to the model’s output, the human can now add value to specific parts of the process instead of reducing value in total.

 

Running FVA Without Getting It Wrong

There are four things that we have to get out of the way to discuss FVA further.

Ensure all data is ‘frozen’ in time. It must be possible to go back to check on what happened and retrieve all data in exactly the same format as it was used for the work. It is all too easy to ‘rework’ data in a spreadsheet and teams have found that they are not able to go back to check on previous work.

Use more than one baseline to calculate FVA. The demand pattern of your historical data can easily be ‘mirrored’ by a simple baseline (like a naive 1) thus potentially making subsequent steps look worse than they really are. Thus check your FVA findings against at least two different baselines (e.g. naive 1 and seasonal naive).

Be sure to match the metrics to your patterns of data. For example, intermittent items will generally have poor MAPE values because of all the errors in between. So good interventions could actually appear to decrease value. In such cases, use WMAPE or MAE for aggregate metrics and track bias as a separate metric.

This should not be used as a scorecard to measure individuals’ performance. People will start to interfere with the process to report better numbers. In extreme cases people might even report worse numbers in order to report improvements in subsequent periods.

The change in error does not necessarily translate into better service or lower inventory levels. This needs to be validated against other relevant KPIs. The FVA also does not necessarily understand the impact of a process step, especially if it is applied differently to the hardest cases.

 

What to Do Before Next Month’s Review

Three checks, each of which takes an afternoon.

Re-calculate your headline accuracy numbers (e.g. MAPE) at the individual SKU-location level. Show these numbers side-by-side with your current numbers to help your stakeholders and yourself see just how much of an illusion you’ve been presenting to them.

Identify the bias for each segment over the last 12 months and merely acknowledge its direction rather than trying to put a figure on it.

Calculate the value of the final consensus forecast against the naive forecast for the same SKUs over the last quarter. It will likely be better than the hard work your team has put into their models. Use this to identify areas of your business where your work is adding value and where it isn’t.

All of this already exists in your tooling. It just needs to be decided that you wish your measures to reflect reality as opposed to the reality that you wish to exist.

← All posts

Want the perspective most relevant to you?

Tell us the decision you’re trying to improve and we’ll share the most relevant KEPLER thinking.

Talk to KEPLER