Every procurement spend analysis project starts the same way. Someone pulls three years of AP and PO data, loads it into a tool, and produces a spend cube sliced by supplier, category and business unit. The cube looks credible. The top ten suppliers are ranked. Category totals add up to the general ledger. Somebody presents it to the CFO.
Then a category manager looks at the supplier list and says the number for their biggest vendor is off by a factor of three.
They are usually right. A spend cube is a pivot, and a pivot inherits every flaw in the data underneath it. The tool did its job perfectly. It aggregated exactly what you gave it. The problem is that spend data cleansing was treated as a preliminary step to get through quickly rather than the bulk of the work, which in practice is what it is.
This article walks through the five cleansing failures that quietly corrupt a spend cube, what each one hides in savings terms, and a sequence for fixing them without stalling the whole programme.
What a Spend Cube Is Supposed to Do
A spend cube organises expenditure along three or more dimensions, typically supplier, category and business unit, with geography, time period or contract added as needed. The point is to answer questions that a single-axis report cannot: how much do we spend with this supplier across every plant, in this category, compared to last year.
Getting there involves four stages that every practitioner describes in roughly the same order. Collect data from AP, purchasing systems, P-cards and expense reports. Cleanse and normalise it. Classify each transaction into a category taxonomy. Then build the cube and analyse it.
Stage two and stage three are where the effort actually sits. Skip past them and you get a fast answer that nobody in the business believes, which is worse than no answer at all, because the credibility does not come back easily.
Failure One: Supplier Names Were Never Normalised
Your ERP holds hundreds of variations of the same vendor. Acme Corp, ACME Corporation, Acme Co. LLC. One plant set up the record in 2011, another in 2019, and neither knew about the other.
The cube shows three suppliers with moderate spend. Reality is one supplier with substantial spend and a negotiating position you never used.
Two reasons this happens, and both are structural rather than careless. Data lives in multiple systems that export in different formats. And vendor naming conventions on purchase orders are rarely standardised across an organisation, particularly one that has grown through acquisition or operates across regions.
The fix is entity resolution: deduplicating, standardising and reconciling supplier records across every source system before any aggregation happens. Fuzzy matching gets you most of the way. Tax IDs, addresses and bank details resolve the ambiguous cases. The last few percent need a human who knows the supply base.
Failure Two: Nobody Built the Parent-Child Hierarchy
This one is subtler and costs more. Even with perfectly deduplicated records, a global supplier appears in your data as many separate legal entities. Regional subsidiaries. Divisions acquired under different names. A distributor reselling the same manufacturer’s parts.
Without a resolved corporate family, every fundamental procurement question returns the wrong answer. A supplier relationship you believe is worth a few million resolves, at the parent level, to several times that once its subsidiaries are rolled up. Your negotiating leverage was always larger than you knew. Spend with three subsidiaries of one parent looks like three unremarkable relationships until you combine them.
Risk exposure breaks the same way. A supplier in financial distress looks minor until you find three of your other vendors are its sister companies.
There is a fast diagnostic here. Ask your current reporting how much you spend with your largest supplier’s entire corporate family, and whether any of its subsidiaries sit in your tail spend. If answering takes a project rather than a query, the hierarchy is not managed.
Failure Three: GL Codes Were Used as Categories
This is the most common shortcut in spend analysis, and the most defensible on the surface. The GL already categorises everything. Why build a separate taxonomy?
Because GL codes and cost centre mappings are built for financial reporting, not commercial decisions. They rarely reflect how a category manager thinks about a market. “Maintenance supplies” as a GL line tells you nothing about whether bearings, hydraulic components and fasteners should be sourced together or separately. Two items that compete for the same supplier sit in different codes. Two items with nothing in common share one.
A commercial taxonomy groups spend by how it is bought and who sells it. Whether you use UNSPSC, eCl@ss or a custom structure matters less than picking one and mapping to it consistently. The practical warning from people who do this repeatedly: do not over-engineer the taxonomy. Depth you cannot maintain degrades faster than a shallower structure you can.
Failure Four: The Tail Was Left Unclassified
Classification is where automation meets its limit, and the numbers here are worth knowing before you set expectations.
Rules-based classification typically reaches around 75 to 85% accuracy on structured, PO-backed spend, and then falls over on P-card transactions, services invoices and tail spend. That leaves a meaningful share of total spend sitting in “miscellaneous” or “other.” Machine learning classification against a standard taxonomy commonly lands at 60 to 70% on the first automated pass, with practitioner review closing the remaining gap. Vendors reporting 95%+ accuracy are generally describing a trained, tuned model after human correction cycles, not a first-pass result.
A workable benchmark for most organisations is 85 to 90% classification coverage at category level, with high-confidence assignments across the top 80% of spend by value. For monitoring, a common rule of thumb treats unclassified spend above 5% as a warning and above 8% as a problem needing attention.
The tail matters more than its share of value suggests. It is where maverick and off-contract buying hides, and it typically involves the majority of your supplier count. It is also, in a lot of engagements, where the first measurable savings appear.
One further warning: a rules-based taxonomy that started at 85% accuracy can degrade towards 60 to 70% within two years as vendors change, business units restructure and acquisitions introduce new spend patterns. Classification is a maintained capability, not a one-off project.
Failure Five: Currency, Units and Timing Were Never Reconciled
The unglamorous one. Multi-country data arrives in mixed currencies, and converting at today’s rate distorts a three-year trend. Use the transaction-date rate, and document the choice.
Units of measure break price comparisons quietly. The same component bought per piece in one plant and per box of fifty in another produces two unit prices that look like a sourcing opportunity and are actually a data error.
Timing mismatches between PO date, invoice date and payment date shift spend across periods and make year-on-year comparisons unreliable. Pick one date convention for the cube and hold to it.
None of this is difficult. It is simply invisible until someone challenges a number in a category review, at which point the whole cube comes under suspicion.
A Sequence That Works
Scope before you cleanse. Decide which spend you are analysing and why. Direct materials for a sourcing wave is a different dataset from all indirect spend for a savings baseline. Cleansing everything at once is how these projects die.
Resolve the top suppliers manually. Take the twenty suppliers that matter most and resolve their corporate families by hand if necessary. Top suppliers hide the largest concentration surprises, and twenty entities is a week of work, not a programme.
Automate classification, then review the exceptions. Let the model handle the bulk. Route low-confidence assignments to a category manager. A classification system that shows its confidence level and flags uncertain assignments is more useful than one that assigns everything with false precision.
Validate against the ledger before anyone sees the cube. Total cube spend should tie to AP within a defined tolerance. Explain every gap. Unexplained variance is what destroys trust in the output.
Build the refresh, not just the file. A one-off cleansing exercise degrades from the day it is delivered. Establish a repeatable process so quality holds as new spend flows in, and monitor the unclassified percentage as a standing metric.
What Clean Spend Data Is Actually Worth
Typical spend analysis programmes are associated with cost reductions in the 5 to 15% range, but that number attaches to organisations that acted on the findings, not to the ones that produced a cube. The realistic value comes from three things.
Consolidation opportunities that were invisible while one supplier looked like four. Price variance you can now see because units and currencies agree across plants. And tail spend you can finally route through preferred suppliers, because you know what is in it.
There is a softer benefit that experienced procurement leaders rate highly. Numbers that survive scrutiny in a category review carry weight in commercial conversations. A business case built on data the finance team has already challenged and accepted moves faster than one that has to be defended from scratch every time.
Where to Start If Your Cube Already Exists
Do not rebuild it. Test it first with three questions.
Ask what you spend with your largest supplier’s full corporate family. Ask what percentage of transactions sit in unclassified, miscellaneous or other. Ask whether the cube total reconciles to the general ledger, and if not, by how much and why.
The answers will tell you which of the five failures you have, and roughly what it will take to fix. In most cases, it is less work than the original build, because the hard part was never the cube. It was always the data underneath it.
