Every operating partner who asks about vendor consolidation wants one number: how much are we spending with this vendor across the portfolio? Getting to that number is harder than it should be.
The problem is not that the data doesn't exist. Every portfolio company has invoices, contracts, and purchase orders that contain the answer. The problem is that the data exists in formats that were not designed to be aggregated, across systems that were not designed to talk to each other, at companies that were not designed with portfolio-level visibility in mind.
Why the number is hard to get
Consider a mid-size enterprise with a Salesforce contract, a MuleSoft contract, and a Tableau contract. At the vendor level, all three are Salesforce. At the invoice level, they may be billed by three different Salesforce entities, with three different cost center codes, through three different procurement channels. Finance sees three line items. Nobody sees one Salesforce number unless someone explicitly builds the mapping.
Now multiply that across eight portfolio companies, each with their own chart of accounts, their own ERP, and their own definition of what goes in the "software" bucket. The number that matters for the consolidation conversation — total portfolio spend with this vendor — does not exist in any system. It has to be constructed.
What construction looks like without a deterministic layer
Without a deterministic layer, construction is a consulting engagement. An analyst pulls the data from each portfolio company, normalizes the vendor names, maps the cost center codes to a common taxonomy, and produces a number in a spreadsheet. The engagement takes four to six weeks. The number is accurate as of the date the spreadsheet was closed.
Six months later, when the operating partner wants to know if the consolidation thesis is tracking, the answer is "we need to update the spreadsheet." The engagement starts again.
What construction looks like with a deterministic layer
With a deterministic layer, the vendor taxonomy is defined once and applied consistently across every data source. The mapping rules are version-controlled. The output is reproducible — the same data in, the same number out, every time. The operating partner can ask the question on any given day and get an answer that is current, consistent with the answer they got last quarter, and traceable back to the source data.
That is the difference between a number you produce and a number you have. The consolidation thesis should rest on a number you have.