Skip to content

Quality over quantity: the art of software data normalisation

Providers compete on the size of their normalisation database. Size is the wrong measure. What decides whether recognition is any good is how each entry got there.

Software normalisation — or recognition, as it is also known — is not new. It became a live topic once ITSM teams started demanding normalised software data to populate the CMDB, and it has stayed one ever since, because the demands on it keep changing.

There are nuances to a normalisation service, and to the database underneath it, that rarely get discussed. Most of the discussion is about size. Size is the wrong thing to discuss.

Why normalisation is needed at all

Take an inventory of every piece of software installed across your organisation and what comes back is a vast, noisy list. Raw inventory data is full of near-duplicates, inconsistent publisher strings, runtimes, drivers, components and versions of the same product recorded four different ways.

Turning that into meaningful information is hard and resource-intensive. Of the many thousands of applications running across a typical organisation, only hundreds carry any real commercial exposure — and knowing which hundreds is what compliance turns on. Building a list of licensable software, with publisher, product, version, edition, release date and the upgrade and downgrade rights attached, is a substantial piece of work.

For almost every organisation, doing that work internally does not make economic sense. Which is why providers built normalisation services underpinned by a reference database, used to decipher inventory data and identify the applications that actually matter.

Database size is not the measure

When providers promote their normalisation capability, most reach for the size of the database. Hundreds of thousands of entries. Millions.

A large proportion of that data is not commercially relevant. It may be useful operationally or technically, but its value for licensing and compliance is limited. A more meaningful measure would be the number of licensable applications in the database. Even that is not sufficient on its own, because it still says nothing about whether the entries are right.

For the record: the Software Recognition Database behind CerteroX SAM holds over 3.5 million titles. We are not going to pretend that number is small, and we are not going to argue it is the reason to choose us. It is a consequence of covering the market properly over eighteen years, not an achievement in itself. A database of three and a half million entries built carelessly would be worse than useless — it would be confidently wrong at scale.

Quality is what counts

So ask the question the size figure is designed to distract from. How were the entries made?

What if the process for identifying, categorising and recording licensable applications is not rigorous and not consistent? What if the work is done by unskilled resource rather than people with real licensing knowledge? The result is inaccurate entries, and a large database full of inaccurate entries is not an asset.

In our experience, the thing that makes a normalisation database dependable is understanding how the software publishers define their own products. Most publishers use SKUs — stock keeping units — as the definitive identifier for each application. Using those SKUs is the only way to remove ambiguity and populate a normalisation database accurately.

Some providers do not use SKUs, and instead create their own definitions for each piece of software. That approach was established when normalisation was new to the market and offered a quick answer to immediate demand. It served its purpose. It is not adequate now, because a bespoke definition drifts from the publisher’s definition over time, and it is the publisher’s definition you will be audited against.

Some publishers do not use SKUs at all. In those cases the provider has to research the product and establish the necessary information directly — which brings you straight back to needing skilled people and a consistent process.

Classification: UNSPSC and SWID

Normalisation is also expected to classify installed software against UNSPSC application categories. Doing that accurately requires the same rigour and the same understanding of the product being classified. We have come across normalisation databases carrying obvious category errors, which tells you the classification was applied mechanically rather than by anyone who knew what the product was.

CerteroX SAM handles Software Identification (SWID) tags with UNSPSC classification alongside publisher normalisation and version recognition, so the classification travels with the recognised title rather than being bolted on afterwards.

Recognition also has to carry lifecycle data. The Software Recognition Service holds release date, end-of-support date and extended-support date against recognised titles, which is what lets you find the software that is still installed, still working, and no longer receiving security updates. That is normalisation earning its keep beyond the licence position.

An example

Consider a customer who only wants a clear, normalised view of their Microsoft software. All they need is assurance that their provider’s database contains what is required to identify and normalise every deployed Microsoft application accurately.

The fact that one provider’s database has millions of entries and another’s has thousands — I exaggerate for effect — is of no relevance whatsoever. Either could be the one with the accurate Microsoft data. The only way to know is to look at the data itself.

The same problem has moved to SaaS

Everything above was written about software installed on machines you own. The problem has since reappeared, unchanged in shape, on the subscription side.

A SaaS application discovered through an identity provider, a vendor API and a browser extension arrives as three different strings describing the same thing. It has to be resolved to one application, with an owner, a category and a set of feature tags, before anyone can say whether it duplicates something you already pay for. That is normalisation, and it needs a catalogue behind it for exactly the same reason installed software does.

CerteroX SaaS Management runs against a catalogue of more than 35,000 applications. It is what makes feature-tag classification possible, which is in turn what makes application rationalisation possible — you cannot rank overlapping applications by recoverable saving until you know which applications actually overlap. It is also how AI tools are identified: classified from feature tags in the catalogue rather than matched against a hardcoded list, so the detection set grows as the catalogue does.

In conclusion

Using database size as a measure of a normalisation provider’s capability is flawed, whoever is quoting the number.

If you want to make sense of the software running across your organisation and get compliance and optimisation right, ask your provider what processes they use to generate and maintain their normalisation database. A consistent, rigorous approach; skilled people; publisher SKUs; correct UNSPSC categorisation; and genuine coverage of commercially licensable applications are the essentials.

Accuracy outweighs size. Work with a provider who understands that — and who will answer the process question without changing the subject back to the headline figure.

Related reading

Other posts covering the same ground.

  • Project Where's My Stuff?

    Before optimisation, rationalisation or audit defence, most organisations need something less glamorous: a trustworthy answer to what they actually own. Why the foundational phase deserves its own name — and what counts as "stuff" now that most of it never touches your network.

    • ITAM
    • SAM
    • SaaS
    5 min
  • Choosing the Right IT Asset Discovery and Inventory Tools

    Discovery and inventory are not the same thing, and most tools that claim the first are only good at the second. Six criteria for choosing between them, and why "good enough" coverage stops being good enough the moment anyone outside SAM uses your data.

    • ITAM
    • SAM
    • SaaS
    8 min
  • Licence Reharvesting: Reclaiming Your IT Assets

    Reclaiming hardware and licences that nobody is using is the cheapest saving available to an IT team. What it takes is metering you can trust, a policy people will accept, and an inventory that actually finds everything.

    • ITAM
    • SAM
    • SaaS
    7 min
From reading to evidence

Put the hardest claim here
to a technical person.

Everything argued above is checkable. Name the publisher, the billing account or the platform you would argue with, and the session is built around it — the reasoning attached, not a summary slide.

No gated download at the end of it.