Product data enrichment is the work of turning a supplier’s raw record into something you can actually sell from. A part number, a trade price and forty words of manufacturer prose go in. A classified, attributed, described and illustrated product record comes out. That is the whole job. What varies enormously is how much of it a given catalogue needs, which is why nobody will quote you a price over the phone.
This page is the definitive version of what we mean by the term, what the work contains, and what moves the cost. It is written from inside the work rather than from a vendor brochure.
What product data enrichment actually means
Two different disciplines use the phrase “data enrichment” and they have nothing to do with each other.
The first appends firmographic and contact fields to a CRM record. You upload a list of companies, a provider matches them, and you get back employee counts and industry codes. That is contact data enrichment. It is not what this page is about.
The second, product data enrichment, operates on the product record. It takes whatever your supplier sent, in whatever state it arrived, and brings it up to a defined standard for a defined set of channels. The unit of work is the SKU. The output is measured against a schema, not against a match rate.
If your search brought you here from a generic data-enrichment result, this is the catalogue version. The two markets share a word and share nothing else.
Enrichment is one stage of four
We split the work into four stages. Collect, structure, enrich, deliver. Enrichment is the third. It is also the stage that everybody wants to buy first, because it is the one whose absence is visible on the website.
Collecting is finding the data. That means chasing supplier spreadsheets, pulling manufacturer PDFs, scraping brand sites, and occasionally photographing a physical sample in a warehouse. We cover that side of it under product data sourcing.
Structuring is deciding where each value belongs. That is your taxonomy and your attribute schema. If the schema does not exist, enrichment has no target and the work cannot be scoped, let alone priced.
Enriching is populating the record to the standard the schema sets.
Delivering is getting it into the PIM, the ERP, the marketplace template or the channel feed in the format each one demands.
The reason this matters commercially is simple. When a buyer asks us what product data enrichment costs, the honest first answer is a question about stage two. A client with a working schema is a different quote from a client with a 200-column spreadsheet and no agreed structure.
What product data enrichment includes
Six distinct kinds of work hide under the single word. Most enrichment briefs we receive ask for two of them and need five.
Attribute population
Filling the structured fields defined by the schema for that category. Voltage, thread pitch, tensile grade, IP rating, pack quantity, hazard classification. This is the part that drives faceted search and filtering, and it is the part most catalogues are worst at. If you want the longer argument for why attributes matter more than prose, we made it in our piece on product attributes.
Value normalisation
Making the populated values comparable. A category where length arrives as “1.5m”, “1500mm”, “1,500 mm” and “59 in” is not filterable, no matter how complete it looks. Normalisation covers units, casing, controlled vocabularies, tolerance formats and range notation. It is unglamorous and it is where most of the actual hours go.
Written content
Product titles built to a naming convention, short descriptions, long descriptions, feature bullets, application notes and SEO metadata. The naming convention is the important bit. A title assembled from brand, range, key attributes and pack size is repeatable across 40,000 SKUs. A title written freehand is not.
Digital assets
Images matched to the right SKU rather than the right range. Renamed to a convention, checked for background and resolution, tagged by asset type. Add datasheets, safety data sheets, installation guides, CAD files and video where the category calls for them. Asset matching is consistently underestimated, particularly in ranges where one photograph has historically served twelve variants.
Relationships
Accessories, spares, consumables, replaces and replaced-by, cross-sells, and kit membership. In automotive this is the parts-to-vehicle application data. In electrical it is the compatible-with chains. Relationship data is the highest commercial value output of an enrichment programme and the one most often cut from scope.
Channel variants
The same record expressed differently per channel. Amazon, a B2B trade portal, a printed catalogue, a marketplace with a 200-character title limit. This is not a second enrichment pass. It is a mapping exercise on top of one enriched master record, and treating it as anything else doubles your cost.
“Enriched” means nothing until you define it per category
The most expensive mistake here is starting work before agreeing what finished looks like.
Completeness is not a single number across a catalogue. A hex bolt might need eight attributes to be sellable. A three-phase inverter drive might need sixty. Scoring both against the same field list produces a percentage that tells you nothing and a remediation plan that fixes the wrong things.
What works is a per-category definition with three tiers. Mandatory attributes without which the product cannot be published. Commercially important attributes that drive filtering and comparison in that category. Nice-to-have attributes that improve the page but do not block it. Define those three lists per category, on one side of A4, before anyone starts populating anything. Our work on taxonomy and attribution exists largely because this step keeps getting skipped.
The tier definitions also give you a cost lever. Mandatory-only enrichment across a whole catalogue is a very different budget from full enrichment of your top three categories. Both are defensible. Choosing between them is a commercial decision, not a data one.
Where the source data actually comes from
Enrichment does not create facts. It finds them, verifies them and structures them. These are the sources, in rough order of how heavily we use them:
- Supplier and manufacturer datasheets. Usually PDF, usually laid out as a table that survives extraction badly.
- Manufacturer product pages. Often better structured than the PDFs, and more current.
- Supplier price files and product feeds. They carry identifiers and pack data, rarely attributes.
- Classification standards such as ETIM, ACES and PIES. They supply the attribute framework, not the values.
- Existing internal content. The legacy website, the printed catalogue, the sales team’s own spreadsheets.
- Physical inspection. This sounds absurd until you meet a forty-year-old own-brand range with no documentation.
Two practical consequences follow. First, coverage is never uniform. Some suppliers give you a clean structured feed and some give you a scanned fax. Your cost per SKU follows that distribution, not an average. Second, provenance matters. Every enriched value should be traceable to where it came from, because the first time a customer disputes a specification, somebody will ask.
What product data enrichment costs
Nobody in this market publishes numbers, which is a poor state of affairs for buyers. Here is how the pricing actually works.
There are four models in common use. Per SKU, which is the most common and the easiest to budget. Per attribute or per field populated, which suits programmes where scope varies wildly by category. Per category or per project, which suits fixed-scope remediation. And day rate or dedicated team, which suits ongoing enrichment where volume is unpredictable.
Per SKU pricing is what most buyers want, and it is fair as long as the scope behind the unit is nailed down. “Cost per SKU” with no attribute count attached is meaningless. Ours starts at £0.30 for a light pass on a well-documented category. It reaches £12 for a full pass on a category with poor source data, including asset matching and relationships.
Seven variables move that number, and they compound:
- Source data quality. A structured supplier feed against a scanned PDF is the single biggest multiplier.
- Attribute count per category. Eight fields against sixty is not eight times the work, but it is not far off.
- Whether the schema exists. If we are defining it as we go, that is a separate piece of work with a separate cost.
- Language and market count. Localisation is not translation. Units, compliance marks and sizing conventions change with the market.
- Asset work. Image matching and renaming is often a third of the total when variant ranges are involved.
- Relationship data. Accessories and fitment are the most labour-intensive output per SKU by a wide margin.
- Verification standard. Sample-checked, fully checked and dual-keyed are three different prices.
The number people forget to budget is their own internal time. Every enrichment programme needs a client-side owner who can answer category questions within a day. Where that person does not exist, throughput halves and the programme runs long. Budget for that role explicitly.
Where AI changes the economics
We use AI extraction on every programme now. It has genuinely changed the cost of two things.
Pulling structured attributes out of unstructured documents is the first. A model reading a datasheet and proposing values against a schema is faster than a person. It is also consistent in a way people are not on hour six.
Drafting written content is the second. Given verified attributes and a tone-of-voice brief, generated descriptions are a good enough starting point that human effort shifts from writing to editing.
What has not changed is verification. A model will produce a confident, well-formatted, wrong value for an IP rating. The format is indistinguishable from a right one. Every attribute that carries commercial or safety consequence still gets checked by somebody who understands the category. We wrote up the trade-off in more detail in our comparison of manual, AI and hybrid description approaches.
The practical effect on cost is that AI compresses the cheap end of the range hard and the expensive end of the range very little. Categories with clean digital source documents get much cheaper. Categories where the answer is not written down anywhere do not.
Three things buyers get wrong about product data enrichment
Treating it as a project rather than a rate. New products arrive every week. A one-off remediation with no operating model behind it decays from the day it finishes. The number to plan around is your monthly new-line volume, not your catalogue size. We made the case for treating product content enrichment as an ongoing discipline in a separate piece.
Buying enrichment to fix a taxonomy problem. If products are in the wrong categories, enriching them puts correct values into the wrong schema. Classification comes first, always.
Expecting the PIM to do it. A PIM stores, governs and distributes product data. It does not know the tensile grade of your fasteners. Buying a PIM and expecting enriched content to appear is the single most common reason we get called into a rescue. The empty PIM six months after go-live is a real and recurring pattern.
Key takeaways
- Product data enrichment means bringing a product record up to a defined standard for defined channels. It is not the CRM enrichment that shares the name.
- The work splits into six types: attributes, normalisation, written content, assets, relationships and channel variants. Most briefs underscope by three of them.
- Nothing can be priced until “complete” is defined per category, in three tiers, in writing.
- Cost per SKU is only meaningful alongside an attribute count and a verification standard.
- AI has made well-documented categories much cheaper and badly documented ones barely cheaper at all.
- Plan for a rate, not a project, and name the internal person who answers category questions.
Want a view on what your own catalogue would take? Send us one category and let us enrich a sample against your schema. You get real records back, plus a cost per SKU based on your data rather than an average. Get in touch. Or read how we run product content enrichment as an operation, not a one-off clean-up.