Product data sourcing sets the ceiling on everything downstream. No amount of cleaning, enrichment, or platform investment recovers information that was never obtained, and most catalogues are limited by acquisition rather than by processing.
This is about the acquisition end specifically. For what happens after the data arrives, we have written on the repeatable supplier onboarding process and on cleaning the data once it arrives.
Where product data sourcing actually starts
Not with the supplier. With a decision about what you need.
Define the attribute standard per category first, then go looking. Teams that request “everything you have” receive whatever the supplier finds convenient, which is usually a marketing PDF and a price list. Teams that request eleven named attributes in a defined format receive most of them.
This makes attribute standards a prerequisite rather than a parallel workstream. Without them there is no way to say whether a supplier file is complete, so nobody can be held to anything.
Keep the initial request short. A list of eleven attributes gets filled. A list of ninety gets ignored, or worse, gets filled with guesses. Ask for the mandatory set first and go back for the rest once the relationship is working.
Ranking your sources by authority
Four sources, and they are not equivalent. Decide the order before you have a conflict, not during one.
The manufacturer is authoritative for technical specifications, compliance data, and dimensions. If you can get a structured feed rather than a catalogue, take it.
Supplier feeds arrive as spreadsheets, XML, CSV, or occasionally an API. Quality varies enormously between suppliers and, more awkwardly, between files from the same supplier over time. A feed that validated cleanly last quarter can fail this quarter because someone changed a column heading.
Aggregators and industry data pools are good for coverage and for filling gaps across long tails. They are second-hand by definition, so treat them as a supplement rather than a source of truth for critical values. In sectors with an established pool, check what your merchants or customers already subscribe to. Matching their source removes an entire category of disagreement.
Scraping is a last resort. It breaks whenever a site changes, and it carries terms-of-use and database-right questions that deserve a conversation with someone qualified before it becomes routine practice. That caution rarely appears in articles on this subject and it should.
Recording where each value came from
This is the step that separates a manageable pipeline from a permanent argument, and almost nobody does it.
Record the source, the date, and the method for every value you ingest. Then when the manufacturer says 500mm and the aggregator says 50cm, you can resolve it by rule rather than by opinion. You can also answer the question that eventually arrives from a customer or an auditor: where did this figure come from.
Set precedence per attribute type rather than per source. Manufacturer wins on technical and compliance values. Your own team wins on marketing copy and categorisation. Aggregators fill gaps but never overwrite a manufacturer value. Written down, that is a page. Undocumented, it is a recurring dispute between merchandising and whoever last loaded a file.
What to do when a supplier cannot supply
Some suppliers will not meet your standard. A few cannot, because they are small, or because the data genuinely does not exist in structured form anywhere in their business.
Dropping them is rarely available to you. Commercial relationships do not turn on attribute completeness, and a buyer will not lose a range over a spreadsheet. So tier them instead, and be realistic about which tier each one belongs in.
Suppliers who send structured data to your template get automated ingestion. Suppliers who send unstructured files get a mapping built once and reused. Suppliers who send nothing usable get their top-selling lines sourced by your own team, and the rest handled as capacity allows.
That last group is a commercial decision, not a data one. Sourcing a hundred SKUs yourself is worth it for a range that sells. It is not worth it for a long tail nobody searches for.
Be explicit about which tier each supplier sits in, and revisit it annually. Suppliers move up when they see that better data gets their products listed faster, and that incentive works better than any amount of chasing.
Meanwhile, make compliance easier than non-compliance. A template with clear examples, a validation response that names the specific problem, and a named contact will move more suppliers than escalation does. Most send poor data because nobody ever told them precisely what good looked like.
From product data sourcing to a usable record
Once data arrives, the sequence is standardise, validate, enrich, and only then publish.
Standardisation is mostly units and vocabulary. Convert to your internal units at ingestion rather than at display, and map supplier terms to your controlled values at the same moment. Doing this on the way in means doing it once. Doing it on the way out means doing it per channel, forever.
Validation should reject rather than warn. A rule that produces a warning nobody reads is not a rule. Incomplete records held in a queue are better than incomplete records published quietly.
Build the exception route at the same time. Rejection without a route for legitimate edge cases turns validation into an obstacle people work around. Usually by loading the data somewhere it will not be checked at all.
Enrichment then fills what sourcing could not reach, which is product content enrichment rather than a sourcing problem. The distinction matters for planning, because the two need different people.
Automation makes this survivable at volume. Mapping, normalising, and validating inbound supplier files is repetitive, rule-based work, which is what was built to handle as part of supplier data onboarding.
Product data sourcing is not a one-off
Treating sourcing as a project is the most expensive mistake available here.
Specifications change. Certificates expire. Suppliers reformulate products without announcing it, and ranges get superseded quietly. A record sourced accurately two years ago may be wrong today, and nothing in your platform will tell you unless someone asked it to.
Set a refresh cadence by risk. Compliance documents and regulated attributes need a review date and an owner. Technical specifications warrant periodic re-verification against the manufacturer, particularly in categories where a wrong figure has consequences on site. Marketing copy can sit until the range changes, since nothing breaks if it is a season out of date.
Then measure the inbound side. What share of supplier files pass validation first time, and which suppliers generate the most rework. Those two numbers tell you where to spend your supplier engagement effort, and they change slowly enough to be worth tracking.
Where this leaves you
Sourcing is the least visible part of product data work and it constrains everything after it. Define the standard first, rank your sources, record provenance, and tier suppliers by what they can realistically deliver.
If supplier data is arriving in more formats than your team can absorb, book a thirty-minute discovery call. We will talk it through against your supplier base. Our product data sourcing service sits alongside wider product data services and PIM and PXM services.