AI in product data has a clean dividing line, and it is not the one most articles draw. The question is not whether a task is complex. It is whether the output can be checked against something.
Mapping a supplier’s “shade” field to your “colour” attribute is verifiable, because a schema either accepts the value or rejects it. Asserting that a coating is suitable for marine use is not, because nothing in the system knows whether that is true. Everything useful about AI in product data follows from that difference.
Where AI in product data already works
Five tasks are past the argument stage and into normal operation.
Attribute mapping. Supplier fields get matched to your internal schema at volume, so “shade”, “colour”, and “finish” all land in the right place. The output is checkable against the model, which is why it works.
Specification extraction. Values pulled out of unstructured PDFs, spreadsheets, and supplier pages into structured fields. This is the highest-value task in most catalogues, because the data already exists and is simply trapped in documents.
Classification. Products assigned to internal categories and to external taxonomies such as GS1 or a marketplace structure. Volume makes this unbearable manually and routine automatically.
First-draft copy. Titles, bullet points, and descriptions generated from an attribute set. Draft is the operative word, and treating it as finished output is the most common mistake we see.
Gap detection. Flagging missing mandatory fields, values outside expected ranges, and records that contradict their category siblings. Quietly the most useful of the five, because it converts data quality from a periodic audit into a continuous check.
The gain is speed on work that was previously bounded by how many people you could put on it. That is the honest benefit, and it is a large one.
Notice what is absent from the list. None of these five decides anything. They propose, extract, and flag, and each output is then either accepted or corrected. That is the shape of the work AI is currently good at.
Why those tasks work and others do not
Look at what the five have in common. Each produces an output that something else can check.
A mapped attribute either fits the schema or does not. An extracted dimension can be validated against a range. A category assignment can be scored against sibling products. A missing field is missing or present. In each case the model is doing structural work, and structure is verifiable without a human knowing anything about the product.
Claims are different. A model writes that a fixing is suitable for exterior use, or that a compound meets a standard. Nothing in the pipeline can confirm either. The sentence is fluent, plausible, formatted correctly, and possibly wrong. No amount of model improvement changes that, because the problem is not capability. It is that the truth lives outside the system.
There is a further consequence worth planning for. Structural tasks get more reliable as your data improves, because there is more signal to learn from. Claim-based tasks do not, because the failure is not statistical. A model with perfect grammar and complete attributes can still assert something untrue about a product it has never encountered.
Sort your intended use cases along that line before you buy anything. It predicts which ones will need review far better than any vendor demo.
The limits of AI in product data
Four limits, each with a practical response.
Input quality sets the ceiling. Inconsistent supplier data produces inconsistent output faster than before. AI does not repair a broken attribute model, and applying it over one multiplies the inconsistency across every channel at once.
Regulated content needs a human. A model can produce an elegant description and silently drop a legally required safety statement. Nothing told it the statement was load-bearing. In building materials, chemicals, healthcare, and food, review is not optional. We have set out the risks of AI-generated content separately.
Opaque decisions propagate errors. If nobody can see why an attribute was mapped a particular way, a systematic mistake spreads unnoticed. The response is provenance: record the source field, the method, and the confidence for every generated value. Then a wrong mapping is traceable and reversible instead of permanent.
Domain nuance defeats general models. A model knows stainless steel is a material. It does not know whether grade 304 or 316 is acceptable in a chlorinated environment, and it will not hesitate before telling you. That judgement belongs to someone who has specified the part.
How to deploy AI in product data with review built in
The operating model matters more than the tool, and this is the part most implementations skip.
Route by confidence rather than by task. High-confidence structural output can post automatically. Low-confidence output goes to a review queue. Set the threshold per category, so fixings and fasteners run looser than anything safety-critical or regulated.
Keep humans on claims, always. Anything a customer could quote back at you stays in review, regardless of confidence score. The same goes for anything a regulator might ask about.
Scope by risk, not by enthusiasm. Start where errors are cheap and volume is high, which is usually supplier data onboarding and classification. Move toward customer-facing content once the review workflow is proven. That sequencing is how our applied AI work usually runs.
Put it inside the platform. AI running as a side process produces output nobody governs. Inside the PIM and PXM services workflow it inherits the same approvals and audit trail as everything else.
What to measure
Three numbers tell you whether it is working, and none of them is a productivity claim. Vendors will offer you percentages saved. Your own operation will tell you more within a month.
Automation rate: the share of records passing without human intervention.
Correction rate: how often accepted output later needs fixing.
Time from supplier file to live product: which is the outcome anyone outside the team cares about.
The second number is the important one. If your correction rate is not falling over successive batches, the model is not learning your categories and the review burden will never reduce. That is a signal to change the approach rather than to push more volume through it.
Baseline all three before switching anything on. A comparison against a remembered version of last quarter convinces nobody, least of all the person approving next year’s spend.
Where this leaves you
AI has not replaced anyone in product data but it has allowed businesses to scale their operations. Effort moves away from mapping and typing, toward defining standards, resolving edge cases, and deciding what is true.
If you want to work out which of your tasks are genuinely automatable, book a thirty-minute discovery call. We will talk it through against your catalogue.