Akeneo’s new piece on preparing product data for agentic commerce opens with a car seat. A sister needs one that fits a specific car, a weight range, a safety standard, a washable cover, and a two-day delivery window. She does not open five retailer tabs. She asks ChatGPT and gets a shortlist.
That is the shopper side of the same story we have been writing. Searchable’s ChatGPT shopping data, which we unpacked in Incomplete Product Data Does Not Rank You Lower. It Drops You From the Answer, showed Target, Walmart and eBay showing up far more often than Amazon, because accessible structured product data beat catalog scale. The car seat is that finding with a person attached. If compatibility lives in a PDF, the cover fabric is marketing copy, and two retailers disagree on the weight range, the model does not guess. It picks a listing it can defend.
Akeneo’s useful contribution is to name what “ready” actually means, and it is not a filled-in PIM completeness bar.
Completeness for a person is not completeness for an agent
A mandatory-field score in a PIM is a team metric. An agent is answering a constraint. “Waterproof backpack suitable for commuting with a 16-inch laptop” needs waterproofing, laptop size, capacity, intended use and materials as facts, not as a vibe in the description. Akeneo’s line is practical: enrich around environments, occasions, compatibility and the problem the buyer is trying to close.
The same goes for adjectives. “Sleek design” is noise. “6-inch footprint that fits beneath standard kitchen cabinets” is something a model can match to a kitchen. Returns, customer questions and search queries already tell you which facts shoppers care about. Those facts often never made the original record because nobody owned them as attributes.
Machine-readable is the other half. Height, width, depth, material, identifiers and relationships as fields, not as a paragraph the agent has to scrape. Beautiful PDPs were built for people who forgive gaps. Agents do not.
Then keep the record the same everywhere. An agent will meet your product on your site, a retailer, a marketplace, a feed or a review page. Conflicting specs are not a channel problem. They are a confidence problem. The agent walks away.
The board thinks it is ready. The project plan says otherwise.
Five days before that blog, Akeneo published a survey of 1,000 senior IT decision makers in the US and Europe. Ninety-five percent say their organisation is ready for AI-driven commerce. Eighty-seven percent expect AI budgets to rise. In the US, 44 percent spend more than half of total AI project effort on data preparation: missing attributes, messy taxonomies, conflicting product information, regional price and compliance cleanup.
That is the data tax. Confidence is high. The work is still reconciling the same catalog for every new agent use case. Romain Fouache puts it plainly: if you clean and restructure the same data for every initiative, you are not building a capability. You are rebuilding the foundation every time.
This is why AEO as a content sprint is the wrong budget conversation. SEO got a page high enough that a human clicked. AEO is whether an engine can use the record in an answer. Ranking does not save you if the underlying information is incomplete, inconsistent or unverifiable. We already argued the purchase half in Crystallize Is Building a Headless Commerce OS: without an order object the agent can create and follow, you have a readable brochure and nowhere to place the order. Akeneo is the catalog half of that OS. Governed product truth, then a transaction the agent can complete.
What we would do with this
We would stop treating “AI-ready catalog” as a checkbox next to mandatory fields.
Pick ten products that actually make money. Run the questions customers already ask, the car-seat kind, through ChatGPT and Perplexity. Write down where the shortlist fails: missing compatibility, vague materials, two channels that disagree, a fact that only exists in a PDF.
Then fix the record the way an agent reads it. Put the constraint in an attribute with a unit. Kill the adjective that cannot be checked. Make the retailer listing match the brand site. Pull return reasons and customer questions into the model, because that is the intent layer most catalogs never stored.
Most companies are still struggling to unlock a single source of truth for product data, let alone put that truth in front of an agent. The car seat shortlist is already happening. The 44 percent data tax is the bill for pretending completeness was a form. Product data is the salesperson the merchandiser does not get to stand next to. Make it impossible to misunderstand, or the agent will simply shop somewhere else.
Source: https://www.akeneo.com/blog/how-to-prepare-your-product-data-for-agentic-commerce/
