See where your products drop out of an AI recommendation — and what changes after your team intervenes.
For ecommerce teams managing considered-purchase products. Verify what assistants say about your priority SKUs, investigate the conditions associated with lost recommendations, and prioritise by assortment, availability and commercial importance.
Start with the failures affecting your hero products, your in-stock inventory, and the merchant routes your business actually wants.
catalog truth: [—] · truth timestamp: [—] · merchant: [—]
The shelf moved into the answer.
The failure we test for
A shopper asks an assistant for a product under a specific budget that meets a specific requirement. Yours qualifies. The assistant names it, quotes a fact it got from an outdated source, concludes the product doesn’t meet the requirement, and recommends a competitor instead.
Three things are now true at once. The product was mentioned. It was represented incorrectly. And the incorrect fact appeared in the explanation used to reject it.
A brand-level mention score can miss that distinction and the team able to correct the product record may never see it.
A price that went stale three months ago puts a product outside the shopper’s stated budget.
A discontinued variant gets recommended, and the buy link goes to a dead page.
A specification carried over from a previous model generation rules a product out of a compatibility question.
of consumers who received irrelevant AI product suggestions abandoned the assistant and searched elsewhere.
Nosto, survey of 2,000 UK and US consumers, 2026higher conversion rate for shoppers arriving at U.S. retail sites from AI assistants in May 2026, with 53% more revenue per visit, versus non-AI traffic.
Adobe Analytics, 2026Much of the evidence assistants cite comes from sources brands manage or influence, which makes product-data accuracy an operational lever — but not the only factor shaping a recommendation.
Illustrative. This describes the class of failure our pilots test for. We publish observed findings only after they have been verified against catalog truth and repeated across multiple runs.
This sits between the systems you already have
Your feed and PIM platforms govern, enrich, distribute and monitor product data across commerce destinations. What they generally cannot show is the product representation and recommendation outcome produced inside each independent AI shopping experience whether the assistant misread the variant, quoted a superseded price, or routed the shopper to a retailer showing out of stock.
Governs and distributes what you submit
Retrieve, interpret, compare, recommend
Observes what supported AI surfaces retrieve, state and recommend; identifies testable interventions; measures post-change outcomes
We measure what happens after your data leaves your systems, prioritise by what it costs you commercially, and hand findings back in the format your team already works in.
Designed to work with product data from Shopify, Google Merchant Center, and major PIM and feed-management platforms.
A mention is only the first checkpoint
Being mentioned tells you the assistant found you. It doesn’t tell you whether the product was described correctly, affirmatively recommended, allowed to satisfy the shopper’s constraints, or purchasable when the shopper clicked. We track two things about every priority SKU: how far it got, and whether the commerce facts held up.
Dimension 1 — Recommendation outcome A genuine sequence. Each stage requires the one before it.
| Stage | Question | Example failure |
|---|---|---|
| Retrieved | Did the product appear at all? | Absent from every answer in the category |
| Recommended | Was it affirmatively suggested? | Named only as a competitor’s alternative |
| Preferred | Did it get the strongest placement? | Third in a list where shoppers read two |
Dimension 2 — Commerce integrity Independent checks. Any can fail at any stage.
| Check | Question | Example failure |
|---|---|---|
| Accurately represented | Was the correct product and variant described correctly? | Spec carried over from a previous generation |
| Qualified for the request | Did it satisfy the shopper’s stated constraints? | Ruled out on a budget using a stale price |
| Purchasable | Was there a valid, available route to buy? | Routed to a retailer showing out of stock |
Read together, these show where a product stopped progressing and what condition was present when it did. A product can be purchasable and never recommended; it can be recommended while described incorrectly.
Every actionable loss is routed to the team best placed to test it
Ranked by what the loss is worth — hero products, in-stock inventory, margin, launch SKUs — not by severity in the abstract.
Catalog eligibility
Could the surface retrieve and correctly identify the product?
Recommendation evidence
Was there enough credible evidence to support it for this use case?
Merchant eligibility
Was there a valid, available, trusted place to buy it?
Preference fit
Did the product actually meet what the shopper asked for?
A recommendation isn’t finished until the change has been re-tested
Finding discrepancies is becoming table stakes. Several tools now do it at SKU level. The harder question is which discrepancies are worth anyone’s time, and whether correcting them changes what the assistant does. Verified recovery is our operating standard: every material intervention we recommend includes a test plan, and we report what happened after the change — including when nothing did.
Repeated observation against catalog truth, ranked by commercial materiality, an evidence-backed change with an owner, and the same journeys re-measured under a documented comparison design. The full loop is documented in the methodology.
- Intervention
- [what changed, which fields, which records]
- Implemented
- [date]
- Observed change in recommendation win rate
- [—]
- Uncertainty range
- [—]
- Treatment observations
- [—]
- Comparison observations
- [—]
- Comparison design
- [—]
- Alternative explanations considered
- model version change · seasonal drift · competitor catalog change · concurrent interventions
Every recovery test resolves to one of three states
| Result | Meaning |
|---|---|
| Recovered | The predefined outcome improved beyond the decision threshold |
| No detected recovery | The observed change was too small or too uncertain to distinguish from ordinary variation |
| Inconclusive | The comparison design, sample size or concurrent changes prevented interpretation |
“No detected recovery” and “inconclusive” are different findings, and you need to be able to tell them apart. Collapsing them is how measurement products quietly become marketing products.
How we measure
Repeated observation
Assistants don’t answer the same way twice — we sample, report distributions, and publish the run count.
Truth with a timestamp
A price is only wrong relative to a specific merchant, market and moment.
Stated uncertainty
We publish ranges. A single number would be easier to sell and harder to defend.
Correlation named as correlation
We say “likely recommendation-affecting” until a test says otherwise.
AI Shelf Recovery Pilot
A fixed-scope engagement that ends with a tested change, not a report.
- Five priority SKUs, one market
- A curated set of high-intent shopping journeys for your category
- Repeated baseline observations across the AI shopping surfaces agreed during scoping
- Product claims verified against your catalog truth
- Ranked diagnostic hypotheses with owners and proposed tests
- One agreed intervention, prepared for implementation and tested after deployment
- Post-change re-measurement with a documented comparison design
- Working sessions with your ecommerce and product-content owners
- Final recovery report
fully credited against your first year if you continue
What we do with your catalog
We process your product data only to provide, secure, support and maintain your service. We do not use confidential catalog or commercial data — margin, inventory, cost, customer data — to train shared models or create cross-customer benchmarks. Public catalog data enters aggregate benchmarks only with your explicit opt-in. We state our retention period, deletion process, model-training policy and subprocessors in full in our privacy policy.