AI shelf intelligence

See where your products drop out of an AI recommendation — and what changes after your team intervenes.

For ecommerce teams managing considered-purchase products. Verify what assistants say about your priority SKUs, investigate the conditions associated with lost recommendations, and prioritise by assortment, availability and commercial importance.

Start with the failures affecting your hero products, your in-stock inventory, and the merchant routes your business actually wants.

shopping journey · constraint: budget + compatibilityrun 07 / n = 12
recommended
alternative
excluded — your product
rejected on a quoted spec from a previous generation
catalog truth: [—] · truth timestamp: [—] · merchant: [—]
Diagram of the failure class we test for — not a screenshot, not a finding.

The shelf moved into the answer.

The failure we test for

A shopper asks an assistant for a product under a specific budget that meets a specific requirement. Yours qualifies. The assistant names it, quotes a fact it got from an outdated source, concludes the product doesn’t meet the requirement, and recommends a competitor instead.

Three things are now true at once. The product was mentioned. It was represented incorrectly. And the incorrect fact appeared in the explanation used to reject it.

A brand-level mention score can miss that distinction and the team able to correct the product record may never see it.

A price that went stale three months ago puts a product outside the shopper’s stated budget.

A discontinued variant gets recommended, and the buy link goes to a dead page.

A specification carried over from a previous model generation rules a product out of a compatibility question.

69%

of consumers who received irrelevant AI product suggestions abandoned the assistant and searched elsewhere.

Nosto, survey of 2,000 UK and US consumers, 2026
+54%

higher conversion rate for shoppers arriving at U.S. retail sites from AI assistants in May 2026, with 53% more revenue per visit, versus non-AI traffic.

Adobe Analytics, 2026

Much of the evidence assistants cite comes from sources brands manage or influence, which makes product-data accuracy an operational lever — but not the only factor shaping a recommendation.

Illustrative. This describes the class of failure our pilots test for. We publish observed findings only after they have been verified against catalog truth and repeated across multiple runs.

This sits between the systems you already have

Your feed and PIM platforms govern, enrich, distribute and monitor product data across commerce destinations. What they generally cannot show is the product representation and recommendation outcome produced inside each independent AI shopping experience whether the assistant misread the variant, quoted a superseded price, or routed the shopper to a retailer showing out of stock.

Your PIM, feed platform, merchant accountsProduct truth

Governs and distributes what you submit

ChatGPT, Perplexity, Gemini, CopilotAI shopping surfaces

Retrieve, interpret, compare, recommend

UsAI shelf intelligence

Observes what supported AI surfaces retrieve, state and recommend; identifies testable interventions; measures post-change outcomes

Findings and tested changes return to the systems your team already runs.

We measure what happens after your data leaves your systems, prioritise by what it costs you commercially, and hand findings back in the format your team already works in.

Designed to work with product data from Shopify, Google Merchant Center, and major PIM and feed-management platforms.

A mention is only the first checkpoint

Being mentioned tells you the assistant found you. It doesn’t tell you whether the product was described correctly, affirmatively recommended, allowed to satisfy the shopper’s constraints, or purchasable when the shopper clicked. We track two things about every priority SKU: how far it got, and whether the commerce facts held up.

recommendation-state profileone SKU · one journey · runs: n = [—]
outcome
Retrievedreached
Recommendedreached
Preferrednot reached
integrity
Accurately representedfailed — spec carried over from a previous generation
Qualified for the requestholds
Purchasableholds
reading: recommended — but on an incorrect representation
Illustrative. Outcome stages are sequential; integrity checks are independent and can fail at any stage.

Dimension 1 — Recommendation outcome A genuine sequence. Each stage requires the one before it.

StageQuestionExample failure
RetrievedDid the product appear at all?Absent from every answer in the category
RecommendedWas it affirmatively suggested?Named only as a competitor’s alternative
PreferredDid it get the strongest placement?Third in a list where shoppers read two

Dimension 2 — Commerce integrity Independent checks. Any can fail at any stage.

CheckQuestionExample failure
Accurately representedWas the correct product and variant described correctly?Spec carried over from a previous generation
Qualified for the requestDid it satisfy the shopper’s stated constraints?Ruled out on a budget using a stale price
PurchasableWas there a valid, available route to buy?Routed to a retailer showing out of stock

Read together, these show where a product stopped progressing and what condition was present when it did. A product can be purchasable and never recommended; it can be recommended while described incorrectly.

Every actionable loss is routed to the team best placed to test it

Ranked by what the loss is worth — hero products, in-stock inventory, margin, launch SKUs — not by severity in the abstract.

path 1 / 4

Catalog eligibility

Could the surface retrieve and correctly identify the product?

Product contentPIMFeed
path 2 / 4

Recommendation evidence

Was there enough credible evidence to support it for this use case?

ContentPRProduct marketing
path 3 / 4

Merchant eligibility

Was there a valid, available, trusted place to buy it?

EcommerceChannelMarketplace
path 4 / 4

Preference fit

Did the product actually meet what the shopper asked for?

MerchandisingCategoryProduct

A recommendation isn’t finished until the change has been re-tested

Finding discrepancies is becoming table stakes. Several tools now do it at SKU level. The harder question is which discrepancies are worth anyone’s time, and whether correcting them changes what the assistant does. Verified recovery is our operating standard: every material intervention we recommend includes a test plan, and we report what happened after the change — including when nothing did.

DetectPrioritiseRecommendImplementRe-measurere-measurementfeeds detection
Every material intervention includes a test plan and a re-measurement.

Repeated observation against catalog truth, ranked by commercial materiality, an evidence-backed change with an owner, and the same journeys re-measured under a documented comparison design. The full loop is documented in the methodology.

Recovery reportresult: recovered / no detected recovery / inconclusive
Intervention
[what changed, which fields, which records]
Implemented
[date]
Observed change in recommendation win rate
[—]
Uncertainty range
[—]
Treatment observations
[—]
Comparison observations
[—]
Comparison design
[—]
Alternative explanations considered
model version change · seasonal drift · competitor catalog change · concurrent interventions
Illustrative report structure. Customer results are published only after a documented intervention and repeated post-change measurement.

Every recovery test resolves to one of three states

ResultMeaning
RecoveredThe predefined outcome improved beyond the decision threshold
No detected recoveryThe observed change was too small or too uncertain to distinguish from ordinary variation
InconclusiveThe comparison design, sample size or concurrent changes prevented interpretation

“No detected recovery” and “inconclusive” are different findings, and you need to be able to tell them apart. Collapsing them is how measurement products quietly become marketing products.

How we measure

Repeated observation

Assistants don’t answer the same way twice — we sample, report distributions, and publish the run count.

Truth with a timestamp

A price is only wrong relative to a specific merchant, market and moment.

Stated uncertainty

We publish ranges. A single number would be easier to sell and harder to defend.

Correlation named as correlation

We say “likely recommendation-affecting” until a test says otherwise.

0105101520product aproduct bproduct cproduct dproduct eruns 01–20 · same journey · no underlying change
Illustrative — the same journey, repeated, returns different result sets. This is why we sample, report distributions, and publish run counts.
Read the full methodology →
The offer

AI Shelf Recovery Pilot

A fixed-scope engagement that ends with a tested change, not a report.

  • Five priority SKUs, one market
  • A curated set of high-intent shopping journeys for your category
  • Repeated baseline observations across the AI shopping surfaces agreed during scoping
  • Product claims verified against your catalog truth
  • Ranked diagnostic hypotheses with owners and proposed tests
  • One agreed intervention, prepared for implementation and tested after deployment
  • Post-change re-measurement with a documented comparison design
  • Working sessions with your ecommerce and product-content owners
  • Final recovery report
$15,000six weeks · six pilots this quarter
fully credited against your first year if you continue
Scope a Recovery Pilot

What we do with your catalog

your catalog
margininventorycostcustomer data
confidential commercial fields
public catalog data
provide · secure · support · maintain
model training · cross-customer benchmarks never
aggregate benchmarks — only with explicit opt-in
your analysisthe only standing destination
shared models & benchmarksconfidential data never enters · public data by opt-in only
Retention period, deletion process, model-training policy and subprocessors are stated in full on the data-use page.

We process your product data only to provide, secure, support and maintain your service. We do not use confidential catalog or commercial data — margin, inventory, cost, customer data — to train shared models or create cross-customer benchmarks. Public catalog data enters aggregate benchmarks only with your explicit opt-in. We state our retention period, deletion process, model-training policy and subprocessors in full in our privacy policy.