product feed article
AI 🤖 eCommerce 🛒

We Rebuilt Our Product Feed for AI Shopping Agents. Here's What Actually Mattered.

Sometime last spring we opened Google Merchant Center for the first time in longer than we'd like to admit, and found ten product listings sitting there. Every one of them had the same little label attached: Click potential: Low.

Ten listings. All low. It's a bit like getting your report card back and every subject is a D.

The thing is, the website was fine. Photography was good, pricing was live, the 3D configurator worked. None of that mattered, because none of that was what Merchant Center was reading. It was reading a table of product data that we'd typed in by hand, once, years earlier, and then never thought about again.

That kicked off about two months of work rebuilding the entire product data layer — for Google, but also for Meta, Pinterest, TikTok, and ChatGPT Shopping. What came out the far end: a generator producing 893 product rows, three feeds that don't agree on which of those rows they want, one design we built most of the way and then threw away, and one Google attribute that got deprecated out from under us six months in.

The parts that went wrong turned out to be the most instructive, so they get the most space.


There is no separate AI channel

One thing is worth establishing before the rest makes sense, because we had it wrong at the start too.

The natural assumption is that getting found by AI shopping agents means doing something new — registering with some fresh surface, separate from the unglamorous product-feed plumbing you've been ignoring for years. That isn't how it works.

Google's Shopping Graph is built out of Merchant Center feeds. It's the structured product layer that grounds shopping answers in AI Mode, in AI Overviews, and in Gemini. Products that aren't in a Merchant Center feed aren't in the Shopping Graph, which means they can't appear in any of those answers at all. Google now ships a Merchant Center report specifically for tracking how your brand surfaces in AI Mode, and its AI-generated ad copy pulls straight from feed attributes like material.

ChatGPT gets there by a different route. Nothing crawls you — you push a structured file to an OpenAI endpoint, as often as every fifteen minutes if your prices move. Different pipe, same principle: a structured product feed is the entry ticket.

So the boring tab-separated file is the thing the agents are reading. There's no AI channel sitting alongside it waiting to be discovered.

That's genuinely good news, because it makes the work finite and legible instead of mystical. It's also why an article about AI shopping agents spends so much of its time on feed mechanics. That's where the agents are.


First, the autopsy

Ten hand-typed listings, and the problems were all the same species.

The titles were near-identical: "10 Custom Printed Mailer Boxes," "100 Custom Printed Mailer Boxes," "1000 Custom Printed Mailer Boxes." They varied by quantity — the one attribute nobody on earth searches by — while size and material, which is how people actually look for boxes, appeared nowhere at all.

A couple of listings had stock icons instead of real product photos. The "last update" timestamps ranged from November 2023 to September 2025, because there was no sync to live pricing; the numbers were whatever someone had typed on the day. And because every quantity of the same box was its own separate listing with no grouping, our own products were bidding against each other and splitting their own signal.

None of that was a judgment on the boxes themselves. The data was simply too thin to match against anything a shopper might plausibly ask for.

If you've got a Merchant Center account you set up once and haven't opened since, go open it. The diagnostics tab is the most honest performance review your product data is ever going to get, and it takes about ninety seconds.


What a machine needs from you

Everything downstream got easier once we started treating the feed as an API rather than a brochure. Working out what that API has to serve took longer.

Three things, roughly. Catalog facts — what you sell, in a form that can be filtered, sorted, and set against a competitor. Trust signals — some reason to believe those facts, ideally originating with someone other than you. And provenance — a path back to a source, so the first two can be checked rather than taken on faith.

Most brands have built a partial version of the first and nothing at all of the other two. We were no exception. What follows is roughly the order we filled them in, mistakes included.

A shopping agent — Google's Shopping surface, ChatGPT, Perplexity, whatever comes next — doesn't experience your catalog the way a visitor does. It ingests rows and columns. Each row is a product, each column is something it can filter, sort, and compare on. Anything that isn't in a column may as well not exist.

Which means narrative selling just evaporates in the transition. "Premium quality, fast turnaround, trusted by thousands" is a string in a description field. You can't filter on it. It has no effect whatsoever on whether you get surfaced. What gets you surfaced is a clean set of checkable facts — exact dimensions, material, unit price, minimum order quantity, availability — because those are the terms a shopper's constraints are written in.

So we went from ten hand-typed listings to 531 generated rows, one for every real combination of box type, dimensions, and material we can actually manufacture. (It's grown to 893 since, as other kinds of row got added — more on that shortly.) They come out of the same pricing engine that powers our live quotes, which means a row can't exist unless it describes something you could genuinely order, at that price, today.

That last bit ended up mattering more than the row count. A sprawling feed full of combinations you can't actually make is worse than a small accurate one — it turns a discovery win into a customer service problem, and you've paid for the click.


The mistake: we almost shipped pack prices

Here's the part we're glad we caught.

The first design enumerated products by quantity tier. Every size, in every material, at every quantity we offer — 10-packs, 100-packs, 500-packs, 1000-packs, all the way up. It came out to 5,811 products. It felt thorough. It mirrored how our own quantity slider works, so it seemed obviously right.

It would have been a disaster.

The problem is what ends up in the price column. Under a pack-total model, a 250-pack of custom mailers lists at something like $312.50. Meanwhile every competitor showing up for "custom mailer boxes" is listing per-unit: $0.50, $0.89, $1.25.

Two things happen, and either one on its own is enough to sink you.

Price filters kill you first. A shopper narrowing to "under $2" never sees a $312.50 listing — it's excluded before ranking is even a consideration. Then, in the results where you do appear, a $312.50 entry sitting in a column of sub-dollar prices reads as insane. Nobody clicks it. Ranking systems interpret that as irrelevance and push you down further, which gets you fewer impressions, which gets you fewer clicks. It compounds.

You're being compared on a number that means something completely different from everyone else's number, and the agent has no way to know that. It sees one column.

So on June 10th we scrapped it and rebuilt around per-unit pricing. The catalog went from 5,811 items down to 531, and got substantially more competitive in the process. We'd built roughly five thousand products that would have actively hurt us.

That's the counterintuitive lesson, and it's the one I'd most want another brand to take: a smaller feed with the right shape beats a bigger one with the wrong shape, and it isn't close.

If you sell in packs, cases, pallets, or anything other than "one item" — go look at what number is currently in your price column. This is far and away the most common way wholesale and B2B sellers make themselves invisible to agentic search, and it's usually a same-week fix.


Which unit price, though?

Switching to per-unit sounds simple until you have to pick a number, because a custom box doesn't have one price. It has a curve. Ten mailers cost a certain amount each; five hundred cost considerably less each; two thousand less again.

You can only put one number in the price column.

Quote the price at your minimum order quantity and you look expensive, because that's the worst point on your own curve. Quote the price at your largest tier and you look great right up until the shopper lands on the page and discovers they'd need to buy two thousand boxes to get it — which is a bait-and-switch, and it'll show up in your bounce rate before it shows up in your conscience.

We settled on a showcase quantity of 500 units. It's a real order size that real customers place, it's roughly where the curve starts to flatten, and it's low enough to stay comparable against competitors quoting their own volume rates. That unit price goes in price.

Then the true minimum goes in min_order_quantity as its own field — 10 units for mailers and shipping boxes, 500 for rigid and folding. It's disclosed separately, in a machine-readable slot, so any agent filtering on "minimum under 50" finds us correctly and nobody discovers the real minimum at checkout.

The last piece closes the loop. Every feed link deep-links into the configurator with the size, material, and quantity already set — and specifically sets quantity to 500, the same showcase number. So the unit price you see on arrival is the exact number that pulled you there. Not approximately. Exactly. A price mismatch between your feed and your landing page is one of the fastest ways to collect item-level disapprovals, and the only reliable fix is to make the two structurally incapable of disagreeing.

There's one deliberate distortion in there worth mentioning. A single mailer sample costs $35, and at very small dimensions the computed per-unit price at low volumes could theoretically exceed that. So the resolver caps any mailer unit price at the sample price. Nobody should ever see a per-box number higher than what one box costs.

The shape of that decision transfers to anything with a conditional price. Free shipping over a threshold, subscription rates versus one-time, tiered service plans — you have one headline field and a condition attached to it. Put the number that makes you genuinely comparable in the price field, and put the condition in whatever machine-readable field the schema gives you for it. Burying the condition in prose and hoping nobody notices works on a landing page. It stops working the moment something is filtering.


The bug you won't notice for six months

Your feed says $0.62. Your product page says $0.68. Someone changed a pricing rule, updated the site, and forgot the feed exists.

This escalates faster than you'd think. Google issues item-level disapprovals for price mismatch. Merchant Center trust degrades quietly. And in an agent-mediated purchase, an agent has now quoted a customer a price your checkout won't honor — which stops being a data problem and becomes a broken promise.

The obvious fix is a reconciliation job that compares the two and alerts on drift. We think that's treating the symptom. It leaves the actual disease in place: two sources of truth that merely happen to agree right now.

We made the parity structural instead. Three surfaces need a price — the merchant feeds, the structured data injected into each variant page, and the live pricing API the configurator calls. All three go through one shared resolver. There's no code path where they can drift apart, because there's no second calculation available to drift from. Same discipline on product IDs: the feed and the structured data pull identifiers from a single shared module, so a product has the same identity everywhere it shows up.

The question to ask about your own setup is pretty simple: if you changed a price in one place, could somewhere else go stale? If yes, you're relying on a coincidence with an expiry date.

The rule: put the number that makes you comparable in price, and put every condition attached to it in its own machine-readable field. Anything living only in prose is invisible to a filter, and any price computed in two places will eventually disagree.


Channel schemas are not stable ground

Six months after we launched all this, part of it stopped working.

Google's bulk_price attribute was what made the per-unit model fully honest. It let you publish the actual volume ladder — 500 units at this price, 1,000 at that, 2,000 at that — as structured, filterable data sitting next to the headline price. We built nine tier columns to fit the widest ladder in our catalog.

On June 11th, Merchant Center started reporting it as an unrecognized attribute. Deprecated, with no particular fanfare. min_order_quantity survived; the tiers didn't.

The volume ladder moved into the description text, where every product now walks through its pricing in plain sentences. The information survived; it just changed shape. We kept computing the tier data anyway — the generator still builds every field, those values feed the descriptions and the ChatGPT Shopping variant, and they're ready for whatever channel supports tiered pricing next.

That's the specific incident. The general version will happen to you, because every channel you publish to is somebody else's product roadmap and none of them owe you stability. Attributes get deprecated. Required fields appear. A format you depend on turns out to mean something different on the platform you added last quarter.

What decides whether that's a bad afternoon or a bad month is whether your generator emits one platform's format natively. If it does, every schema change is a rewrite. If it builds a neutral internal representation of a product and serializes that per channel, a deprecation means deleting a column from one serializer while everything else carries on untouched.

Worth putting that indirection in before you think you need it. We happened to have it because we were publishing to three channels from the start, which made the neutral layer obvious. Had we built Google-first and bolted the others on afterward, June 11th would have cost us a week.

There's a quieter point here too. When the structured field went away, prose absorbed the data — and prose turned out to be a perfectly good container, because the systems reading it can parse sentences now. That's genuinely new, and it changes what a description field is for.


Three feeds, because one wasn't enough

The instinct is to build a single feed and point everyone at it. That works right up until it doesn't.

We ended up publishing three serializations of the same underlying catalog:

google-merchant.txt — read by Google Shopping and Bing. The full Google attribute set, repeated columns and all.

products-universal.txt — read by Meta, Pinterest and TikTok. A conservative subset, because those platforms reject or mangle Google-specific structures.

openai-products.txt — read by ChatGPT Shopping. A different schema and a different writing register, which gets its own section below.

The split comes down to a concrete format disagreement.

Google's tab-delimited format lets you repeat a column header to express a multi-value attribute. You want five product images? Five separate columns, all headed additional_image_link. Six selling points? Six columns headed product_highlight. It's an odd-looking file, and easier to show than describe — here's the real header row of ours, with the tabs broken onto separate lines:

id
title
description
link
image_link
additional_image_link     <- same name
additional_image_link     <- five times
additional_image_link
additional_image_link
additional_image_link
availability
price
min_order_quantity
...

That's how Google expresses a list in a flat format, and it works fine.

Meta, Pinterest, and TikTok don't do that. Depending on the platform, repeated headers get rejected outright or silently collapsed, which is worse, because you don't find out until you notice half your images are missing. Those platforms want a single column with comma-separated values.

There's no clever encoding that satisfies both. So the universal feed joins the lists into single columns and drops the Google-only wholesale attributes entirely, stating the minimum order in the description text instead. Same products, same underlying data, two different shapes.

The product highlights are worth a note of their own, since they're the one place in a Merchant Center listing where you get to make an argument. Google renders them as bullets in the Shopping product panel and accepts up to ten. We emit six: one row-specific lead bullet, plus five standing differentiators — that you design it yourself in a 3D editor, that there are no setup fees or plate charges, full-color digital printing edge to edge, curbside-recyclable board, printed and shipped in the USA.

Those five were chosen against the competition rather than in a vacuum. We sit in Google product category 973, "Moving & Shipping Boxes," which is dominated by plain brown commodity boxes. Every bullet is a thing those listings can't say.

That move is worth stealing whatever you sell. Look at what the listings sitting beside you in your category are unable to claim, and spend your bullets there. Highlights are one of the few slots in a machine-read listing where you get to make an argument rather than state a fact, and most merchants either leave them empty or fill them with things every competitor could say word for word.

The wider point is to budget for translation, not just publication. Adding a sales channel is rarely a matter of pointing it at the feed you already have. Each one is a dialect with its own opinions about how a list gets expressed and which attributes it will tolerate, and you find out which by having something break.

The rule: build a neutral internal representation of a product and serialize it per channel. A generator that emits one platform's format natively turns every deprecation into a rewrite.


Same catalog, different guest lists

Format turned out to be the easy half. The harder question was which products belong on which channel at all.

For the first stretch, everything went everywhere — all 893 rows to Google, to the universal feed, to ChatGPT. Then we looked at what they were doing.

Between July 7th and August 3rd, the specific-size rows — the 528 that name exact dimensions, the ones we'd been proudest of — took about 22% of reported impressions and returned 12% of the clicks. A 1.06% click-through rate against 2.24% for the head-term listings. Worse, 596 of 799 products fell below Merchant Center's roughly 24-impression reporting floor altogether, meaning they were served almost nothing and converted nothing.

They weren't sitting there harmlessly either. On a channel that ranks by engagement, a big tail of listings nobody clicks drags on the whole account's quality signal. You're patiently teaching the system that your listings don't get clicked.

So they came out of Google, and out of Meta, Pinterest and TikTok. What's left there is 365 rows — the quantity products, which cover size-qualified searches with a correct per-unit price anyway, plus the head-term heroes and the samples.

They stayed in the ChatGPT feed, though, and that's the more interesting half of the decision.

Retrieval doesn't behave like an auction. Nothing on OpenAI's side is ranking our rows against each other by click-through, so a long tail costs approximately nothing to carry — and it buys real coverage, because the size rows are the only ones carrying structured length, width and height values. That's 187 distinct dimension sets in the ChatGPT feed against 19 on Google. If somebody asks an agent for a 14x10x6 mailer, the row that answers them exists in exactly one of our feeds.

So the three feeds disagree twice over. They disagree about format, and they disagree about which products should be in them at all.

The rule: on a channel ranked by engagement, a long tail nobody clicks is a liability rather than dead weight. On a channel that retrieves, breadth is close to free. Work out which kind you're publishing to before deciding how much catalog to send.


The ChatGPT feed asks different questions

The third serialization is the one most brands haven't run into yet, and it's the most interesting of the three, because OpenAI's Agentic Commerce schema asks for genuinely different information.

Some of the divergence is cosmetic. item_id rather than id, url rather than link, group_id rather than item_group_id. Renaming, essentially.

The substantive differences are more revealing. Here's a real row from ours, trimmed to the fields Google has no equivalent for:

is_eligible_search      true
is_eligible_checkout    false
item_id                 mailer-box-10x8x4-kraft-board
price                   2.95 USD
length                  10
width                   8
height                  4
dimensions_unit         in
material                Kraft Board
group_id                mailer-box-10x8x4
listing_has_variations  true
seller_name             Packwire
seller_url              https://packwire.com
target_countries        US

Three of those are worth dwelling on.

It asks for explicit eligibility flagsis_eligible_search and is_eligible_checkout, as two separate booleans. That distinction doesn't exist in Google's format, and it's a thoughtful one: it lets a merchant be discoverable without yet being transactable. Ours ship with search true and checkout false, which is an honest description of where we are. A custom box gets designed in a configurator, so there's no meaningful "buy now" for an agent to execute yet.

It asks for dimensions as real numeric fields — separate length, width, height, and dimensions_unit columns — rather than accepting them baked into a title string. For a business that sells boxes this is a gift, because the single most important attribute of our product finally has a machine-readable home instead of living inside prose that has to be parsed back out.

And it asks for seller_name and seller_url, which tells you something about the purchase model it anticipates. In an agent-mediated transaction the shopper may never land on your domain at all, so provenance has to travel attached to the data rather than being implied by the page it sits on.

Then there's the description field, which is where the two schemas diverge most and where the lesson generalizes furthest.

Google's rewards keyword-dense retail copy, because it's feeding a matching system. OpenAI's rewards conversational copy, because the description is likely to be summarized, paraphrased, or read aloud by a model answering somebody's question.

So ours read like answers rather than ads. Each one gives the material and exact dimensions, explains that you design the box yourself in a 3D configurator, and then walks through the full price ladder in plain sentences — what a unit costs at the showcase quantity, what it costs at the minimum order, what it drops to at the largest tier.

Putting the ladder in prose was deliberate. If an agent is answering "what's the cheapest way to get 2,000 custom mailers," we want that answer sitting in a sentence it can lift directly, rather than locked in a structured field it may or may not have parsed.

Which is the same lesson the bulk_price deprecation taught from the other direction. Write for the model that's going to paraphrase you, not just the crawler that's going to index you. Those are different readers with different needs, and increasingly it's the paraphraser your customer actually hears.


The image problem nobody warns you about

Every product needs an image URL. With 893 generated rows and a photo library of maybe thirty lifestyle shots, you're assigning images programmatically whether you planned to or not.

The naive approach is to give every mailer box the same photo. It works, and it's dull, and it means every listing you have looks identical in a grid of search results.

The next idea is to assign randomly for variety. This one is actively harmful, and the reason is non-obvious: feeds get regenerated. Ours regenerate whenever pricing changes. If image assignment is random, every regeneration reshuffles the deck, and Merchant Center sees hundreds of products whose images just changed. It re-crawls and re-processes all of them, and you've manufactured a pile of churn that buys you nothing.

So assignment is deterministic. Each product's ID gets hashed, and the hash picks its photo out of the pool for that box type. Same product, same photo, forever — but different products get different photos, so the catalog has variety. Regenerate the feed as often as you like and nothing appears to change, because nothing has.

Each listing carries six images total: one primary and five additional, rotating through the pool from the product's own hash offset. Six because Merchant Center's own scorecard puts five to eight images per offer in its top bucket, and there's no reason to leave that on the table when the photos already exist.

We also had no idea which kind of photo should lead — the clean hero shot, or a lifestyle image that's more arresting in a grid — so we split the catalog on hash parity and tagged each half in custom_label_2. That test is still running and we don't have a verdict yet.

The tagging is the part worth copying even if you never run that particular experiment. Custom labels are free-form segmentation slots that most merchants leave completely empty, and anything you write into one becomes a dimension you can break a standard Merchant Center report out by. Ours carry the row class, a quantity cohort, and the image cohort. The practical effect is that an experiment costs no tooling at all — you tag the cohort at generation time and read the result out of a report you already have. Experiments that cheap are experiments that actually get run.


Trust signals: the reviews file we wrote for machines

Everything above is layer one — catalog facts, and the mechanics of making them legible. An agent comparing three suppliers has a second question, and it's the harder one: why should it believe any of you.

Product data is self-reported. Every merchant in that comparison says their boxes are good and their turnaround is fast. The tiebreaker is evidence that didn't come from the merchant — which is to say, from customers.

So alongside the HTML reviews page we publish a plain-text file at /reviews.txt. The two do different jobs on purpose. The HTML page behaves like every testimonial page ever built and leads with the five-star ones. The text file takes the opposite approach: every Google review on record, newest first, with dates and ratings, linking back to Google so any of it can be checked at source.

That includes the unflattering ones. Of the 101 reviews in the file, three are one-star, and there's a two, a three and a four in there as well. The reasoning is written into the code: transparency reads better to a machine cross-checking the aggregate rating. If you claim 4.8 out of 5 and publish only glowing reviews, a system evaluating that claim has nothing to reconcile it against, and an unverifiable claim is worth roughly nothing. Publish the full distribution — including the customer who had a genuinely bad experience — and the 4.8 becomes checkable. It stops being a marketing number and becomes a fact.

There's an instinct to resist this, and it's worth naming. Curating your best reviews for a landing page is normal, and every brand does it. But an agent arrives with a different job. It's assessing a claim, and it has other sources to check it against. The failure mode is subtler than losing a sale to a bad review: an unsubstantiated aggregate gets quietly discounted, and you never find out.

The same logic drives /llms.txt, which is a plain-text summary of the whole business — products, materials, pricing structure, deep links, and pointers to the feeds and the reviews. It carries the top-line rating, and then explicitly instructs any model reading it to cite from the full reviews file rather than from the summary. It's a strange sentence to write. You're addressing a reader who isn't a person, telling it where the better data lives.

The rule: an aggregate rating is worth only what a machine can reconcile it against. Publish the distribution, link the source, and let the number be checkable.


The UGC question we haven't solved

Which brings up the obvious gap.

Reviews are the structured, tractable end of customer-generated evidence. The rest of it — unboxing videos, customer photos, the Instagram post where someone's product looks genuinely great in a box we printed — is enormously persuasive to humans and almost entirely illegible to machines.

There's no ugc_link attribute in any of the three feed schemas we publish to. Nothing in Merchant Center's spec, nothing in OpenAI's Agentic Commerce fields, nothing in the universal format. An agent comparing packaging suppliers has no channel through which to receive "here are four hundred customers who posted photos of this product."

We haven't solved that, and we're not going to pretend otherwise. What we can see is the shape of the problem. The evidence is real and it's genuinely differentiating, and right now the pipe to carry it doesn't exist — so it converts to zero in exactly the comparisons that are growing fastest.

Some of it can be dragged into machine-readable territory today. Review platforms that emit Review and AggregateRating structured data get their content parsed. Testimonials with named attribution on a crawlable page are legible. Case studies with real numbers survive summarization, where a wall of Instagram embeds does not.

The general principle, as far as we can tell: social proof only counts to the extent it exists as text a machine can attribute to a person. An Instagram grid is invisible to that process. A quoted customer with a name, a date, and a source that can be checked survives it.

If you're investing in UGC right now — and you probably should be, because humans still buy things — it's worth asking what fraction of it survives translation into a format an agent can read. For most brands the honest answer is nearly none, and that's a strategic problem that's going to get more expensive to ignore.


What comes after feeds: MCP

Back to that is_eligible_checkout: false flag in the ChatGPT feed. It sits at the honest edge of what a feed can do, and it points squarely at what comes next.

A feed is a snapshot. You generate it, you push it, and it describes what was true at generation time. Ours carries 893 rows at its widest because that's how many combinations we chose to enumerate — but the real product space is effectively unbounded. Any dimensions, any material, any quantity, any artwork. Someone who wants 350 boxes at 9x7x3 in kraft is asking about a product that genuinely exists, that we can price in milliseconds, and that appears nowhere in the feed.

The Model Context Protocol is the shape of the answer to that. It's an open standard Anthropic published in late 2024, now supported by Google, Microsoft, OpenAI and Shopify, and the analogy everyone reaches for is USB-C — one connector instead of a custom cable for every pairing. Where a feed is a file you publish, an MCP server is an interface you operate and an agent calls. Rather than reading a row, the agent asks a question and gets an answer computed on the spot.

The division of labor looks roughly like this:

  agent needs to FIND you
            |
            v
  +----------------------+
  | FEED                 |
  | a file you publish   |
  | fixed rows           |
  | refreshed on a clock |
  | "what we listed"     |
  +----------------------+
            |
  agent needs a SPECIFIC answer
            |
            v
  +----------------------+
  | MCP SERVER           |
  | an interface you run |
  | any configuration    |
  | computed on demand   |
  | "what we can make"   |
  +----------------------+

For us that's uncomfortably close to shipping, because the thing such a server would call already exists. Our pricing API answers exactly that question for any valid configuration, and it's the same resolver behind the feed and the landing pages. Exposing it as a callable tool would mostly be wrapping something we already run.

The general version is worth thinking about whatever you sell, because the gap between what your feed says and what you could answer if someone asked is usually enormous:

  • Configurable or made-to-order products — live quoting for combinations no feed could ever enumerate. Every configurator business has this problem and most haven't noticed it yet.
  • Inventory-heavy retail — real stock and a real delivery date for a specific address, instead of an availability column that was accurate at last sync.
  • Fitment-driven catalogs — auto parts, filters, cartridges. "Will this fit my 2019 model" is a lookup, and a list of filterable attributes approximates it badly.
  • Services and appointments — actual availability windows rather than a price range.
  • Subscriptions — comparing plans against the usage a customer just described, which is arithmetic no static listing can do.

In each case the feed carries the general shape and the server answers the specific question.

So why haven't we built one? A few honest reasons.

The client side isn't there yet for shopping. Product discovery in ChatGPT still runs on the pushed feed, and Google's AI surfaces still run on the Shopping Graph built from Merchant Center. Standing up a server that nothing currently calls is an excellent way to feel productive.

That's changing quickly, though. In March 2026 Shopify and Google announced the Universal Commerce Protocol, built on MCP, aimed at letting agents complete real purchases inside Google Search, the Gemini app and Copilot. Shopify now ships MCP servers to every store on the platform, switched on by default — so if you're on Shopify, you have some of this already and may not know it. We run a custom stack, which means building and operating it ourselves against currently negligible agent traffic.

And checkout is genuinely harder for us than for someone selling t-shirts. A custom box needs artwork, a proof, and an approval before anything gets printed. There's no coherent one-click purchase for an agent to execute, which is precisely why that eligibility flag reads the way it does. The quoting half we could do fairly quickly; the transacting half needs the product to change, not just the plumbing.

For most merchants the reasonable position is that this is an awareness item now and an integration item within a year or two. Worth understanding what it is, worth watching what your platform gives you for free, and worth noticing that competitors on Shopify had it switched on without asking.

One thing it won't do is replace the feed. An agent has to find you before it has any reason to call you, and discovery still starts with the boring file. MCP is depth. The feed is the front door.


What to actually do about it

Roughly in order of return on effort:

Open Merchant Center and read the diagnostics. "Click potential: Low" is a data problem with a data fix, and it's almost always the cheapest win sitting in front of you. It's also your AI Mode visibility, since the Shopping Graph is built from that same feed.

Submit a feed to ChatGPT Shopping. It's a separate push to a separate endpoint with its own schema, and right now the merchants who've bothered are competing against far fewer rivals than they will be in a year.

Check your price column. If you sell in anything other than single units, make sure you're publishing per-unit pricing with min_order_quantity next to it, rather than a pack total. This one change can move you from filtered-out to comparable.

Get searchable attributes into your titles — size, material, dimensions, the things a constraint actually gets expressed in. Quantity belongs in min_order_quantity where it can be filtered on. "Custom Kraft Mailer Box – 10x8x4in" describes a findable object; "100 Custom Boxes" describes almost nothing.

Build a neutral internal representation and serialize per channel. Attributes get deprecated. Formats disagree. If your generator emits one platform's exact format natively, every schema change is a rewrite.

Make image assignment deterministic. If regenerating your feed reshuffles images, you're generating pointless re-crawls every time you touch pricing.

Use your custom labels. They're free segmentation, they cost nothing to populate, and an experiment you can read straight out of a standard report is an experiment you'll actually run.

Publish your reviews in full, in plain text, including the bad ones. An aggregate rating nobody can check against anything is just a number on a page.

Generate the feed, don't type it. Anything maintained by hand goes stale, and stale data in an agentic channel becomes wrong information delivered confidently.


The part that takes getting used to

A growing share of purchase decisions are being made by something that will never see your brand.

It won't notice the photography or be charmed by the copy. It reads a table, applies constraints, returns what fits.

That's bad news if your edge was presentation. It's quite good news if your edge is the product, because a comparison stripped down to checkable facts rewards whoever's genuinely winning on the facts. Low minimums, honest per-unit pricing, no setup fees, accurate dimensions — those win in that environment precisely because the environment ignores everything else.

The brands that struggle here will mostly have perfectly good products. They'll just have described them in a way no machine could verify.

Worth walking through your own data the way a customer would. Or, more to the point, the way a parser would.


If you want to see the real numbers on your own box, the 3D configurator shows live per-unit pricing at every quantity — no quote form, no sales call. And if you're not sure what size you need yet, the Box Size Optimizer works it out from your product dimensions.