Does AI Search Actually Recommend Small Shopify Stores? An Open Experiment
Pre-registered September 5, 2026. The hypotheses, design, primary outcome and decision rule below are locked as of this date. Any change to them is appended in the Changes section with a date and a reason, and never overwrites the original. Typos and contact details may be fixed without a log entry.
Why this exists
Merchants on r/shopify and r/ecommerce keep asking the same three questions:
- Is "GEO" (getting your products into ChatGPT or Gemini answers) real, or is it a scam?
- Can I check whether my products show up in AI answers without paying an agency $3,000 for an audit?
- If it is real, how much of it can I do myself?
Every existing answer comes from someone selling GEO services. This is an independent test, run at cost, with all data published, including a null result if that is what we find.
What we already know (and how little that is)
One pilot observation (September 5, 2026, single model, single run): for a very narrow query ("slow feeder bowl for a flat-faced pug with a sensitive stomach"), a web-grounded Gemini session retrieved and read several independent Shopify stores. For the broad version of the same question, it read only Wirecutter, AKC and PetMD.
This proves nothing yet. The narrow query was chosen after seeing which stores ranked in web search. The stores it read may be old, well-reviewed survivors. And the session ran inside an agent tool that does its own site: searches, which a consumer using the Gemini app does not get. We list this so you can see exactly what we are trying to correct for.
Hypotheses (fixed before data collection)
- H1 (existence): For product questions that no authority site (Wirecutter, NYT, AKC, PetMD and similar) has answered, consumer-facing AI search cites or recommends small independent stores at a materially higher rate than for questions authority sites have covered.
- H2 (intervention): Adding one verifiable, question-specific page to a small store's own domain increases that store's inclusion in AI answers for that question, relative to a matched control product in the same store that gets no new page.
- H0 for both: no difference beyond what 5 repeated runs would show by chance.
Design
Participants. 3 to 5 existing Shopify stores, already indexed and already selling. Not a new store. Each store nominates two comparable products: one treatment, one control. Assignment is by coin flip, recorded on this page.
Queries. Pre-registered per store, written before any search is run, in a 2×2 grid:
| Authority site has covered it | No authority coverage | |
|---|---|---|
| Broad | Q1 | Q2 |
| Narrow | Q3 | Q4 |
Authority coverage is checked and logged after the queries are written, never before.
Instruments. Consumer entry points only: the Gemini app with web grounding, and ChatGPT with search. No agent tools, no custom plugins, no site: operators. Fresh session per run, US locale, logged out where possible. 5 independent runs per query per instrument, spread across at least 3 days.
Baseline, then intervention. Two weeks of baseline measurement for all products. Then treatment products get one new page, drafted with the merchant and published on the merchant's domain. Control products get nothing. Measurement continues for four more weeks. Stores start the intervention on staggered dates.
Primary outcome (the only number that decides)
The target product appears in the top 3 of the AI's final answer with a clickable link to the merchant's domain.
Everything else, such as the page being retrieved, cited, mentioned without a link, or ranked 4th, is recorded as a secondary signal and cannot be used to claim success.
Decision rule. H2 is supported only if treatment products' primary-outcome rate rises by at least 20 percentage points over control across 3 or more stores. Anything less is reported as "no detectable effect at this scale."
What we publish
- Every query, every run, timestamped, with the AI's full final answer (screenshots and text)
- The raw data table, downloadable
- Which store did what and when (stores can opt for anonymization; products and pages are described, not named)
- All failed runs, tool errors, and anything we got wrong, appended with dates
What we ask of participating merchants
- Your store URL and two comparable products (one becomes treatment, one control, by coin flip)
- Permission to publish results (anonymized if you prefer)
- For the treatment product: agree to publish one new page on your domain (we draft it, you approve it, you own it)
- One 15-minute call at the start and one at the end
- Optional: Google Search Console read access for the two product URLs
What you get
- A baseline report of where your two products currently stand in AI answers, before anything changes
- The full dataset at the end, and a plain-language read of what it means for your store
- No fee. No upsell. We are not an agency and do not sell GEO services.
To take part, email hello@shelfseen.com with your store URL. Recruitment closes September 12, 2026, or when 5 stores are in.
Who is running this
Derek. One independent operator, not an agency, not selling anything. Contact: hello@shelfseen.com.
Timeline
- September 5, 2026: pre-registration published (this page)
- September 12: merchant recruitment closes (minimum 3 stores, or the experiment pauses and says so here)
- September 26: baseline complete, published
- October 24: intervention period ends, full results published
Changes to this page
None yet. Any change to the hypotheses, design, primary outcome or decision rule after September 5, 2026 appears here with a date and a reason.