Blog

Monitoring

How to Build an AI Shopping Query Set

Build a repeatable AI shopping query set across discovery, use case, budget, attributes, comparison, availability, and purchase intent.

A general AI search query set can tell you whether a brand is mentioned or cited for buyer Questions. Shopping introduces more observable layers.

An answer can name a brand without displaying a product. A product carousel can show an item without linking to the manufacturer. A product detail can list several merchants. A source citation can support a specification without producing a product result. A labeled ad can appear beside an otherwise independent answer. An eligible merchant may offer checkout inside ChatGPT while another sends the user to its site.

If all of those events become one field called “AI visibility,” the resulting trend cannot explain what changed.

Short answer: Build an AI shopping query set as a versioned measurement contract. Define the catalog and market, classify Questions across seven declared shopping intents, freeze exact wording and relevant context, and run each Question on a named product surface under a documented retry policy. Code product appearance, brand mention, merchant appearance, displayed citation, labeled ad, and checkout separately. Preserve raw evidence and completed denominators. Thirty Questions repeated three times are a useful operational starting point, not an OpenAI recommendation, statistical minimum, or guarantee of representative results.

This guide extends the general AI search query set workflow for ecommerce teams that need product-level evidence.

Shopping Questions Need Their Own Measurement Contract

A normal brand-monitoring Question might ask:

Which project management tools are suitable for a small agency?

A shopping Question can add product attributes, price, inventory, seller choice, delivery, visual similarity, and a transaction path:

Find a quiet cordless vacuum under $400 that works on pet hair and is available from a merchant that delivers to Boston this week.

The second Question creates several possible observations:

  • Was a specific product displayed?
  • Was its governed brand named accurately?
  • Which merchant offers were shown?
  • Were price and availability stated?
  • Was a supporting source displayed?
  • Was any result labeled as an ad?
  • Did the result provide an external purchase link or an in-chat checkout path?

Those fields have different owners and different denominators. Product data may belong to merchandising. Brand identity may belong to marketing. Merchant price and availability may change independently. Citation quality may require editorial review. Ads and organic product results belong in separate reporting lanes.

Use the ChatGPT shopping visibility guide for the surface-level distinction. Use this article to build the Question set that measures those distinctions consistently.

What OpenAI Currently Documents

The product surface is evolving, so record the documentation review date beside the baseline. The summary below reflects OpenAI material reviewed on August 9, 2026.

OpenAI's Shopping with ChatGPT Search help page says a shopping-intent question can produce product options with imagery, details, and links to learn more or purchase. It also says some eligible products and merchants may expose an Instant Checkout option.

The same page makes three boundaries important for measurement:

  • Product results are selected independently and are not ads.
  • A product can appear when ChatGPT perceives it as relevant to the user's intent and context.
  • That context can include Memory or Custom Instructions.

OpenAI's product discovery announcement describes visual browsing, side-by-side comparisons, conversational refinement, and product details such as price, reviews, and features. Its shopping research documentation describes a deeper interactive flow for comparisons, trade-offs, and multiple constraints, including follow-up questions and real-time preference feedback.

These are documented product behaviors, not a promise that every account, Question, category, country, or Run will expose the same format.

OpenAI also says product results and shopping research are separate from ads. Its current product-feed advertising guide states that products uploaded to Ads Manager during the described beta are eligible for ads and do not enter organic ChatGPT conversations through that feed. Therefore, code a labeled ad as an ad observation. Do not relabel it as organic product visibility.

The ChatGPT ads-versus-organic guide provides the wider reporting boundary.

The Seven Shopping Intents

The following seven classes are an AEO Table operational framework. They are not official OpenAI query categories, ranking factors, or reporting dimensions.

Their purpose is to stop a team from filling its baseline with one kind of Question, such as branded product lookups, and then presenting the result as complete shopping visibility.

Intent ClassBuyer NeedExample QuestionMain Evidence
DiscoveryExplore a category without a fixed product“What kind of air purifier works for a studio apartment with a cat?”Products and brands entering consideration
Use caseSolve a defined job or situation“Find a carry-on backpack for weekly business travel and a 16-inch laptop”Fit reasons, constraints, and exclusions
BudgetStay within a price or value boundary“Which espresso machines under $600 have an integrated grinder?”Price qualification and merchant options
AttributesMatch specifications, compatibility, or materials“Show waterproof hiking shoes in wide sizes with a rock plate”Product facts and source support
ComparisonEvaluate named or discovered alternatives“Compare these three robot vacuums for pet hair and obstacle avoidance”Product differences, trade-offs, and citations
AvailabilityFind a valid offer in a place and time“Where can I buy this monitor in Singapore with delivery this week?”Merchant, location, price, stock, and handoff
PurchaseComplete or approach a transaction“Can I buy this model now, and what are the return options?”Merchant selection, external link, or checkout

One Question can express more than one intent. Preserve one primary class for reporting and optional secondary classes for analysis. Do not change the primary class after seeing which result appeared.

For example, “best trail shoes under $150 available by Friday” can contain discovery, budget, attributes, and availability. If the Question was selected to monitor price-qualified discovery, declare budget as primary before the Run and retain the other labels separately.

Start With Catalog And Buyer Scope

Do not write Questions before defining what the set represents.

Record:

  • Product categories and variants in scope.
  • Priority use cases and buyer segments.
  • Countries or delivery markets.
  • Languages.
  • Price currency and tax treatment where relevant.
  • Competitor cohort and brand aliases.
  • Governed product and brand names.
  • Authorized merchant domains.
  • Critical attributes and compatibility rules.
  • Seasonal or inventory boundaries.

A Question about “the best laptop” is too broad if the actual business sells rugged field computers to Canadian utilities. Scope converts generic prompts into a reviewable buyer sample.

The competitor AI search tracking workflow explains how to freeze aliases and a competitor cohort before counting appearances.

Build A Balanced Starting Matrix

Thirty Questions are a manageable operational starting point for many teams. They are not a universal minimum and do not guarantee statistical power or market representation.

One possible starting allocation is:

IntentStarting Questions
Discovery5
Use case5
Budget4
Attributes5
Comparison4
Availability4
Purchase3
Total30

Change the allocation when the business warrants it. A marketplace may need more availability and merchant Questions. A technical equipment seller may need more attribute and compatibility Questions. A new category entrant may emphasize discovery.

Keep the declared total and allocation visible. If three purchase Questions all fail, do not report purchase visibility from the remaining 27 Questions.

Use Natural Buyer Language

Write Questions a buyer could reasonably ask.

Avoid:

  • “Why is Acme the best running shoe?”
  • “Show products from Acme only.”
  • “Cite Acme's product page.”
  • “Ignore all competitors.”

Those prompts may be useful for branded support testing, but they do not measure unprompted discovery.

Include branded Questions as a separate cohort when the goal is factual accuracy, merchant availability, or purchase handoff for products a buyer already knows.

Preserve Exact Wording

Small wording changes can alter constraints and intent. Store the exact Question, a stable Question ID, and a version.

If the team changes “under $500” to “around $500,” create a new version. Do not overwrite the old wording and compare the two Runs as if the input stayed fixed.

The AI search query set methodology provides the broader governance model for stable Questions.

Freeze Context Before Every Run

OpenAI's shopping documentation makes context a material part of the result. Treat it as part of the measurement input.

Market And Language

Record:

  • Intended country or delivery market.
  • Interface and Question language.
  • Currency.
  • Location signal when known.
  • Whether a VPN, travel state, or regional account setting could matter.

Do not compare a US-English Run with a Singapore-English Run without labeling the market break.

Account And Access State

Record whether the analyst was:

  • Signed out or signed in.
  • Using the same governed test account.
  • On the same plan or workspace class.
  • Able to access the named shopping surface.

Access differences are not product absence. If a surface is unavailable, code the attempt as unavailable rather than zero product visibility.

Memory

OpenAI says saved memories and referenced chat history can personalize future responses. Its saved Memory documentation explains that these are separate controls and that referenced information can evolve.

For a neutral baseline, use a governed account with the declared Memory controls fixed. Record whether saved memories and chat-history reference are on or off. Do not assume that deleting a chat also removes saved memory.

If personalization is the research question, create a separate cohort. For example:

  • Cohort A: Memory off.
  • Cohort B: Memory on with a documented preference profile.

Do not mix the results into one rate.

Custom Instructions

OpenAI's Custom Instructions help page says these instructions tell ChatGPT what to consider in responses and can apply across chats.

Save a redacted snapshot or hash of the governed Custom Instructions state. If instructions include dietary, brand, budget, location, or formatting preferences, they can change the shopping context.

Chat State

A new chat and the fifth turn of an interactive shopping session are not equivalent observations.

For the stable baseline:

  • Start each Question in a new chat.
  • Do not provide follow-up feedback before coding the first result.
  • Record any automatic clarification request.
  • Use a predeclared response policy for clarifications.

For journey testing, create a separate multi-turn cohort and preserve every turn. OpenAI's shopping research material describes follow-up questions and live preference feedback as part of the experience, so conversational refinement should be measured deliberately rather than leaking into a single-turn baseline.

Product Surface

Name the observed surface instead of writing only “ChatGPT.”

Examples include:

  • A regular response with shopping product results.
  • Shopping research.
  • A product detail and merchant list.
  • An eligible checkout path.
  • A labeled ad unit.

Do not merge them unless the analysis explicitly defines a cross-surface summary.

Time Window

Price, availability, merchant ordering, interfaces, and inventory can change. Run matched observations in a declared window and record the timezone.

OpenAI warns that merchant price or shipping updates can take time to appear in product information. A later difference may reflect fresher commerce data rather than a content intervention.

Retry Policy

Set the policy before running:

  • Maximum attempts.
  • Retryable errors.
  • Wait rule.
  • Whether a clarification counts as completed.
  • Which result is retained.
  • How refusals and unavailable surfaces are coded.

Do not keep retrying until the target brand appears. Do not count every retry as an independent shopper.

Use One Experimental Unit

Use:

Question × named shopping surface × Run

Each scheduled unit should end in one governed status:

  • Completed and eligible.
  • Completed but ineligible for the named metric.
  • Clarification required.
  • Failed.
  • Refused.
  • Surface unavailable.

Three repeats per Question are a practical starting point for estimating obvious answer variation. They are not an OpenAI rule, a confidence guarantee, or a substitute for a sample-size design.

If the initial set has 30 Questions and three repeats, the schedule contains 90 planned units. Report all 90 statuses. A denominator of only the best completed results hides execution quality.

The AI search volatility guide explains why repeated Runs and visible completed denominators matter.

Code Six Signals Separately

Do not store one visible field.

SignalCoding QuestionWhat It Does Not Prove
ProductDid the named product or governed variant appear in a product result?Brand recommendation, merchant availability, or sale
BrandDid answer text or product UI name the governed brand?Owned citation or product-level visibility
MerchantWas a merchant option shown for the product?Inventory accuracy, lowest price, or conversion
CitationWas a source displayed, and what claim did it appear to support?That every answer claim is supported
AdWas a result visibly labeled as advertising?Organic product selection or answer influence
CheckoutWas an external purchase handoff or eligible in-chat checkout path exposed?Completed order, revenue, or customer satisfaction

Add product rank or merchant order only when the surface exposes a stable, reviewable order. Do not invent a rank from visual prominence.

For citations, preserve the displayed URL, source title, position when available, and the claim being reviewed. The cited-but-not-mentioned funnel shows why source appearance and brand naming must remain separate.

For ads, use the visible label and preserve a screenshot. OpenAI says ads are separate from answers and product results. Do not infer that the ad changed the answer.

For checkout, distinguish:

  • Product link to a merchant.
  • Merchant-selection panel.
  • Explicit Instant Checkout eligibility.
  • Checkout initiated.
  • Purchase confirmed in a governed commerce system.

The first three are interface observations. The last two require transaction evidence and appropriate privacy controls.

A Reproducible Execution Workflow

Step 1: Version The Contract

Assign a query-set version and freeze:

  • Scope.
  • Seven intent definitions.
  • Exact Questions.
  • Providers and surfaces.
  • Market and language.
  • Account and personalization state.
  • Chat-state policy.
  • Time window.
  • Retry and eligibility rules.
  • Product, brand, merchant, citation, ad, and checkout coding.

Step 2: Run A Small Pilot

Pilot one or two Questions from each intent.

Look for ambiguity:

  • Does the Question require a location?
  • Does a product family need variant-level coding?
  • Is “available” defined as displayed, in stock, or deliverable?
  • Does the surface request clarifying information?
  • Can two reviewers identify the same merchant and product?

Fix the contract before the baseline, then lock it.

Step 3: Execute Every Planned Unit

Run the Questions in the declared window. Preserve:

  • Raw response.
  • Product cards or product details.
  • Brand and merchant names.
  • Displayed sources.
  • Ad labels.
  • Checkout or handoff state.
  • Screenshot or approved artifact.
  • Timestamp and reviewer.

Step 4: Review Accuracy

For high-priority products, verify:

  • Product identity and variant.
  • Price and currency.
  • Availability.
  • Material specifications.
  • Compatibility.
  • Merchant identity.
  • Return or delivery claim.

OpenAI's shopping research announcement says shopping research can make mistakes about details such as price and availability and advises users to visit the merchant site for the most accurate information. Preserve this limitation instead of presenting the observed answer as catalog truth.

Step 5: Calculate Native Rates

Calculate each signal with its own eligible denominator:

product appearance coverage = eligible answers showing a governed product ÷ completed eligible answers

brand mention coverage = eligible answers naming the governed brand ÷ completed eligible answers

merchant coverage = eligible answers showing an authorized merchant ÷ completed eligible answers

owned citation coverage = eligible answers displaying a governed owned source ÷ completed eligible answers

ad appearance coverage = eligible observations with a labeled target ad ÷ ad-eligible completed observations

checkout-path coverage = eligible product observations exposing the declared checkout or handoff state ÷ eligible product observations

Show raw counts beside rates. Keep intent and surface breakdowns before calculating an overall total.

Step 6: Repeat Without Moving The Contract

Repeat the matched schedule and compare:

  • Completion.
  • Product appearance.
  • Brand mention.
  • Merchant options.
  • Citations.
  • Ads.
  • Checkout paths.
  • Accuracy errors.

If the Question set, surface, market, personalization, or coding changes, label a contract break.

Step 7: Test One Change

Use a predeclared hypothesis:

Correcting accurate size and compatibility data across the catalog, product page, and supported structured fields will remove a documented source inconsistency. We will observe whether product accuracy and appearance change across later matched Runs.

Do not promise that a feed or page edit will cause ChatGPT visibility. The AEO content experiment guide provides a stronger pre/post and control design for focused interventions.

An Auditable CSV Schema

Use one row per planned Question-surface-Run, with child rows or structured fields for multiple products, merchants, and citations.

query_set_version, question_id, question_text, primary_intent, secondary_intents, catalog_scope, market, language, currency, account_state, plan_state, memory_saved_state, memory_history_state, custom_instructions_version, chat_state, surface, run_id, attempt_number, attempt_status, metric_eligibility, observed_at, product_id, product_name, variant_id, product_appeared, brand_id, brand_mentioned, merchant_id, merchant_appeared, displayed_price, availability_claim, citation_url, citation_role, ad_labeled, checkout_state, external_handoff_url, accuracy_verdict, raw_artifact_url, reviewer, coding_version, notes

Do not put personal account data, payment details, private chat content, or customer information into a shared monitoring export. Use governed test accounts and access controls.

Common Mistakes

Do not call the seven intent classes an OpenAI taxonomy.

Do not call 30 Questions or three repeats a statistical requirement.

Do not mix new chats with personalized multi-turn sessions.

Do not leave Memory or Custom Instructions undocumented.

Do not compare markets, languages, currencies, or surfaces as if they were identical.

Do not count a failed or unavailable Run as product absence.

Do not treat a product card as a brand recommendation.

Do not treat a merchant option as proof of inventory or lowest price.

Do not treat a citation as support for every claim in the answer.

Do not count a labeled ad as organic product visibility.

Do not treat an external purchase link as completed checkout.

Do not keep retrying until the desired product appears.

Do not attribute a later visibility change to one catalog or content edit without a design that supports that claim.

The Bottom Line

An AI shopping query set is not a list of ecommerce keywords copied into a chatbot. It is a governed sample of buyer needs connected to explicit product evidence.

Start with a catalog and market boundary. Use discovery, use case, budget, attributes, comparison, availability, and purchase as declared operating classes, not provider labels. Freeze wording, market, language, account, Memory, Custom Instructions, chat state, surface, time, and retries. Then code product, brand, merchant, citation, ad, and checkout separately.

Thirty Questions and three repeats can make the first workflow manageable. They do not make the sample universal. The value comes from preserving every planned unit, showing denominators, reviewing raw evidence, and keeping the contract stable long enough to distinguish ordinary answer variation from a repeatable pattern.

Create a free AEO Table account to organize stable shopping Questions into Tasks, preserve repeatable Runs, compare brand and competitor visibility, review product and merchant appearances in the preserved evidence, and inspect citations without collapsing every shopping signal into one score.

FAQ

What is an AI shopping query set?

An AI shopping query set is a governed collection of buyer Questions used to observe how a named AI shopping surface returns products, brands, merchants, citations, ads, and checkout paths. It preserves exact wording, intent, market, language, account and personalization state, chat state, execution policy, and evidence so later Runs are comparable.

Which shopping intents should the query set include?

A practical set can cover discovery, use case, budget, attributes, comparison, availability, and purchase. These seven classes are an AEO Table operating framework, not an OpenAI taxonomy. Teams should adapt the mix to their catalog and buyer journey while keeping category definitions stable within a baseline.

How many shopping Questions should I monitor?

There is no universal minimum. Thirty Questions are a manageable operational starting point for many teams, not a statistical standard. The set should be large enough to cover priority categories and intents but small enough to rerun, review, and preserve consistently.

Why should Memory and Custom Instructions be fixed?

OpenAI says shopping product selection can consider context such as Memory and Custom Instructions. If those settings differ between observations, a change in the result may reflect personalization rather than a durable visibility change. Record the state and keep it fixed, or test it as a separate declared cohort.

Are product results, citations, ads, and checkout the same signal?

No. Code product appearance, brand mention, merchant option, displayed citation, labeled ad, and checkout availability as separate observable fields. One can occur without the others, and none by itself proves ranking, traffic, conversion, or revenue.