Blog

Measurement

Multi-Turn AI Visibility: Test the Full Conversation

Measure AI visibility across multi-turn conversations with session contracts, turn evidence, shortlist changes, constraints, and citations.

A buyer rarely begins with a complete brief.

They may first ask for a category of tools, then add a budget, reject one option, require an integration, and finally ask, "Which one would you choose?" The last sentence contains almost none of the decision context. It is meaningful because the conversation carries the rest.

An AI visibility test that extracts only that final prompt and replays it in a fresh chat has changed the request. It may still produce a useful isolated-prompt observation, but it is no longer measuring the original conversation.

Short answer: Treat a multi-turn conversation as a versioned measurement object. Preserve every turn, the request-state changes introduced at that turn, the history supplied to the model, the shortlist before and after each new constraint, and the sources displayed with each answer. Report turn-level and session-level outcomes separately. Do not infer that history caused a brand or citation change unless a paired experiment actually holds the final turn and model conditions fixed while varying the supplied history.

This guide builds that contract without pretending a conversational assistant exposes a stable rank.

Why The Final Prompt Is Not The Whole Query

In keyword search, one submitted string is usually a defensible input unit. In conversational search, the current message can be a state update rather than a self-contained specification.

Consider this constructed buyer journey:

  1. "What tools can monitor how AI assistants describe a B2B SaaS brand?"
  2. "We are a five-person marketing team and need something under $200 a month."
  3. "Exclude tools that only generate content. We need preserved answers and citations."
  4. "Does either option support weekly competitor tracking?"
  5. "Which one would you choose?"

Turn five does not restate the category, team size, budget, exclusion, evidence requirement, cadence, or competitor need. An isolated replay of "Which one would you choose?" is under-specified. Rewriting it into a complete prompt is a different treatment because the analyst decides which prior details remain active and how to phrase them.

A stable AI search query set remains useful for repeatable category, comparison, pricing, proof, and technical monitoring. The isolated baseline and conversational journey simply answer different questions and need separate labels.

What The July 2026 Study Actually Found

The preprint The Prompt Is Not the Query analyzed user turns from two sources. The commercial source included 670 English multi-turn conversations in a discovery-and-replication design. The public PRISM source included 7,463 eligible conversations from 1,389 participants. The study analyzed 8,133 conversations in total and excluded assistant text from its outcomes.

The study measured unique user-side content vocabulary and nine transparent request-state cue families: price or budget, location or proximity, persona or use case, attribute requirement, time, alternatives, correction or redirect, comparison or evaluation, and explanation or evidence.

ResultCommercial conversationsPRISM conversationsCorrect interpretation
Conversations analyzed6707,463Two separately reported sources, not one representative sample of all AI use
Median final-prompt share of unique user-side content vocabulary35.6%36.4%Local information availability, not semantic loss
Final prompt contains at most half of that vocabulary68.4%74.3%The endpoint is often short relative to the session
At least one detected dimension appears in history but not the final prompt50.3%44.8%Explicit request evidence often remains outside the endpoint
Final prompt reproduces the complete detected dimension set26.1%26.2%Among dimension-bearing conversations only: 456 commercial and 4,534 PRISM conversations
Final prompt adds a previously unseen dimension17.9%19.3%The endpoint can add state rather than summarize prior state

The vocabulary result needs special care. Length-matched null models reproduced most of the low lexical coverage, so the study interprets it as information availability caused largely by turn length, not evidence of semantic drift. The cue rules are also conservative pattern detectors, not a model of a person's latent intent.

Most importantly, no prompt was rerun. The study did not compare answers with full history, user-only history, and no history. It therefore did not estimate how history changes facts, brands, recommendations, sources, confidence, or constraint satisfaction. Its evidence supports treating session state as a measurement object; it does not prove that history causes a particular visibility outcome.

Multi-Turn Surfaces Exist, But There Is No Special Ranking Shortcut

Google says AI Mode supports nuanced questions, further exploration, and complex comparisons. Google also says AI Mode and AI Overviews may use query fan-out to issue related searches across subtopics and sources.

Those statements do not disclose how a prior turn affects source selection, establish a universal conversation length, or create a special optimization requirement. Google says normal Search eligibility and people-first SEO practices still apply, with no additional technical requirement for inclusion in AI Mode or AI Overviews.

Use product documentation to define the surface and captured answers to describe observations. Neither reveals an undisclosed ranking formula.

Define The Session-Level Measurement Unit

Use this as the primary unit for a governed conversational test:

Session script version × answer surface × state profile × session execution

The session script version defines the planned buyer journey and allowed branches. The answer surface identifies the actual product interface. The state profile captures context that can change eligibility or output. The session execution is one complete conversation attempt during a declared window.

Within that session, the smallest observable unit is:

Turn × cumulative request state × answer evidence

This separation prevents analysts from counting turns as independent buyers, treating one session as stable market share, or ignoring inherited and superseded constraints.

Session-Level Fields

FieldWhat to preserveWhy it matters
session_idStable identifier for one executed conversationConnects every turn and answer without pooling sessions
script_versionGoverned scenario, branch rules, and revision dateMakes material prompt changes visible
buyer_jobThe outcome the constructed buyer needsKeeps the conversation commercially relevant
surfaceNamed product and interface, such as AI Mode or ChatGPT webPrevents unlike experiences from being pooled
market_languageDeclared market and languagePreserves locale conditions
state_profilePublic or logged-in status, device, plan, workspace, memory, and personalization when relevant and observableExposes context differences without guessing
history_policyFull interleaved history, user-turn-only history, isolated endpoint, or another declared policyDefines what the model received
execution_identityDate, time, product version or model when exposed, and search or retrieval mode when exposedHelps diagnose runtime changes
completion_policyCompleted, failed, refused, blocked, ineligible, or abandonedPreserves the denominator
privacy_ruleRetention, redaction, consent, and access policyProtects conversational data

Do not fill unavailable identity fields with an analyst's guess. Record not exposed or unknown and keep the limitation visible.

Turn-Level Fields

FieldWhat to preserveQuestion it answers
turn_indexOrdered user and assistant turn numberWhere did the state or answer change?
user_textExact submitted text before any rewriteWhat did the user actually add or revise?
state_deltaNew, changed, negated, or superseded dimensionsWhat changed at this turn?
active_stateCumulative constraints believed active under a declared coding ruleWhat specification was available for review?
assistant_textComplete displayed answer where collection is permittedHow was the decision framed?
shortlistGoverned brands or products named, compared, or recommendedWho entered, survived, or left consideration?
claim_evidenceMaterial product claims and supporting passagesWere constraints described accurately?
citationsDisplayed URL, domain, placement, and cited claimWhich sources accompanied the answer?
tool_or_product_eventProduct card, app suggestion, invocation, or other separately defined eventDid the interface do more than mention a brand?
review_statusReviewer, code version, disagreement, and adjudicationCan another analyst audit the classification?

The active_state field is an analyst construct, not access to the model's hidden state. A buyer can revoke a constraint, and an assistant can misunderstand one. Keep observed user evidence, analyst coding, and model behavior in separate columns.

If tool_or_product_event is a private plugin selection rather than ordinary answer evidence, use the ChatGPT plugin activation test guide to score eligible selection, false activation, arguments, and completion without mixing those events into public brand visibility.

Build Conversation Scripts That Change One Decision Dimension At A Time

A useful script resembles a real evaluation journey but remains reviewable.

Start with an unbranded discovery prompt. Then add one material constraint per turn. Good dimensions for B2B SaaS include buyer role, company size, budget, integration, security, data location, workflow, switching cost, and proof requirement. The SaaS and Plugin visibility guide provides a fuller buyer-job and account-state framework.

Use a five-turn core such as:

TurnUser operationExampleIntended observation
1Discover"Which tools help a SaaS team monitor brand visibility in AI answers?"Initial category shortlist
2Constrain"We have five marketers and a $200 monthly budget."Budget and team-fit survival
3Require evidence"We need the raw answers and cited URLs, not only a score."Evidence-feature accuracy
4Compare risk"Which options support competitor tracking without a long contract?"Competitive framing and commercial claim accuracy
5Decide"Which two should we evaluate first, and why?"Final shortlist and rationale

This table is a constructed example, not customer data and not evidence of how any provider will answer.

Avoid a script that names the target brand in turn one if the purpose is unprompted discovery. Use separate branded controls to test factual understanding. Keep competitor aliases and coding rules frozen, as described in the competitor AI search tracking workflow.

Predeclare branches. If the assistant asks a clarifying question, use one governed reply; if it recommends an irrelevant category, use one neutral correction. Record any branch as its own treatment.

Track Shortlist Entry, Survival, Exit, And Return

A single final mention rate hides the decision path. Code each governed brand at every turn.

Shortlist stateOperational definitionExample interpretation
EnteredFirst turn where the brand is named as a relevant optionThe brand joined consideration after a use-case cue
SurvivedBrand remains relevant after a new material constraintThe answer still presents the product as fitting the stated budget
ExitedBrand is removed or explicitly described as unsuitableA security or integration requirement excludes it
ReturnedBrand reappears after an additional change or correctionA revised budget or requirement changes fit
RecommendedAnswer affirmatively proposes the brand for the active stateStronger than a passing mention
Mentioned onlyBrand is named without positive selection framingAssociation, not endorsement

Useful session summaries include initial shortlist coverage, first-entry turn, constraint-survival count, final shortlist coverage, and recommendation framing. Always show counts beside percentages and keep the session denominator visible.

An exit is not automatically an AEO failure. Exclusion can be accurate when a product lacks a required integration. Use the AI share-of-voice framework only after the cohort, eligible answers, and recommendation coding are stable.

Track Constraints As A State Ledger

For every constraint, record the turn where it was introduced, its current lifecycle state (active or superseded), and its answer verdict (satisfied, violated, unverified, ambiguous, or not applicable).

Constraint recordExample value
DimensionBudget
Introduced at turn2
User evidence"under $200 a month"
Current stateActive
Answer treatmentProduct described as starting at $149 per month
Verification statusNeeds current vendor pricing source
VerdictSatisfied, violated, unverified, ambiguous, or not applicable

Do not equate a fluent response with constraint satisfaction. Verify material prices, features, integrations, security claims, and availability against current governed sources. If an answer silently violates an inherited constraint, preserve the exact conflict rather than editing the prompt after the fact.

For "Not Europe; we need US data residency," mark the earlier location assumption as superseded and the new one as active. Do not let the analyst create an impossible request by retaining both.

Preserve Citation Movement At Every Turn

Sources can change before the shortlist changes. A documentation page may appear when the buyer asks about an integration. A pricing page may appear after a budget constraint. A review may support a comparison while the vendor's domain disappears.

At each turn, preserve:

  • Displayed citation URLs and final resolved URLs.
  • Source domain and declared source type.
  • Citation placement and the claim it appears to support.
  • First appearance, persistence, disappearance, and return.
  • Owned, competitor-owned, official, primary research, editorial, marketplace, and community source classes.
  • Access, freshness, and claim-support status as separate fields.

A URL appearing at turn four does not prove that the turn-four wording caused retrieval. It is an observed sequence. Apply the AI citation accuracy audit when a high-risk claim needs a claim-to-source verdict, and use the cited-but-not-mentioned workflow when the source appears without visible brand credit.

Run A Paired History Test When You Need An Answer-Effect Claim

Descriptive session monitoring answers, "What happened along this observed conversation?" A paired history experiment asks, "How did the supplied history change the answer under this test design?"

Hold the final user turn, named surface, timing window, market, language, and exposed model conditions as stable as practical. Compare at least three declared treatments:

  1. Full interleaved history: prior user and assistant turns plus the final turn.
  2. User-turn-only history: prior user turns plus the final turn, without assistant messages.
  3. Isolated endpoint: the final turn in a new session.

Repeat each treatment because generated answers vary. Preserve failures and do not select the most favorable response. Compare brands, recommendation framing, constraint satisfaction, factual claims, and displayed sources separately.

This design can estimate a difference under the tested treatments. It does not reveal a ranking system or prove that one source caused a recommendation. Consult the AI search volatility guide when defining repeated-observation ranges, and use the AEO content experiment protocol when testing a content intervention.

Use A Seven-Step Operating Workflow

1. Define The Buyer Decision

Choose one category and one meaningful decision, such as selecting a support analytics tool for a regulated mid-market company. Do not combine unrelated journeys into one session.

2. Freeze The Script And Branch Rules

Version exact user turns, permitted clarifications, constraint order, branded controls, and stopping rules. Mark constructed scenarios as test data.

3. Declare The State Profile

Record product surface, public or logged-in state, market, language, device, and observable personalization or workspace conditions. Use accounts and data only when authorized.

4. Execute Without Silent Repairs

Submit the planned turns. Preserve refusals, failures, clarifying questions, and irrelevant answers. If a branch is triggered, record the branch identifier.

5. Code The Turn Evidence

Update request state, shortlist, recommendation framing, constraint treatment, claims, and citations. Use two reviewers for material or ambiguous classifications.

6. Repeat Matched Sessions

Use the same script, state profile, and cadence. Treat a provider interface or history-policy change as a contract break rather than continuing one trend line.

7. Route The Finding

Send inaccurate owned claims to the governed source owner, unsupported third-party claims to the appropriate correction path, genuine product gaps to product, and unstable observations back to measurement. The wrong AI answer correction workflow helps keep the evidence and remediation paths separate.

Keep Isolated Monitoring And Session QA Separate

A stable isolated-Question panel and a multi-turn session test are complementary.

ProgramBest useMain limitation
Isolated Question baselineRepeatable category, proof, comparison, and competitor monitoringDoes not reproduce an evolving conversation unless the full state is written into each Question
Governed session QAShortlist and source movement as constraints accumulateMore expensive to execute, code, and repeat
Paired history experimentAnswer differences between declared history treatmentsNarrow experimental result, not a universal causal law
Webmaster and analytics dataDiscovery, impressions, clicks, sessions, and downstream behaviorDoes not expose the complete prompt-level conversation

AEO Table's current Task-and-Run workflow can support stable Questions and preserved answer evidence for supported channels, as described in the AI search monitoring operating guide. This article does not claim that AEO Table automatically orchestrates multi-turn sessions, controls consumer account state, or captures every conversational product surface. Run session tests manually or through an authorized QA system unless current product documentation explicitly says otherwise.

Protect Conversational Data

Conversation transcripts can contain personal details, confidential requirements, and sparse language that remains identifying after superficial redaction. Prefer constructed scenarios. For real transcripts, establish consent, purpose limitation, retention, access control, deletion, and review rules before collection. Remove secrets and customer data from editorial artifacts, and never move a sales, support, or procurement transcript into a testing surface without authorization.

Limitations To Put In Every Report

  • The script is a governed scenario, not a representative sample of every buyer conversation.
  • Request-state coding observes explicit evidence; it does not reveal private intent or hidden model state.
  • Later user turns may respond to assistant framing, so state evolution is not user-only causation.
  • Product surfaces, models, retrieval behavior, account state, and sources can change during the study.
  • One session does not establish a stable brand rank or market share.
  • A sequence between a new constraint and a shortlist change is not proof that the constraint caused the change.
  • Citation presence is not claim support, and source disappearance is not proof of deindexing.
  • Isolated endpoint and rewritten full-context prompts are constructed treatments, not replicas of the original interaction.

The July 2026 paper has additional source boundaries: the commercial corpus was governed rather than randomly sampled from global assistant use; PRISM was recruited and protocol-driven; cue rules can miss or misclassify state; and no answer effects were measured. Preserve those limits whenever its percentages are reused.

Report The Conversation, Not A Synthetic Winner

A decision-ready report should include:

  • Session script version, purpose, branch, and execution window.
  • Surface, market, language, state profile, and history policy.
  • Attempted, completed, failed, refused, blocked, ineligible, and abandoned sessions.
  • Turn-by-turn state deltas and cumulative active constraints.
  • Shortlist entry, survival, exit, return, and final recommendation evidence.
  • Constraint satisfaction and material factual-accuracy verdicts.
  • Citation appearance, persistence, claim mapping, and source class.
  • Variation across matched session executions.
  • Contract breaks, coding disagreements, privacy controls, and limitations.
  • One proportional next action with an owner and validation plan.

Avoid one blended "conversation visibility score." A final recommendation rate, early shortlist entry, constraint survival, and citation accuracy answer different business questions. The AI search visibility report template can hold the executive decision, while the turn ledger remains attached as auditable evidence.

Common Mistakes

  • Replaying only the final prompt and calling it the original query.
  • Rewriting a conversation into one polished prompt without labeling a new treatment.
  • Counting turns as independent sessions or retaining revoked constraints.
  • Mixing public, logged-in, personalized, and managed-workspace states in one denominator.
  • Treating a passing mention as a recommendation or a citation as factual verification.
  • Hiding failures and refusals or inferring a ranking formula from one answer order.
  • Claiming a history effect without a matched repeated experiment.
  • Implying that a product automates session behavior it has not documented.

The Bottom Line

A final prompt is one message. A conversational decision can be distributed across the session.

Keep isolated Questions for a repeatable baseline, then add governed session tests for the buyer journeys where constraints, corrections, comparisons, and evidence requests develop over several turns. Preserve the history policy, request-state ledger, shortlist, claims, citations, failures, and account conditions. When causality matters, run a paired repeated experiment instead of treating sequence as proof.

That approach produces defensible evidence about when a brand enters consideration, survives a constraint, is accurately excluded, and receives source support.

Create a free AEO Table account to build the stable isolated-Question baseline that your governed conversation tests can extend.

FAQ

What is multi-turn AI visibility?

Multi-turn AI visibility is the observed way a brand, product, competitor, claim, and source appear or disappear as a buyer develops one request across several conversational turns. It should be measured at both the turn level and the complete-session level.

Why is the final prompt not enough for AI visibility testing?

A final prompt may rely on a budget, use case, location, alternative, correction, or evidence request stated earlier. Replaying only the last sentence changes the available request state and may test a different object from the original conversation.

What should a multi-turn AI visibility test record?

Record the session script and version, every user and assistant turn, cumulative active constraints, superseded constraints, shortlist membership and framing, citations, failures, surface, market, language, account state, history policy, and timestamp.

Does research prove that conversation history changes brand recommendations?

No. The July 2026 study discussed in this guide measured where explicit user request evidence appeared across observed conversations. It did not rerun final prompts with and without history, so it did not estimate a causal effect on answers, brands, rankings, or citations.

Can AEO Table automatically run multi-turn conversation tests?

This guide does not claim that capability. AEO Table can support a stable isolated-Question baseline and preserve answer evidence for its current supported channels. Session orchestration, account-state testing, and turn-by-turn conversation capture should be treated as a separate manual or authorized QA workflow unless the product explicitly documents automation for them.