Measurement
Multi-Turn AI Visibility: Test the Full Conversation
Measure AI visibility across multi-turn conversations with session contracts, turn evidence, shortlist changes, constraints, and citations.
A buyer rarely begins with a complete brief.
They may first ask for a category of tools, then add a budget, reject one option, require an integration, and finally ask, "Which one would you choose?" The last sentence contains almost none of the decision context. It is meaningful because the conversation carries the rest.
An AI visibility test that extracts only that final prompt and replays it in a fresh chat has changed the request. It may still produce a useful isolated-prompt observation, but it is no longer measuring the original conversation.
Short answer: Treat a multi-turn conversation as a versioned measurement object. Preserve every turn, the request-state changes introduced at that turn, the history supplied to the model, the shortlist before and after each new constraint, and the sources displayed with each answer. Report turn-level and session-level outcomes separately. Do not infer that history caused a brand or citation change unless a paired experiment actually holds the final turn and model conditions fixed while varying the supplied history.
This guide builds that contract without pretending a conversational assistant exposes a stable rank.
Why The Final Prompt Is Not The Whole Query
In keyword search, one submitted string is usually a defensible input unit. In conversational search, the current message can be a state update rather than a self-contained specification.
Consider this constructed buyer journey:
- "What tools can monitor how AI assistants describe a B2B SaaS brand?"
- "We are a five-person marketing team and need something under $200 a month."
- "Exclude tools that only generate content. We need preserved answers and citations."
- "Does either option support weekly competitor tracking?"
- "Which one would you choose?"
Turn five does not restate the category, team size, budget, exclusion, evidence requirement, cadence, or competitor need. An isolated replay of "Which one would you choose?" is under-specified. Rewriting it into a complete prompt is a different treatment because the analyst decides which prior details remain active and how to phrase them.
A stable AI search query set remains useful for repeatable category, comparison, pricing, proof, and technical monitoring. The isolated baseline and conversational journey simply answer different questions and need separate labels.
What The July 2026 Study Actually Found
The preprint The Prompt Is Not the Query analyzed user turns from two sources. The commercial source included 670 English multi-turn conversations in a discovery-and-replication design. The public PRISM source included 7,463 eligible conversations from 1,389 participants. The study analyzed 8,133 conversations in total and excluded assistant text from its outcomes.
The study measured unique user-side content vocabulary and nine transparent request-state cue families: price or budget, location or proximity, persona or use case, attribute requirement, time, alternatives, correction or redirect, comparison or evaluation, and explanation or evidence.
| Result | Commercial conversations | PRISM conversations | Correct interpretation |
|---|---|---|---|
| Conversations analyzed | 670 | 7,463 | Two separately reported sources, not one representative sample of all AI use |
| Median final-prompt share of unique user-side content vocabulary | 35.6% | 36.4% | Local information availability, not semantic loss |
| Final prompt contains at most half of that vocabulary | 68.4% | 74.3% | The endpoint is often short relative to the session |
| At least one detected dimension appears in history but not the final prompt | 50.3% | 44.8% | Explicit request evidence often remains outside the endpoint |
| Final prompt reproduces the complete detected dimension set | 26.1% | 26.2% | Among dimension-bearing conversations only: 456 commercial and 4,534 PRISM conversations |
| Final prompt adds a previously unseen dimension | 17.9% | 19.3% | The endpoint can add state rather than summarize prior state |
The vocabulary result needs special care. Length-matched null models reproduced most of the low lexical coverage, so the study interprets it as information availability caused largely by turn length, not evidence of semantic drift. The cue rules are also conservative pattern detectors, not a model of a person's latent intent.
Most importantly, no prompt was rerun. The study did not compare answers with full history, user-only history, and no history. It therefore did not estimate how history changes facts, brands, recommendations, sources, confidence, or constraint satisfaction. Its evidence supports treating session state as a measurement object; it does not prove that history causes a particular visibility outcome.
Multi-Turn Surfaces Exist, But There Is No Special Ranking Shortcut
Google says AI Mode supports nuanced questions, further exploration, and complex comparisons. Google also says AI Mode and AI Overviews may use query fan-out to issue related searches across subtopics and sources.
Those statements do not disclose how a prior turn affects source selection, establish a universal conversation length, or create a special optimization requirement. Google says normal Search eligibility and people-first SEO practices still apply, with no additional technical requirement for inclusion in AI Mode or AI Overviews.
Use product documentation to define the surface and captured answers to describe observations. Neither reveals an undisclosed ranking formula.
Define The Session-Level Measurement Unit
Use this as the primary unit for a governed conversational test:
Session script version × answer surface × state profile × session execution
The session script version defines the planned buyer journey and allowed branches. The answer surface identifies the actual product interface. The state profile captures context that can change eligibility or output. The session execution is one complete conversation attempt during a declared window.
Within that session, the smallest observable unit is:
Turn × cumulative request state × answer evidence
This separation prevents analysts from counting turns as independent buyers, treating one session as stable market share, or ignoring inherited and superseded constraints.
Session-Level Fields
| Field | What to preserve | Why it matters |
|---|---|---|
session_id | Stable identifier for one executed conversation | Connects every turn and answer without pooling sessions |
script_version | Governed scenario, branch rules, and revision date | Makes material prompt changes visible |
buyer_job | The outcome the constructed buyer needs | Keeps the conversation commercially relevant |
surface | Named product and interface, such as AI Mode or ChatGPT web | Prevents unlike experiences from being pooled |
market_language | Declared market and language | Preserves locale conditions |
state_profile | Public or logged-in status, device, plan, workspace, memory, and personalization when relevant and observable | Exposes context differences without guessing |
history_policy | Full interleaved history, user-turn-only history, isolated endpoint, or another declared policy | Defines what the model received |
execution_identity | Date, time, product version or model when exposed, and search or retrieval mode when exposed | Helps diagnose runtime changes |
completion_policy | Completed, failed, refused, blocked, ineligible, or abandoned | Preserves the denominator |
privacy_rule | Retention, redaction, consent, and access policy | Protects conversational data |
Do not fill unavailable identity fields with an analyst's guess. Record not exposed or unknown and keep the limitation visible.
Turn-Level Fields
| Field | What to preserve | Question it answers |
|---|---|---|
turn_index | Ordered user and assistant turn number | Where did the state or answer change? |
user_text | Exact submitted text before any rewrite | What did the user actually add or revise? |
state_delta | New, changed, negated, or superseded dimensions | What changed at this turn? |
active_state | Cumulative constraints believed active under a declared coding rule | What specification was available for review? |
assistant_text | Complete displayed answer where collection is permitted | How was the decision framed? |
shortlist | Governed brands or products named, compared, or recommended | Who entered, survived, or left consideration? |
claim_evidence | Material product claims and supporting passages | Were constraints described accurately? |
citations | Displayed URL, domain, placement, and cited claim | Which sources accompanied the answer? |
tool_or_product_event | Product card, app suggestion, invocation, or other separately defined event | Did the interface do more than mention a brand? |
review_status | Reviewer, code version, disagreement, and adjudication | Can another analyst audit the classification? |
The active_state field is an analyst construct, not access to the model's hidden state. A buyer can revoke a constraint, and an assistant can misunderstand one. Keep observed user evidence, analyst coding, and model behavior in separate columns.
If tool_or_product_event is a private plugin selection rather than ordinary answer evidence, use the ChatGPT plugin activation test guide to score eligible selection, false activation, arguments, and completion without mixing those events into public brand visibility.
Build Conversation Scripts That Change One Decision Dimension At A Time
A useful script resembles a real evaluation journey but remains reviewable.
Start with an unbranded discovery prompt. Then add one material constraint per turn. Good dimensions for B2B SaaS include buyer role, company size, budget, integration, security, data location, workflow, switching cost, and proof requirement. The SaaS and Plugin visibility guide provides a fuller buyer-job and account-state framework.
Use a five-turn core such as:
| Turn | User operation | Example | Intended observation |
|---|---|---|---|
| 1 | Discover | "Which tools help a SaaS team monitor brand visibility in AI answers?" | Initial category shortlist |
| 2 | Constrain | "We have five marketers and a $200 monthly budget." | Budget and team-fit survival |
| 3 | Require evidence | "We need the raw answers and cited URLs, not only a score." | Evidence-feature accuracy |
| 4 | Compare risk | "Which options support competitor tracking without a long contract?" | Competitive framing and commercial claim accuracy |
| 5 | Decide | "Which two should we evaluate first, and why?" | Final shortlist and rationale |
This table is a constructed example, not customer data and not evidence of how any provider will answer.
Avoid a script that names the target brand in turn one if the purpose is unprompted discovery. Use separate branded controls to test factual understanding. Keep competitor aliases and coding rules frozen, as described in the competitor AI search tracking workflow.
Predeclare branches. If the assistant asks a clarifying question, use one governed reply; if it recommends an irrelevant category, use one neutral correction. Record any branch as its own treatment.
Track Shortlist Entry, Survival, Exit, And Return
A single final mention rate hides the decision path. Code each governed brand at every turn.
| Shortlist state | Operational definition | Example interpretation |
|---|---|---|
| Entered | First turn where the brand is named as a relevant option | The brand joined consideration after a use-case cue |
| Survived | Brand remains relevant after a new material constraint | The answer still presents the product as fitting the stated budget |
| Exited | Brand is removed or explicitly described as unsuitable | A security or integration requirement excludes it |
| Returned | Brand reappears after an additional change or correction | A revised budget or requirement changes fit |
| Recommended | Answer affirmatively proposes the brand for the active state | Stronger than a passing mention |
| Mentioned only | Brand is named without positive selection framing | Association, not endorsement |
Useful session summaries include initial shortlist coverage, first-entry turn, constraint-survival count, final shortlist coverage, and recommendation framing. Always show counts beside percentages and keep the session denominator visible.
An exit is not automatically an AEO failure. Exclusion can be accurate when a product lacks a required integration. Use the AI share-of-voice framework only after the cohort, eligible answers, and recommendation coding are stable.
Track Constraints As A State Ledger
For every constraint, record the turn where it was introduced, its current lifecycle state (active or superseded), and its answer verdict (satisfied, violated, unverified, ambiguous, or not applicable).
| Constraint record | Example value |
|---|---|
| Dimension | Budget |
| Introduced at turn | 2 |
| User evidence | "under $200 a month" |
| Current state | Active |
| Answer treatment | Product described as starting at $149 per month |
| Verification status | Needs current vendor pricing source |
| Verdict | Satisfied, violated, unverified, ambiguous, or not applicable |
Do not equate a fluent response with constraint satisfaction. Verify material prices, features, integrations, security claims, and availability against current governed sources. If an answer silently violates an inherited constraint, preserve the exact conflict rather than editing the prompt after the fact.
For "Not Europe; we need US data residency," mark the earlier location assumption as superseded and the new one as active. Do not let the analyst create an impossible request by retaining both.
Preserve Citation Movement At Every Turn
Sources can change before the shortlist changes. A documentation page may appear when the buyer asks about an integration. A pricing page may appear after a budget constraint. A review may support a comparison while the vendor's domain disappears.
At each turn, preserve:
- Displayed citation URLs and final resolved URLs.
- Source domain and declared source type.
- Citation placement and the claim it appears to support.
- First appearance, persistence, disappearance, and return.
- Owned, competitor-owned, official, primary research, editorial, marketplace, and community source classes.
- Access, freshness, and claim-support status as separate fields.
A URL appearing at turn four does not prove that the turn-four wording caused retrieval. It is an observed sequence. Apply the AI citation accuracy audit when a high-risk claim needs a claim-to-source verdict, and use the cited-but-not-mentioned workflow when the source appears without visible brand credit.
Run A Paired History Test When You Need An Answer-Effect Claim
Descriptive session monitoring answers, "What happened along this observed conversation?" A paired history experiment asks, "How did the supplied history change the answer under this test design?"
Hold the final user turn, named surface, timing window, market, language, and exposed model conditions as stable as practical. Compare at least three declared treatments:
- Full interleaved history: prior user and assistant turns plus the final turn.
- User-turn-only history: prior user turns plus the final turn, without assistant messages.
- Isolated endpoint: the final turn in a new session.
Repeat each treatment because generated answers vary. Preserve failures and do not select the most favorable response. Compare brands, recommendation framing, constraint satisfaction, factual claims, and displayed sources separately.
This design can estimate a difference under the tested treatments. It does not reveal a ranking system or prove that one source caused a recommendation. Consult the AI search volatility guide when defining repeated-observation ranges, and use the AEO content experiment protocol when testing a content intervention.
Use A Seven-Step Operating Workflow
1. Define The Buyer Decision
Choose one category and one meaningful decision, such as selecting a support analytics tool for a regulated mid-market company. Do not combine unrelated journeys into one session.
2. Freeze The Script And Branch Rules
Version exact user turns, permitted clarifications, constraint order, branded controls, and stopping rules. Mark constructed scenarios as test data.
3. Declare The State Profile
Record product surface, public or logged-in state, market, language, device, and observable personalization or workspace conditions. Use accounts and data only when authorized.
4. Execute Without Silent Repairs
Submit the planned turns. Preserve refusals, failures, clarifying questions, and irrelevant answers. If a branch is triggered, record the branch identifier.
5. Code The Turn Evidence
Update request state, shortlist, recommendation framing, constraint treatment, claims, and citations. Use two reviewers for material or ambiguous classifications.
6. Repeat Matched Sessions
Use the same script, state profile, and cadence. Treat a provider interface or history-policy change as a contract break rather than continuing one trend line.
7. Route The Finding
Send inaccurate owned claims to the governed source owner, unsupported third-party claims to the appropriate correction path, genuine product gaps to product, and unstable observations back to measurement. The wrong AI answer correction workflow helps keep the evidence and remediation paths separate.
Keep Isolated Monitoring And Session QA Separate
A stable isolated-Question panel and a multi-turn session test are complementary.
| Program | Best use | Main limitation |
|---|---|---|
| Isolated Question baseline | Repeatable category, proof, comparison, and competitor monitoring | Does not reproduce an evolving conversation unless the full state is written into each Question |
| Governed session QA | Shortlist and source movement as constraints accumulate | More expensive to execute, code, and repeat |
| Paired history experiment | Answer differences between declared history treatments | Narrow experimental result, not a universal causal law |
| Webmaster and analytics data | Discovery, impressions, clicks, sessions, and downstream behavior | Does not expose the complete prompt-level conversation |
AEO Table's current Task-and-Run workflow can support stable Questions and preserved answer evidence for supported channels, as described in the AI search monitoring operating guide. This article does not claim that AEO Table automatically orchestrates multi-turn sessions, controls consumer account state, or captures every conversational product surface. Run session tests manually or through an authorized QA system unless current product documentation explicitly says otherwise.
Protect Conversational Data
Conversation transcripts can contain personal details, confidential requirements, and sparse language that remains identifying after superficial redaction. Prefer constructed scenarios. For real transcripts, establish consent, purpose limitation, retention, access control, deletion, and review rules before collection. Remove secrets and customer data from editorial artifacts, and never move a sales, support, or procurement transcript into a testing surface without authorization.
Limitations To Put In Every Report
- The script is a governed scenario, not a representative sample of every buyer conversation.
- Request-state coding observes explicit evidence; it does not reveal private intent or hidden model state.
- Later user turns may respond to assistant framing, so state evolution is not user-only causation.
- Product surfaces, models, retrieval behavior, account state, and sources can change during the study.
- One session does not establish a stable brand rank or market share.
- A sequence between a new constraint and a shortlist change is not proof that the constraint caused the change.
- Citation presence is not claim support, and source disappearance is not proof of deindexing.
- Isolated endpoint and rewritten full-context prompts are constructed treatments, not replicas of the original interaction.
The July 2026 paper has additional source boundaries: the commercial corpus was governed rather than randomly sampled from global assistant use; PRISM was recruited and protocol-driven; cue rules can miss or misclassify state; and no answer effects were measured. Preserve those limits whenever its percentages are reused.
Report The Conversation, Not A Synthetic Winner
A decision-ready report should include:
- Session script version, purpose, branch, and execution window.
- Surface, market, language, state profile, and history policy.
- Attempted, completed, failed, refused, blocked, ineligible, and abandoned sessions.
- Turn-by-turn state deltas and cumulative active constraints.
- Shortlist entry, survival, exit, return, and final recommendation evidence.
- Constraint satisfaction and material factual-accuracy verdicts.
- Citation appearance, persistence, claim mapping, and source class.
- Variation across matched session executions.
- Contract breaks, coding disagreements, privacy controls, and limitations.
- One proportional next action with an owner and validation plan.
Avoid one blended "conversation visibility score." A final recommendation rate, early shortlist entry, constraint survival, and citation accuracy answer different business questions. The AI search visibility report template can hold the executive decision, while the turn ledger remains attached as auditable evidence.
Common Mistakes
- Replaying only the final prompt and calling it the original query.
- Rewriting a conversation into one polished prompt without labeling a new treatment.
- Counting turns as independent sessions or retaining revoked constraints.
- Mixing public, logged-in, personalized, and managed-workspace states in one denominator.
- Treating a passing mention as a recommendation or a citation as factual verification.
- Hiding failures and refusals or inferring a ranking formula from one answer order.
- Claiming a history effect without a matched repeated experiment.
- Implying that a product automates session behavior it has not documented.
The Bottom Line
A final prompt is one message. A conversational decision can be distributed across the session.
Keep isolated Questions for a repeatable baseline, then add governed session tests for the buyer journeys where constraints, corrections, comparisons, and evidence requests develop over several turns. Preserve the history policy, request-state ledger, shortlist, claims, citations, failures, and account conditions. When causality matters, run a paired repeated experiment instead of treating sequence as proof.
That approach produces defensible evidence about when a brand enters consideration, survives a constraint, is accurately excluded, and receives source support.
Create a free AEO Table account to build the stable isolated-Question baseline that your governed conversation tests can extend.
FAQ
What is multi-turn AI visibility?
Multi-turn AI visibility is the observed way a brand, product, competitor, claim, and source appear or disappear as a buyer develops one request across several conversational turns. It should be measured at both the turn level and the complete-session level.
Why is the final prompt not enough for AI visibility testing?
A final prompt may rely on a budget, use case, location, alternative, correction, or evidence request stated earlier. Replaying only the last sentence changes the available request state and may test a different object from the original conversation.
What should a multi-turn AI visibility test record?
Record the session script and version, every user and assistant turn, cumulative active constraints, superseded constraints, shortlist membership and framing, citations, failures, surface, market, language, account state, history policy, and timestamp.
Does research prove that conversation history changes brand recommendations?
No. The July 2026 study discussed in this guide measured where explicit user request evidence appeared across observed conversations. It did not rerun final prompts with and without history, so it did not estimate a causal effect on answers, brands, rankings, or citations.
Can AEO Table automatically run multi-turn conversation tests?
This guide does not claim that capability. AEO Table can support a stable isolated-Question baseline and preserve answer evidence for its current supported channels. Session orchestration, account-state testing, and turn-by-turn conversation capture should be treated as a separate manual or authorized QA workflow unless the product explicitly documents automation for them.