Blog

Measurement

How to Run an AEO Content Experiment: Pre/Post Runs and Controls

Use matched AI search Runs, predeclared hypotheses, control Questions, and answer evidence to test whether a page change is associated with better visibility.

You update a comparison page on Monday. On Friday, your brand appears in an AI answer that omitted it the week before.

Did the page work?

The honest answer is that you observed a change after an edit. You have not yet separated the edit from normal answer variation, different completed results, a search or model update, new competitor content, or a source-discovery delay.

Short answer: Treat an AEO content test as a predeclared, repeated measurement study. Choose one decision and primary metric, freeze the Task, learn its baseline range, change one defined content unit, preserve publication and discovery evidence, then run matched post-change windows with concurrent controls. Inspect the answers and citations behind the aggregate. Unless assignment was randomized or the design supports stronger inference, report that the intervention was associated with the observed movement rather than claiming it caused the movement.

This protocol will not turn an answer engine into a laboratory instrument. It will make your content decision more auditable and much harder to cherry-pick.

Monitoring, Rollouts, And Experiments Are Different

Teams often call all three activities a "test."

ActivityPrimary QuestionMinimum EvidenceAppropriate Claim
MonitoringWhat did selected AI surfaces return?Timestamped Questions, answers, channels, sources and outcomes"This Run observed..."
Content rolloutWas the page published and discoverable?Deployment, rendered page, canonical, sitemap and available crawl/index evidence"The change is live and eligible..."
Content experimentDid the selected outcomes move after the intervention under a predefined design?Repeated matched Runs, frozen definitions, controls, raw evidence and confounder log"The post window was associated with..."

A rollout can pass while answer visibility stays unchanged. Monitoring can find a movement without explaining it. An experiment connects the two with a declared comparison and explicit limits.

Why One Before/After Screenshot Fails

One pre-change answer and one post-change answer confound several events:

  1. Answer variance: the same Question can return different brands and citations across repeats.
  2. Completion variance: failures and ineligible results can change the denominator.
  3. Retrieval variance: the answer surface may select different sources.
  4. Time: news, product releases, reviews, competitor edits, and index changes continue during the test.
  5. Platform changes: the provider, model, search mode, or visible product can change without your Task changing.
  6. Question context: wording, conversation history, locale, and account state can alter the observed result.
  7. Coding drift: a reviewer can change what counts as a mention, recommendation, citation, or correct claim.

The 2026 preprint Don't Measure Once: Measuring Visibility in AI Search argues that AI visibility should be treated as a distribution across repeated observations rather than one fixed ranking. Another preprint, Quantifying Uncertainty in AI Visibility, found substantial citation variation in repeated samples across three answer systems and three consumer-product topics. Its scope is narrow, but it demonstrates why some precise-looking domain differences can fall inside an observed noise range.

Those papers do not supply one universal sample size for every AEO experiment. They support a more basic rule: measure the variability of your own Questions and channels before assigning meaning to a change.

Define The Experimental Unit

The most useful raw unit is:

Question × channel × Run

Each unit needs an attempted status and, when completed, the answer, citations, brand coding, framing, market, language, execution identity, and timestamp.

Do not treat every answer as statistically independent merely because it occupies a separate row. Answers to the same Question or from the same channel can share sources and behavior. Pooled answer counts are useful operational denominators, but high-stakes statistical inference should account for repeated Questions, channels, and time rather than pretending the rows are unrelated coin flips.

The AI visibility measurement methodology defines the Task and Run evidence needed to make this unit reviewable.

Start With A Decision, Not A Tactic

"Add FAQs and see what happens" is not an experiment brief.

Begin with the decision the result will change:

  • Keep, revise, or revert a new evidence section.
  • Extend a page pattern to the rest of a content cluster.
  • Invest in original research rather than another comparison page.
  • Fix entity and source clarity before producing more articles.
  • Stop an ineffective tactic before scaling it.

Then state a minimum meaningful change. A two-percentage-point movement can be statistically interesting in a large, stable dataset but commercially irrelevant. A ten-point movement can still be untrustworthy when the baseline routinely swings by fifteen points.

Predeclare both:

  • Decision threshold: the smallest movement that would justify action.
  • Evidence threshold: the persistence and data quality required before using that movement.

Write A Falsifiable Hypothesis

Use this template:

If [one content intervention] on [specific canonical page or page cohort] becomes publicly discoverable, then [primary outcome] for [target Question class and channels] will change by at least [minimum meaningful amount] relative to [baseline and controls] during [predeclared post window], while [guardrail outcome] remains within [declared boundary].

Example:

If the security comparison page adds a reviewed SAML SSO evidence table with links to canonical documentation, then the owned-citation rate for 12 unbranded enterprise-authentication Questions will increase by at least eight percentage points relative to its pre window and concurrent control Questions across the same three channels. The claim-accuracy rate must not decline.

This hypothesis can fail. That is a feature.

Avoid hypotheses such as "make the page more AI-friendly" or "increase authority." They do not define an intervention, observable result, comparison, or decision.

Use An Evidence Ladder

Not every team can run the strongest design. Label the design you actually have.

DesignWhat It AddsMain Limitation
One pre and one post observationA documented anecdoteCannot estimate normal variation
Repeated matched pre/post RunsBaseline range and persistenceTime-varying confounders remain
Concurrent control Questions or pagesComparison with outcomes not expected to respondControls may not share the same shocks or retrieval dynamics
Staggered rollout across matched page groupsMultiple intervention times and stronger counterfactual evidenceRequires comparable pages, longer study and disciplined rollout
Randomized platform-controlled assignmentDirect treatment/control assignmentRarely available to publishers measuring external answer engines

Most publisher-led AEO studies are observational or quasi-experimental. Calling them A/B tests does not create random assignment.

The original Generative Engine Optimization paper reported sizable visibility improvements for some methods in its benchmark, with effects varying by domain. That is useful experimental research in a defined environment, not a guarantee that copying one tactic onto one live page will produce the same lift across commercial answer engines. Test the decision in your own measurement contract.

The Twelve-Step AEO Content Experiment Protocol

1. Choose One Primary Decision And Metric

Select one primary metric before seeing the post data:

  • Brand mention rate.
  • Recommendation rate.
  • Owned citation rate.
  • Claim accuracy rate.
  • Supported-answer rate.

Keep other outcomes secondary or guardrails. If you test five tactics against eight metrics and promote whichever cell turns green, the reported win reflects the search process as much as the content.

2. Define The Intervention Precisely

Record:

  • Canonical URL or page cohort.
  • Exact sections, claims, tables, source links, schema, or technical access controls changing.
  • Before-state evidence.
  • Intended mechanism stated as a hypothesis, not a fact.
  • Author, reviewer, approval, merge or deployment reference, and publication time.

Prefer one coherent intervention. If you change the title, URL, template, internal links, body evidence, schema, and ten supporting pages at once, the bundle may be worth shipping, but the test cannot isolate which component mattered.

3. Freeze The Task

Before the baseline, freeze or record:

LayerRequired Contract
BrandCanonical name, aliases and owned domains
CompetitorsFixed cohort and aliases
QuestionsExact text, order, buyer stage and brand class
ScopeMarket, language and device or account state when applicable
ChannelsExplicit answer-engine products or surfaces
RuntimeProvider, requested and returned model, search mode and version where exposed
EligibilityCompleted, failed, blocked and ineligible rules
CodingMention, recommendation, citation, support and accuracy definitions
EvidenceAnswer, citations, source URLs, timestamp and reviewer

If one of these changes during the study, log the break. A material contract change usually requires a new baseline.

The AI search query set guide explains how to separate unbranded discovery, branded validation, comparison, problem, and buying-stage Questions.

4. Run A Repeated Pre-Change Window

Use the complete frozen Task on a fixed cadence. Do not stop the baseline after a conveniently low Run.

The number of Runs depends on:

  • Baseline variability for the primary outcome.
  • Number and diversity of Questions and channels.
  • Completion and failure rates.
  • Expected discovery and category-change speed.
  • Minimum meaningful change.
  • Cost and consequence of a wrong decision.

Report individual Runs and their range. An average alone can hide a channel outage or one unusual spike.

5. Select Controls Before The Intervention

Controls should be measured concurrently and should not be expected to respond directly to the edited page.

Useful options include:

  • Control Questions: similar buyer intent but about an unaffected topic.
  • Control pages: a comparable page or page cohort not receiving the change.
  • Negative-control outcome: a metric the intervention should not plausibly move.
  • Positive process check: a stable Question and source that should remain detectable, helping reveal execution or coding failures.
  • Competitor and channel context: concurrent movement that could indicate a broader platform or market shift.

Controls need a rationale. Easy branded Questions are poor controls for difficult unbranded discovery Questions. A stable control does not eliminate all confounding, and a moving control does not automatically invalidate the study; it changes the interpretation.

6. Publish And Preserve The Deployment Evidence

After the baseline is complete:

  1. Publish the predeclared change.
  2. Confirm the final public URL, rendered text, response status, canonical, internal links, and sitemap.
  3. Preserve the deployment time and before/after diff.
  4. Record supported discovery actions without treating them as indexing guarantees.
  5. Avoid unrelated edits to the target page during the post window.

If an emergency correction is required, make it. Log the deviation rather than preserving a bad customer experience for experimental purity.

7. Use Two Clocks And A Fixed Start Rule

Track:

  • Deployment clock: when the content became public.
  • Discovery clock: when available evidence first shows a relevant crawler, index, or source surface recognized the updated URL.

Discovery evidence is often incomplete. Therefore predeclare a post-window start rule such as:

"Begin matched post Runs at the later of 72 hours after deployment or the first available index confirmation, capped at seven days after deployment."

That rule is only an example. Choose one appropriate to the site and cadence before the intervention. Do not keep delaying the window until a favorable answer appears.

Google's recrawl guidance says discovery can take days to weeks and is not guaranteed. Other answer surfaces can use different retrieval and update paths. A fixed rule makes the uncertainty visible; it does not solve it.

8. Run The Matched Post-Change Window

Repeat:

  • The same target and control Questions.
  • The same channels and locale.
  • The same cadence and number of planned Runs.
  • The same completion and retry policy.
  • The same brand, citation, framing, and accuracy coding.

Preserve failures. If the post period has fewer completed answers, use completed eligible answers as the rate denominator and show attempted and failed counts beside it.

9. Score Outcomes Separately

For completed eligible answers:

Brand mention rate

Answers naming the governed brand or alias ÷ completed eligible answers

Owned citation rate

Answers displaying at least one owned-domain citation ÷ completed eligible answers

Recommendation rate

Answers affirmatively recommending the brand for the defined need ÷ completed eligible answers

Claim accuracy rate

Answers stating the governed claim correctly, including required qualifications ÷ completed eligible answers

Supported-answer rate

Answers whose relevant claim is supported by an inspectable displayed source ÷ answers where source support can be reviewed

Do not merge mention, citation, recommendation, and accuracy into one "visibility" score during the experiment. A page can gain citations without a brand mention, or gain mentions while an answer becomes less accurate.

10. Maintain A Confounder And Deviation Log

During both windows, record:

  • Provider, model, search-mode, or product-surface changes.
  • Channel incidents, rate limits, access failures, and retry deviations.
  • Major search news or category events.
  • Competitor launches, acquisitions, rebrands, and content changes.
  • New third-party reviews, directories, research, or press coverage.
  • Changes to the site template, robots rules, canonical, sitemap, or internal links.
  • Reviewer or coding-rule changes.
  • Any target-page edits after the intervention.

A confounder does not automatically erase the result. It limits the claim and can explain why channels diverged.

11. Compare Ranges, Controls, And Raw Evidence

Start with a descriptive review:

  1. Show attempted, completed, failed, and ineligible counts.
  2. Plot or list the primary metric for every pre and post Run.
  3. Compare the post values with the observed baseline range.
  4. Compare target movement with concurrent controls.
  5. Break results down by Question and channel.
  6. Inspect the answers and source URLs responsible for the movement.
  7. Check secondary and guardrail metrics.

A simple control-adjusted descriptive contrast is:

(Target post rate − Target pre rate) − (Control post rate − Control pre rate)

This resembles a difference-in-differences contrast, but the arithmetic does not validate its causal assumptions. Target and control trends may have differed before the edit, repeated answers may be dependent, and platform shocks may affect the groups differently.

When the design and sample support it, report uncertainty intervals or use resampling that respects repeated Questions and channels. The 2026 paper Auditing Citation Behavior in AI-Generated Search Summaries offers a useful observational framework built around query-document pairs and source provenance. It is an audit framework, not permission to treat every answer row as independent.

For executive, contractual, or public benchmark claims, involve someone qualified to review the sampling unit, dependence, missingness, uncertainty method, and causal language.

12. Apply The Predeclared Decision Rule

An illustrative rule might be:

Ship the pattern to the next page cohort only if the owned-citation rate improves by at least eight percentage points, the improvement appears in at least three of four post Runs, the target moves beyond its baseline range, the control remains within its declared range, and answer review confirms that citations support the intended claim. Revert or revise if claim accuracy falls.

That is an operating heuristic, not a universal statistical threshold. Your rule should reflect the observed baseline, cost, downside, and decision.

Possible dispositions:

  • Extend: evidence threshold and guardrails pass.
  • Hold: direction is promising, but persistence or data quality is insufficient.
  • Revise and retest: answer evidence reveals the wrong mechanism or an implementation problem.
  • Revert: the change harms accuracy, usability, or another guardrail.
  • Inconclusive: the design, failures, or confounders do not support a decision.

Worked Example: Evidence Links On A Security Comparison Page

The following numbers are illustrative sample data, not AEO Table customer results.

The team measures:

  • 12 target Questions about enterprise authentication.
  • 8 concurrent control Questions about data export, an unaffected topic.
  • 3 answer-engine channels.
  • 4 pre-change and 4 post-change Runs.

That creates 144 attempted target opportunities and 96 attempted control opportunities per window.

Group And WindowAttemptedCompletedOwned-Citation AnswersOwned-Citation RateBrand-Mention AnswersBrand-Mention Rate
Target pre1441322015.2%3828.8%
Target post1441353525.9%4130.4%
Control pre96901314.4%2527.8%
Control post96911415.4%2628.6%

The target owned-citation rate rose by about 10.8 percentage points. The control rose by about 0.9 point, producing a 9.8-point control-adjusted descriptive contrast when calculated from the unrounded rates.

That table is not enough to declare victory.

The per-Run review shows that the target rate improved in all four post Runs and exceeded the highest pre Run in three of four. Nine of the fifteen additional citation answers display the edited comparison page; six display the canonical SSO documentation linked from it. Claim accuracy stays inside its predeclared guardrail. One channel produces most of the remaining uncited answers, and the control stays within its baseline range.

A defensible report says:

"After the reviewed evidence section and documentation links became public, matched post Runs showed a persistent increase in owned citations for the target authentication Questions. The target movement exceeded the concurrent control movement, and answer review connected most additional citations to the edited page or linked documentation. The quasi-experimental design supports an association, not proof that the page change caused every observed answer."

The next decision could be to extend the pattern to one matched page cohort, not to rewrite the entire site.

Use A Pre-Registered Experiment Brief

You do not need an academic registry. You need a timestamped plan that cannot be silently rewritten after the result.

FieldWhat To Record
Experiment ID and ownerStable name and accountable decision-maker
DecisionExtend, revise, revert, hold or stop
HypothesisIntervention, outcome, comparison, window and guardrail
Target unitURL or page cohort plus exact Question class
Primary metricOne metric and its denominator
Secondary metricsMention, recommendation, citation, accuracy or support
Minimum meaningful changePractical threshold
Task contractBrands, Questions, channels, locale, runtime and coding rules
Pre windowCadence, planned Runs and stop rule
ControlsQuestions, pages, outcomes and selection rationale
InterventionExact diff and intended mechanism
Discovery ruleFixed post-window start boundary
Post windowCadence, planned Runs and stop rule
Decision rulePersistence, range, control and guardrail requirements
ConfoundersKnown events and deviation process
AnalysisBreakdowns, uncertainty method and reviewer
Claim languageMaximum conclusion the design can support

Lock the brief before the intervention. Append deviations; do not overwrite the original plan.

Common AEO Experiment Mistakes

Choosing The Metric After Seeing The Result

If owned citations stay flat but a passing brand mention rises, the team cannot quietly rename the experiment a mention test. Report both and keep the original primary outcome.

Changing Multiple Content Systems At Once

A site migration, template release, internal-link overhaul, research launch, and page rewrite may be a valuable program. It is not an isolated page experiment.

Using Branded Questions To Prove Discovery

"Is Acme Cloud good?" primes the brand. It cannot substitute for unbranded category or problem Questions when the decision concerns discovery.

Hiding Failed Answers

A rising success rate with a shrinking denominator may reflect missing difficult observations. Show attempted, completed, failed, and ineligible outcomes.

Waiting Until The Result Looks Good

Starting the post window after the first citation, stopping after a peak, or extending only a disappointing window biases the comparison.

Treating A Control As Magic

An unaffected topic may have different sources, variance, seasonality, and channel behavior. Explain why it is informative and inspect its pre trend.

Scaling From An Aggregate Without Reading Answers

Additional citations can be irrelevant, unsupported, negative, or attached to the wrong entity. Source and framing review determines whether the metric moved in the intended way.

Calling Association Causation

Pre/post order, a control-adjusted contrast, or a small p-value cannot by itself rule out platform, market, source, timing, and measurement confounders.

How AEO Table Fits The Protocol

Use a stable AEO Table Task for the predeclared brand, competitor, Question, channel, market, and language scope. Each Run preserves a measurement snapshot so the team can compare completed outcomes, answers, mentions, and citations without relying on one screenshot.

The product is the evidence layer, not a causality machine. Keep the intervention diff, deployment proof, discovery rule, control rationale, decision threshold, and confounder log in the experiment brief. Then use AI search reports to review matched Runs and AI citation tracking to inspect the source movement behind the aggregate.

Start with the AI search volatility guide if you have not established a stable baseline. An experiment becomes useful only after the team knows what it is comparing and what decision the evidence is allowed to change.

FAQ

How many Runs does an AEO content experiment need?

There is no universal number. Use repeated baseline Runs to estimate the normal variation of the exact Task, then choose the pre and post windows based on that variation, the minimum change worth acting on, and the stakes of the decision.

Is a before-and-after AEO test a true A/B test?

Usually not. Most teams cannot randomly assign an answer engine's retrieval or generated output to treatment and control conditions. Matched pre/post Runs with controls are generally observational or quasi-experimental and should be reported as associations unless the design supports a stronger causal claim.

Which metrics should an AEO content experiment measure?

Use completed eligible answers as the denominator and score brand mention, recommendation framing, owned citation, total citation, answer accuracy, and source support separately. Choose one primary metric before the experiment and keep the others as secondary or guardrail metrics.

How long should an AEO team wait after publishing a content change?

There is no guaranteed discovery window. Predeclare a waiting rule using available crawl or index evidence and a fixed elapsed-time boundary, preserve both clocks, and do not move the start date until the first favorable answer appears.