Blog

AI Citations

AI Citation Accuracy Audit: Does the Source Support the Claim?

Audit whether each source supports the exact claim an AI answer makes, with reproducible verdicts, evidence fields, and review rules.

An AI answer can display a credible-looking citation and still attribute the wrong fact to it.

The link may work. The page may discuss the right topic. The publisher may be reputable. None of those observations proves that the cited passage supports the exact claim shown in the answer.

That is the job of a citation accuracy audit.

Short answer: Audit at the level of atomic claim × cited source × answer Run. Preserve the answer and citation mapping, retrieve the cited content, and assign one factual verdict: Supported, Partially supported, Conflicting, Contradicted, or Not found. Record accessibility, freshness, and source type in separate fields. An inaccessible page is an access exception, not an unsupported claim. Report eligible counts beside rates, retain repeated-Run evidence, and route factual errors according to their risk and source ownership.

This is narrower than an AI citation gap analysis. A gap analysis asks which sources or brands repeatedly win a defined answer context. An accuracy audit asks whether a displayed source actually carries the evidentiary weight the answer places on it.

A Citation Is Not A Verification Result

Citation presence answers a structural question: did the answer show or associate a source with some text?

Citation accuracy answers an evidentiary question: does that source support the exact factual claim, with the same entity, scope, time period, geography, product version, denominator, and qualification?

Those questions often diverge. A pricing page can be relevant to a pricing claim while documenting a different plan. A research abstract can mention an outcome while studying a different population. A current documentation page can support a feature today but not prove it existed when an older answer Run was captured. A review can correctly describe a category while failing to support a specific vendor comparison.

This is also different from the cited-but-not-mentioned pattern. There, the source may be used without the brand being named. Here, the concern is whether the answer's claim-to-source attribution is faithful, regardless of whether a brand receives visible credit.

Treat every citation marker as a hypothesis to test, not a quality seal.

Use Atomic Claim × Cited Source × Answer Run

The smallest auditable record has three parts.

Atomic claim

An atomic claim is one independently checkable assertion. Split compound sentences before judging them.

Consider this answer sentence:

Acme launched in 2024, supports SSO on every paid plan, and reduced customer onboarding time by 40%.

That sentence contains at least three claims:

  1. Acme launched in 2024.
  2. Acme supports SSO on every paid plan.
  3. Acme reduced customer onboarding time by 40%.

One company page might support the launch year, a pricing page might limit SSO to an enterprise plan, and a case study might report the 40% result for one customer rather than all customers. A sentence-level verdict would conceal those differences.

Make each claim self-contained. Replace pronouns with the named entity, retain important qualifiers, and do not silently turn an estimate into a fact or a case study into a universal result.

Cited source

One claim can have zero, one, or several displayed sources. Create a separate row for every claim-source relationship. If three sources are attached to one claim, judge all three; do not let one supporting page excuse two irrelevant citations.

Preserve both the displayed URL and the final resolved URL. Record redirects, canonical URL, cited title, displayed snippet, and the passage actually retrieved. URL normalization helps analysis, but the raw citation is part of the evidence.

Answer Run

The same Question can generate different wording, sources, and citation placement across time, providers, markets, languages, accounts, or personalization states. The Run therefore belongs in the unit.

Capture provider and surface, Question text, timestamp, market, language, device or client, account state, Task version, answer text, citation positions, and a screenshot or raw response where permitted. The AI search volatility guide explains why one answer should not be treated as a durable platform state.

If the same claim-source pair appears in five Runs, preserve five observations. For summaries, show both Run-level recurrence and deduplicated claim-source counts so repeated answers do not create false precision.

Keep The Factual Verdict Separate From Source Conditions

Use one primary factual verdict, then add independent fields for accessibility, freshness, and source type.

VerdictUse it whenDo not confuse it with
SupportedThe retrieved source directly supports the complete atomic claim with its material qualifiers intactThe page merely discussing the same topic
Partially supportedThe source supports a meaningful part of the claim, but a material scope, number, date, condition, or causal statement is missing or narrowerA stylistic paraphrase that preserves the full meaning
ConflictingThe cited source contains materially incompatible statements or evidence that prevents one stable reading for the claimTwo independent sources disagreeing; record each claim-source pair separately first
ContradictedThe source explicitly supports an incompatible proposition, value, condition, or outcomeThe source saying nothing about the claim
Not foundThe accessible content captured for review does not address the atomic claimAn inaccessible page, a failed extraction, or proof that the claim is false

Not found means the reviewer could not locate support in the content that was successfully captured. It is an attribution failure for that citation observation, not necessarily a real-world falsehood. The fact could be true and documented elsewhere.

Conflicting is useful when a single cited artifact is internally unstable: for example, a product page says a feature is included while its plan table excludes it, or a report's narrative and appendix use incompatible denominators. When two different sources disagree, retain two rows and judge each source against the claim before escalating the cross-source conflict.

Accessibility is a separate field

Use statuses such as accessible, paywalled, login required, blocked, timeout, removed, or extraction failed. Record HTTP status and retrieval time where available.

Do not convert inaccessible into Not found, Contradicted, or a generic unsupported result. A crawler failure describes your observation boundary, not the source's contents. Leave the factual verdict unassessed, exclude the row from the factual-verdict denominator, and place it in an access-exception queue. Retry with a permitted browser, licensed archive, publisher-provided copy, or contemporaneous snapshot when policy allows.

Freshness is a separate field

Record source publication date, visible update date, answer Run date, retrieval date, and whether the cited fact is time-sensitive. Useful statuses include current for the claim, dated but applicable, stale for the claim, undated, and changed since Run.

A stale source may still accurately support a historical claim. A newly updated page may no longer show what the answer used at Run time. Preserve snapshots and avoid treating updatedAt as proof that every sentence was reviewed.

Source type is a separate field

Classify source ownership and role without using it as a verdict shortcut. A practical taxonomy is first-party product or documentation, official or regulatory, primary research, secondary editorial, marketplace or directory, partner or reseller, and community or user-generated content.

An official source can be misquoted. A community source can accurately support an observation. Source type helps route remediation and assess risk; it does not determine claim fidelity.

What Recent Research Does And Does Not Establish

Research supports a multidimensional audit, but published percentages must stay attached to their samples and methods.

CiteEval, published at ACL 2025, argues that binary or ternary natural-language-inference labels are an incomplete proxy for citation evaluation. Its framework considers the user query, generated text, cited source, and wider retrieval context, then builds a human-annotated multi-domain benchmark. The practical lesson is not to copy one benchmark label set blindly. It is to preserve enough context for a reviewer to understand what role the citation was supposed to play.

A 2026 preprint, Cited but Not Verified, separates URL accessibility, topical relevance, and factual consistency when evaluating cited deep-research reports. In its experiment across 14 closed- and open-source models, the strongest frontier models exceeded 94% link validity and 80% relevance, while reported factual-accuracy results ranged from 39% to 77%. Those figures belong to that paper's prompts, systems, retrieval setup, rubric-based model judges, and human calibration. They are not a universal accuracy range for every ChatGPT, Google, Perplexity, or enterprise RAG answer.

The same study reports an ablation across two frontier models in which its Fact Check measure declined as tool calls increased from 2 to 150, while its access and relevance measures stayed comparatively stable. That is useful evidence against assuming that more citations automatically mean better synthesis. It does not prove that every longer research process reduces accuracy.

The Google AI Overview study Measuring Google AI Overviews needs equally careful framing. The authors issued 55,393 trending queries across 19 categories during a 40-day collection window in March and April 2026. They captured 7,583 AI Overviews and treated 7,491 as verifiable after excluding 80 with no extractable claims and 12 with only social-media references. Their pipeline decomposed those verifiable AIOs into 98,020 atomic claims.

Within that Google AIO sample, the paper labeled 1.4% of claims Ambiguous, 2.7% Incorrect, and 7.0% Omitted: 11.0% combined Ambiguous/Incorrect/Omitted under the study's verification taxonomy. The displayed component percentages are rounded and therefore add to 11.1%; the paper's combined value is calculated from the underlying counts. This is not an 11% error constant for AI search. It should not be generalized to ChatGPT, Perplexity, other Google surfaces, commercial buyer queries, or your own Task. The preprint was under review, used an automated claim-extraction and verification pipeline with manual validation, and documented retrieval limitations including paywalled or unavailable content.

The defensible takeaway is methodological: claim fidelity and source reputation are different variables, omission can be more common than explicit contradiction, and an audit must publish its denominator and exclusions.

Build A Reproducible Citation Accuracy Audit

1. Predeclare the decision and sample

State what the audit can change. Possible decisions include correcting a product fact, updating an owned source, contacting a publisher, changing a campaign claim, or monitoring a recurring provider pattern.

Build the sample before reading convenient examples. Stratify by buyer intent, provider, market, source ownership, and risk. Include high-intent comparison, pricing, security, compliance, and integration Questions where an attribution error could affect a purchase. Add a random sample of ordinary Questions so the audit is not only a collection of known failures.

Freeze inclusion rules, date window, minimum Run count, retry policy, and escalation threshold. If the Question set is still unstable, use the AI search query set guide before calculating a rate.

2. Preserve the original answer evidence

Store the complete answer, not only the sentence that looks wrong. Preserve citation markers, source panels, displayed titles, snippets, raw URLs, provider surface, and Run metadata. A citation marker can apply to one clause, one sentence, or a paragraph; record that observed placement before interpreting it.

Hash or version the evidence record where practical. If the interface later changes, the audit should still show what the reviewer saw.

3. Atomize claims without changing meaning

Split conjunctions, lists, comparisons, numbers, causal statements, and time qualifiers. Keep normative opinions separate from verifiable claims. “Acme is the best choice” may not be directly verifiable, while “Acme is the least expensive option at $20 per month” contains testable price and comparison claims.

Have a second reviewer inspect claim splitting for high-stakes answers. A poor split can manufacture support by dropping a qualifier or manufacture contradiction by removing context.

4. Retrieve and snapshot every cited source

Attempt retrieval under a documented policy. Record redirect chain, response status, final URL, canonical URL, retrieval timestamp, page title, visible dates, and extracted passage. Save a permitted snapshot or content hash so later edits can be distinguished from the state reviewed.

Do not bypass access controls. Put inaccessible citations in the separate exception queue and retry according to policy. Report the accessible share before any factual rate.

5. Judge the claim against the source

Read the smallest passage that can resolve the claim, then inspect surrounding context, tables, footnotes, definitions, and date boundaries. Compare every named entity, number, unit, population, geography, plan, version, and causal verb.

Record a short rationale and quote only the minimum source fragment needed for internal evidence. Assign the factual verdict only after confirming that extraction did not omit a table, script-rendered block, PDF page, or relevant continuation.

Use automation to propose claim splits, locate passages, or flag mismatched numbers. Do not let one uncalibrated model silently make final decisions about consequential claims. The AEO content experiment protocol offers a useful pattern: predeclare the rule, preserve raw evidence, and limit the claim to what the design supports.

6. Resolve disagreement

Double-code a calibration sample before scaling. Track reviewer agreement by verdict and discuss recurring boundary cases such as narrower populations, approximate numbers, implied causality, and updated product documentation.

Send Contradicted, Conflicting, and high-risk Partially supported records to a second reviewer. If reviewers cannot resolve a case, retain both rationales and mark it for adjudication rather than forcing consensus invisibly.

Use An Evidence Table, Not A Screenshot Folder

At minimum, preserve these fields:

Field groupRequired fields
RunTask version, Question, provider and surface, timestamp, market, language, account state
ClaimAtomic claim ID, exact claim, surrounding answer text, risk class
CitationDisplayed URL and title, citation position, final URL, canonical URL
RetrievalAccess status, HTTP status, retrieval time, snapshot or hash, extracted passage
Source conditionSource type, publisher, publication and update dates, freshness status
DecisionFactual verdict, rationale, reviewer, review time, adjudication state
ActionOwner, correction path, due date, verification Run

This structure keeps evidence available when an aggregate changes. It also prevents source availability or prestige from leaking into the factual verdict.

Report Denominators Before Percentages

Start with a flow count:

  1. Answer Runs captured.
  2. Answers with displayed citations.
  3. Atomic claims extracted.
  4. Claim-source-Run rows created.
  5. Sources successfully retrieved.
  6. Rows eligible for factual verdicts.
  7. Verdict counts.

Then calculate separate rates.

  • Citation access rate: accessible claim-source-Run rows divided by attempted rows.
  • Strict support rate: Supported divided by factually assessed rows.
  • Qualified support rate: Supported plus Partially supported divided by factually assessed rows, always reported beside the component counts.
  • Contradiction rate: Contradicted divided by factually assessed rows.
  • Not-found rate: Not found divided by factually assessed rows.
  • Fresh-evidence share: assessed rows whose sources are current for the claim divided by rows with a determinable freshness status.

Do not hide Conflicting records inside a support bucket. Do not include inaccessible rows in a factual denominator. Do not call the resulting number “platform accuracy” unless the sampling design genuinely represents that platform and surface.

Break results down by Question intent, provider surface, source type, risk, and owned versus third-party source. Show unique claim-source pairs beside Run-level rows. For broader reporting design, connect the audit to the AI visibility dashboard guide without collapsing citation presence and citation correctness into one score.

Turn Verdicts Into Owned Actions

The same verdict can require different work depending on source ownership and business risk.

PatternLikely next action
Supported, current owned sourcePreserve the evidence, improve discoverable context only if needed, and monitor recurrence
Partially supported by owned documentationAdd missing qualifiers or split broad marketing language from documented facts
Not found in an owned pageCreate or update the canonical evidence page; do not manufacture proof for a claim the business cannot substantiate
Contradicted by current first-party factsEscalate as an answer correction and check whether other public pages create the conflict
Third-party source is stale or wrongPrepare a source-backed correction request and track whether the publisher updates it
Inaccessible sourceRetry or obtain an authorized snapshot; do not label it unsupported
Conflicting sourceResolve the underlying documentation conflict before trying to influence the AI answer

When the answer itself is materially wrong, follow the wrong AI answer correction workflow. When competitors repeatedly receive better source support, connect the verified findings back to the citation gap backlog. When internal discovery is the problem, use the internal link audit guide to strengthen legitimate paths to the canonical evidence.

Prioritize by claim risk, buyer intent, recurrence, source ownership, and feasibility. A false security certification or dosage statement deserves a different response from an outdated office location. Keep legal, medical, financial, safety, and regulatory decisions with qualified reviewers.

Re-Run The Exact Question After Remediation

Publishing a correction does not prove that an answer engine discovered it, and a changed answer does not prove that your edit caused the change.

Record the source change, publication time, crawl or indexing evidence available to you, and the first eligible post-change Run. Repeat matched Runs under the same Task contract. Review the underlying answer and citation mapping, not only a summary score.

If the corrected source appears and supports the claim, record an observed post-change association. If another source appears, treat that as new evidence. The AEO ROI measurement guide applies the same discipline downstream: join evidence only through observed or explicitly modeled links.

The Bottom Line

A citation accuracy audit is not a link checker and not a hunt for embarrassing screenshots. It is a reproducible comparison between one atomic claim and one cited source in one captured answer Run.

Keep the factual verdict narrow. Keep access, freshness, and source type separate. Preserve denominators and exclusions. Repeat important Questions. Then route each verified problem to the team that can correct the underlying evidence.

Create a free AEO Table account to capture repeatable Tasks and Runs, review the citations behind AI answers, and build an evidence-backed accuracy backlog instead of relying on isolated prompt checks.

FAQ

What is an AI citation accuracy audit?

An AI citation accuracy audit checks whether the exact cited source supports each atomic factual claim in a captured AI answer. It preserves the answer Run, claim, citation mapping, retrieved source passage, verdict, and separate source-condition fields so another reviewer can reproduce the decision.

What is the correct unit for auditing an AI citation?

Use atomic claim × cited source × answer Run. A sentence may contain several claims, one claim may cite several sources, and the same Question may produce different answers across Runs, so sentence-level or page-level scoring hides attribution errors.

Does an inaccessible citation count as unsupported?

No. Record accessibility separately and exclude an inaccessible source from the factual-verdict denominator until its content can be retrieved from an authorized snapshot or another valid capture. A timeout, paywall, robots restriction, or login requirement is not evidence that the source contradicts or omits the claim.

Which citation accuracy verdicts should a team use?

Use Supported, Partially supported, Conflicting, Contradicted, and Not found for claim-to-source review. Keep accessibility, freshness, and source type in separate fields so source condition does not get mistaken for factual support.

How often should AI citation accuracy be audited?

Audit high-risk or high-intent Questions on every material change and review a stable stratified sample on a regular cadence. Preserve repeated Runs because an accurate citation once does not prove durable accuracy, and one inaccurate Run does not establish a persistent platform pattern.