Measurement & Research
AEO Reporting: Turn Evidence into Decisions
This guide is part of the King of AEO learning library.
The short answer
An AEO report should explain what changed, what was observed and what decision the evidence supports. Separate completed work from platform visibility and business outcomes. Give every percentage a denominator, describe changes to the observation method, and attach source records. Finish with a few actions whose owners, rationale and verification criteria are clear.
In this guide
Start with the decision the reader must makeKeep three kinds of evidence distinctUse native reports without double countingMake the denominator visibleAttach evidence people can inspectClose with choices and thresholdsSourcesStart with the decision the reader must make
An AEO report becomes useful when its reader can choose between competing actions. A marketing director might need to decide whether to expand a product documentation programme. An editor might need to choose between repairing existing comparisons and commissioning research. Put that decision near the beginning. Then select evidence that changes the choice. A screenshot of a favourable answer may be interesting, but it does not establish the coverage, reliability or commercial value needed to justify a larger programme. Define those separately before making the recommendation.
Write a short opening assessment that includes the strongest finding and its practical limit. For example, an illustrative report might say that revised compatibility pages now answer the intended questions, while repeated observations still show mixed citation coverage. That supports completing the remaining documentation, especially if support teams need it too. It does not support claiming that the rewrite caused revenue growth. Use a small set of AEO KPIs agreed before the reporting period, so the narrative cannot quietly replace an inconvenient target with an easier success measure.
Keep three kinds of evidence distinct
Delivery evidence establishes work completed: a corrected canonical, a published explanation, or a repaired link. Visibility evidence establishes an observed appearance within a stated product and test context. Business evidence concerns visits, enquiries, purchases or another agreed outcome. These layers can reinforce one another, but none automatically proves the next. A page can be technically sound without appearing in an answer. A citation can appear without sending a visit. A visit can help a future purchase without receiving conversion credit in the chosen analytics model.
Give each layer its own compact presentation. Describe the change and affected URLs in the delivery section. Show comparable observations in the visibility section. Show the relevant downstream outcomes in the business section, with the attribution rule next to the number. Readers should be able to trace a claim across these sections without mistaking sequence for causation. The guide to conversion attribution covers those business measurement choices in detail. Here, the reporting task is to preserve their meaning rather than compress all evidence into one opaque visibility score.
Delivery: Verified changes
Visibility: Comparable observations
Outcomes: Defined business measures
Decision: Action with evidence
Delivery, visibility and business evidence each contribute to a decision without proving one another.
Use native reports without double counting
Google now documents dedicated generative AI performance reports for Search and Discover. Its announcement describes impressions and page, country, device and date views, with device information specified for Search. Label the exact native metric and product surface you export. Do not rename an impression as a citation, a visit or a lead. A platform measure reflects its own counting rules. A third-party prompt observation reflects your selected questions and test conditions. Both can be useful, but they answer different questions.
The same announcement says generative AI data also remains within overall performance reporting. Therefore, adding the dedicated figure to an overall total can count overlapping activity twice. Keep the native AI view as a breakdown where that relationship applies, and document the reporting source. Compare like periods and note any gap in availability. If the reporting interface changes, preserve the original export and update the metric definition before continuing the chart. Otherwise, a new measurement boundary can look like a sudden improvement in content performance when the underlying change is administrative.
Make the denominator visible
An illustrative citation panel with eight appearances across forty completed observations should say eight of forty, alongside the percentage. If ten other attempts failed, disclose them separately and explain whether they were retried. A failed collection is not evidence that the brand was absent. Break results down by meaningful question group when a single total hides a different pattern. Buying questions and implementation questions might behave differently, and improving the wrong group can make the headline rise while the commercial problem remains unresolved.
Do not create precision that your method cannot support. Repeated runs of one prompt are not forty independent customer journeys. A handpicked panel is not a population estimate of all questions people ask. Explain those limitations beside the result using ordinary language. The detailed treatment of sampling bias can help define the panel, while prompt tracking explains its records. In the report, show whether the panel changed, which questions were added or removed, and whether a comparable subset gives the same direction of movement.
Attach evidence people can inspect
A defensible report lets another person follow an important number back to its source. Preserve observation identifiers, collection times, product modes, prompts, cited URLs and stored answer records where permitted. Link completed changes to their release or content records. Keep confidential customer material out of broad circulation, while retaining an authorised route for review. The goal is not to overwhelm executives with raw exports. It is to ensure that a concise claim can be checked when the decision is expensive, disputed or difficult to reverse.
Treat unusual observations as investigation leads. If a key page disappears from a sample, first check the page and collection method before naming a cause. Google's Search Console guidance describes tools for examining search performance and indexing. Use those diagnostics to support specific technical findings, while recognising that an indexed page is not promised an answer appearance. Record confirmed faults separately from hypotheses. This prevents the report turning every noisy observation into an urgent technical ticket, or overlooking a real site problem because average visibility remains healthy.
Close with choices and thresholds
Each recommended action should name the problem, the supporting evidence, the owner and the condition for considering the work complete. In an illustrative example, a product manager could confirm missing regional restrictions before an editor updates three affected guides. Completion would mean the restrictions are accurate across those pages and their linked product records. Continued observation could then assess visibility, but that is a separate follow-up. A good recommendation makes progress possible even when the platform outcome is uncertain, because the underlying customer information becomes more reliable.
Include a deliberate decision to wait when evidence does not justify intervention. State what further observation would change that decision, such as a persistent loss across comparable runs combined with an access fault. Revisit earlier recommendations and explain what happened after implementation. This creates an honest learning record rather than a succession of unrelated monthly claims. When a change needs stronger causal evidence, route it into an AEO experiment. Reporting should make the next decision clearer, including when the available evidence is insufficient to claim that the last decision worked.
When moving an action into the AEO backlog, preserve the report finding that justified it. Otherwise, a concrete recommendation can become a vague ticket such as improve AI visibility, with no reliable completion condition. Include the reporting period and evidence identifier so a future review can distinguish the original problem from later changes. If the action depends on another team, name the dependency rather than presenting the task as ready to execute. This handover keeps reporting connected to actual improvements and makes it possible to explain why a sensible recommendation was delayed, revised or abandoned when better evidence arrived.
Sources and further reading
- Google generative AI reportingDedicated generative AI impression reporting exists and overlaps overall performance data.
- Google Search Console guideSearch Console provides search performance and indexing diagnostics.