Measurement & Research
Prompt Tracking: Build a Reproducible Answer Baseline
This guide is part of the King of AEO learning library.
The short answer
Prompt tracking records answers to a defined set of questions under documented conditions. Keep a stable core panel, preserve the exact prompt and output, record the platform and relevant settings, and repeat observations to reveal variability. Separate exploratory questions from the baseline so changes in wording or sample composition do not masquerade as performance gains.
In this guide
Turn a question list into an observation protocolFreeze the wording that defines the baselineRecord the conditions that can change an answerPreserve outputs before reducing them to scoresRepeat observations without pretending they are independent peopleHandle unavailable runs as data about coverageUse the baseline to identify questions worth investigatingSourcesTurn a question list into an observation protocol
A spreadsheet of interesting prompts is a starting point, not a reproducible tracking system. Each question needs an identifier, exact wording and an intended task. Record whether it asks for a recommendation, a factual explanation, a comparison or troubleshooting help. Those categories help interpret later differences without changing the prompt itself. Also record whether the company is explicitly named. A branded question tests a different situation from a neutral question in which the system may or may not introduce the company on its own.
Build the panel from real information needs rather than prompts designed to produce praise. A question such as why is our company the leading choice already supplies the conclusion it appears to test. An illustrative neutral comparison might instead ask which capabilities matter for a small team choosing appointment software. Use question research to identify useful tasks, then use sampling bias to define what the panel does and does not represent. A stable panel can support trend analysis even when it is not a probability sample of every user question.
Freeze the wording that defines the baseline
Small wording changes can alter the task. Asking for the cheapest option differs from asking for the best value option. Adding a location, budget or industry can change the eligible set substantially. Keep the exact core wording fixed across reporting periods unless there is a documented reason to retire or replace it. If a question becomes obsolete, preserve its history and introduce a new identifier for the replacement rather than silently editing the old row and presenting the series as continuous.
Exploratory prompts remain useful. They can uncover missing facts, misunderstood product features or new ways people phrase a need. Store them in a separate panel so that discoveries do not reshape the baseline after each run. When an exploratory question deserves ongoing tracking, add it through a versioned panel update. Report the old stable subset separately during the transition if needed. This protects the comparison from a subtle form of selection: continually retaining prompts that produce favourable answers and discarding the questions on which the site remains absent.
Fixed prompt panel: Stable identifiers and wording
Documented conditions: Surface, context and timing
Raw observations: Answers, sources and failures
Versioned grading: Apply explicit classifications
Comparable baseline: Retain coverage and variability
Prompt tracking separates collection from interpretation. Raw answers and documented conditions make later comparisons reviewable.
Record the conditions that can change an answer
Preserve the platform, product surface, date, time and any visible model or mode designation. Record relevant search settings, location assumptions, language and whether the run begins in a fresh conversation. If the product uses account-level preferences or conversation context, note the conditions you can actually control. Do not claim two runs are identical merely because the typed sentence matches. The environment may differ in ways that matter, and some platform behaviour may remain unobservable to the tester.
A fresh session is particularly important when the aim is to compare independent answers to the same opening question. A follow-up in an existing discussion tests a different context. If conversational journeys are the intended subject, define that sequence explicitly and preserve every preceding turn. The multi-platform testing guide covers comparisons between surfaces that expose different controls. The tracking record should be honest about missing settings rather than inventing a level of reproducibility that the consumer interface or available tooling does not provide.
Preserve outputs before reducing them to scores
Save the full answer, visible citations and collection status. Where practical, retain a screenshot or export alongside structured text so that later reviewers can see how sources were presented. A classification such as brand present is a derived judgement, not a substitute for the evidence. If the classifier changes or a company alias is discovered, preserved outputs let you reassess past observations consistently. Without them, a historical score may be impossible to audit or compare under a revised rule.
Anthropic's evaluation explanation describes evaluation through inputs and grading. Apply that basic separation to tracking: the prompt and answer are observations; mention detection and quality labels are assessments. Keep the raw record unchanged when correcting an assessment. The citation tracking guide adds the URL-level detail needed when sources are a key outcome. A response can mention a company without citing its site, so preserve enough information to distinguish those events rather than collapsing them into a single visibility flag.
Repeat observations without pretending they are independent people
A single run can capture an unusual answer. Repeated runs help reveal how stable the observed behaviour is under the chosen conditions. Decide the repetition schedule in advance and apply it consistently across the core panel. An illustrative protocol might collect several independent sessions per question during a defined weekly window. The exact schedule should reflect the cost and intended decision, not a universal claim that a particular number of runs makes the result statistically representative of the market.
NIST's sampling guidance connects precision with sampling design and independent replication. For prompt tracking, repeated answers to one question should not be mistaken for many distinct user needs. They provide evidence about variability around that question. Preserve prompt identity when summarising so that a question with extra retries does not accidentally receive more weight. If failures trigger retries, record both the original failure and the retry rule. Otherwise collection reliability can become entangled with the apparent rate of successful brand appearances.
Use the baseline to identify questions worth investigating
Tracking can show that a pattern changed; it usually cannot explain the cause by itself. A drop in owned citations could reflect source changes, platform behaviour, a content release or ordinary variation. Inspect the underlying answers and affected prompts before attributing the movement to an optimisation. The citation volatility guide covers interpretation of changing source appearances. Keep release dates and known platform changes as contextual annotations, but do not turn temporal proximity into proof that one event caused the other.
When the baseline reveals a persistent problem, formulate a bounded next question. Perhaps a product limitation is consistently misstated, or one topic rarely surfaces an available explanation. That can lead to a content correction, an access investigation or an experiment. Preserve the original tracking protocol while doing that exploratory work. A useful baseline is a stable instrument for noticing change, not a set of prompts rewritten until they confirm the desired narrative. Its value grows when another reviewer can reconstruct exactly what was asked, what appeared and how the result was classified.
Sources and further reading
- Anthropic evaluation methodsEvaluation combines inputs with defined grading of outputs.
- NIST sampling schemesSampling design affects precision and systematic error.