Methodology
How SeenAndCited Measures AI Visibility
AI answers are variable, contextual and constantly changing. SeenAndCited uses repeated, evidence-backed measurement across relevant questions and AI engines to build a more reliable view of how a business is being understood, cited and recommended.
This page explains what we measure, what each measurement means, and — just as importantly — what the evidence does not prove.
Why AI visibility needs a different measurement approach
Traditional search produces a stable, inspectable list of results. AI answers do not work that way. The same question can be answered differently on different days, different engines answer from different data, and the sources shown alongside an answer vary from run to run.
One prompt run is therefore weak evidence. Anything useful has to come from repeated observation of relevant questions, recorded with enough context to be compared later.
Measurement starts with the questions people are likely to ask
We do not simply check whether a brand appears somewhere in AI. We measure a portfolio of questions that reflect how customers actually describe their problem — informed by the business proposition, its services and products, the market and customer intent, and competitor context where relevant.
The quality of that question portfolio materially affects how useful the measurement is. A portfolio of questions nobody asks will produce numbers nobody should act on.
Automatically generated questions start out as provisional. They are proposed, reviewed and confirmed before they become part of the monitored portfolio that ongoing measurement is built on.
What a probe is
A probe is one measurement attempt: one monitored question, asked of one AI engine, once. Asking five questions of two engines is ten probes. Repeating that observation to reduce variance multiplies the probes again.
This is why plan capacity is expressed in probes per month rather than "questions" — it is the unit of measurement work actually performed.
Different AI engines can produce different answers
Answer engines retrieve, rank and cite differently, so being cited by one says little about the others. We measure across the engines included in a plan; engine coverage is one of the things that differs between plans.
Where a result comes from a narrower engine set, we say so rather than presenting it as a view of AI as a whole.
One answer is not enough
Because answers vary, higher measurement tiers repeat the same observation several times before drawing a conclusion. Repeated sampling reduces the risk of treating a single, unrepresentative answer as the truth.
Lighter sampling still produces useful directional evidence; deeper sampling produces more stable evidence and is what we rely on for confident movement claims. See current engine coverage and sampling depth by plan.
A measured zero is not the same as no measurement
Every number we show carries an explicit state. A result is either measured, not yet measured, based on insufficient evidence, still in progress, or unavailable because a provider or source could not be reached.
This matters more than any single metric. If an engine fails, or evidence cannot be retrieved, we do not record that as "not cited". Absence of evidence is reported as absence of evidence, with a plain-language reason, and it is excluded from rates rather than silently counted as a zero.
Citation Rate — what it does and does not mean
Citation Rate is the share of your measured monitored questions where your site was cited with evidence linking the answer to a page you own.
Questions with no valid observation in the period are not part of that calculation. They are reported separately as measurement coverage, so a partially completed measurement run can never depress the rate.
Citation Rate is a metric over your measured question, engine and sample portfolio. It is not a measure of "how visible you are in AI" overall, and it should never be read that way.
We distinguish a citation from a mention
Not every source that appears alongside an answer is evidence about your business. We separate sources you own from third-party sources, and we only treat a source as supporting evidence when the underlying response actually supports that interpretation.
Sources that merely appeared while an engine answered a question in your market are recorded as market context, not as evidence about you.
Independent Sources are earned, not inferred
A third-party site only counts as an Independent Source once a page on that site has been checked and found to reference the business — by linking to it, naming it, or identifying it in structured data.
Appearing in an AI answer is not enough to qualify. Sources awaiting verification are shown as awaiting verification rather than being counted early, which means this number is usually lower — and more defensible — than a raw count of hosts seen in answers.
AI Understanding
AI Understanding examines how a business appears to be interpreted — its category, audience, services and products, positioning, and any ambiguity or missing signals — based on public evidence and measured AI responses.
A confident interpretation is not the same as a correct one. The business owner remains the authority on whether the interpretation is accurate and complete, which is why the product asks you to confirm or correct it rather than treating it as settled.
AI Readiness
AI Readiness covers machine-verifiable technical and public signals that affect how accessible and interpretable a site is to AI systems — things that can be checked directly rather than estimated.
Readiness is a supporting condition, not a guarantee. A technically excellent site can still go uncited, and improving readiness removes obstacles rather than producing citations by itself.
Change only matters when the measurements are comparable
We do not compare arbitrary historical observations. A previous measurement only becomes a baseline when it was an authoritative measurement of the same kind, covering a sufficiently similar question set and a comparable set of engines, taken far enough apart to represent a genuinely earlier period.
When no comparable baseline exists, we say "starting position" rather than inventing a change against an implied zero. We also suppress apparent movement caused by a definition correction on our side, so a correctness fix can never be presented as a real-world loss.
Provisional and orientation samples — including the AI Visibility Snapshot — are never used as the authoritative historical baseline for ongoing measurement.
Why the AI Visibility Snapshot is different
The Snapshot is a one-time orientation exercise: a small set of automatically generated questions, run once against a single answer engine, combined with public evidence from the website.
It reports which of its sample questions were successfully measured and which were not, and it is intentionally provisional. It is not a Citation Rate, not a confirmed monitored portfolio, and it is kept isolated from ongoing measurement.
The Snapshot is designed to provide useful initial evidence with minimal input, so the relevance of its sample questions depends on how clearly the site communicates the business, its services and its market. The full Evaluation develops a richer business context and a broader, confirmed measurement portfolio.
Measuring what happens after work is completed
When eligible work is completed, we schedule follow-up measurement at roughly one week and one month afterwards, and compare it against the pre-change baseline for the same questions. We require several post-change observations before stating an outcome, for the same reason we sample repeatedly in the first place.
Outcome measurement shows what changed after the work was completed. It provides evidence of subsequent movement, but AI visibility is influenced by multiple external factors and should not be interpreted as proof of sole causation.
Model and provider changes, competitor activity, changes to third-party sources, and general movement on the web all affect results. We present outcomes as evidence of what followed, alongside the context needed to judge it.
What this methodology does not claim
- That one AI answer represents what every person will see.
- That an answer engine will produce the same answer twice.
- That technical readiness guarantees citations.
- That every third-party source appearing in an answer is authority evidence.
- That Snapshot results are canonical ongoing measurement.
- That movement after a completed task proves sole causation.
- That visibility on one engine implies visibility on another.
These limits are the reason for the design. Because AI measurement is inherently variable, we use explicit evidence states, repeated measurement and comparable baselines rather than presenting the data as more precise than it is.
Methodology summary
- 1Relevant questions
- 2Multiple AI engines
- 3Repeated sampling
- 4Verified evidence
- 5Comparable baselines
- 6Change over time
- 7Work completed
- 8Outcome re-measurement