2026-10-05

Live vs sampled: how we label every AI answer

An AI-visibility metric is only trustworthy if its provenance is. Here is exactly when a run is called live, when it is sampled, and why aggregates blend them.

If a tool tells you "ChatGPT mentions you 38% of the time", the first question should be: how do you know? Asking an LLM to simulate an answer engine is a reasonable measurement technique — it is not the same as querying the consumer product, and pretending otherwise is how dashboards start lying.

So AIEO.ee labels every single run, and the label is the product.

The two labels

live is earned, not configured. Three engines support a real AI-search backend — ChatGPT via OpenAI web search, Perplexity via the direct Sonar API, Google AI Overview via Gemini grounding. But configuring a key isn't enough: a run only counts as live when the backend actually returned search evidence for that run. A live call that comes back without grounding gets demoted to sampled, automatically. A live run's citations are real sources the backend retrieved.

sampled is everything else: engines without a live backend, a live engine whose key isn't configured, or a grounded call with no search evidence. An LLM produces its best estimate of what the answer engine would say. That is a directional signal — useful, trackable over time, and clearly not a measurement of the consumer product. We never present it as one.

Where you can verify the label

  • On every run in the run list — the raw answers with their provenance. That list is the audit trail: any aggregate number can be traced back to labeled runs.

The aggregates — the summary meters, share of voice, the trend series and the per-engine breakdowns — currently include live and sampled runs together; there is no source filter on them today.

Where the labels blend

If your mix is mostly sampled, a 38% SOV is a weaker claim than one built on live runs, and we think you should know that before quoting it internally.

Why not filter silently? Deciding for you what counts would be the same dishonesty in the other direction. Per-source views are on the roadmap; until then, verify anything that matters against the labeled run list.

What this means in practice

  1. Configure live keys for the engines that support them — the numbers get materially stronger.
  2. When a run matters (a report, a decision), open it and check the label.
  3. Treat sampled trends as direction, and treat "live" as the only number you call a measurement.

The data sources page documents the full matrix, engine by engine.