TL;DR

The short version.

  • Measure access, eligibility, retrieval and representation as separate stages.
  • A mention, a citation, a recommendation and a referral are different outcomes.
  • Use a fixed prompt sample for trend measurement and a separate exploratory set for discovery.
  • Treat third-party source gaps as a diagnosis—not permission to chase irrelevant placements.
MEASUREMENT DIAGNOSTIC

Find the stage where visibility breaks.

Choose a stage. The lab shows the metric, owner and next action that belongs there.

MEASUREReachable URLs
PRIMARY OWNERTechnical SEO
NEXT ACTIONTest status, robots, rendering, canonicals and internal discovery.
Download the AI visibility audit tracker
01

One visibility score can hide four different failures

A brand can have technically healthy pages and still be absent from category answers. It can also be mentioned often while the answer cites weak, outdated or unrelated sources. Compressing those outcomes into a single percentage makes the dashboard tidy and the diagnosis nearly useless.

Google's guidance says a page must be indexed and eligible to appear with a snippet before it can be shown as a supporting link in AI Overviews or AI Mode. OpenAI separately tells publishers to allow OAI-SearchBot if they want content considered for ChatGPT search summaries and snippets. Access and eligibility are therefore prerequisites, not visibility wins.

The funnel below is a working model: access, eligibility, retrieval and representation. Count each stage separately, then look for the largest loss between adjacent stages. That loss tells the team what kind of work to investigate next.

FIGURE 01The AI visibility funnelIllustrative cohort of 100 useful, public URLs. This is a diagnostic model, not a platform benchmark.
100AccessibleHealthy URL + allowed crawl
78EligibleIndexed and snippet-eligible
44RetrievedRelevant to the prompt set
21RepresentedMentioned or cited accurately
02

Stage one: prove that important evidence is reachable

Begin with the URLs that contain your strongest category evidence: product facts, service definitions, methodology, original research, comparison criteria, pricing context and accountable authorship. Test the live response, not only the source repository.

For each URL, record status code, canonical target, robots access, index directive, renderability and internal-link path. A sitemap can help discovery, but it does not repair a blocked, redirected or low-value page. IndexNow can notify participating engines that a URL changed; it still does not guarantee indexing or visibility.

The access rate is the share of priority URLs that a chosen crawler can retrieve successfully. Keep a separate rate by crawler when policies differ. Do not average Googlebot, OAI-SearchBot and training crawlers into one number because they do not serve the same outcome.

  • Test public responses after CDN and bot-management rules are applied.
  • Keep canonical, noindex and sitemap signals consistent.
  • Record the crawler, URL, result, test date and owner.
  • Escalate policy decisions instead of silently copying a competitor's robots file.
03

Stage two: separate eligibility from ranking or answer inclusion

Eligibility means a system is allowed to consider the page under its documented rules. It does not mean the page will rank, be retrieved or appear in a generated answer. Google explicitly says that meeting requirements does not guarantee crawling, indexing or serving.

Use Search Console index and performance data for Google. In June 2026, Google announced dedicated generative-AI performance views for a subset of sites, while keeping the data included in overall Search reporting. Where the dedicated view is unavailable, preserve the limitation instead of estimating it from screenshots.

Structured data can help systems understand visible facts, but it must match the page. Course, Article, Organization and Breadcrumb markup describe real content; they are not switches that manufacture authority. Validate the markup and the underlying claims together.

04

Stage three: measure retrieval with a stable prompt sample

Retrieval asks whether a relevant owned or third-party page appears in the evidence environment for a buyer question. Because generated responses can vary, one prompt run is an anecdote. Use a stable panel of natural prompt variations across the platforms and geographies that matter to the audience.

Build the sample from buyer decisions: finding options, comparing alternatives, checking trust, evaluating fit and resolving risk. Keep 25 to 40 baseline prompts unchanged for monthly measurement. Maintain a second exploratory set for discovering new language and sources without contaminating the trend line.

For every run, save platform, model or interface, date, exact prompt, brands named, URLs cited and a short accuracy note. A retrieval rate can then be calculated as runs with at least one relevant source divided by total eligible runs. Report sample size beside the percentage.

  • Do not rewrite the baseline after every disappointing result.
  • Separate a cited URL from an uncited brand mention.
  • Capture the full source URL, not only the publisher domain.
  • Mark volatile or unrepeatable observations instead of smoothing them away.
05

Stage four: classify representation and source support

Representation is what the answer says about the brand. Source support is the evidence it exposes. A favorable mention with no useful source, an accurate citation without a brand mention and a recommendation supported by an outdated directory are not equivalent outcomes.

Use four fields for every brand observation: presence, position, claim accuracy and supporting URL. Add a fifth field for recommendation context when the answer explicitly frames the brand as a fit for a use case. This preserves the evidence needed to correct errors and prioritize placement gaps.

The original GEO research established visibility in generated responses as a measurable optimization problem, but results varied by method and domain. Use research as a reason to test disciplined, evidence-based changes—not as a guaranteed uplift claim.

FIGURE 02Do not count every appearance as the same outcomeClassify each observed answer by brand representation and source support.
STRONGER REPRESENTATION ↑
Named, not supported

The brand appears, but the answer cites no useful evidence. Track accuracy risk.

Named + supported

The brand appears accurately and the answer points to a relevant source.

Absent + unsupported

No clear category evidence appears. Improve the information environment.

Absent, source present

A relevant source is retrieved but omits the brand. Investigate the placement gap.

STRONGER SOURCE SUPPORT →
06

Turn third-party source gaps into qualified opportunities

When a relevant comparison, directory, review or trade article repeatedly supports category answers but omits the brand, log a source gap. Then qualify it before outreach. The page should reach the right audience, use defensible criteria, stay reasonably current and offer a truthful reason the brand belongs.

A source gap is not permission to buy every link or force inclusion. Publishers control their pages, and platforms control their answers. The useful action is to prepare accurate evidence, explain the fit and make a respectful request that an editor can independently evaluate.

Track outreach separately from visibility. Record contactability, status, editorial response, placement result and live URL. If the page changes, annotate the date and observe future prompt runs without claiming causation from a single result.

07

Build the dashboard around decisions, not decoration

A practical monthly dashboard needs counts, denominators and movement by stage. Show reachable priority URLs, eligible URLs, prompt runs, relevant retrieved sources, accurate mentions, citations, explicit recommendations and verified referral sessions. Keep the raw observation log behind the summary.

OpenAI says referral traffic from ChatGPT can be tracked in analytics when OAI-SearchBot access is allowed; its referral URLs include attribution that can support analysis. Referral sessions are valuable, but they represent visits—not every unseen mention or recommendation.

Review the funnel with the people who can act on it. Technical owners handle access and eligibility. Content owners improve useful evidence. Brand and earned-media teams address representation and qualified source gaps. One dashboard can coordinate the work without pretending one team controls the entire system.

FIGURE 03A useful dashboard preserves the stagesIllustrative monthly observations from a fixed prompt set; counts are examples, not promised results.
Retrieved sourcesAccurate representations
FEB
MAR
APR
MAY
JUN
JUL
08

Use a 90-day cadence that can survive platform change

In weeks one and two, define the priority URLs, buyer intents, prompt sample, platforms and observation rules. In weeks three and four, establish the technical baseline and first source map. During months two and three, improve owned evidence, pursue qualified outside corrections or inclusions and rerun the fixed sample.

At every review, ask where the largest stage loss occurs, what changed in the evidence environment and which observation is strong enough to act on. Document platform changes and reporting limitations. A smaller repeatable system is more credible than a complex score nobody can reproduce.

The downloadable workbook includes the prompt register, source log, representation rubric, stage dashboard and 90-day review sheet. Use it to keep the measurements auditable while the interfaces continue to evolve.