What is AI visibility?
AI visibility is the observed presence and treatment of a brand, product, person, or source inside generated search answers. It includes mentions, linked citations, recommendation prominence, factual accuracy, sentiment, and the visits or conversions that follow. Because outputs vary, AI visibility should be measured across a fixed prompt cohort and repeated samples—not a single “rank.”
A useful measurement system answers three different questions:
- Exposure: Does the brand appear in relevant generated answers?
- Evidence: Do the answers cite owned or trusted third-party sources that support the brand's claims?
- Business impact: Do those appearances produce qualified visits, assisted journeys, leads, or revenue?
Do not combine the three into one opaque score. They diagnose different problems.
Why traditional rank tracking is not enough
Traditional search results usually present an ordered set of URLs for a query. Generated answers may synthesize multiple subtopics, cite several passages, mention a brand without a link, or produce different wording on another run.
Visibility can vary by:
- platform and search mode;
- model or product version;
- country and language;
- personalization or session context;
- prompt wording and follow-up questions;
- date and index freshness;
- whether live search was used.
That makes “we rank number one in ChatGPT” an unstable claim unless the provider defines the exact prompt, mode, market, time, and observation protocol. Even then, it is a recorded result—not ownership of a position.
Build the measurement cohort
1. Choose business topics
Start with categories tied to revenue or strategic reputation. A B2B engineering studio might use:
- generative engine optimization agency;
- AI visibility audit;
- enterprise RAG development;
- headless Shopify Plus agency;
- technical SEO relaunch.
2. Create prompt families
| Family | Example | Buyer stage |
|---|---|---|
| Definition | “What is GEO and how is it different from SEO?” | Awareness |
| Problem | “How do I measure whether my brand appears in ChatGPT?” | Awareness/consideration |
| Comparison | “GEO agency vs AI visibility software” | Consideration |
| Recommendation | “Recommend GEO agencies for a DACH B2B company” | Decision |
| Implementation | “Which crawler controls ChatGPT Search inclusion?” | Implementation |
| Branded accuracy | “What does AppWebSeo specialize in?” | Validation |
Include prompts where the brand should not be mentioned. A monitoring system that expects the brand in every answer encourages irrelevant optimization.
3. Freeze observation conditions
Record:
- exact prompt text;
- language and target country;
- platform, mode, and model label where available;
- logged-in or neutral session state;
- date and time;
- number of repetitions;
- whether follow-up context was used.
Run prompts in a consistent order or randomize the order and record it. Preserve the full answer and source URLs for auditability.
Core AI visibility metrics
Mention rate
Mention rate shows how often a brand appears in valid responses for the selected cohort.
Mention rate = responses mentioning the brand / valid responses × 100Report it by platform, topic, market, and buyer stage. A global average can hide that a brand is visible for definitions but absent from high-intent recommendations.
Owned citation rate
Owned citation rate shows how often a response links to the brand's own domain.
Owned citation rate = responses linking to an owned URL / valid responses × 100A mention without a citation and a citation without a brand mention are different observations. Keep both.
Citation source coverage
Source coverage counts the distinct owned pages cited during the period. It reveals whether visibility depends on one URL or whether the whole topic cluster contributes.
Track:
- unique owned URLs cited;
- citation frequency per URL;
- prompt families served by each URL;
- stale or redirected URLs still appearing;
- third-party pages that cite or describe the brand.
Competitive share of voice
Define a closed competitor set before collecting the data.
Share of voice = brand mentions / mentions of all tracked brands × 100State whether multiple mentions can occur in one response. Do not silently change the competitor set between reporting periods.
Recommendation prominence
Not every mention has equal value. Use a documented rubric:
| Label | Definition |
|---|---|
| Recommended | Brand is explicitly suggested for the use case |
| Included | Brand appears in a relevant list without strong endorsement |
| Referenced | Brand is mentioned as an example or source |
| Negative | Brand appears with a material warning or criticism |
| Irrelevant | Mention does not answer the target intent |
Avoid invented precision. Human classification with clear examples is more defensible than a 97.3 “prominence score” with no validation.
Factual accuracy
Create a list of verifiable brand facts: company name, services, locations, pricing model, certifications, named team members, and product capabilities. Mark each observed claim as accurate, outdated, unsupported, or contradictory.
Accuracy rate can be reported only against facts that actually appear in the sample. Preserve the evidence for every error.
Sentiment
Use a simple rubric—positive, neutral, mixed, or negative—and define it before review. For high-stakes analysis, have a second reviewer classify a sample and resolve disagreements.
Sentiment is context, not business value. A neutral citation in a technical guide may be more valuable than a positive but irrelevant mention.
Platform data and what it proves
Google Search Console
Google's Generative AI performance report is being rolled out to eligible properties. It reports impressions from AI Overviews and AI Mode and supports page, country, device, and date dimensions.
Use it to measure whether owned URLs appeared in Google's generative features. It does not provide a universal cross-platform citation score, and availability may vary by property.
Keep ordinary Search metrics beside it:
- non-brand impressions and clicks;
- landing pages;
- country and device;
- conversions and assisted conversions;
- indexed coverage and technical errors.
ChatGPT referrals
OpenAI says publishers can track ChatGPT Search referrals because outbound URLs include utm_source=chatgpt.com. Preserve query parameters and classify the source in analytics.
Measure:
- sessions and users;
- landing pages;
- engaged sessions;
- lead or purchase conversion;
- qualified pipeline and revenue where CRM attribution allows it.
Referral traffic is lower-funnel evidence than a manual mention observation, but it still captures only visits that click through.
Perplexity and other sources
Track referrers that appear in your analytics and verify the current platform documentation before creating parsing rules. Platform behavior and referral labels can change.
Manual or third-party monitoring is still needed for unclicked mentions and citations outside Google. Treat vendor metrics as that vendor's observations, not internal platform ranking data. Google explicitly warns that third-party tools do not have access to its internal ranking systems.
Connect visibility to revenue
Use a simple evidence ladder:
| Level | Evidence | Interpretation |
|---|---|---|
| 1 | Brand mentioned | Exposure observed |
| 2 | Owned or trusted source cited | Evidence selected |
| 3 | Referral session | User visited the site |
| 4 | Engaged or assisted session | Visit contributed to evaluation |
| 5 | Qualified conversion | Commercial intent confirmed |
| 6 | Pipeline or revenue | Business outcome recorded |
Do not claim revenue impact from a change in mention rate alone. Use analytics and CRM evidence, state the attribution model, and acknowledge assisted journeys.
A minimum viable dashboard
| Section | Metric | Segment |
|---|---|---|
| Cohort | Valid responses and repetitions | Platform, market, language |
| Exposure | Mention rate and prominence | Topic, buyer stage |
| Evidence | Owned citation rate and source coverage | URL, source type |
| Competition | Share of voice | Fixed competitor set |
| Trust | Accuracy and sentiment | Fact category, reviewer |
| Search | Google generative impressions | Page, country, device, date |
| Traffic | AI referral sessions and engagement | Source, landing page |
| Business | Qualified leads, pipeline, revenue | Service, market, attribution model |
Add a methodology tab with prompt text, run dates, platform labels, classification rules, and known limitations.
How often should you measure?
- Technical and referral monitoring: continuous where logs and analytics allow it.
- Prompt cohort: monthly for a stable strategic view; more often during a controlled experiment.
- Executive reporting: monthly or quarterly, depending on sales cycle.
- Content experiments: baseline before the change, then a defined follow-up window such as 30, 60, and 90 days.
Avoid changing prompts every week. Add new prompts as a separate cohort so the trend line remains interpretable.
Experiment design for GEO changes
- State the hypothesis: for example, “publishing a sourced comparison page will increase owned citation coverage for comparison prompts.”
- Freeze the prompt cohort and observation conditions.
- Capture the baseline across repeated runs.
- Record the exact pages and technical changes deployed.
- Wait for crawling and indexing; check logs and platform controls.
- Re-run the same cohort.
- Compare platform-specific results and business metrics.
- Report alternative explanations and sample limitations.
Do not attribute every change to the implementation. Models, indices, competitors, and answer formats can change during the same period.
Reporting mistakes to avoid
- Calling a proprietary visibility score an official platform metric.
- Mixing mentions, citations, impressions, and clicks into one number.
- Reporting one screenshot as a persistent ranking.
- Hiding the prompt list, country, language, or sample size.
- Changing prompts or competitors between periods without restating the baseline.
- Treating sentiment as conversion.
- Counting
llms.txtdeployment or schema validation as an outcome. - Claiming causation without a controlled comparison.
- Promising a fixed uplift before baseline measurement.
Starting template
For each observation, store these fields:
date | platform | mode/model | country | language | session
prompt_id | prompt_text | valid_response
brand_mentioned | prominence | owned_url_cited
all_source_urls | competitors_mentioned
accuracy_label | sentiment_label | reviewer | notesFor analytics, add:
source | landing_page | sessions | engaged_sessions
conversions | qualified_leads | pipeline | revenue
attribution_model | reporting_periodBottom line
AI visibility is not one rank. Measure exposure, evidence selection, traffic, and commercial impact as separate layers. Use a stable prompt cohort, retain the full answers and citations, connect observations to first-party analytics, and state uncertainty. That produces a dashboard leaders can act on and analysts can reproduce.
The GEO audit checklist shows how to collect the required technical and editorial evidence. The complete GEO guide explains which improvements to test after the baseline exists.