What a GEO audit should answer
A GEO audit should show whether priority content can be accessed, whether it is useful and verifiable enough to serve as a source, how the brand currently appears in generated answers, and whether that visibility produces qualified business activity. It should not promise future citations or compress every finding into a mysterious “AI ranking score.”
The audit has seven parts:
- Goals and prompt baseline.
- Search eligibility and crawler access.
- Content usefulness and extractability.
- Evidence, authorship, and entity accuracy.
- Structured data and machine-readable surfaces.
- Citation and competitor patterns.
- Analytics, conversion, and experiment design.
Use this checklist for Google AI Overviews and AI Mode, ChatGPT Search, and Perplexity. Record platform-specific evidence rather than assuming one engine's behavior applies to the others.
Before the audit: define the scope
Do not begin by scanning every page. Begin with the decisions the content is expected to influence.
- [ ] Identify 3–5 business-critical services, products, or categories.
- [ ] Define the countries and languages that matter.
- [ ] Select 3–5 direct competitors for each category.
- [ ] Identify the primary conversion: audit request, demo, qualified lead, purchase, or another event.
- [ ] Choose the owned pages that should support each category.
- [ ] Agree on a review window and who will validate technical, editorial, and business findings.
Build a prompt cohort
Create 20–50 prompts that reflect real research behavior. Use the same wording for every baseline and follow-up run.
| Prompt type | Example | What it tests |
|---|---|---|
| Category | “What is generative engine optimization?” | Definition and topical authority |
| Problem | “How can a B2B brand measure visibility in ChatGPT?” | Problem-solution relevance |
| Comparison | “GEO vs SEO for an international SaaS company” | Decision support |
| Recommendation | “Which GEO agencies work with DACH enterprise companies?” | Brand consideration |
| Implementation | “How should OAI-SearchBot be configured?” | Technical expertise |
| Branded accuracy | “What services does AppWebSeo provide?” | Entity and factual consistency |
For each observation, record platform, model or mode where available, language, location context, date, session state, answer, linked sources, brand mention, competitors, and factual errors.
A single output is not a rank. Generated answers can vary between runs. Use repeated observations and report the sample size.
Part 1: Google Search eligibility
Google states that its generative AI features use the core Search index and quality systems. Start with ordinary Search requirements.
- [ ] Priority URLs return
200and are not blocked byrobots.txt. - [ ] Pages are not marked
noindexand are eligible to show a snippet. - [ ] Canonicals point to the intended indexable URL.
- [ ] Pages appear in the XML sitemap with accurate modification dates.
- [ ] Main content is visible in rendered HTML.
- [ ] Internal links connect hubs, supporting articles, and commercial pages.
- [ ] Google Search Console shows the expected indexed canonical.
- [ ] The Search generative AI control is not excluding the property when inclusion is desired.
- [ ] The dedicated Generative AI performance report is reviewed if the property has access.
Google's AI optimization guide says there is no special AI markup, required text length, or chunking formula.
Part 2: AI search crawler access
Crawler names have different purposes. Audit the actual documented search crawler rather than allowing every bot under an “AI” heading.
ChatGPT Search
- [ ]
OAI-SearchBotis not blocked when ChatGPT Search inclusion is desired. - [ ] WAF, CDN, and bot management rules allow OpenAI's published searchbot IP ranges.
- [ ] The policy for
GPTBotis documented separately because it relates to potential model training. - [ ] The team understands that
ChatGPT-Usermay be used for user-initiated visits and is not the automatic Search crawler.
Source: OpenAI crawler documentation.
Perplexity
- [ ]
PerplexityBotis not blocked when Perplexity search inclusion is desired. - [ ] Published Perplexity IP ranges are not rejected by infrastructure.
- [ ] Server logs are checked for successful crawler responses and unexpected
403,429, or5xxpatterns.
Source: Perplexity crawler documentation.
Policy check
- [ ] Search inclusion, model training, and user-initiated fetching are treated as separate decisions.
- [ ] Legal, privacy, and content teams approve the policy.
- [ ]
robots.txtand edge security implement the same decision. - [ ] Changes are dated and documented.
Crawler access supports discovery. It does not guarantee selection or citation.
Part 3: Content usefulness and extractability
Review the priority pages as a skeptical buyer and as an editor.
- [ ] The page directly answers its primary question near the beginning.
- [ ] The content adds first-hand experience, original data, a method, a template, or a decision framework.
- [ ] Important terms are defined in plain language.
- [ ] Headings reflect reader tasks rather than keyword variations.
- [ ] Tables are used for genuine comparisons.
- [ ] Processes use ordered steps.
- [ ] Each key paragraph can be understood without an unexplained pronoun or missing qualifier.
- [ ] The page covers limitations and cases where the recommendation does not apply.
- [ ] The content has a visible publication or update date in metadata.
- [ ] The next step is appropriate for the reader's buying stage.
Do not optimize for an invented token window. Google explicitly says it can understand multiple topics on a page and does not require publishers to divide content into tiny AI-oriented chunks.
Part 4: Evidence and claim review
Create a claim ledger for every priority page.
| Claim | Source | Date checked | Scope/conditions | Action |
|---|---|---|---|---|
| A platform uses a named crawler | Platform documentation | YYYY-MM-DD | Platform and crawler purpose | Keep and cite |
| A client result improved | First-party analytics or case record | YYYY-MM-DD | Client, period, sample | Publish with permission |
| A tactic increases citations by a fixed percentage | No verified source | — | Unknown | Remove or test |
- [ ] Every statistic links to an original or authoritative source.
- [ ] The source actually supports the wording used on the page.
- [ ] Dates, products, crawler names, and standards are current.
- [ ] Case results include baseline, period, scope, and material changes.
- [ ] Correlation is not described as causation.
- [ ] Superlatives such as “best,” “leading,” or “top 1%” have evidence or are removed.
- [ ] Guarantees concern deliverables under the provider's control, not third-party search outcomes.
- [ ] Authors are real people or an accurately identified organization.
Part 5: Entity and structured data review
Structured data should reflect visible facts. It can help systems understand page entities and qualify pages for supported Search features, but Google says there is no special schema required for generative AI search.
- [ ] Organization name, legal identity, URL, logo, and contact facts are consistent.
- [ ] Service names and descriptions agree across navigation, landing pages, articles, and public profiles.
- [ ] Author markup matches the visible author.
- [ ]
ArticleorTechArticledata uses accuratedatePublishedanddateModifiedvalues. - [ ] Breadcrumb markup matches visible navigation.
- [ ] Structured data validates and does not describe hidden content.
- [ ] Only supported, relevant Schema.org types and properties are used.
- [ ] Non-versioned URLs such as
https://schema.org/Organizationare used. - [ ] “Schema 3.0” or other invented AI schema layers are removed.
Schema.org's public release history lists version 30.0 as the current release at the time of this audit template. Publishers generally use the ordinary non-versioned vocabulary URLs.
Part 6: llms.txt review
Treat llms.txt as an optional navigational file for agents that choose to use the proposal—not as a ranking factor.
- [ ] The file describes the business accurately.
- [ ] Links resolve to canonical public pages.
- [ ] Claims match visible website content.
- [ ] Stale services, prices, authors, and results are removed.
- [ ] The file is not presented as a Google visibility factor.
- [ ] Deployment is not counted as a citation result.
Google says Search ignores llms.txt. The specification calls itself a proposal. A valid file may still be useful for documentation or agent workflows that deliberately request it.
Part 7: Citation and competitor analysis
Run the agreed prompt cohort and create a source map.
- [ ] Record owned citations separately from unlinked brand mentions.
- [ ] Record third-party pages that mention the brand.
- [ ] Identify recurring competitor sources.
- [ ] Classify each cited source: official documentation, publisher, directory, review, forum, social, academic, or other.
- [ ] Note whether the cited passage directly supports the generated claim.
- [ ] Record factual inaccuracies and outdated descriptions.
- [ ] Identify questions for which no current source gives a strong answer.
- [ ] Prioritize gaps by buyer importance, not by citation count alone.
Do not copy a competitor's format blindly. Determine what uncertainty its page resolves and whether you can contribute stronger evidence.
Part 8: Measurement and conversion
- [ ] Google Generative AI impressions are exported where Search Console provides the report.
- [ ] AI referral sources and UTM parameters are preserved in analytics.
- [ ] ChatGPT referrals using
utm_source=chatgpt.comare monitored. - [ ] Landing page, engagement, conversion, and qualified lead data are connected.
- [ ] Brand mention rate and owned citation rate are separate metrics.
- [ ] Sentiment and factual accuracy use a documented classification rubric.
- [ ] The same prompt cohort is rerun on a fixed cadence.
- [ ] Reports state platform, market, sample size, and limitations.
Use the AI visibility measurement guide for formulas and a reporting template.
How to prioritize findings
Classify each finding by evidence, impact, and control.
| Priority | Definition | Examples |
|---|---|---|
| P0 | Blocks eligibility or creates material trust risk | noindex, wrong canonical, blocked search crawler, fabricated statistic, false guarantee |
| P1 | Prevents a priority page from becoming a strong source | Commodity content, missing evidence, inconsistent entity facts, weak comparison |
| P2 | Improves clarity or coverage after fundamentals are fixed | Better table, additional FAQ, secondary profile update |
| Monitor | Evidence is insufficient or platform behavior is changing | Experimental file formats, undocumented crawler assumptions |
An audit score can help organize work, but it is not an AI-platform score and should never be presented as a prediction of citations.
Recommended audit deliverables
A useful audit should produce:
- A documented scope, prompt cohort, competitors, markets, and date.
- A technical eligibility report with reproducible evidence.
- A page-level claim ledger.
- A source and citation map.
- An entity inconsistency list.
- A prioritized 30/60/90-day backlog with owners.
- A measurement specification and baseline export.
- A list of assumptions that still require testing.
Bottom line
A GEO audit is a disciplined evidence review, not a magic score. Fix eligibility first, unsupported claims second, and source quality third. Then measure whether visibility, referrals, and qualified business outcomes change under a repeatable protocol.
For the strategic context behind this checklist, read the complete GEO guide. AppWebSeo can also run the technical, editorial, and measurement review as one scoped diagnostic.
Primary sources
- Google: Optimizing for generative AI features
- Google Search Console: Generative AI performance report
- Google: Search generative AI control
- OpenAI: Overview of OpenAI crawlers
- OpenAI: Publishers and Developers FAQ
- Perplexity: Crawler documentation
- Google: Structured data guidelines
- Schema.org: Releases
- The `llms.txt` proposal