A practical 2026 framework for auditing how a brand is discovered, described, mentioned, cited and recommended across ChatGPT, Google AI features, Perplexity and other answer engines—without pretending one prompt is a ranking report.
What is an AI search visibility audit?
An AI search visibility audit is a structured review of whether—and how—your brand, website, products, people and evidence appear inside generative search experiences. Depending on the market and buyer journey, that can include ChatGPT Search, Google AI Overviews and AI Mode, Perplexity, Gemini-connected experiences, Copilot and other answer surfaces.
The audit should answer more than “Are we mentioned?” A useful audit investigates a chain of questions:
- Eligibility: Can the relevant search and retrieval systems access and index your content?
- Understanding: Do they correctly understand who you are, what you sell and who it is for?
- Retrieval: Does your content or third-party evidence enter the candidate source set?
- Selection: Are you cited, mentioned or recommended for prompts that matter commercially?
- Accuracy: Are the facts, pricing, positioning and comparisons correct?
- Competition: Which competitors appear instead, and which sources support them?
- Stability: Does the observation persist across repeated runs and time?
- Outcome: Can any of this be connected to visits, branded demand, assisted conversions or sales?
This is why an AI visibility audit sits above a basic AEO readiness checker. A checker can inspect your site. An audit must inspect the answer environment around your category.
AI search audit vs traditional SEO audit
A traditional SEO audit asks whether pages can be crawled, indexed, ranked and converted. Those fundamentals remain essential. In fact, Google explicitly states that there are no additional technical requirements for appearing as a supporting link in AI Overviews or AI Mode beyond being indexed, eligible for Search and eligible for snippets. Google’s 2026 generative AI Search guidance also says these experiences remain rooted in core Search ranking and quality systems.
That means AEO/GEO does not replace SEO. A serious AI audit starts with SEO eligibility, then measures layers classic rank tracking does not capture well: answer inclusion, recommendation framing, citation sources, competitor co-occurrence, factual representation and run-to-run variability.
| Dimension | Traditional SEO audit | AI search visibility audit |
|---|---|---|
| Primary object | URLs and rankings | Entities, answers, sources and recommendations |
| Technical eligibility | Crawl, index, canonical, rendering | Same baseline + surface-specific crawler/search eligibility |
| Measurement | Rank, clicks, impressions, conversions | Mention rate, citation rate, recommendation rate, source share, factual accuracy, referrals |
| Query model | Keyword set | Prompt/intention set with paraphrases and buyer contexts |
| Competition | Who outranks you? | Who is recommended, cited or used as evidence instead? |
| Stability | Rankings fluctuate, but snapshots are often useful | Repeated sampling is essential because identical prompts can produce different answers |
| Off-site evidence | Links, authority, reputation | Also asks which third-party pages the answer engine is actually using |
Google does not require a secret “AI SEO” file
Google’s current Search documentation says no special markup, AI text file or llms.txt file is required to appear in Google Search generative AI features. For Google specifically, foundational SEO, crawlability, index eligibility, visible high-quality content and accurate structured data remain the core technical baseline. If an audit marks “missing llms.txt” as a critical Google AI error, challenge the methodology.
Why one ChatGPT prompt is not an audit
Generative search is probabilistic. Identical or near-identical prompts can produce different brands, sources, answer order and citations across repeated runs. Surface changes, search activation, context and time can all affect the response.
A 2026 study titled Don’t Measure Once: Measuring Visibility in AI Search evaluated repeated AI-search observations and concluded that visibility should be treated as a distribution rather than a single-point outcome. Another 2026 paper, Quantifying Uncertainty in AI Visibility, found substantial citation variability across repeated samples and warned that single-run metrics can look more precise than the underlying behavior supports.
This changes the audit methodology. A screenshot saying “ChatGPT did not mention you” is an observation. It is not evidence that your brand has zero ChatGPT visibility.
The 12 layers of a real AI visibility audit
1. Search and crawler eligibility
Before measuring recommendations, verify that the relevant surfaces can access the content they may need.
For ChatGPT Search, OpenAI’s current publisher documentation says publishers should allow OAI-SearchBot if they want their content included in ChatGPT summaries and snippets. OpenAI separates this from GPTBot, which is related to potential training use. For Perplexity, official crawler documentation says PerplexityBot is designed to surface and link websites in Perplexity search results, while Perplexity-User can access pages in response to user actions.
For Google AI Overviews and AI Mode, the baseline is different: Google says pages must be indexed and eligible to appear in Search with a snippet, with no additional AI-specific technical requirement.
Your audit should therefore test actual crawler and robots.txt policy rather than applying one generic “AI bots allowed” checkbox.
2. Index and discoverability health
AI visibility cannot be separated cleanly from search visibility. Check whether key commercial, comparison, documentation and research pages are indexable, internally linked, canonicalized correctly and discoverable through the search systems that feed generative experiences.
For Google, use Search Console data. For Bing-dependent ecosystems, include Bing index health where relevant. The audit should distinguish not indexed, indexed but not retrieved, and retrieved but not selected. Those are different problems.
3. Entity clarity
Audit whether the open web consistently answers basic questions about the entity:
- What is the brand’s canonical name?
- What category does it belong to?
- What products or services does it offer?
- Who is the founder, author or responsible organization?
- Which geography does it serve?
- What are the current pricing, features and policies?
- Which external profiles corroborate those facts?
Then compare those facts with what AI systems actually say. Entity consistency is not useful if the model confidently repeats an outdated price or describes the wrong target customer.
4. Content extractability
Check whether important facts exist in visible, accessible text and are easy to verify. Tables, direct answer paragraphs, labelled comparisons, dates, authorship, definitions and clearly separated claims can make content easier for retrieval systems and humans to interpret.
Do not turn this into a ritual of arbitrary chunk sizes. Google’s 2026 guidance explicitly warns site owners to prioritize valuable, non-commodity content rather than generic “AEO hacks.” For Google specifically, special AI files or forced chunking are not required.
5. Structured data accuracy
Structured data can reinforce entity relationships and factual consistency when it matches visible content. Audit the existing JSON-LD graph for contradictions, duplicated entities, stale offers, mismatched organization names, orphaned authors and unsupported properties.
The goal is not to maximize schema count. The goal is to make the machine-readable representation agree with the page humans see. See the JSON-LD @graph method.
6. Prompt-set design
A serious audit does not use ten random prompts invented by the analyst. Build the prompt set from the buyer journey and real information demand.
A practical prompt taxonomy can include:
| Intent | Example pattern | What it tests |
|---|---|---|
| Category discovery | “Best [category] tools/services for…” | Whether the brand enters the consideration set. |
| Problem discovery | “How should I solve [problem]?” | Whether your expertise/content is used before vendor selection. |
| Shortlisting | “Which providers should I consider for…?” | Recommendation presence and relative prominence. |
| Comparison | “Brand A vs Brand B for [context]” | Feature accuracy, framing and comparative evidence. |
| Pricing / fit | “What does [service] cost and who is it for?” | Commercial fact accuracy and suitability descriptions. |
| Trust | “Is [brand] reputable / experienced in…?” | Off-site consensus, credentials and reputation sources. |
| Implementation | “How do I implement [technical task]?” | Whether docs, guides or technical authority are cited. |
Commercial providers already vary widely here: some audits use around 30–40 reviewed prompts, while deeper services publicly describe 50–100 prompts across several engines. The number matters less than whether the prompt set covers the decisions your buyers actually make.
7. Multi-engine coverage
Do not merge all AI surfaces into one universal ranking. ChatGPT, Google AI features and Perplexity have different retrieval paths, interfaces, source pools and citation behaviors.
At minimum, record results by engine separately. If the target audience uses multiple surfaces, your report should show both the cross-platform pattern and the platform-specific differences.
8. Mention, citation and recommendation separation
These outcomes are not interchangeable.
- Mention: the brand appears in answer text.
- Citation: your site or another page about you is used as a linked source.
- Recommendation: the model actively positions the brand as a suitable choice for the user’s need.
- Prominence: where and how strongly the brand is surfaced inside the answer.
9. Repeated sampling and uncertainty
For each important prompt, record more than one observation where budget and tooling allow. The audit should document:
- Engine and product surface.
- Date and time.
- Exact prompt.
- Run number.
- Whether web/search mode was active where observable.
- Brands mentioned.
- Brands recommended.
- Sources cited.
- Position/prominence.
- Incorrect or stale facts.
Then report rates with the sample size attached. “Mention rate: 42% (21/50 observations)” is more honest than “AI Visibility Score: 42” with no method.
10. Competitive source analysis
When competitors are repeatedly recommended, inspect the evidence layer behind them.
Ask:
- Are their own product/service pages cited?
- Are review sites doing the heavy lifting?
- Are Reddit/forum discussions shaping the answer?
- Are industry publications validating the category position?
- Do they have comparison pages, docs, studies or original data you lack?
- Are the same third-party domains recurring across multiple engines?
This is where AI visibility auditing becomes strategic. The fix may not be “rewrite our landing page.” It may be “we have no independent evidence in the source ecosystem that answer engines use for this buying question.”
11. Factual accuracy and brand representation
Visibility is not automatically good. An answer can mention you frequently while describing the wrong pricing, outdated product name, incorrect location, unsupported capability or a weakness you no longer have.
Audit the factual fields that matter commercially:
- Price and pricing model.
- Service/product category.
- Target customer.
- Geography.
- Features and exclusions.
- Founder/leadership.
- Certifications and credentials.
- Support and policy information.
- Competitor comparisons.
For every incorrect claim, record the answer, source if present, canonical fact and likely correction path.
12. Referral and business measurement
Not every AI exposure produces a measurable click. But when referral data is available, capture it.
OpenAI’s current publisher FAQ says ChatGPT referral links automatically include utm_source=chatgpt.com, allowing publishers to track inbound ChatGPT search traffic in analytics platforms such as Google Analytics. Google reports AI Overviews and AI Mode traffic within Search Console and has introduced dedicated generative-AI reporting in its current documentation.
Use referral traffic as one outcome—not the definition of visibility. A recommendation can influence a later branded search, direct visit or offline buying decision without a clean last-click trail.
What metrics should an AI visibility audit report?
A defensible audit should provide raw counts and rates before compressing anything into an executive score.
| Metric | Example definition | What it tells you |
|---|---|---|
| Mention rate | Brand mentioned ÷ total observations | How often you enter the answer at all. |
| Citation rate | Your domain cited ÷ total observations | How often your site becomes explicit evidence. |
| Recommendation rate | Brand recommended ÷ relevant commercial observations | How often you enter the buyer shortlist. |
| Prompt coverage | Prompt clusters with ≥1 meaningful visibility event | Where visibility exists across the buyer journey. |
| Source share | Your domain citations ÷ all observed citations in scope | Your evidence presence relative to the source market. |
| Prominence | Position / first-mentioned / shortlist inclusion | Whether you are central or incidental. |
| Fact accuracy | Correct audited claims ÷ audited claims | Whether visibility represents the real business. |
| Volatility | Run-to-run change across repeated observations | How stable the apparent visibility is. |
| AI referrals | Tracked sessions / conversions from identifiable AI sources | Observable traffic and revenue contribution. |
Always publish the denominator
“We appeared 18 times” is not interpretable without knowing whether the audit ran 20, 200 or 2,000 observations. Every visibility percentage should be traceable to the underlying prompt set, engine set, run count and date range.
How many prompts should an AI visibility audit test?
There is no universal number because the right sample depends on category breadth, markets, products, buyer roles and how many engines you are testing.
Public commercial audits currently illustrate the range: Tilio describes a reviewed 30–40 prompt set across ChatGPT, Google AI Overviews and Perplexity, while Veza Digital describes 50–100 prompts across six AI platforms with a 30-day tracking period. Those are examples of service design, not scientific minimums.
For a focused service business, I would rather test 30 carefully designed buyer prompts with repeat observations than 500 generic prompts that nobody commercially cares about.
For a broad SaaS or ecommerce category, the prompt matrix may need to grow across:
- Products.
- Features.
- Personas.
- Industries.
- Competitors.
- Use cases.
- Pricing tiers.
- Geographies.
- Languages.
Free AEO checker vs professional AI visibility audit
A free tool can be extremely useful if its promise is narrow. For example, a checker can inspect:
- robots.txt and crawler directives;
- indexability;
- structured data presence;
- basic entity signals;
- content structure;
- technical readiness.
That is why the AI Search Optimization Tools on Maksut.net are useful as a first layer.
A professional audit should go further:
- Custom prompt research.
- Multiple AI engines.
- Repeated observations.
- Competitor benchmarking.
- Citation/source extraction.
- Fact checking.
- Technical + content review.
- Prioritized implementation plan.
- Raw data and methodology.
The difference is simple: the checker looks at your site; the audit looks at the market’s AI answers and then explains why your site is or is not winning inside them.
What should the audit deliverable contain?
If a client is paying for an audit, they should receive more than a dashboard login and a branded score.
-
Executive baseline
What is visible, what is not, where the brand is misrepresented and which competitors dominate the highest-value prompt clusters?
-
Methodology
Exact engines, prompts, markets, dates, run count, matching rules and known limitations.
-
Raw prompt-level evidence
Every prompt, observed answer, brands, citations, recommendation outcome and fact error should be exportable.
-
Technical eligibility findings
Crawlability, indexing, robots, rendering, structured data and discoverability barriers.
-
Entity and factual findings
What AI systems think the brand is, where that differs from canonical truth and which sources appear to create the mismatch.
-
Competitive source map
Which third-party domains, review sites, communities, documentation or media pages repeatedly support competitors.
-
Content gap map
Which buyer questions and evidence types you do not currently answer well enough.
-
90-day action plan
Prioritized by expected impact, implementation effort, dependency and confidence—not a flat list of 80 SEO tasks.
How much does an AI search visibility audit cost in 2026?
There is no standardized market price because “AI visibility audit” currently describes very different products.
At the lower end, automated snapshots may cost little or be free. Public specialist examples in 2026 include Tilio’s £800 fixed-price audit using 30–40 reviewed prompts across three engines, and Veza Digital’s $4,500 fixed-price audit that publicly describes 30 days of tracking, 50–100 prompts across six platforms, competitor benchmarking, technical/content review and a 90-day plan.
Those two examples demonstrate why audit pricing should be compared by sample design and analyst work, not the word “audit.”
| Audit level | Typical scope | What to verify before buying |
|---|---|---|
| Free / automated scan | Technical readiness, a few sample prompts, basic score. | Is it clearly labelled as a snapshot rather than “market visibility”? |
| Focused audit | Reviewed prompts, 2–3 engines, competitor comparison, analyst recommendations. | Do you get raw data and the approved prompt set? |
| Deep visibility audit | More prompts, repeat runs/time series, technical/content/entity review, source mapping. | Does methodology explain sampling, variability and matching? |
| Enterprise program | Ongoing monitoring, markets/languages, many products, implementation and executive reporting. | Is the monitoring linked to real optimization and business outcomes? |
Red flags in an AEO/GEO audit
- One prompt per topic. A single observation is presented as a ranking.
- No raw data. You receive a score but cannot inspect the answers that created it.
- No methodology. The report does not state engines, dates, run count or prompt set.
- Guaranteed ChatGPT rankings. Organic generative answers are not sold as guaranteed placements.
- “Missing llms.txt” marked as a Google AI blocker. Google’s current documentation says Google Search ignores llms.txt for visibility/ranking purposes.
- Schema volume as the score. More schema types do not equal more AI visibility.
- Mentions and citations merged. These are different outcomes.
- No competitor evidence. You are told what is wrong with your site but not why competitors are being selected.
- No fact checking. The audit counts visibility even when the model describes the business incorrectly.
- No uncertainty. Every number is presented as exact despite stochastic output.
What happens after an AI visibility audit?
The audit is useful only if findings become implementation priorities.
A practical post-audit sequence is:
- Fix eligibility blockers first. Crawl/index problems invalidate downstream work.
- Correct canonical facts. Fix pricing, product, organization and entity inconsistencies.
- Close high-value content gaps. Prioritize prompts closest to revenue or strategic authority.
- Strengthen missing evidence. Build documentation, original data, credible comparisons and third-party proof where competitors dominate.
- Improve extractability where needed. Make critical facts clear in visible text and accurate structured data.
- Re-run the same baseline. Use the original prompt set and methodology before claiming improvement.
- Expand only after the baseline is stable enough. Add markets, products and prompt families without destroying comparability.
How I would structure an AI visibility audit for Maksut.net clients
My preferred structure is intentionally transparent:
-
Business intake
Define products/services, markets, buyer roles, commercial competitors and the questions prospects ask before purchase.
-
Technical eligibility pass
Check search/index eligibility, crawler policies, important rendering barriers, structured-data accuracy and internal discoverability.
-
Prompt map
Create an approved set across discovery, comparison, pricing, trust, alternatives and implementation intents.
-
Multi-surface sampling
Capture observations separately by platform and repeat high-value prompts enough to expose obvious instability.
-
Evidence extraction
Log mentions, citations, recommendations, competitor positions, source domains and factual errors.
-
Gap analysis
Separate technical blockers, content gaps, entity gaps and third-party evidence gaps.
-
Prioritized roadmap
Rank actions by buyer value, evidence strength, implementation effort and confidence.
-
Re-measure
Use the same prompt baseline after implementation so “improvement” means more than a new screenshot.
If you only want a fast technical baseline before commissioning a full audit, start with the free AI search tools. If you need implementation rather than measurement alone, see SEO & AI Search Optimization.
Frequently asked questions
- What is an AI search visibility audit?
-
An AI search visibility audit measures how a brand is discovered, described, mentioned, cited and recommended across generative search experiences. A complete audit combines technical eligibility, prompt-based observation, competitor benchmarking, source analysis, factual accuracy and a prioritized implementation plan.
- Can one ChatGPT search tell me whether my brand is visible?
-
No. It can give you one observation, but generative search responses vary across runs, time, search conditions and product surfaces. A useful baseline uses a reviewed prompt set and repeated observations for important prompts rather than treating one answer as a fixed ranking.
- How many prompts should an AEO audit use?
-
There is no universal minimum. Focused commercial audits may use roughly 30–40 carefully reviewed prompts, while deeper multi-product programs may use 50–100 or many more. Prompt quality, buyer-intent coverage, repeated sampling and platform coverage matter more than raw count.
- What is the difference between an AI mention and an AI citation?
-
A mention means the brand appears in generated text. A citation means a page is explicitly linked or attributed as supporting evidence. A recommendation goes further: the model actively positions the brand as a suitable choice for the user’s need. A good audit tracks these separately.
- Does Google require special AEO schema or llms.txt?
-
No. Google’s current 2026 Search documentation says there are no additional technical requirements to appear in AI Overviews or AI Mode beyond normal Search eligibility, and that Google Search does not use llms.txt as a special visibility or ranking file. Accurate structured data can still help Search understand content when it matches the visible page.
- Can ChatGPT referral traffic be tracked in GA4?
-
Some of it can. OpenAI’s current publisher documentation says referral links from ChatGPT Search include
utm_source=chatgpt.com, which can be measured in analytics platforms. That does not reveal every AI exposure or every no-click recommendation. - How much does an AI visibility audit cost?
-
Pricing varies because the scope is not standardized. Public 2026 specialist examples range from focused audits around the high hundreds to deep multi-engine, multi-week audits costing several thousand dollars. Compare the approved prompt set, platform coverage, repeat sampling, analyst work, raw-data access and implementation plan before comparing price.
- Can an AEO agency guarantee a ChatGPT ranking?
-
No credible organic audit should promise a guaranteed position in ChatGPT or another generative answer engine. These systems are probabilistic and their retrieval, model and product behavior changes. The useful promise is transparent measurement, prioritized improvements and repeatable re-testing.
Measure the answers your buyers actually see—not a vanity AI score.
Start with a technical readiness check if you need a baseline. If you need to know why competitors are being cited or recommended instead, build a reviewed prompt set, measure the answer environment, inspect the evidence and turn the findings into a prioritized implementation plan.
Discuss an AI Search Visibility AuditStart free: AI Search Tools · Learn: AEO Guide · Implement: SEO & AI Search Optimization
Discussion
0 comments
No comments yet.
Have a technical question, correction or a different interpretation? Add to the discussion.