AI Search 19–20 min read

AI Search Visibility Audit: What a Real AEO/GEO Audit Should Measure in 2026

A practical 2026 framework for auditing how a brand is discovered, described, mentioned, cited and recommended across ChatGPT, Google AI features, Perplexity and other answer engines—without pretending one prompt is a ranking report.


What is an AI search visibility audit?

An AI search visibility audit is a structured review of whether—and how—your brand, website, products, people and evidence appear inside generative search experiences. Depending on the market and buyer journey, that can include ChatGPT Search, Google AI Overviews and AI Mode, Perplexity, Gemini-connected experiences, Copilot and other answer surfaces.

The audit should answer more than “Are we mentioned?” A useful audit investigates a chain of questions:

  • Eligibility: Can the relevant search and retrieval systems access and index your content?
  • Understanding: Do they correctly understand who you are, what you sell and who it is for?
  • Retrieval: Does your content or third-party evidence enter the candidate source set?
  • Selection: Are you cited, mentioned or recommended for prompts that matter commercially?
  • Accuracy: Are the facts, pricing, positioning and comparisons correct?
  • Competition: Which competitors appear instead, and which sources support them?
  • Stability: Does the observation persist across repeated runs and time?
  • Outcome: Can any of this be connected to visits, branded demand, assisted conversions or sales?

This is why an AI visibility audit sits above a basic AEO readiness checker. A checker can inspect your site. An audit must inspect the answer environment around your category.

The seven layers of an AI search visibility audit A seven-layer stack moves from crawlability to entity understanding, extractability, off-site evidence, prompt visibility, mentions and citations, and finally business outcomes. 1 · Eligibility crawl · index · robots · snippet eligibility 2 · Entity understanding brand · product · person · category · facts 3 · Extractability answers · tables · structured data · visible text 4 · Evidence & consensus reviews · docs · media · forums · references 5 · Prompt visibility intent coverage · engines · competitors · reruns 6 · Answer outcome mention · citation · recommendation · framing 7 · Business outcome referrals · brand demand · pipeline · revenue visibility becomes more commercially meaningful
A useful audit moves from technical eligibility to observed answer behavior and finally to business impact. Stopping at schema or robots.txt is not a visibility audit.

AI search audit vs traditional SEO audit

A traditional SEO audit asks whether pages can be crawled, indexed, ranked and converted. Those fundamentals remain essential. In fact, Google explicitly states that there are no additional technical requirements for appearing as a supporting link in AI Overviews or AI Mode beyond being indexed, eligible for Search and eligible for snippets. Google’s 2026 generative AI Search guidance also says these experiences remain rooted in core Search ranking and quality systems.

That means AEO/GEO does not replace SEO. A serious AI audit starts with SEO eligibility, then measures layers classic rank tracking does not capture well: answer inclusion, recommendation framing, citation sources, competitor co-occurrence, factual representation and run-to-run variability.

Dimension Traditional SEO audit AI search visibility audit
Primary object URLs and rankings Entities, answers, sources and recommendations
Technical eligibility Crawl, index, canonical, rendering Same baseline + surface-specific crawler/search eligibility
Measurement Rank, clicks, impressions, conversions Mention rate, citation rate, recommendation rate, source share, factual accuracy, referrals
Query model Keyword set Prompt/intention set with paraphrases and buyer contexts
Competition Who outranks you? Who is recommended, cited or used as evidence instead?
Stability Rankings fluctuate, but snapshots are often useful Repeated sampling is essential because identical prompts can produce different answers
Off-site evidence Links, authority, reputation Also asks which third-party pages the answer engine is actually using

Google does not require a secret “AI SEO” file

Google’s current Search documentation says no special markup, AI text file or llms.txt file is required to appear in Google Search generative AI features. For Google specifically, foundational SEO, crawlability, index eligibility, visible high-quality content and accurate structured data remain the core technical baseline. If an audit marks “missing llms.txt” as a critical Google AI error, challenge the methodology.

Why one ChatGPT prompt is not an audit

Generative search is probabilistic. Identical or near-identical prompts can produce different brands, sources, answer order and citations across repeated runs. Surface changes, search activation, context and time can all affect the response.

A 2026 study titled Don’t Measure Once: Measuring Visibility in AI Search evaluated repeated AI-search observations and concluded that visibility should be treated as a distribution rather than a single-point outcome. Another 2026 paper, Quantifying Uncertainty in AI Visibility, found substantial citation variability across repeated samples and warned that single-run metrics can look more precise than the underlying behavior supports.

This changes the audit methodology. A screenshot saying “ChatGPT did not mention you” is an observation. It is not evidence that your brand has zero ChatGPT visibility.

One prompt screenshot versus repeated AI visibility measurement The left side shows one prompt producing one answer and a misleading zero visibility conclusion. The right side shows a reviewed prompt set, multiple engines, repeated runs and time producing a distribution-based baseline. Snapshot audit 1 prompt × 1 run × 1 engine Brand absent in this answer Wrong conclusion: “0% AI visibility” Measurement audit Prompt set Multiple engines Repeated runs Time series Visibility distribution + evidence An audit should tell you what was sampled, when, where and how often—not only the score.
Observation is not measurement. Repeated sampling does not make AI search deterministic; it makes your uncertainty visible.

The 12 layers of a real AI visibility audit

1. Search and crawler eligibility

Before measuring recommendations, verify that the relevant surfaces can access the content they may need.

For ChatGPT Search, OpenAI’s current publisher documentation says publishers should allow OAI-SearchBot if they want their content included in ChatGPT summaries and snippets. OpenAI separates this from GPTBot, which is related to potential training use. For Perplexity, official crawler documentation says PerplexityBot is designed to surface and link websites in Perplexity search results, while Perplexity-User can access pages in response to user actions.

For Google AI Overviews and AI Mode, the baseline is different: Google says pages must be indexed and eligible to appear in Search with a snippet, with no additional AI-specific technical requirement.

Your audit should therefore test actual crawler and robots.txt policy rather than applying one generic “AI bots allowed” checkbox.

2. Index and discoverability health

AI visibility cannot be separated cleanly from search visibility. Check whether key commercial, comparison, documentation and research pages are indexable, internally linked, canonicalized correctly and discoverable through the search systems that feed generative experiences.

For Google, use Search Console data. For Bing-dependent ecosystems, include Bing index health where relevant. The audit should distinguish not indexed, indexed but not retrieved, and retrieved but not selected. Those are different problems.

3. Entity clarity

Audit whether the open web consistently answers basic questions about the entity:

  • What is the brand’s canonical name?
  • What category does it belong to?
  • What products or services does it offer?
  • Who is the founder, author or responsible organization?
  • Which geography does it serve?
  • What are the current pricing, features and policies?
  • Which external profiles corroborate those facts?

Then compare those facts with what AI systems actually say. Entity consistency is not useful if the model confidently repeats an outdated price or describes the wrong target customer.

4. Content extractability

Check whether important facts exist in visible, accessible text and are easy to verify. Tables, direct answer paragraphs, labelled comparisons, dates, authorship, definitions and clearly separated claims can make content easier for retrieval systems and humans to interpret.

Do not turn this into a ritual of arbitrary chunk sizes. Google’s 2026 guidance explicitly warns site owners to prioritize valuable, non-commodity content rather than generic “AEO hacks.” For Google specifically, special AI files or forced chunking are not required.

5. Structured data accuracy

Structured data can reinforce entity relationships and factual consistency when it matches visible content. Audit the existing JSON-LD graph for contradictions, duplicated entities, stale offers, mismatched organization names, orphaned authors and unsupported properties.

The goal is not to maximize schema count. The goal is to make the machine-readable representation agree with the page humans see. See the JSON-LD @graph method.

6. Prompt-set design

A serious audit does not use ten random prompts invented by the analyst. Build the prompt set from the buyer journey and real information demand.

A practical prompt taxonomy can include:

Intent Example pattern What it tests
Category discovery “Best [category] tools/services for…” Whether the brand enters the consideration set.
Problem discovery “How should I solve [problem]?” Whether your expertise/content is used before vendor selection.
Shortlisting “Which providers should I consider for…?” Recommendation presence and relative prominence.
Comparison “Brand A vs Brand B for [context]” Feature accuracy, framing and comparative evidence.
Pricing / fit “What does [service] cost and who is it for?” Commercial fact accuracy and suitability descriptions.
Trust “Is [brand] reputable / experienced in…?” Off-site consensus, credentials and reputation sources.
Implementation “How do I implement [technical task]?” Whether docs, guides or technical authority are cited.

Commercial providers already vary widely here: some audits use around 30–40 reviewed prompts, while deeper services publicly describe 50–100 prompts across several engines. The number matters less than whether the prompt set covers the decisions your buyers actually make.

7. Multi-engine coverage

Do not merge all AI surfaces into one universal ranking. ChatGPT, Google AI features and Perplexity have different retrieval paths, interfaces, source pools and citation behaviors.

At minimum, record results by engine separately. If the target audience uses multiple surfaces, your report should show both the cross-platform pattern and the platform-specific differences.

8. Mention, citation and recommendation separation

These outcomes are not interchangeable.

  • Mention: the brand appears in answer text.
  • Citation: your site or another page about you is used as a linked source.
  • Recommendation: the model actively positions the brand as a suitable choice for the user’s need.
  • Prominence: where and how strongly the brand is surfaced inside the answer.
Mention, citation and recommendation are different AI visibility outcomes Three cards distinguish brand mention, linked citation and active recommendation, followed by a fourth layer for prominence and accuracy. Mention “Maksut.net is one provider in this space.” Brand appears Citation A page is linked as supporting evidence. Source is attributed Recommendation “Choose Maksut.net if you need custom Woo…” Commercial endorsement Audit also scores accuracy prominence A brand can be frequently mentioned but rarely cited—or cited without ever being recommended.
Do not collapse every outcome into one score. Mentions, citations, recommendations and accuracy answer different business questions.

9. Repeated sampling and uncertainty

For each important prompt, record more than one observation where budget and tooling allow. The audit should document:

  • Engine and product surface.
  • Date and time.
  • Exact prompt.
  • Run number.
  • Whether web/search mode was active where observable.
  • Brands mentioned.
  • Brands recommended.
  • Sources cited.
  • Position/prominence.
  • Incorrect or stale facts.

Then report rates with the sample size attached. “Mention rate: 42% (21/50 observations)” is more honest than “AI Visibility Score: 42” with no method.

10. Competitive source analysis

When competitors are repeatedly recommended, inspect the evidence layer behind them.

Ask:

  • Are their own product/service pages cited?
  • Are review sites doing the heavy lifting?
  • Are Reddit/forum discussions shaping the answer?
  • Are industry publications validating the category position?
  • Do they have comparison pages, docs, studies or original data you lack?
  • Are the same third-party domains recurring across multiple engines?

This is where AI visibility auditing becomes strategic. The fix may not be “rewrite our landing page.” It may be “we have no independent evidence in the source ecosystem that answer engines use for this buying question.”

11. Factual accuracy and brand representation

Visibility is not automatically good. An answer can mention you frequently while describing the wrong pricing, outdated product name, incorrect location, unsupported capability or a weakness you no longer have.

Audit the factual fields that matter commercially:

  • Price and pricing model.
  • Service/product category.
  • Target customer.
  • Geography.
  • Features and exclusions.
  • Founder/leadership.
  • Certifications and credentials.
  • Support and policy information.
  • Competitor comparisons.

For every incorrect claim, record the answer, source if present, canonical fact and likely correction path.

12. Referral and business measurement

Not every AI exposure produces a measurable click. But when referral data is available, capture it.

OpenAI’s current publisher FAQ says ChatGPT referral links automatically include utm_source=chatgpt.com, allowing publishers to track inbound ChatGPT search traffic in analytics platforms such as Google Analytics. Google reports AI Overviews and AI Mode traffic within Search Console and has introduced dedicated generative-AI reporting in its current documentation.

Use referral traffic as one outcome—not the definition of visibility. A recommendation can influence a later branded search, direct visit or offline buying decision without a clean last-click trail.

What metrics should an AI visibility audit report?

A defensible audit should provide raw counts and rates before compressing anything into an executive score.

Metric Example definition What it tells you
Mention rate Brand mentioned ÷ total observations How often you enter the answer at all.
Citation rate Your domain cited ÷ total observations How often your site becomes explicit evidence.
Recommendation rate Brand recommended ÷ relevant commercial observations How often you enter the buyer shortlist.
Prompt coverage Prompt clusters with ≥1 meaningful visibility event Where visibility exists across the buyer journey.
Source share Your domain citations ÷ all observed citations in scope Your evidence presence relative to the source market.
Prominence Position / first-mentioned / shortlist inclusion Whether you are central or incidental.
Fact accuracy Correct audited claims ÷ audited claims Whether visibility represents the real business.
Volatility Run-to-run change across repeated observations How stable the apparent visibility is.
AI referrals Tracked sessions / conversions from identifiable AI sources Observable traffic and revenue contribution.

Always publish the denominator

“We appeared 18 times” is not interpretable without knowing whether the audit ran 20, 200 or 2,000 observations. Every visibility percentage should be traceable to the underlying prompt set, engine set, run count and date range.

How many prompts should an AI visibility audit test?

There is no universal number because the right sample depends on category breadth, markets, products, buyer roles and how many engines you are testing.

Public commercial audits currently illustrate the range: Tilio describes a reviewed 30–40 prompt set across ChatGPT, Google AI Overviews and Perplexity, while Veza Digital describes 50–100 prompts across six AI platforms with a 30-day tracking period. Those are examples of service design, not scientific minimums.

For a focused service business, I would rather test 30 carefully designed buyer prompts with repeat observations than 500 generic prompts that nobody commercially cares about.

For a broad SaaS or ecommerce category, the prompt matrix may need to grow across:

  • Products.
  • Features.
  • Personas.
  • Industries.
  • Competitors.
  • Use cases.
  • Pricing tiers.
  • Geographies.
  • Languages.

Free AEO checker vs professional AI visibility audit

A free tool can be extremely useful if its promise is narrow. For example, a checker can inspect:

  • robots.txt and crawler directives;
  • indexability;
  • structured data presence;
  • basic entity signals;
  • content structure;
  • technical readiness.

That is why the AI Search Optimization Tools on Maksut.net are useful as a first layer.

A professional audit should go further:

  • Custom prompt research.
  • Multiple AI engines.
  • Repeated observations.
  • Competitor benchmarking.
  • Citation/source extraction.
  • Fact checking.
  • Technical + content review.
  • Prioritized implementation plan.
  • Raw data and methodology.

The difference is simple: the checker looks at your site; the audit looks at the market’s AI answers and then explains why your site is or is not winning inside them.

What should the audit deliverable contain?

If a client is paying for an audit, they should receive more than a dashboard login and a branded score.

  1. Executive baseline

    What is visible, what is not, where the brand is misrepresented and which competitors dominate the highest-value prompt clusters?

  2. Methodology

    Exact engines, prompts, markets, dates, run count, matching rules and known limitations.

  3. Raw prompt-level evidence

    Every prompt, observed answer, brands, citations, recommendation outcome and fact error should be exportable.

  4. Technical eligibility findings

    Crawlability, indexing, robots, rendering, structured data and discoverability barriers.

  5. Entity and factual findings

    What AI systems think the brand is, where that differs from canonical truth and which sources appear to create the mismatch.

  6. Competitive source map

    Which third-party domains, review sites, communities, documentation or media pages repeatedly support competitors.

  7. Content gap map

    Which buyer questions and evidence types you do not currently answer well enough.

  8. 90-day action plan

    Prioritized by expected impact, implementation effort, dependency and confidence—not a flat list of 80 SEO tasks.

How much does an AI search visibility audit cost in 2026?

There is no standardized market price because “AI visibility audit” currently describes very different products.

At the lower end, automated snapshots may cost little or be free. Public specialist examples in 2026 include Tilio’s £800 fixed-price audit using 30–40 reviewed prompts across three engines, and Veza Digital’s $4,500 fixed-price audit that publicly describes 30 days of tracking, 50–100 prompts across six platforms, competitor benchmarking, technical/content review and a 90-day plan.

Those two examples demonstrate why audit pricing should be compared by sample design and analyst work, not the word “audit.”

Audit level Typical scope What to verify before buying
Free / automated scan Technical readiness, a few sample prompts, basic score. Is it clearly labelled as a snapshot rather than “market visibility”?
Focused audit Reviewed prompts, 2–3 engines, competitor comparison, analyst recommendations. Do you get raw data and the approved prompt set?
Deep visibility audit More prompts, repeat runs/time series, technical/content/entity review, source mapping. Does methodology explain sampling, variability and matching?
Enterprise program Ongoing monitoring, markets/languages, many products, implementation and executive reporting. Is the monitoring linked to real optimization and business outcomes?

Red flags in an AEO/GEO audit

  • One prompt per topic. A single observation is presented as a ranking.
  • No raw data. You receive a score but cannot inspect the answers that created it.
  • No methodology. The report does not state engines, dates, run count or prompt set.
  • Guaranteed ChatGPT rankings. Organic generative answers are not sold as guaranteed placements.
  • “Missing llms.txt” marked as a Google AI blocker. Google’s current documentation says Google Search ignores llms.txt for visibility/ranking purposes.
  • Schema volume as the score. More schema types do not equal more AI visibility.
  • Mentions and citations merged. These are different outcomes.
  • No competitor evidence. You are told what is wrong with your site but not why competitors are being selected.
  • No fact checking. The audit counts visibility even when the model describes the business incorrectly.
  • No uncertainty. Every number is presented as exact despite stochastic output.

What happens after an AI visibility audit?

The audit is useful only if findings become implementation priorities.

A practical post-audit sequence is:

From free AI readiness check to audit, implementation and ongoing monitoring Four connected stages show free readiness checking, professional visibility audit, implementation, and ongoing measurement and optimization. 1 · Free checker crawlability schema · structure technical baseline 2 · Visibility audit prompts · engines citations · competitors priority plan 3 · Implementation technical fixes content · entities off-site evidence 4 · Monitor repeat measurement new competitors business outcomes The audit is a baseline. Optimization starts when findings change the web evidence and the measurement is repeated.
The clean funnel is readiness → measurement → implementation → repeated measurement. Selling endless monitoring without fixing anything is not optimization.
  1. Fix eligibility blockers first. Crawl/index problems invalidate downstream work.
  2. Correct canonical facts. Fix pricing, product, organization and entity inconsistencies.
  3. Close high-value content gaps. Prioritize prompts closest to revenue or strategic authority.
  4. Strengthen missing evidence. Build documentation, original data, credible comparisons and third-party proof where competitors dominate.
  5. Improve extractability where needed. Make critical facts clear in visible text and accurate structured data.
  6. Re-run the same baseline. Use the original prompt set and methodology before claiming improvement.
  7. Expand only after the baseline is stable enough. Add markets, products and prompt families without destroying comparability.

How I would structure an AI visibility audit for Maksut.net clients

My preferred structure is intentionally transparent:

  1. Business intake

    Define products/services, markets, buyer roles, commercial competitors and the questions prospects ask before purchase.

  2. Technical eligibility pass

    Check search/index eligibility, crawler policies, important rendering barriers, structured-data accuracy and internal discoverability.

  3. Prompt map

    Create an approved set across discovery, comparison, pricing, trust, alternatives and implementation intents.

  4. Multi-surface sampling

    Capture observations separately by platform and repeat high-value prompts enough to expose obvious instability.

  5. Evidence extraction

    Log mentions, citations, recommendations, competitor positions, source domains and factual errors.

  6. Gap analysis

    Separate technical blockers, content gaps, entity gaps and third-party evidence gaps.

  7. Prioritized roadmap

    Rank actions by buyer value, evidence strength, implementation effort and confidence.

  8. Re-measure

    Use the same prompt baseline after implementation so “improvement” means more than a new screenshot.

If you only want a fast technical baseline before commissioning a full audit, start with the free AI search tools. If you need implementation rather than measurement alone, see SEO & AI Search Optimization.

Frequently asked questions

What is an AI search visibility audit?

An AI search visibility audit measures how a brand is discovered, described, mentioned, cited and recommended across generative search experiences. A complete audit combines technical eligibility, prompt-based observation, competitor benchmarking, source analysis, factual accuracy and a prioritized implementation plan.

Can one ChatGPT search tell me whether my brand is visible?

No. It can give you one observation, but generative search responses vary across runs, time, search conditions and product surfaces. A useful baseline uses a reviewed prompt set and repeated observations for important prompts rather than treating one answer as a fixed ranking.

How many prompts should an AEO audit use?

There is no universal minimum. Focused commercial audits may use roughly 30–40 carefully reviewed prompts, while deeper multi-product programs may use 50–100 or many more. Prompt quality, buyer-intent coverage, repeated sampling and platform coverage matter more than raw count.

What is the difference between an AI mention and an AI citation?

A mention means the brand appears in generated text. A citation means a page is explicitly linked or attributed as supporting evidence. A recommendation goes further: the model actively positions the brand as a suitable choice for the user’s need. A good audit tracks these separately.

Does Google require special AEO schema or llms.txt?

No. Google’s current 2026 Search documentation says there are no additional technical requirements to appear in AI Overviews or AI Mode beyond normal Search eligibility, and that Google Search does not use llms.txt as a special visibility or ranking file. Accurate structured data can still help Search understand content when it matches the visible page.

Can ChatGPT referral traffic be tracked in GA4?

Some of it can. OpenAI’s current publisher documentation says referral links from ChatGPT Search include utm_source=chatgpt.com, which can be measured in analytics platforms. That does not reveal every AI exposure or every no-click recommendation.

How much does an AI visibility audit cost?

Pricing varies because the scope is not standardized. Public 2026 specialist examples range from focused audits around the high hundreds to deep multi-engine, multi-week audits costing several thousand dollars. Compare the approved prompt set, platform coverage, repeat sampling, analyst work, raw-data access and implementation plan before comparing price.

Can an AEO agency guarantee a ChatGPT ranking?

No credible organic audit should promise a guaranteed position in ChatGPT or another generative answer engine. These systems are probabilistic and their retrieval, model and product behavior changes. The useful promise is transparent measurement, prioritized improvements and repeatable re-testing.

Measure the answers your buyers actually see—not a vanity AI score.

Start with a technical readiness check if you need a baseline. If you need to know why competitors are being cited or recommended instead, build a reviewed prompt set, measure the answer environment, inspect the evidence and turn the findings into a prioritized implementation plan.

Discuss an AI Search Visibility Audit

Start free: AI Search Tools · Learn: AEO Guide · Implement: SEO & AI Search Optimization

Share this article

Related posts

Discussion

0 comments

No comments yet.

Have a technical question, correction or a different interpretation? Add to the discussion.

Leave a reply