AI Visibility Audit: What to Check, in What Order
Most AI visibility audits produce a list of thirty findings with no order. Everything on the list is true. Perhaps four items matter this quarter, and nothing on the page tells you which four.
The ordering is not a matter of taste. There is a dependency chain, and work done out of sequence cannot land. This is the sequence we use, and the reason each stage gates the next.
In this article:
The principle
Check in the order an engine encounters you: can it fetch the page, can it extract a claim from it, does it know who you are, does it have a reason to pick you, and did any of it change the numbers. Each stage is wasted effort until the one before it is sound.
The practical consequence is blunt. If AI crawlers cannot fetch your pages, your content strategy is irrelevant this quarter. If the content has no extractable claim, your authority is irrelevant. Teams routinely start at stage four because it is the interesting one, and get nothing for it. Improving content while AI crawlers are blocked produces nothing measurable. Taking a visibility baseline before the blockers are cleared simply records the blockers.
Stage 1: Access
Can an AI crawler fetch the page at all? Cheapest stage to check, most likely to contain a hard blocker, and the one nearly every audit skips because it looks fine in a browser.
- Request your key pages with the user agent strings of GPTBot, PerplexityBot, ClaudeBot and Google-Extended. Confirm a 200 and content in the body, not a challenge page.
- Compare against the same request with a browser user agent. A difference in status code is your finding.
- Read robots.txt for AI-specific directives, including ones added by a plugin or a security vendor without anyone deciding to.
- Check firewall, CDN and bot-protection rules for user-agent filtering. A ruleset tightened against scrapers blocks AI crawlers just as effectively.
- Confirm a missing URL returns a real 404 rather than a soft 404 serving a page of content.
- Check whether
llms.txtis served, and whether it is accurate if it is.
A single blocking rule here invalidates every other stage, which is why it goes first. More on the detail in llms.txt and AI search readiness.
Stage 2: Extraction
Having fetched it, can the engine lift an answer out? This is where most firms with genuinely good content lose, and the failure is invisible to a human reader because a person reads around the problem.
- View the raw HTML, not the rendered page. Content behind JavaScript, a consent wall, a tab or an accordion may not be in the response body.
- For each priority question, check there is a direct answer within the first sentence or two beneath a heading that matches the question.
- Check headings are phrased as the question, not as a clever label. “What does probate cost?” is extractable. “Demystifying probate” is not.
- Check claims are self-contained. A sentence beginning “As we saw above” cannot be quoted in isolation.
- Verify structured data is present, valid and consistent with the visible content.
- Check tables and lists are real markup rather than images of tables.
The test we use: could you lift one sentence from this page and have it stand alone as a correct, attributable answer? If not, an engine cannot either. See structured data for AI search and what AI reads on your website.
Stage 3: Identity
Does the engine know, with confidence, who you are? Engines are cautious about naming organisations they cannot resolve, and ambiguity reads as risk.
- Run branded prompts. Ask each engine what your firm does and whether it is any good. If you do not appear for your own name, stop and fix this before anything else.
- Check the organisation name, address, phone and company registration number are identical in your schema, your footer and your directory listings. Inconsistency is the commonest cause of entity confusion.
- Confirm one canonical organisation entity rather than several competing ones across language or regional versions of the site.
- Check authors are real, named, and have a page establishing why they are credible on the subject.
- Check your own description of what you do matches how the engines describe you. A mismatch means the market’s understanding of you is not yours.
The branded-prompt test is the fastest diagnostic in the whole audit and almost nobody runs it. Background in entity-driven SEO.
Stage 4: Evidence
Given a choice, why would an engine name you? Only now does this question become worth answering, because the previous three stages are the mechanism through which any answer to it operates.
- Map your prompt set against your content. Which buyer questions have no page that answers them directly?
- For prompts where competitors appear instead of you, read what was cited. Usually it is more specific than your equivalent page, not better written.
- Count unclaimed prompts, where no brand is named by anyone. These are the cheapest openings you have.
- Check whether you hold anything genuinely citable: original data, a named method, a published position. Engines cite sources, and a page that restates the consensus gives them no reason to pick it.
- Check third-party corroboration. Directory profiles, review platforms and trade press carry disproportionate weight because they are independent of you.
- Check topical breadth. One good page is fragile; a cluster that covers a subject properly is a position. See building topic clusters.
Stage 5: Measurement
Did any of it work? Last, because without the first four stages there is nothing to measure, and a baseline taken before the blockers are cleared just records the blockers.
- Freeze the prompt set and record it with a date.
- Record appearances per engine separately, never blended.
- Record how many distinct pages the citations draw on. Breadth is the measure of durability.
- Record characterisation, not just presence.
- Record who appeared instead of you.
- Fix the interval and the conditions, then repeat identically.
The measure itself is defined in what an AI Visibility Score is, the procedure in full is in how to measure your AI visibility score, and how to read the result in what a good AI visibility score looks like.
Where this becomes a paid engagement
Worth being straight about this, because the list above is the actual sequence and not a teaser.
Stage 1 is a morning’s work and you should do it yourself this week regardless of who you hire. It is the highest-value hour in digital marketing right now and it needs curl and patience, nothing else.
Stages 2 and 3 are a few days for someone who has done them before and considerably longer for someone learning on your site, mostly because the failures are subtle. A heading that reads beautifully and cannot be extracted looks like good work.
Stages 4 and 5 are where outside help earns its fee, for one reason: both depend on comparison. Knowing which of your competitors the engines treat as authoritative, and whether 11 of 48 is good, requires having seen the same measurement across many sites. That is the part you cannot generate from inside your own organisation, and it is the part we sell.
Everything before it, do yourself. You will understand your own position better for it, and you will be a considerably harder client to mislead.
FAQ
What should an AI visibility audit check first?
Whether AI crawlers can fetch your pages at all. Request your key URLs with the user agent strings of GPTBot, PerplexityBot, ClaudeBot and Google-Extended and confirm you get a 200 with real content in the body, then compare against the same request with a browser user agent. Firewall rules, CDN bot protection and robots directives frequently block AI crawlers while the site looks perfect to a human, and a single blocking rule makes every other finding in the audit irrelevant until it is cleared.
Why does the order of an AI visibility audit matter?
Because the stages form a dependency chain and work done out of sequence cannot take effect. An engine has to fetch the page before it can extract a claim from it, has to resolve who you are before it will name you, and only then weighs whether you are the best source. Improving content while AI crawlers are blocked produces nothing measurable, and taking a visibility baseline before the blockers are cleared simply records the blockers. Order the findings by what gates everything else, not by effort.
Can I run an AI visibility audit myself?
Much of it, yes, and you should. The crawler access checks are a morning’s work with curl and are the single highest-value hour available. Extraction and entity checks are achievable but the failures are subtle, since a heading can read well and still be impossible to extract from. The parts that genuinely need outside help are the competitive comparison and the benchmarking, because knowing which competitors the engines treat as authoritative and whether your number is healthy requires having run the same measurement across many sites.
Want the five stages run on your site?
Our SEO Intelligence Report works through these stages in order. Stage 1 is a per-crawler access table. Stages 2 and 3 sit inside the technical, on-page and schema scoring. Stages 4 and 5 are the competitive comparison and the AI visibility baseline across ChatGPT, Perplexity and Gemini. Six scored dimensions, one overall figure, and a prioritised plan ordered by what gates everything else.
Related
Articles