---
title: "How to Measure Your AI Visibility Score"
description: "The whole method, in the order we run it, including the tedious parts. Follow it and you will have a defensible AI visibility baseline by the end of a working day."
url: https://www.agiledigitalagency.com/blog/how-to-measure-ai-visibility-score/
date: 2026-09-11
modified: 2026-10-07
author: "The Agile Team"
image: https://www.agiledigitalagency.com/wp-content/uploads/2026/09/how-to-measure-ai-visibility-score.avif
type: blog
lang: en
---

# How to Measure Your AI Visibility Score

There is an obvious objection to publishing this. If the method is written down, why would anyone pay for it? The answer is that the method was never the scarce part. Running it consistently every month, on a prompt set that survives contact with how buyers actually ask, is the scarce part. Most teams who try it get two months in and quietly stop.

So here is the whole thing, in the order we do it, including the bits that are tedious and the point at which doing it yourself usually stops being worth the hours. If you follow it you will have a defensible number by the end of a working day.

## In this article:

- [The method in one paragraph](#the-method-in-one-paragraph)
- [Step 1: Build the prompt set](#step-1-build-the-prompt-set)
- [Step 2: Fix the conditions](#step-2-fix-the-conditions)
- [Step 3: Run it enough times to count](#step-3-run-it-enough-times-to-count)
- [Step 4: Record four things, not one](#step-4-record-four-things-not-one)
- [Step 5: Set the baseline](#step-5-set-the-baseline)
- [Where doing it yourself breaks down](#where-doing-it-yourself-breaks-down)
- [FAQ](#faq)

## The method in one paragraph

**Fix a list of buyer-intent prompts. Run each one several times, on each engine, under controlled conditions. Record whether your brand appeared, which page was cited, and how you were characterised. Count appearances per engine as a share of prompts run. Repeat identically next month and compare.**

That is it. The reason it still takes skill is that every one of those five sentences contains a decision that determines whether the resulting number means anything. We define the measure itself in [what an AI Visibility Score is](/blog/ai-visibility-score/); this is the procedure for producing one.

## Step 1: Build the prompt set

This is the step that decides everything, and the step most people rush. A prompt set of your service names will tell you almost nothing, because that is not how anyone talks to an assistant. People arrive with a situation, not a keyword.

Aim for 40 to 60 prompts, drawn from five categories:

- **Category discovery.** “Who are the best commercial property solicitors in Manchester?” The question a buyer asks before they know any names.
- **Problem-first.** “My landlord is withholding my deposit, who do I speak to?” No category named, high intent.
- **Comparison.** “How do I choose between a high street firm and a specialist?” The deliberation stage, where being named as an option is the whole prize.
- **Branded.** “What does do?” and “Is any good?” You should appear here. If you do not, the problem is entity clarity, not content.
- **Competitor-branded.** “Alternatives to [competitor].” Where you discover whether you are in the consideration set at all.

Pull the raw material from places where real phrasing already exists: your enquiry inbox, sales call notes, the questions clients ask in first meetings, and the People Also Ask boxes for your core terms. Resist writing them yourself. Prompts invented by marketers sound like marketing, and they produce flattering, useless results.

Then freeze the list. The prompt set is your instrument, and an instrument you adjust every month cannot measure change. Write it down, date it, and change it only at planned intervals with the old set kept running alongside for a period.

## Step 2: Fix the conditions

AI answers are personalised, localised and non-deterministic. Any of those will corrupt a comparison if you let them vary.

- **Engines.** Decide which you are tracking and report each separately. ChatGPT, Perplexity and Gemini cover most of the ground for UK professional services, and Google AI Overviews, AI Mode and Copilot are worth adding if the budget runs to it. A single blended score hides the only thing worth knowing, which is where you are absent.
- **Account state.** Use a logged-out session, or a clean account used for nothing else. Your own history will show you your own site.
- **Location.** Set it explicitly and record it. “Best accountant near me” from two cities is two different measurements.
- **No follow-ups.** Each prompt is asked cold, in a fresh conversation. A second question inherits context from the first and inflates your appearance rate.

Record all four alongside the results. In six months you will not remember, and without them the baseline is not comparable.

## Step 3: Run it enough times to count

One run is not a measurement. Ask the same engine the same question three times and you can get three different answers, with your brand in one of them. Reporting that as “we appear in ChatGPT” is wrong in both directions depending on which run you saw.

Run each prompt at least three times per engine, spaced across several days rather than in one sitting. The output you want is not yes or no but a frequency: appeared in 2 of 3 runs. That frequency is the honest unit of AI visibility, and it is why serious reporting talks about appearance rates rather than positions.

At 50 prompts, 5 engines and 3 runs that is 750 queries. This is the point at which most in-house attempts fail, and the point at which an index-based platform such as Ahrefs Brand Radar earns its cost, because it is querying a standing corpus of AI responses rather than making you generate one.

## Step 4: Record four things, not one

For every run, capture:

1. **Appearance.** Named, or not.
2. **Source.** Which URL was cited, if any. A mention with no link still counts as an appearance, and the distinction matters more than it sounds: see [the citation gap](/blog/citation-gap-ai-mentions-brand-no-link/).
3. **Characterisation.** Positive, neutral or negative. Being named as the expensive option is an appearance and a problem at once.
4. **Who else appeared.** The competitors named instead of you. This is frequently the single most useful column in the sheet, because it tells you who the engine currently considers the authority.

Keep the verbatim answer text too. Scores summarise; the raw answers are where you see that three engines are all citing one page of yours and none of the others, or that an engine is describing a service you stopped offering in 2023.

## Step 5: Set the baseline

Your first run is a baseline, not a verdict. Write it up as appearances out of prompts run, per engine, with the date and conditions attached. Something like: 11 of 48 on Perplexity, 6 of 48 on ChatGPT, 2 of 48 in AI Overviews, 0 of 48 on Gemini, drawing on 5 distinct pages.

Then do nothing with it for a month except act on the diagnosis. Re-measure at a fixed interval, identically. The second measurement is the first one that tells you anything, because AI visibility is only meaningful as a direction of travel. We go further into reporting cadence in [how to measure GEO and AI SEO success](/blog/measure-geo-ai-seo-success/).

One warning about expectations. A baseline of zero on an engine is common and is not a crisis. Absence frequently means your category is not being asked about on that platform yet, which is a different problem from being passed over.

## Where doing it yourself breaks down

Honestly: not at the first measurement. A competent marketing person can produce a credible baseline in a day using nothing but a spreadsheet and a clean browser session. If that is all you need, do that and keep the money.

It breaks down in three specific places.

**Month three.** 750 manual queries is a day of work. Doing it twelve times, identically, while the rest of your job continues, is where the discipline fails. The measurement stops and the baseline becomes a document nobody updates.

**The diagnosis.** The sheet tells you that you are absent from 37 of 48 prompts. It does not tell you why. Working out whether the cause is crawlability, extractable structure, entity ambiguity or a stronger competitor requires separate technical work on each candidate cause, and the four have nothing in common. That is the part that is genuinely an audit.

**Knowing what to fix first.** Once you have thirty findings, the ordering is worth more than the list. We cover the sequence in [the GEO-ready website checklist](/blog/geo-ready-website-checklist/) and the underlying technical work in [structured data for AI search](/blog/structured-data-for-ai-search-professional-services/).

If you take one thing from this: measure it yourself once, this week, badly. A rough number you produced is worth more than a polished one you were sent, because you know exactly what it counted.

## FAQ

### How many prompts do I need to measure AI visibility?

Between 40 and 60 for most professional services firms, spread across category discovery, problem-first questions, comparison questions, your own brand and competitor brands. Fewer than about 30 and a single answer swings the percentage too far to be a trend. Many more than 60 and the manual running cost becomes the reason you stop. What matters more than the count is that the list is frozen, so the same instrument is used every period.

### How often should I re-measure?

Monthly is the useful interval for most firms, and quarterly is defensible if the category moves slowly. Weekly is too often: AI answers are probabilistic, so short-interval changes are mostly noise rather than progress. The important discipline is that the interval is fixed and the conditions are identical, because the comparison is the finding. A single measurement is a baseline and tells you almost nothing on its own.

### Do I need a paid tool to measure AI visibility?

Not for a first baseline. A spreadsheet, a logged-out browser session and a day of work will give you a defensible number you fully understand. Paid platforms earn their place at scale and over time: 50 prompts across 5 engines with 3 runs each is 750 queries a month, and index-based tools such as Ahrefs Brand Radar query a standing corpus of AI responses rather than requiring you to generate one. Start manual, then buy when the repetition is what is failing.

### Want a baseline without building the whole instrument?

Our SEO Intelligence Report gives you a dated baseline: we ask ChatGPT, Perplexity and Gemini directly about your brand and your category, and report discoverability, mention rate and sentiment per engine, the competitors named instead of you, and whether each AI crawler can reach your site at all. It sits inside a six-dimension audit, so the absences arrive with the technical diagnosis behind them.

Being straight about the scope: that is a point-in-time sample, and we label it as one in the report. Continuous tracking across a frozen prompt set of 40 to 60, of the kind described above, is an index-based programme on top and it costs accordingly. Worth it for some firms and not for others, and we will tell you which you are.

[**See what the report covers →**](/services/seo/seo-audit/)
