Key takeaways
- Measure four things: how often you're mentioned (mention rate), how often your site is cited as a source (citation rate), how you compare with competitors (share of voice), and on how many buyer questions you appear at all (query coverage).
- One prompt tells you almost nothing. Answers differ between systems, between runs of the same question, and over time.
- Use a fixed set of 20 to 30 real buyer questions, run each several times on each system, and repeat monthly.
- Report trends and ranges, not single scores. Month-over-month movement on the same question set is the signal.
- The numbers show visibility, not revenue, and no one can guarantee that an assistant will mention or cite you.
the four metrics
AI search visibility is how often, and how, AI assistants bring up your brand when people ask the questions your buyers ask. These four metrics cover it. All of them are simple ratios you can calculate in a spreadsheet.
| Metric | What it answers | How to calculate |
|---|---|---|
| Mention rate | Does the assistant name us? | Answers that mention your brand ÷ all answers collected |
| Citation rate | Do they link to our site as a source? | Answers citing a page on your domain ÷ answers that show sources |
| Share of voice | How often are we named compared with competitors? | Your mentions ÷ mentions of all brands you track |
| Query coverage | On how many buyer questions do we appear at all? | Questions where you're mentioned at least once ÷ all questions |
why one prompt is not enough
Asking ChatGPT once whether it recommends you is a natural first test, and a misleading one. AI answers are generated fresh every time, so a single answer is a sample of one.
- Between systems: ChatGPT, Perplexity, Gemini and Google's AI Overviews draw on different sources and search indexes, and some answer from the model's training data when they don't search.
- Between runs: the same question asked twice can produce different brands, in a different order.
- Over time: models are updated, new pages get published and indexed, and the sources an assistant trusts change.
- Between people: language, location, account history and the exact wording of the question all shift the answer.
Treat every answer as one data point. Visibility is what you see across many of them.
build a repeatable query set
The query set is the backbone of the method. If it changes every month, your numbers can't be compared.
- Start from real buyer questions: sales calls, support emails, Search Console queries and "People also ask".
- Cover the buying stages: category questions ("best accounting software for freelancers"), comparisons ("X vs Y"), problems ("how do I reduce churn") and brand questions ("is X any good").
- Aim for 20 to 30 questions. Fewer makes the numbers jumpy; many more makes the work hard to repeat.
- Write each question down exactly as it will be asked, including language and country.
- List the competitors you'll count in share of voice, and keep that list fixed.
- When you add questions later, add them as a new version and keep reporting the old set too, so trends stay comparable.
how to run the measurement
- Pick the systems your buyers use, typically ChatGPT, Perplexity, Gemini and Google's AI Overviews.
- Use clean sessions where possible: logged out or a fresh profile, the same country and language each time.
- Run every question several times on every system. Three runs per question per system is a practical minimum.
- Log one row per answer (see the table below).
- Repeat in the same week each month, and note any model or product changes you're aware of.
| Field | Example |
|---|---|
| Date | 2026-09-26 |
| System | Perplexity |
| Question (from the fixed set) | Q07: best solar installer for small businesses |
| Run | 2 of 3 |
| Brands mentioned, in order | Competitor A, Example Co, Competitor C |
| Your brand mentioned? | Yes (position 2) |
| Cited URLs | example.com/business-solar, a trade publication |
| Described correctly? | Partly: wrong service area |
a worked example
Example Co is a fictional solar installer. The numbers below are invented to show the arithmetic, not real results. Its query set has 20 questions, run 3 times on 4 systems: 240 answers in total.
| System | Answers | Mentions of Example Co | Mention rate |
|---|---|---|---|
| ChatGPT | 60 | 14 | 23% |
| Perplexity | 60 | 15 | 25% |
| Gemini | 60 | 8 | 13% |
| AI Overviews | 60 | 5 | 8% |
| All systems | 240 | 42 | 18% |
| Brand | Mentions | Share of voice |
|---|---|---|
| Competitor A | 96 | 42% |
| Competitor B | 61 | 27% |
| Example Co | 42 | 18% |
| Competitor C | 30 | 13% |
- Mention rate: 42 of 240 answers, 18%.
- Citation rate: of the 150 answers that showed sources, 9 cited example.com, 6%.
- Share of voice: 42 of 229 brand mentions, 18%, behind two competitors.
- Query coverage: mentioned at least once on 7 of the 20 questions, 35%.
What Example Co would take from this: it's visible on some questions but absent on most, it's rarely the cited source even when it's named, and Competitor A dominates. The next step is to look at which sources the assistants cite on the questions where Example Co is missing.
what you can and can't conclude
| You can reasonably conclude | You can't conclude |
|---|---|
| Whether you're mentioned for your buyers' questions | How much traffic or revenue AI answers bring you |
| Whether that's improving month over month | Your exact "rank", as if answers were a fixed list |
| Which competitors assistants recommend instead | That one change caused one month's movement |
| Which sources assistants cite in your category | What every user sees: answers are personalised and vary |
Small differences between months are often noise. Look for movement that holds across systems and over several months before drawing conclusions.
what tends to move the numbers
The same work that makes a brand easy to find and trust in search: pages that answer buyer questions directly, consistent facts about your brand everywhere, structured data, and mentions on the sources assistants cite for your category. It can improve your visibility over time; it can't guarantee that any assistant will mention or cite you.
questions
Is there a standard AI visibility score?
No. Tools calculate their own scores in different ways. The four ratios here are transparent, so you can compare them over time and explain them to others.
How many prompts do I need?
A practical minimum is 20 to 30 questions, each run several times on each system. Fewer and the numbers swing too much from month to month.
Can I use a tool instead?
Yes, tools automate the collection. Check how they sample (which systems, how many runs, which location), because the same method questions apply.
How often should I measure?
Monthly is enough for most brands. Answers change too much day to day for weekly numbers to mean much.
Does GEO guarantee that AI assistants will recommend us?
No. No one controls what an assistant says. GEO improves the sources and signals assistants draw on, and measurement shows whether that's working.
Sources and data
- Search demand: Google Ads data via DataForSEO, US and UK, pulled 26 September 2026
- AI Overview frequency: our own samples of Google results pages for marketing searches (139 US and UK, 112 Dutch), September 2026
- OpenAI crawlers: OpenAI platform documentation