Skip to content
All guides

How to track your brand's visibility in ChatGPT and AI search

What to measure when you track your brand in ChatGPT, Google AI Overviews, Perplexity and other AI engines, how to collect the answers people see, and why every rate needs an interval.

By Klaus Byskov Pedersen, founder of sirefs

The short answer

Pick the questions your buyers ask, check them on every engine people use, and record four things per answer: whether your brand is mentioned, where it ranks, who's named instead and which pages are cited. Collect what the consumer apps show a logged-out visitor rather than what a developer API returns, spread your samples over days instead of firing them at once, and read every rate with its sample size and a confidence interval. Otherwise you'll be reacting to noise.

Why track AI answers at all?

Buyers now ask AI engines the questions they used to type into a search box: "best AI chatbot for a webshop", "Tool A vs Tool B", "how much does X cost". The engine answers with a few names and a few links. If you're not among them, you lose that buyer before they reach any list of search results, and nothing in your analytics shows it.

The gap can be large. When we first ran sirefs on our own products, one SaaS brand appeared in 1 of 132 answers across seven engines, while its top competitor appeared in 70% of them. You only find that out by asking.

Which questions should you track?

The ones buyers ask before they know your name. Cover each intent:

  • Best: "best invoicing app for freelancers"
  • For a use case: "invoicing app for a small agency"
  • Comparison: "Tool A vs Tool B"
  • Alternatives: "Tool A alternatives"
  • Pricing: "how much does invoicing software cost"
  • How to: "how to send recurring invoices"

Write them the way people talk to ChatGPT, in each market's language. Leave your own brand name out of most prompts: an answer to a question that names you always mentions you. Somewhere between 25 and 50 prompts covers most categories.

Which engines should you track?

The ones your buyers use, which is usually more than one. They answer differently, cite different sources and favour different brands:

Engine What you see
ChatGPT OpenAI's assistant; answers with sources when it searches the web.
Google AI Overviews The AI summary above Google's results, with linked references.
Google AI Mode Google's conversational search mode, with linked references.
Gemini Google's assistant app.
Perplexity An answer engine that cites its sources prominently.
Microsoft Copilot Microsoft's assistant, which searches the web with Bing.
Claude Anthropic's assistant.
Grok xAI's assistant.

App answers or API answers?

Measure the apps where you can. The consumer apps decide for themselves what to search for, which pages to read and how to phrase the answer. A developer API with web search switched on makes its own choices, so for the same question it often cites different pages and names different brands. Your buyers see the app.

Collect the apps logged out, too. A logged-in answer is shaped by that account's history and memory: useful for nobody but its owner. A logged-out answer, requested for a set country and language, is the closest you get to what a new buyer sees.

Where an app can't be read logged out (Claude and Grok, today), the API is the only option. Use it, but label those numbers and never put them on the same trend line as app answers. How sirefs collects answers shows the method for each engine.

What should you record for each answer?

Measure What it tells you
Visibility The share of answers that mention you. The headline number.
Share of voice Your mentions against your competitors' mentions. Who owns the category.
Recommendation rate How often the answer recommends you, not just names you.
Rank Where you appear in list-style answers. Noisy: read it as a tendency.
Citation rate How often your own pages are cited as a source.
Cited sources Which pages the engines cite for each question: yours, competitors', third parties'. Your to-do list.
Who's named instead The competitors, including ones you didn't know about.

How many answers do you need before a number means something?

More than most dashboards admit. Engines rarely give the same answer twice, so a single answer is an anecdote, and a rate from a few answers has a wide margin of error. Here's the 95% interval around the same rate at different sample sizes:

Answers Mentioned Rate 95% interval
6 2 33% 10% to 70%
15 6 40% 20% to 64%
170 68 40% 33% to 48%

At six answers, "33%" could be anything from 10% to 70%. That's why a drop "from 40% to 33%" on a small sample is usually noise. With 170 answers on each side, a drop from 40% to 33% still isn't significant (p = 0.18), while a drop from 40% to 26% is (p = 0.008).

Three habits keep you honest:

  1. Spread samples over time. One answer per prompt every day or week beats ten answers fired at once: the exact list of brands changes from run to run, but how often a brand appears over many answers is far steadier.
  2. Pool where you can. Share of voice across all your prompts settles much sooner than any single prompt's rate.
  3. Test before you react. Only call a change a change when a significance test says so. Reading the numbers shows the test sirefs uses.

What do you do with the results?

Turn them into pages and placements:

  • For every prompt where you're missing, look at what the engines cite instead. Those pages show the format and the facts that get quoted.
  • Write the page that answers the question: short answer first, "best for" lines, a comparison table, sourced facts and an FAQ.
  • Get onto the third-party sources the engines keep citing: review sites, forums, industry lists.
  • Fix your site's foundations so engines can read and trust it. The GEO audit covers that.
  • Keep measuring to see whether it worked, with intervals, so you know when it did.

How can you start?

Run the free check to see your GEO score and whether ChatGPT, Google AI Overviews and Perplexity mention you for three buyer questions. To track your own prompts on all eight engines over time, sirefs plans start free.

Frequently asked questions

How often should I check my AI visibility?

Weekly is enough to see where you stand and spot real shifts for most brands. Daily checks give larger samples, so intervals narrow faster and real changes show sooner, which matters when you're actively publishing to close a gap.

Do AI engines show everyone the same answer?

No. Answers vary from run to run, and they depend on the country and language, and on the person's account, history and saved memory when they're logged in. A logged-out answer collected for a set market is the most neutral baseline you can get.

Should I track my brand name as a prompt?

Not for visibility. An answer to a question that names you always mentions you, so it says nothing about whether engines bring you up on their own. Track the category questions buyers ask before they know your name, and use a separate check of what engines say about you to catch wrong facts.

Can I track AI visibility by hand?

For a few prompts, yes: ask each engine in a private browser window, logged out, and note who's mentioned and what's cited. It gets hard quickly: eight engines, several markets, enough answers per prompt to mean something, and the same conditions every time.

How we wrote this guide

sirefs wrote this guide from how our own GEO audit and monitoring work. Where we state facts about AI engines and crawlers, we cite the vendors' own documentation, checked on 6 October 2026. Spotted something out of date? Write to [email protected].

How are you being referenced by SIs?

See your GEO score and whether ChatGPT, Google AI Overviews and Perplexity mention you, free.