How sirefs collects answers from AI engines
sirefs collects the answers people actually see: what each consumer app shows a logged-out visitor, collected by our servers through data vendors. Where an app can't be read logged out, we use the provider's official API and label those answers api.
Which method does each engine use?
| Engine | Method | How we collect it |
|---|---|---|
| ChatGPT | ui |
The answer chatgpt.com gives a logged-out visitor, with its sources and the searches it ran, through a data vendor. |
| Google AI Overviews | serp |
The AI Overview on Google's results page for the prompt, with its references. |
| Google AI Mode | serp |
Google's AI Mode answer page for the prompt, with its references. |
| Gemini | ui |
The answer the Gemini app gives a logged-out visitor, through a data vendor. |
| Perplexity | ui |
The answer perplexity.ai gives a logged-out visitor. Until that collector is connected, answers come from Perplexity's Sonar API and are labelled api. |
| Microsoft Copilot | ui |
Coming next: the answer copilot.microsoft.com gives a logged-out visitor, through a data vendor. Our vendor's Copilot collector currently hits a sign-in wall, so Copilot answers aren't collected yet. |
| Claude | api |
Anthropic's API with web search, located in the prompt's market. claude.ai can't be used logged out, so there is no app answer to collect. |
| Grok | api |
xAI's API with web search. No data vendor collects Grok's app answers today. |
Every answer stores its method, the vendor and the model, and the dashboard shows the method next to every number.
What do ui, serp and api mean?
ui: The consumer app, as a logged-out visitor sees it: the same answer a person gets on the website, collected by our servers through a data vendor.serp: Google's own results page for the prompt: the AI Overview or the AI Mode answer, with the pages it links to.api: The provider's developer API with web search switched on. Used only where the app can't be read logged out. API answers differ from app answers, so they are labelled and never mixed into the same trend.
Why collect logged out, on our servers?
- No personal account. Answers aren't shaped by anyone's chat history, saved memory or preferences, so you see what a new buyer sees.
- The right market. Each prompt carries a country and a language, and the request is made for that market.
- Nothing to keep running. There's no browser extension and no laptop that has to stay awake. Checks run on schedule whether or not you're online.
- The same conditions every time. Comparable conditions are what make a trend line mean something.
Why do API answers differ from the apps?
The consumer apps decide for themselves what to search for, which pages to read and how to answer, with instructions and models the public API doesn't share. An API call with web search makes its own choices, so it often cites different pages and names different brands than the app does for the same question.
That's why sirefs collects the apps wherever a logged-out version can be read, labels API answers as api, and never mixes methods in one trend line.
What happens when a collector fails?
- A failed request is retried with backoff.
- If it keeps failing, the check fails over to a secondary collector where one exists.
- An answer from a failover with a different method is stored and shown, but kept out of the trend line, because the methods answer differently.
- A check that still fails is recorded as failed. It's never counted as "not mentioned".
Google sometimes returns an AI Overview that hasn't finished loading. Those are retried, never stored as an answer without a mention.
When are checks run?
Every check (one prompt on one engine in one market) has a due date set by your plan's frequency. Once an hour, sirefs takes the checks that are due and runs each at a random moment within that hour. Spreading samples over days gives a steadier picture than firing many at once: engines rarely give the same list of brands twice, but how often a brand appears over many answers is far more stable.
How is each answer analysed?
Two passes read every answer:
- Matching, on every answer. Brand and competitor names and aliases are matched in the text (whole words, ignoring case and accents). Cited URLs are cleaned of tracking parameters and tagged as your own, a competitor's or a third party's. Numbered and bulleted lists give each brand its rank.
- Reading, once per check per week. A language model (Claude Haiku) reads the answer and lists every brand it mentions, including ones you don't track, whether each is recommended, its rank and the sentiment, and what the answer claims about your brand.
A brand counts as mentioned when either pass finds it. The model catches name variants your aliases miss; the matching catches what the model overlooks. Brands named in the prompt itself are left out of that prompt's numbers: every answer to "alternatives to X" mentions X.
When you change a brand's aliases or the competitor list, stored answers are matched again and the numbers rebuilt, at no cost to you.
Do you scrape the engines yourselves?
No. Consumer-app answers come from data vendors that collect them logged out (DataForSEO and Bright Data), and Claude and Grok come from Anthropic's and xAI's official APIs. We hold no consumer accounts with any engine.