# Reading the numbers: visibility, share of voice and confidence intervals

What each sirefs metric means, why every rate comes with a sample size and a 95% Wilson interval, and when a change counts as real.

Last updated: 2026-10-06. HTML version: https://sirefs.com/docs/reading-the-numbers

Every rate in sirefs is a share of answers, shown with the number of answers behind it and a 95% confidence interval. AI engines vary from run to run, so the interval is what tells you whether a number is solid or a coin toss.

## What does each metric mean?

| Metric | What it measures |
|---|---|
| Visibility | The share of answers that mention your brand, per prompt, topic, engine or market. |
| Share of voice | Your brand's mentions divided by the mentions of every tracked brand (yours and your competitors'). |
| Average rank | Your position in answers that list brands ("1. X, 2. Y"). Only list-style answers count. |
| Citation rate | The share of answers that cite at least one of your own pages as a source. |
| Recommendation rate | The share of answers that actively recommend your brand, not just name it. |
| Sentiment | The average tone of the mentions, from -1 (negative) to 1 (positive). |
| Cited sources | The domains and pages the engines cite, tagged as yours, a competitor's or a third party's. |

## What counts as a mention?

A brand is mentioned when its name or one of its aliases appears in the answer, or when the language model reading the answer finds it under another name. Brands named in the prompt itself don't count for that prompt. See [how each answer is analysed](https://sirefs.com/docs/how-we-collect#how-is-each-answer-analysed).

## Why does every rate have an interval?

Because a rate from a handful of answers can't be told apart from luck. The same 40% means very different things depending on how many answers it comes from:

| Answers | Mentioned | Rate | 95% interval |
|---|---|---|---|
| 6 | 2 | 33% | 10% to 70% |
| 15 | 6 | 40% | 20% to 64% |
| 30 | 12 | 40% | 25% to 58% |
| 170 | 68 | 40% | 33% to 48% |

sirefs uses the Wilson score interval, which stays sensible with small samples and with rates close to 0% or 100%. Read it as: the true rate is very likely somewhere in this range. When two intervals overlap a lot, the difference between the two rates may be noise.

## When does a change count as real?

sirefs calls something a change only when a two-proportion test says so: at least 20 answers on each side and a p-value under 0.05. Two examples with the same number of answers:

- Visibility goes from 68 of 170 answers (40%) to 56 of 170 (33%). p = 0.18: **not a change.** It's within what chance produces.
- Visibility goes from 68 of 170 (40%) to 45 of 170 (26%). p = 0.008: **a real drop.** This one sends an alert.

Alerts compare the last 7 days with the 7 days before, per engine. The weekly digest uses the same test and only calls significant movements changes.

## Which time windows do the numbers use?

Per-prompt rates are shown over rolling 7 and 28-day windows. Share of voice and overall visibility pool every prompt, which gives larger samples: a project with 40 weekly prompts collects about 160 answers per engine over 28 days. Trend lines use daily rollups, grouped by week for longer ranges.

## Why isn't rank the headline number?

Because it barely repeats. Engines rarely list the same brands in the same order twice, so a single rank says little. Rank is shown for list-style answers, averaged, but visibility and share of voice lead.

## Why are some answers missing from the trend?

Each engine's trend line uses one method: its primary collector. When a check fails over to a collector with a different method (for example, an API instead of the app), the answer is stored and shown in the answer list, but it's kept out of the trend, because the methods answer differently. See [how we collect answers](https://sirefs.com/docs/how-we-collect).
