How we measure AI visibility
A scan records a defined sample of AI answers. Here is what we ask, how we score completed answers, and what a change can tell you.
Buyer questions, without naming your business
We generate category and location questions such as “best plumber in Sacramento” and “where should we stay in Mendocino”. These are generated buyer-intent prompts, not measured search-volume data. You can add or edit your tracked questions. Custom questions come first in the priority set, followed by the oldest active questions.
The initial free scan plans a small question set. Local tracks up to 100 questions; Growth tracks up to 120 per business. Each question and assistant is a separate observation.
What gets checked, and when
- Daily priority checks
- your top 25 questions on Local or 30 per business on Growth, on ChatGPT and Gemini.
- Weekly full sweep
- all 100 tracked questions on Local or 120 per business on Growth, across ChatGPT, Claude, Gemini, Perplexity, Grok and Copilot. Sweeps start Monday.
- Google AI answers
- Up to 6 priority questions on paid scans. Results may be up to 7 days old; an AI answer is not available for every question.
Email alerts after a mention change holds across two consecutive checks of the same question and assistant. Models, freshness and measurement limits
Which models produce the answers?
We query model APIs through OpenRouter. Copilot is the exception: it runs on an Azure OpenAI model with Grounding with Bing Search, the same pairing Microsoft Copilot runs on, and uses the same deployment on every tier. Grok runs xAI's grok-4.3 model with xAI's own web search, also the same model on every tier. The free scan and paid monitoring otherwise use different model tiers. These observations can differ from the consumer ChatGPT, Claude, Gemini, Perplexity, Grok and Copilot apps, which may use other models, browsing settings, personal history or location context.
| Family | Free scan | Paid monitoring |
|---|---|---|
| openai/gpt-4o | openai/gpt-4o-miniDaily priority checks and weekly sweep | |
| anthropic/claude-sonnet-4.5 | anthropic/claude-haiku-4.5Weekly sweep | |
| google/gemini-2.5-pro | google/gemini-2.5-flashDaily priority checks and weekly sweep | |
| perplexity/sonar-pro | perplexity/sonarWeekly sweep | |
| x-ai/grok-4.3 | x-ai/grok-4.3Weekly sweep | |
| Azure OpenAI + Grounding with Bing Search | Azure OpenAI + Grounding with Bing SearchWeekly sweep |
Free scans and manual rescans request live web grounding. Scheduled web grounding is enabled on Mondays (UTC) by default; the other scheduled daily checks do not request the additional web search. Sonar and Copilot are web-grounded whenever they run; Grok searches the web whenever grounding is requested. Scheduled checks may reuse a recent answer for the same model, question and grounding setting.
A manual paid rescan checks your 25 or 30 priority questions across all six paid model families, up to 10 rescans per user per day. Its coverage is broader than the scheduled daily check. Full sweeps start Monday and may finish later if queued.
Google freshness and unavailable answers
Paid scans request Google AI Overviews and AI Mode through SerpApi for the first six priority questions, once per full sweep. The default cache retains each question-and-surface result for seven days. A scan may therefore reuse an earlier result instead of fetching Google again. Provider availability and monthly capacity can also limit collection. An absent AI Overview is not a negative mention result, and quota skips are labeled. A scan timestamp is not a promise that every underlying Google answer was fetched at that moment.
How the score is calculated
Mention rate is the fraction of completed answers that name your business. Citation rate is the fraction that link to your site. Returned Google AI answers are included on paid scans. Failed checks and Google surfaces that return no answer do not enter that denominator. New scans retry token-limited answers once with more room; answers that still do not complete are excluded as failed checks, never counted as missing mentions. List position comes from detected provider lists; prose mentions have no list position.
Score = round(min(100, 80 × mention rate + 15 × citation rate + rank bonus))
Rank bonus = max(0, 8 − average list position), using only answers that mention you and have a detected position. It is zero when none has a position. Sentiment and competitor share of voice are reported separately.
Worked example · illustrative
Of 20 completed answers, 14 mention the business and 4 cite its site. Its average detected list position is 2. The score is 80 × 0.70 + 15 × 0.20 + (8 − 2) = 65, grade C. If two additional checks failed, the denominator remains 20, not 22.
Grades: A ≥ 85, B ≥ 70, C ≥ 50, D ≥ 30, F below 30. With no completed provider answers, we do not publish a zero score. Partial coverage is labeled; more than 20% failed chat checks marks a scan as degraded and excludes it from change alerts.
Business matching and answer receipts
We match names on word boundaries, normalize accents and punctuation, treat “&” and “and” alike, and allow common legal suffixes to be omitted. Similar business names can still need review. Receipts show the prompt, captured excerpt, detected mention, list position and cited sources so you can inspect the evidence.
Paid answer exports include date, prompt, assistant, mention, rank, citation domains and excerpt. Failed chat checks are exported as failed with blank mention fields. This preserves the difference between “not mentioned” and “not measured”.
What changes mean, and what they do not prove
- Answers vary between runs. A rate summarizes this question set; it is not a statistically representative estimate of all customer searches.
- Compare the same prompts, model coverage and grounding settings. Free scans, daily checks, manual rescans and full sweeps have different coverage and can produce different scores for that reason alone.
- Alerts require a change to persist across two consecutive checks of the same question and assistant, relative to an earlier baseline. Missing checks cannot count as a drop. Confirmation reduces noise; it cannot eliminate it.
- We measure observed answers and cited sources, not the model's hidden reasoning. A citation does not prove why a recommendation was made, and a rising score alone does not prove that a particular edit caused it.
- These measurements do not establish leads, bookings or revenue. Recommendations are actions to evaluate; no assistant recommendation or ranking is guaranteed.