The two clocks
Training data is frozen at a point in time and updates only when a new model ships. It is good at broad associations, which is why assistants know the big brands in every category. It is unreliable for a nine-person HVAC company that changed its name last year.
Retrieval is live. OpenAI, Google, Anthropic, and Perplexity all publish documentation on the crawlers and search tools their assistants use to fetch pages at answer time. That is the half you can affect this month.
No, it is not reading Google Maps
This is the most common misconception, and it has an expensive consequence. Assistants do not have a live feed of the map. What they have is the enormous sediment of web pages that were themselves written from map data: directory pages, roundups, review sites, aggregators, and your own site. Your Business Profile matters enormously, but it matters through those pages, not directly.
Google's own AI surfaces are the exception worth knowing. AI Overviews and AI Mode sit inside Google Search and do draw on Google's index and local data. If you care about those specifically, they behave more like search than like a chat assistant.
What our own corpus shows
The table below is aggregated from every source domain the assistants actually pointed at across the scans stored in this product, refreshed a few times a day. It is not a survey and not an estimate. Two things tend to surprise owners: how large the long tail of ordinary business websites is, and how far down the list the big directories sit.
Retrieval-readiness check
Ten things that decide whether a page about you can be fetched and used. All ten are checkable in an hour and none require a developer for more than one of them.
Can it be fetched
- Your robots.txt does not block AI crawlers you want reading you. Check for GPTBot, OAI-SearchBot, Google-Extended, ClaudeBot, and PerplexityBot by name.
- Key pages return normal HTML without requiring JavaScript to show their text. Disable JavaScript in your browser and reload your top service page.
- No login, no interstitial, no cookie wall in front of the content.
- Pages load in a couple of seconds. Slow pages get abandoned by fetchers as well as by people.
Can it be understood
- The business name appears as text on every page, not only in the logo image.
- Each service page has one clear subject and says the city in the first hundred words.
- Hours, phone, and service area are text.
- The page has a visible last-updated or published date.
Is it worth quoting
- There is at least one number, range, or specific commitment on the page that nobody else in your town publishes.
- There is a sentence that would survive being pasted into an answer with no surrounding context.
23,926
citations recorded in AI answers
798
distinct source domains cited
85%
of them point at a business's own website
The non-business domains assistants cited most
A business’s own website is the single largest bucket, at 85% of all citations. These are the platforms behind the rest.
- 01angi.com595
- 02yelp.com484
- 03reddit.com468
- 04bbb.org299
- 05homeadvisor.com222
- 06thumbtack.com218
- 07consumeraffairs.com211
- 08tripadvisor.com190
- 09cslb.ca.gov127
- 10irs.gov69
Computed live from 182 stored scans covering 23,926 citations across 798 distinct domains. Assistants were asked real buyer questions unprimed; every source domain they cited or linked was recorded and classified. Full method and sample.
You are looking for the term
Grounding and retrieval
The live-search half of this has a name: retrieval-augmented generation, usually shortened to RAG, and the practice of tying an answer to fetched sources is called grounding. When a vendor tells you your content needs to be "groundable", this is the thing they mean, and the checklist above is the whole of it.
- Should I block AI crawlers?
- Not if you want to be recommended. Blocking GPTBot or PerplexityBot removes your pages from the pool an assistant can retrieve, which is the opposite of the goal for a local business that needs to be found.
- Does ChatGPT use Yelp?
- Sometimes, among many other sources. In our corpus the review platforms are a real but minority share of citations, well behind business and product websites.
- 01Overview of OpenAI crawlersOpenAI
- 02Web search toolOpenAI
- 03Web search toolAnthropic
- 04Perplexity crawlersPerplexity
- 05AI features and your websiteGoogle Search Central
Every link above was opened and confirmed reachable on September 1, 2026.
The automated version
See what ChatGPT, Claude, Gemini and Perplexity say about your business
The free scan runs the buyer questions for your trade and city across all four assistants, shows who gets named instead of you, and lists the sources those answers came from. No signup, about thirty seconds.
Run my free scan