Skip to content

Where does ChatGPT get its local recommendations from?

The short answer

From two places. Some of it is baked in from training, which is why an assistant can name a national chain with no lookup at all. For anything local and current it runs a live search, retrieves a small number of pages, and writes the answer out of those. For local questions the retrieved pages dominate, which is why coverage of your name across ordinary web pages is what decides the answer.

Last reviewed: September 1, 2026Skip to the checklist

01The detail

The two clocks

Training data is frozen at a point in time and updates only when a new model ships. It is good at broad associations, which is why assistants know the big brands in every category. It is unreliable for a nine-person HVAC company that changed its name last year.

Retrieval is live. OpenAI, Google, Anthropic, and Perplexity all publish documentation on the crawlers and search tools their assistants use to fetch pages at answer time. That is the half you can affect this month.

No, it is not reading Google Maps

This is the most common misconception, and it has an expensive consequence. Assistants do not have a live feed of the map. What they have is the enormous sediment of web pages that were themselves written from map data: directory pages, roundups, review sites, aggregators, and your own site. Your Business Profile matters enormously, but it matters through those pages, not directly.

Google's own AI surfaces are the exception worth knowing. AI Overviews and AI Mode sit inside Google Search and do draw on Google's index and local data. If you care about those specifically, they behave more like search than like a chat assistant.

What our own corpus shows

The table below is aggregated from every source domain the assistants actually pointed at across the scans stored in this product, refreshed a few times a day. It is not a survey and not an estimate. Two things tend to surprise owners: how large the long tail of ordinary business websites is, and how far down the list the big directories sit.

02Use this

Retrieval-readiness check

Ten things that decide whether a page about you can be fetched and used. All ten are checkable in an hour and none require a developer for more than one of them.

Can it be fetched

  • Your robots.txt does not block AI crawlers you want reading you. Check for GPTBot, OAI-SearchBot, Google-Extended, ClaudeBot, and PerplexityBot by name.
  • Key pages return normal HTML without requiring JavaScript to show their text. Disable JavaScript in your browser and reload your top service page.
  • No login, no interstitial, no cookie wall in front of the content.
  • Pages load in a couple of seconds. Slow pages get abandoned by fetchers as well as by people.

Can it be understood

  • The business name appears as text on every page, not only in the logo image.
  • Each service page has one clear subject and says the city in the first hundred words.
  • Hours, phone, and service area are text.
  • The page has a visible last-updated or published date.

Is it worth quoting

  • There is at least one number, range, or specific commitment on the page that nobody else in your town publishes.
  • There is a sentence that would survive being pasted into an answer with no surrounding context.
03Our data

23,926

citations recorded in AI answers

798

distinct source domains cited

85%

of them point at a business's own website

The non-business domains assistants cited most

A business’s own website is the single largest bucket, at 85% of all citations. These are the platforms behind the rest.

  1. 01angi.com595
  2. 02yelp.com484
  3. 03reddit.com468
  4. 04bbb.org299
  5. 05homeadvisor.com222
  6. 06thumbtack.com218
  7. 07consumeraffairs.com211
  8. 08tripadvisor.com190
  9. 09cslb.ca.gov127
  10. 10irs.gov69

Computed live from 182 stored scans covering 23,926 citations across 798 distinct domains. Assistants were asked real buyer questions unprimed; every source domain they cited or linked was recorded and classified. Full method and sample.

04What this is called

You are looking for the term

Grounding and retrieval

The live-search half of this has a name: retrieval-augmented generation, usually shortened to RAG, and the practice of tying an answer to fetched sources is called grounding. When a vendor tells you your content needs to be "groundable", this is the thing they mean, and the checklist above is the whole of it.

05Also asked
Should I block AI crawlers?
Not if you want to be recommended. Blocking GPTBot or PerplexityBot removes your pages from the pool an assistant can retrieve, which is the opposite of the goal for a local business that needs to be found.
Does ChatGPT use Yelp?
Sometimes, among many other sources. In our corpus the review platforms are a real but minority share of citations, well behind business and product websites.
06Sources
  1. 01Overview of OpenAI crawlersOpenAI
  2. 02Web search toolOpenAI
  3. 03Web search toolAnthropic
  4. 04Perplexity crawlersPerplexity
  5. 05AI features and your websiteGoogle Search Central

Every link above was opened and confirmed reachable on September 1, 2026.

The automated version

See what ChatGPT, Claude, Gemini and Perplexity say about your business

The free scan runs the buyer questions for your trade and city across all four assistants, shows who gets named instead of you, and lists the sources those answers came from. No signup, about thirty seconds.

Run my free scan
07Keep going

Other questions owners ask

All questions