How ChatGPT Decides Whom to Cite: The Four Layers Behind Every Recommendation

How ChatGPT chooses sources: retrieval, trust, extractability and brand layers, what Reddit's purge changed, a 12-point self-audit and a 20-prompt test.

MOVE Agency graphic cover for the article: How ChatGPT Decides Whom to Cite: The Four Layers Behind Every Recommendation

In short

ChatGPT doesn’t rank websites; it composes an answer and picks the sources it trusts enough to rely on. That choice runs through four layers: retrieval (can the page be found in Bing and fetched by OAI-SearchBot and GPTBot), trust (do third parties mention the brand consistently, with reviews and matching facts), extractability (does the page give a structured passage with a date and a direct answer the model can lift) and brand entity (is it clear what the company is, with the same name and description everywhere). Most brands lose already at the first layer, since about 73% of sites have technical barriers for AI crawlers per Otterly.ai, and at the second, since about 85% of AI visibility is built off-site.

Below is the mechanism, a 12-point self-audit and a 20-prompt test you can run this week.

Four layers, one decision

When someone asks ChatGPT “which payroll service is best for a 50-person company in Poland,” the answer that comes back isn’t a search results page. It’s a synthesis: the model reads what it found, combines it with what it already knows, and writes two or three recommendations with links where it drew on a specific source.

Nobody outside OpenAI knows the exact formula, and it changes. But every audit we run points to the same four layers, and the order matters: failing at an earlier layer makes the later ones irrelevant.

LayerQuestion the model is effectively askingWhat blocks you
1. RetrievalCan I find and read this page right now?Not in Bing, bots blocked, JavaScript-only content, stale pages
2. TrustDo other sources agree this brand deserves trust for this task?Few third-party mentions, conflicting facts, no reviews
3. ExtractabilityIs there a passage I can quote as the answer?Long narratives, no definitions or numbers, no date, no author
4. Brand entityDo I clearly understand what this company is?Inconsistent name spellings, vague descriptions, no structured public record

Layer 1: retrieval, or whether you exist at all

ChatGPT Search relies on Bing as a third-party search data provider, and also on OpenAI’s own crawlers: OAI-SearchBot (for search retrieval) and GPTBot (for broader crawling). Three things decide whether a page can be pulled in:

  • Crawler access. About 73% of sites have technical barriers for AI crawlers (GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot): robots.txt blocks, CDN or WAF restrictions, or JavaScript-only rendering, often without anyone deciding this on purpose. Default CMS settings and CDN “AI protection” toggles do it silently. Open the file and look.
  • Bing indexation. A site well indexed in Google can be thinly indexed in Bing. Register in Bing Webmaster Tools and submit pages through IndexNow so Bing finds new and updated pages faster. OpenAI doesn’t confirm this directly speeds up appearing in ChatGPT, but without Bing indexation the page can’t be pulled in at all.
  • Rendering and freshness. Critical content needs to be in the HTML, not loaded only through JavaScript. Pages need a visible, current date. Freshness matters more than length: a 600-word page updated this quarter beats a 3,000-word page from three years ago.

The full technical checklist is in our guide to getting your site indexed by ChatGPT. Fix this layer first: it’s the cheapest and the most common failure.

Layer 2: trust, or what everyone else says about you

Here’s the part most companies get backward. They polish their own site and wonder why a competitor with a worse one gets recommended. The answer is that about 85% of a brand’s visibility in AI answers is built away from its own site: in media, directories, review platforms, forums and YouTube.

Three trust signals stand out:

  1. Third-party mentions. Named brand mentions on other sites correlate with appearing in Google AI Overviews about three times more strongly than backlinks. The model is effectively counting how often trusted sources bring you up in the context of a task. A comparison article that lists you among five options counts; a bare link in a footer barely does.
  2. Consistency of facts across sources. If your site says the company was founded in 2015, a directory says 2017, and a review says you’re based in a different city, the model’s confidence drops and it reaches for a competitor whose story holds together. Audit the top twenty places your brand appears and align the basic facts.
  3. Reviews. Review platforms are among the sources most often pulled in for “best X” queries. Volume, recency and specificity of reviews all feed the trust layer. A hundred reviews naming the specific service you want to be recommended for beat a thousand generic five-star ratings.

Layer 3: extractability, or whether there’s anything to quote

Once a page is found and the brand is trusted, the model still needs a passage it can use. Pages built for extraction share a pattern:

  • Answer first. The first paragraph under a heading answers the question in that heading. No warm-up.
  • Definitions. “X is Y” sentences get lifted into answers verbatim more than any other structure.
  • Tables. Comparisons, price ranges and spec lists in tables are easy for the model to parse and cite.
  • Numbers with a source and a date. “14.2% conversion, Opollo, 2026” is citable; “much higher conversion” isn’t.
  • A named author with a role. Expertise the model can attribute to someone is expertise it can trust.
  • A visible publication or update date. Feeds the freshness signal from layer 1.

What doesn’t help here, despite being sold heavily: per Ahrefs’ research from May 2026, on 1,885 pages that added JSON-LD markup against 4,000 control pages, citations in Google AI Overviews fell 4.6%, while Google AI Mode and ChatGPT rose 2.2-2.4%, statistically close to zero. llms.txt remains optional hygiene with no confirmed effect on visibility. Structure the visible text; don’t expect hidden markup to do the work.

Layer 4: brand entity, or whether the model knows who you are

Models reason in entities, not URLs. To recommend “Acme Clinic,” the model needs a stable idea of what Acme Clinic is: a dermatology clinic in Warsaw, founded in a given year, with certain procedures and a certain price level.

Entity clarity comes from repetition and consistency:

  • Use one canonical brand name everywhere (not “Acme,” “Acme Clinic” and “Acme Medical” interchangeably).
  • Keep the same one-sentence description on your site, in directories, in social bios, in your press kit.
  • Address, phone number, founding year and specialties should match across every listing.
  • Where your category has structured public registries, like Wikidata or established industry directories, make sure your entry exists and matches your site.
  • Publish an “About” page that states the facts plainly, so a crawler can confirm what third parties say.

What Reddit’s purge means for you

Reddit is the most-cited domain in AI answers, which made it an obvious target for manipulation. Reddit itself has reported removing an average of 100,000 automated accounts a day, and separately blocking about 23 million spam views a day, as part of its fight against bots. A permanent ban for astroturfing isn’t a platform-wide policy: it’s a rule only certain communities enforce on their own.

Two consequences follow. First, seeded threads and fake “who would you recommend” posts are increasingly detected and removed along with the accounts that created them. Second, genuine participation is worth more now, because there’s less noise around it. An employee who openly discloses where they work and answers a real question with something useful and specific: that’s exactly the kind of content that survives the purges, and exactly what the model retrieves.

A 12-point self-audit

Score yourself honestly. Each row is a yes or a no.

#LayerCheckYes/No
1Retrievalrobots.txt allows OAI-SearchBot and GPTBot
2RetrievalSite is registered in Bing Webmaster Tools and submits via IndexNow
3RetrievalKey pages render their main content in HTML without JavaScript
4RetrievalKey pages show a visible date updated within the last 12 months
5TrustBrand is named in at least one comparison or “best X” article per core service
6TrustTop 20 third-party listings show identical name, address, founding year, description
7TrustReviews on the main platforms for your category are recent and mention specific services
8ExtractabilityEach core page opens with a direct answer or definition under its H1
9ExtractabilityPrices, comparisons or specs appear in tables
10ExtractabilityArticles carry a named author with a role
11Brand entityOne canonical brand name and one-sentence description used everywhere
12Brand entityStructured public records (Wikidata, industry directories) exist and match the site

Fewer than six “yes” answers means the model has little reason to cite you and not enough material to do it even if it wanted to. Six to nine means the foundation is there and off-site work will move the needle. Ten or more means you’re competing on trust and freshness, which is where share of voice gets won.

How to test with 20 buying prompts

You don’t need a tool for a first read. You need an hour and a spreadsheet.

  1. Write 20 prompts a real buyer would type. Mix them: 5 category prompts (“best CRM for a small sales team”), 5 local prompts (“dermatology clinic in Warsaw with English-speaking doctors”), 5 comparison prompts (“X vs Y for a company like ours”), 5 problem prompts (“we keep losing leads between marketing and sales, what should we use”).
  2. Run each prompt in ChatGPT with search turned on, then repeat in Copilot and Google AI Mode. Use a clean session for each.
  3. Record four things per answer: which brands got named, in what order, which sources were cited, and whether you showed up at all.
  4. Repeat the same 20 prompts a week later. Since only about 30% of brands stay visible in consecutive answers, one run is just noise. Two runs give you a rough share of voice.
  5. Read the cited sources. This is the real output. The directories, review sites and articles that keep coming up are exactly where you need a presence. That list is your off-site plan.

Established trackers (Otterly from $29 a month, Peec from €89, Profound from $99) automate steps 2 through 4 once you want monthly tracking against competitors.

Six mistakes that keep brands out of ChatGPT answers

  1. Auditing Google and assuming Bing will follow on its own. It won’t. Check Bing indexation separately.
  2. Blocking AI crawlers “just in case” through robots.txt, CDN rules or JavaScript-only rendering, and forgetting about it. About 73% of sites have technical barriers for AI crawlers. Check yours today.
  3. Spending the whole budget on-site. At most 15% of the outcome lives there.
  4. Letting facts drift between listings. Every inconsistency is a small vote against you in the trust layer.
  5. Writing for length instead of for citation. A tight, dated 800 words with sources beats a sprawling 3,000.
  6. Measuring referral clicks only. ChatGPT often names a brand without a link. If clicks are your only metric, you’ll write off a channel that’s quietly shaping shortlists. Track mentions and assisted conversions.

Get the audit done for you

At MOVE we run exactly this process as a paid AI visibility audit starting at $1,500: your category’s buying prompts across ChatGPT, Copilot and Google AI Mode, a score on all four layers, the list of cited sources turned into an off-site plan, and a baseline share of voice you can hold us to.

Where the market allows it, we pair this with paid placements through ChatGPT Ads and Microsoft AI Max, so the brand shows up in those conversations while the organic work builds up (see our guide to launching ChatGPT Ads step by step). We don’t promise rankings: about 30% persistence makes that a promise nobody can keep. Book an AI visibility audit on the AI advertising service page or contact us with your site and three competitors, and we’ll send back the 20 prompts we’d test first.

Frequently asked questions

Does ChatGPT use Google to find sources?
Not exactly. Per [OpenAI's documentation](https://developers.openai.com/api/docs/bots), ChatGPT's search relies on Bing as a third-party search data provider, and it also crawls and indexes sites directly through its own crawler, OAI-SearchBot. If your site is missing from Bing or blocks these bots in robots.txt, it can't be pulled in as a source, whatever your Google rankings look like.
Why does ChatGPT recommend a competitor with a worse website?
Because most of the decision happens off your site. Per [AirOps'](https://www.airops.com/report/the-influence-of-offsite-signals-in-ai-search) analysis, about 85% of brand mentions in AI answers come from third-party sources rather than the brand's own site, and per [Ahrefs'](https://ahrefs.com/blog/ai-overview-brand-correlation/) research, such mentions correlate with appearing in Google AI Overviews about three times more strongly than backlinks. A competitor with a plain site but consistent, frequent mentions in trusted directories and reviews gets the recommendation.
Do backlinks still matter for ChatGPT citations?
Indirectly, yes, through Bing rankings and overall authority, but brand mentions are the stronger signal for AI visibility. Shift part of your link-building budget toward named mentions in the sources ChatGPT already cites for your category.
How stable is a ChatGPT citation once you get it?
Not very. Per the [AirOps](https://www.airops.com/report/the-2026-state-of-ai-search) report The 2026 State of AI Search, only about 30% of brands stay visible in ChatGPT's next answer to the same query. Treat visibility as a share of mentions you track monthly, not a position you win once and for all.
Can I just post about my brand on Reddit to get cited?
Not safely. Reddit is the most-cited domain in AI answers per [Peec AI's](https://peec.ai/blog/top-domains-cited-by-ai-search-analysis-based-on-30m-sources) analysis, and Reddit itself has reported removing an average of 100,000 automated accounts a day as part of its fight against bots. Only open, genuinely helpful participation is worth doing.
Add MOVE to your Google sources