"Where does ChatGPT get its information" has two honest answers, and most articles only give the first one. The first is the training corpus — the frozen snapshot of text the model learned from, which explains why it can define photosynthesis instantly and cannot tell you today's Fed rate. The second is retrieval: a search that happens while you wait, whose results the model reads and links.
Which one you get decides whether the answer can be checked, whether it can be current, and — if you run a website — whether you can do anything about it. So rather than describe the plumbing from OpenAI's documentation and stop, I logged what actually came out.
The short answer, in one table
| Supply | What it is | Can you see the source? | Can you influence it? |
|---|---|---|---|
| Pretraining data | Publicly available internet content, licensed third-party data, and material from users, human trainers and researchers[1] | No — the model stores parameters, not documents | Only by opting out of GPTBot before the next training run[3] |
| Live web retrieval | Rewritten queries sent to third-party search providers, pages read at answer time[2] | Yes — cited as links | Yes — allow OAI-SearchBot, be the page that answers[4] |
| Licensed partner content | Deals with roughly 20 media organisations covering 160+ outlets[9], plus Reddit's Data API[8] | Sometimes, as a normal citation | Only by being one of them |
| Product & shopping data | Structured metadata from first- and third-party providers; Shopify Catalog is wired in directly[5] | Partly — merchant links | Yes — product feeds, accurate catalog data |
| Your own context | Memory, uploaded files, connectors, custom instructions[2] | Yes, to you | Not applicable — it is yours |
Everything below is about which of these fires, and when.
How we measured it
I wanted the thing a normal person sees, not an API simulation. So the prompts went through the chatgpt.com interface itself, captured with DataForSEO's ChatGPT scraper[13], US locale, English, logged out.
- 144 questions across 12 intent types — stable facts, current events, consumer "best X", B2B software, comparisons, pricing, health, personal finance, how-to, local, brand reputation, and AI/marketing.
- Each question run twice, minutes apart, to measure how stable the sources are. That gives 288 attempts; 287 returned an answer.
- 894 cited links captured with their titles and domains.
- For each question we also pulled Google's organic top 10 for the identical query text, same locale, same day, to test the "rank in Google, get cited by ChatGPT" assumption.
- Dates: 17–18 September 2026.
Limits worth stating up front: this is one locale, one language, one two-day window, a logged-out session with no memory, and a sample of 144 questions I wrote. It measures what ChatGPT cited, which is not the same as everything it consulted — a page can inform an answer and never appear as a link. Read the percentages as the shape of a system, not as constants.
ChatGPT does not search for most of what it is asked — it searches for two-thirds
Across 287 answers, 195 (68%) showed cited sources and 92 showed none at all. The pattern in which questions trigger a search is the most useful thing in this dataset.
Source: Attrifast, 17–18 September 2026 — 144 US-English prompts run twice on chatgpt.com, 287 answers
Read it as a spectrum of volatility. Questions whose answers change — what a tool costs, what opened downtown, which laptop is best this year — get searched almost every time. Questions whose answers have not changed since the model was trained get answered from memory:
- "What causes the seasons on Earth?" — no search, no sources, 12 of 12 stable-fact prompts, both runs.
- "How much does Shopify cost per month?" — searched in 96% of pricing prompts, and 96% of those citations pointed at the vendor's own pricing page.
- "What is the best CRM for a small business?" — searched every single time.
Two categories break the pattern in ways worth knowing. Health questions searched only 17% of the time — ChatGPT answered "how much sleep does an adult need" from training — but when a health question did trigger a search, it cited 7.2 sources, more than any other category, and 100% of those were government, academic or nonprofit domains: CDC, NIH, the American Heart Association. The model appears to treat medical claims as needing institutional backing when it reaches for backing at all. How-to questions behaved similarly: searched only 38% of the time, but 6.9 sources when they were, overwhelmingly official documentation.
This maps onto how people actually use the product. OpenAI's own research with NBER found "Seeking Information" grew from 14% to 24% of all ChatGPT messages in a year, and describes that category as "a very close substitute for web search"[12]. A quarter of usage is the part where citations — and therefore your website — can exist at all.
What ChatGPT cites: brands' own pages, then institutions
Here is where the popular advice diverges hardest from the measurement.
Source: Attrifast, 17–18 September 2026 — 894 cited links, domains hand-classified
Nearly half of every link ChatGPT showed us pointed at a company's own website. Not a review of the company — the company. shopify.com/pricing. support.google.com help articles. tesla.com. apple.com. When the question was about what something costs, brand-site dominance hit 96%.
The most-cited domains across the whole run:
| Domain | Answers citing it | Type |
|---|---|---|
google.com (support, workspace, store) | 17 | Brand / vendor |
openai.com | 12 | Brand / vendor |
hubspot.com | 9 | Brand / vendor |
theinfatuation.com | 9 | Local guide |
reuters.com | 8 | News |
apple.com | 7 | Brand / vendor |
rtings.com | 7 | Review publisher |
shopify.com | 7 | Brand / vendor |
tomsguide.com | 6 | Review publisher |
technologyadvice.com | 6 | Review publisher |
And the number that surprised me most: across 287 answers and 894 links, Wikipedia was cited zero times. Reddit was cited zero times. YouTube once.
That is worth handling carefully, because it contradicts a lot of published research — including studies of Google's AI surfaces, where Reddit and YouTube are consistently the two most-cited domains. Three explanations, and I cannot fully separate them with this data:
- Wikipedia is a training source, not a citation. It was 3% of the GPT-3 training mix[7]. A model that absorbed Wikipedia does not need to cite it to repeat it — which is precisely the invisible-contribution problem.
- Different surface, different behaviour. Google's AI Overviews cite the open web because they are grounded in Google's index. ChatGPT's search stack, as we will see, overlaps with Google's top 10 only about a tenth of the time.
- Question mix. My 144 questions skew commercial and practical. Ask "what is the Treaty of Westphalia" and the encyclopedia would likely reappear — although note that stable factual questions triggered no search at all in this sample.
The practical reading is the same in all three cases: for the commercial questions where a business hopes to be mentioned, ChatGPT prefers primary sources — the vendor's own page — and institutional ones, and only then the review press.
Ranking in Google does not get you cited
For every question we pulled Google's organic top 10 for the same text on the same day, then checked each ChatGPT citation against it.
Source: Attrifast, 17–18 September 2026 — ChatGPT citations vs Google organic top 10 for the identical query
10.6% of cited URLs were in Google's top 10 for the same question. Another 32.1% came from a domain that ranked in the top 10 but on a different page — the brand ranks, and ChatGPT cites a deeper page like its pricing or docs. And 57.3% came from domains that were nowhere in Google's first page at all.
This lines up almost exactly with the largest independent study of the question: Ahrefs analysed 15,000 long-tail queries and found ChatGPT's in-text citations matched Google's top 10 about 8% of the time, with a cross-assistant average of 12%[11]. Two different methods, two different samples, the same conclusion — so treat the ~10% figure as reasonably solid.
The gap is largest exactly where marketers care most: for B2B software and consumer "best X" questions, three-quarters of citations came from domains absent from Google's top 10. It is smallest for pricing questions, where 68% of citations were at least the right domain — because there is only one authoritative page for what a product costs, and both systems find it.
We wrote up the mechanism behind this — query fan-out, where one question becomes several different searches — in how AI engines choose which sources to cite and ChatGPT query fan-out explained. The short version: you are not being retrieved for the question the user typed.
Ask the same question twice, get a different bibliography
Every question ran twice. The search decision itself was stable — 94.4% of pairs either both searched or both did not. The sources were not.
Source: Attrifast, 17–18 September 2026 — mean overlap between two runs of each prompt, 93 prompts where both runs searched
Averaged across every repeated pair: 56% of domains and just 33% of exact URLs repeated. A third of pairs returned an identical set of domains; 8.6% shared no domain at all between two runs of the same question minutes apart.
Stability tracks how many valid answers exist. "How much does Google Workspace cost" has one right source, and pricing questions held 88% domain overlap. "What are the best noise cancelling headphones" has fifty defensible sources, and consumer "best X" questions dropped to 30%.
If you check ChatGPT once to see whether it mentions your brand, you are sampling a distribution with a single draw. That is the entire argument for tracking prompts on a schedule rather than by hand, which we made with different data in AI visibility prompts.
The training layer: what is actually in there
For everything ChatGPT answers without searching, the source is a corpus nobody outside OpenAI has seen in full. The company describes three inputs: publicly available internet content, information licensed from third parties, and information users, human trainers and researchers provide or generate[1]. It also says it increasingly uses synthetic data, excludes sources that aggregate large amounts of personal data, and respects GPTBot in robots.txt as the opt-out for training[1][3].
The last time OpenAI published an actual recipe was GPT-3, in 2020:
Source: Brown et al., Language Models are Few-Shot Learners (arXiv:2005.14165), Table 2.2 — weight in training mix
Common Crawl — a public archive of scraped web pages — was 60% of the training mix, WebText2 22%, two book corpora 8% each, and Wikipedia 3%[7]. Note that these are sampling weights, not raw size: Wikipedia is small but was sampled 3.4 times, while the far larger Common Crawl was sampled less than once, because OpenAI weighted what it considered higher quality more heavily. No comparable table has been published for GPT-4 or anything after it. Anyone who tells you the current percentage breakdown is guessing.
Since then the visible additions have been licensing deals rather than crawling: Reddit's Data API for "real-time, structured, and unique content"[8], and content partnerships with around 20 media organisations spanning more than 160 outlets[9].
Why the cutoff still matters
Every model carries a date after which it knows nothing, and OpenAI publishes them per model[6].
Source: OpenAI model documentation, read 17 September 2026 — months from published knowledge cutoff
On the day we ran this, the newest model in the lineup — GPT-6 Astra, cutoff 30 April 2026 — was four and a half months behind the world, and models still in wide use were one to three years behind. When ChatGPT searches, that gap closes to minutes. When it does not, the gap is the answer's age.
This is the single most useful thing a reader can take away for everyday use: an answer with no citations is an answer from memory, and its freshness is the model's cutoff, not today.
So does ChatGPT use Google, or Bing?
The most common version of this question deserves a direct answer, in three parts.
What OpenAI says. Its help centre states that ChatGPT search "sometimes partners with other search providers," rewrites your question into one or more targeted queries, and may send follow-up queries to other providers after reading the first results[2]. The only two providers whose privacy policies it links are Microsoft and Shopify, and it says it may share disassociated queries "with search engines like Bing"[2]. It never names Google.
What was reported. In August 2025, The Information reported that OpenAI had been using SerpApi — a scraping service that packages Google's results — to help answer real-time questions on news, sports and finance, and that OpenAI had asked Google to license search data and been refused[10].
What we measured. If ChatGPT were simply reading Google's first page back to you, the overlap would be high. It is 10.6%. Whatever mixture of Bing, its own OAI-SearchBot index, partner feeds and scraped results is in there, the output is not Google's top 10 — a conclusion that holds regardless of which provider contract is in force this quarter.
Compare that with the sibling question for Google's own AI, where the grounding is the Search index: where Google AI gets its information.
If you own a website: what this data tells you to do
Four things follow directly from the numbers above, in order of how much they matter.
1. Allow OAI-SearchBot, and know it is not GPTBot. These are separate robots.txt tokens with separate jobs: OAI-SearchBot decides whether you can appear in ChatGPT's search answers, GPTBot decides whether your content may be used for training, and ChatGPT-User is the agent that fetches a page because a user asked[3]. Blocking the training bot while allowing the search bot is a coherent, supported position. Blocking OAI-SearchBot removes you from the only layer you could have influenced. Our AI crawler directory has the user-agent strings and verification method for each.
2. Be the primary source for facts about yourself. Half of all citations were brands' own pages, and pricing questions cited the vendor's own page 96% of the time. If your price lives inside an image, behind a "contact sales" form, or in a JavaScript widget, you have removed yourself from the one category where you are the undisputed authority. Put the number in text.
3. For "best X" questions, your own site is not the lever. Consumer and B2B recommendation questions searched nearly every time, and leaned on review and comparison publishers — rtings.com, tomsguide.com, technologyadvice.com, nerdwallet.com — plus a long tail of small niche sites I had never heard of. Being in those roundups is a different job from publishing your own page, and it is the one that moves this needle.
4. Measure the clicks, because they are unusually easy to measure. Every one of the 894 cited links in our sample carried utm_source=chatgpt.com — 100%, no exceptions. OpenAI documents this in its publisher FAQ[4]. That means ChatGPT referrals are identifiable in analytics with none of the guesswork that AI traffic normally involves — provided your analytics does not fold them into Direct, which is the usual failure.
One small footnote on method, offered as evidence rather than a boast: of the 894 links, exactly one pointed at this site — for the question "How much website traffic comes from ChatGPT?" One citation out of 894 is roughly what a site our size should expect, and it landed on the page that answers that question with original data rather than on our homepage. That is the shape of the opportunity: specific questions, answered with numbers, on a page the search bot can read.
What this does not tell you
- It is not a ranking factor list. Nothing here reveals why one page was chosen over another — only what was chosen.
- Citations are not the whole input. Pages can shape an answer without being linked, and training data never shows up as a link at all.
- The sample is one slice. US English, logged out, no memory, two days in September 2026, 144 questions of my choosing. A different question mix would move every percentage, though probably not the direction of the findings.
- It changes. OpenAI has swapped search providers, added shopping feeds and signed licensing deals within single quarters. Numbers in this article are dated deliberately so you can tell when they have aged.
FAQ
Where does ChatGPT get its information?
From two separate places. The first is training: OpenAI says its foundation models are developed from publicly available internet content, information licensed from third parties, and information users, human trainers and researchers provide or generate. That knowledge is frozen at the model's cutoff date and carries no per-fact source. The second is retrieval: when ChatGPT searches, it rewrites your question, sends it to third-party search providers, reads the pages that come back, and cites them as links. In our test of 287 answers it searched for 68% of questions, and cited an average of 4.6 links when it did.
Does ChatGPT get its information from Google?
Partly, and OpenAI has never said so directly. Its help center names search providers without listing Google, linking to Microsoft's and Shopify's privacy policies. In August 2025 The Information reported that OpenAI had used SerpApi, a scraping service that packages Google results, to help answer real-time questions about news, sports and finance. What we can measure is the overlap: only 10.6% of the links ChatGPT cited were also in Google's top 10 for the same question, so whatever it queries, it does not simply read Google's first page back to you.
Does ChatGPT use Bing?
Microsoft is the only search engine whose privacy policy OpenAI links from its ChatGPT search help article, and OpenAI states it may share disassociated search queries with search engines like Bing to return web results. So Bing is in the pipeline, but it is not the whole pipeline: OpenAI describes sending rewritten queries to more than one provider, and it operates its own crawler, OAI-SearchBot, to build a search index of its own.
Does ChatGPT search the internet for every question?
No, and the split is sharp. Across 287 answers, ChatGPT searched for 68% of questions. It searched for every B2B software recommendation and every local question we asked, 96% of current-events and pricing questions — and 0% of stable factual questions like what causes the seasons or who wrote Pride and Prejudice. Those it answered from training alone, with no sources shown. Health was the surprise: only 17% of health questions triggered a search, but when one did, it cited more sources than any other category.
How current is ChatGPT's information?
It depends entirely on whether it searched. When it does not search, you get knowledge frozen at the model's cutoff: OpenAI's documentation puts GPT-6 Astra at 30 April 2026 and the GPT-5.6 family at 16 February 2026, so on 17 September 2026 the freshest model in the lineup was running on knowledge between four and seven months old. When it does search, it reads pages published minutes ago. The failure mode to watch for is a confident answer with no citations on a question whose answer changed after the cutoff.
How do I get my website cited by ChatGPT?
Allow OAI-SearchBot in robots.txt — it is the only crawler that governs whether you appear in ChatGPT's search answers, and it is separate from GPTBot, which governs training. Then make the page the kind of thing ChatGPT actually cites: 48% of the links in our sample were brands' own sites, most often a pricing page, a documentation page or a help-center article that states a fact in plain text. For questions where a buyer is comparing options, the citations go to review and comparison publishers instead, so your presence in those roundups matters more than your own copy.
Will ChatGPT cite the same sources if I ask the same question twice?
Usually not. We ran all 144 questions twice, minutes apart. On average only 33% of the exact URLs repeated between the two runs, and 56% of the domains. Pricing questions were the most stable at 88% domain overlap; consumer 'best X' questions were the least at 30%. A single check of what ChatGPT says about your brand is an anecdote, not a measurement — you need repeated runs before a change in citations means anything.
Continue the research path
- Where does Google AI get its information? — the same question for AI Overviews and AI Mode, where the grounding is Google's own index
- How AI engines choose which sources to cite — the retrieval mechanics behind the citation
- How much traffic actually comes from ChatGPT — the measured other end of the pipe
- How to track ChatGPT traffic in 2026 — the setup, including the Direct-traffic trap
- AI citation tracking, joined to revenue
Sources
Primary sources for the claims in this article. Numbered citations in the text link to the matching entry.
- [1]How ChatGPT and our foundation models are developed — OpenAI Help Center (2026).
- [2]Searching the web with ChatGPT — OpenAI Help Center (2026).
- [3]Overview of OpenAI Crawlers — OpenAI (2026).
- [4]Publishers and Developers - FAQ — OpenAI Help Center (2026).
- [5]Shopping with ChatGPT Search — OpenAI Help Center (2026).
- [6]Models — knowledge cutoff dates — OpenAI (2026).
- [7]Language Models are Few-Shot Learners (GPT-3), Table 2.2 — Brown et al., arXiv (2020).
- [8]OpenAI and Reddit Partnership — OpenAI (2024).
- [9]Partnering with Axios expands OpenAI's work with the news industry — OpenAI (2025).
- [10]ChatGPT's answers came from Google Search after all: Report — Search Engine Land (2025).
- [11]Only 12% of AI Cited URLs Rank in Google's Top 10 for the Original Prompt — Ahrefs (2025).
- [12]How People Use ChatGPT (NBER Working Paper 34255) — Chatterji et al., NBER (2025).
- [13]ChatGPT LLM scraper API (data collection method) — DataForSEO (2026).

