Tellr · AI Search

ChatGPT Citations: How ChatGPT Chooses Which Sites to Cite

How does ChatGPT decide which sites to cite? Learn when citations are real, why some are wrong, and what makes pages more likely to get linked.

By Tellr Editorial TeamPublished 7 October 2026

ChatGPT citations come from live retrieval. When ChatGPT searches the web, it pulls pages that match the query, ranks them for relevance, authority, freshness and accessibility, and links the pages whose passages it actually used to write the answer. When it answers from model memory alone, with no search, it has no pages to point to, and any "references" it produces may be invented.

That split explains most of the confusion. A researcher who asks for sources with search turned off can get a confident list of papers that do not exist. A marketer who sees a competitor linked in a search-enabled answer is looking at a different system, one that selects real pages through a retrieval pipeline. As of October 2026, both behaviors exist side by side, and how ChatGPT "chooses" a site depends on which one is running.

Key takeaways

  • ChatGPT only cites real, clickable pages when it retrieves content through search, connected data or uploaded documents, and answers from model memory have no verifiable source behind them.
  • In search mode, a page has to be crawlable, relevant to the exact query and easy to extract a passage from before its authority matters at all.
  • Pages that answer the question in the first few sentences, under clear headings, with visible dates and authors, are easier for retrieval systems to quote.
  • A linked citation does not guarantee the page supports the claim, so every source should be opened and checked against the sentence it is attached to.

When does ChatGPT cite sources, and when does it not?

ChatGPT cites real sources only when its answer is grounded in content it retrieved during the conversation. Without retrieval, it generates text from patterns learned in training and cannot trace that text back to a page.

Three terms get mixed up here. A citation is a marker tied to a specific claim. A linked source is a page listed with the answer, which may or may not support a given sentence. A reference is a bibliography-style entry the model writes as text, and it can be fabricated when no retrieval happened.

ModeWhere the answer comes fromWhat citations look likeReliability
No web access (model memory)Training dataTyped references, if requestedLow, since references may be invented
Search / browsingLive web pages fetched for the queryInline links to retrieved pagesHigher. The pages exist, but support must be checked
Connected data sourcesLinked drives, mailboxes, appsLinks to internal filesDepends on the files and their freshness
Enterprise / document-groundedUploaded or indexed company documentsPointers to specific documents or passagesHigh, but limited to what was indexed

How does ChatGPT choose which sites to cite in search mode?

In search mode, ChatGPT selects sites through retrieval first and generation second. A search step decides which pages are candidates, and the model cites the ones whose passages it used to write the answer.

  1. Query rewriting. The model turns the prompt into one or more search queries, often several narrower ones for a broad question.
  2. Candidate retrieval. A search index returns matching pages. Blocked, unindexed or login-gated pages never enter the pool.
  3. Fetching and chunking. Pages are fetched and split into passages, so the model works with specific sections rather than whole pages.
  4. Ranking. Passages are scored on relevance to the query, source authority, freshness for time-sensitive topics and how cleanly they answer the question.
  5. Grounded writing. The model writes from the top passages and attaches citations to the pages they came from.

Authority alone is not enough. A highly trusted site that is not crawlable, or that buries the answer in paragraph twelve, loses to a less famous page that states it plainly. Our guide on how large language models decide what to cite covers the ranking side in more depth.

What makes a site more likely to be cited?

A site is more likely to be cited when its pages are open to crawlers, match the query closely, state the answer near the top and show that the content is current and written by someone credible.

  • Crawl access: robots.txt allows OpenAI's search crawler (OAI-SearchBot), key content renders in HTML rather than only through client-side JavaScript, and pages are not gated.
  • Answer near the top: the first two or three sentences under each heading answer it directly and make sense when quoted alone.
  • Question-shaped headings: H2s and H3s mirror how buyers phrase queries, so passages map cleanly to sub-questions.
  • Tight topical focus: one page per question beats a long page covering ten loosely related topics.
  • Original evidence: the page offers first-party data, test results, pricing details or named examples that other pages lack.
  • Visible author and date: a named expert and a clear "updated" date help on topics where freshness matters.
  • Structured comparisons: tables with consistent attributes are easy to extract for "X vs Y" and "best X" prompts.
  • Third-party presence: reviews, Reddit threads and comparison pages that mention you increase the number of retrieved pages that can name your brand.

Rewrite the opening of your ten most commercial pages so the first sentence answers each page's main query on its own. It is the cheapest change with the most direct effect on extractability.

Our walkthrough on showing up in ChatGPT answers covers the on-page work in more detail.

Why does ChatGPT sometimes cite the wrong source or invent one?

ChatGPT cites the wrong source or invents one because retrieval and generation are separate steps. The model can write claims its retrieved pages do not support, or write references when it retrieved nothing at all.

This has been known since launch. SANS noted that ChatGPT will give you references when asked, but they are not necessarily the sources from which it is formulating its answer. These are the common failure modes.

  • Hallucinated references: plausible titles, authors and DOIs for papers that do not exist, typical with search off.
  • Misattribution: a real link attached to a claim the page never makes.
  • Aggregator citations: a listicle cited instead of the primary study or vendor page.
  • Stale information: an older page outranks the current one, so outdated pricing or specs get repeated.
  • Inaccessible sources: the best page is paywalled or blocked, so a weaker one gets cited.
  • Manipulated pages: retrieved content can carry hidden instructions. The Cloud Security Alliance describes ChatGPhish as a Cross-Site Prompt Injection Attack disclosed by Permiso Security on May 29, 2026 that exploits how ChatGPT processes web pages.

How does citation selection change by query type?

Each query type weights relevance, authority and freshness differently, so the same site can be cited for one prompt and ignored for the next.

Query typeExampleSources that tend to winMain selection factor
Factual"What is SOC 2 Type II?"Standards bodies, reference sites, clear explainersDefinition near the top
Product comparison"Best CSPM tools for AWS"Comparison pages, review sites, Reddit threads, vendor pagesStructured attributes and third-party mentions
Breaking news"Latest ransomware attack on hospitals"News outlets, security advisoriesFreshness
YMYL (health, finance, legal)"Is this drug safe with alcohol?"Government and institutional sourcesAuthority and expert signals
Niche B2B"How to scope IAM for a multi-cloud org"Vendor documentation, practitioner blogs, forumsSpecific, detailed coverage

Niche B2B gives mid-market and enterprise brands the most room. Few pages answer the exact question, so a focused, well-structured page can become the default source. Our breakdown of which sources ChatGPT and Perplexity quote shows the patterns by category.

How do you verify a ChatGPT citation and cite ChatGPT itself?

To verify a ChatGPT citation, open the source, confirm it exists and says what the answer claims, and check that it is primary and current. To cite ChatGPT itself, treat it as software rather than a publication.

  1. Confirm it exists: click the link or search the exact title and DOI.
  2. Match the claim: find the specific sentence or figure on the page.
  3. Check primary vs secondary: trace summaries back to the original study, filing or vendor page.
  4. Check the date: confirm the page is current enough for the topic.
  5. Save the transcript: keep the full prompt and response, since ChatGPT does not provide a reliable stored list of cited sources you can retrieve later.

How do I cite ChatGPT in APA? Use OpenAI as the author, then the year, the tool name, a description and the URL, for example: OpenAI. (2026). ChatGPT [Large language model]. https://chatgpt.com/. ChatGPT can format APA or MLA citations when you give it the source details, but it is not a reliable automatic citation generator, so check every field.

How Tellr helps brands earn ChatGPT citations

Tellr runs a managed earned-visibility program that puts enterprise brands inside the pages AI answers retrieve, including Reddit threads and comparison content, so buyers see the brand when they ask ChatGPT about the category. Trackers show where you are missing. A senior Tellr team does the work that changes it, as one governed program with approvals and an audit trail. It is built for marketing teams spending $10k+ a month at companies worth $500M+ or with 200+ employees, so smaller teams will get more from self-serve tools.

  • Tellr writes answer-shaped articles, comparison pages and reviews meant to be quoted by ChatGPT, Perplexity and Google AI Overviews, and publishes them to your CMS.
  • Every week, Tellr tracks who Google and its AI Overviews cite for your category's queries.
  • Tellr maps relevant subreddits, runs a daily thread radar and drafts guideline-checked Reddit replies that go out only after approval.

ChatGPT citations reward pages that are accessible, focused, current and easy to quote, and they punish anyone who trusts a link without checking it. Build pages retrieval systems can extract, earn mentions on the third-party sources they favor, and verify every source before you repeat it.

FAQ

When does ChatGPT cite real sources?

ChatGPT cites real, clickable sources only when its answer is grounded in retrieved content, such as live web search, connected data sources or uploaded documents. When it answers from model memory alone, it has no verifiable page behind the response.

How does ChatGPT choose which sites to cite in search mode?

In search mode, ChatGPT rewrites the query, retrieves candidate pages from a search index, splits them into passages, ranks those passages by relevance, authority, freshness and answer clarity, then cites the pages whose passages it actually used to write the answer.

What makes a page more likely to be cited by ChatGPT?

Pages are more likely to be cited when they are crawlable, closely match the query, answer the question near the top, use clear question-shaped headings, show visible author and update details, and include original evidence or structured comparisons that are easy to extract.

Why does ChatGPT sometimes cite the wrong source or invent one?

This happens because retrieval and generation are separate steps. ChatGPT can attach a real link to a claim the page does not support, cite an aggregator instead of the primary source, repeat stale information, or invent bibliography-style references entirely when no retrieval happened.

How do you verify a ChatGPT citation?

Open the linked source, confirm it exists, check that it actually supports the specific claim, trace secondary summaries back to the primary source, and confirm the page is current enough for the topic. Saving the full prompt and response also helps because cited sources may not be easy to retrieve later.