How to get cited by ChatGPT
Most advice about getting cited by ChatGPT describes how to format a page. That work matters, but it decides the last step of four. Before any of it applies, a search has to fire, your URL has to enter through the right retrieval channel, and your title has to match a question you never see.
The numbers behind those three gates are unforgiving. Profound found that roughly 18% of ChatGPT conversations trigger a web search at all. Ahrefs found that 88.46% of URLs arriving through the general search channel get cited, against 1.93% for Reddit. Formatting cannot rescue a page that never reaches the shortlist.
Does ChatGPT search the web at all?
Usually not. Profound measured web searches in about 18% of conversations, and the rate held steady across three months of data.
That single figure reframes the whole exercise. Four conversations in five are answered from the model's own parameters, with no retrieval and therefore no citation available to win. When a search does fire, it fires early: Profound's turn-by-turn breakdown shows citations decaying sharply as a conversation continues.
| Turn number | Share of turns carrying citations |
|---|---|
| 1 | 12.6% |
| 2 | 8.98% |
| 3 | 7.53% |
| 5 | 6.2% |
| 10 | 4.5% |
| 20 | 3.0% |
The practical reading: opening questions need factual grounding, while later turns tend to be clarifications, deeper dives or creative work that no longer needs the web. The queries worth targeting are the ones that start a research session, not the ones that continue it. Note the scope before leaning on these figures - the sample is around 730,000 conversations from US-based, English-language users on ChatGPT.com over October to December 2025, and every one of them already contained at least one citation.
Which retrieval channel are you in?
The one that matters is plain search. Ahrefs found citation rates differ by more than a factor of forty across ChatGPT's internal retrieval channels.
When ChatGPT retrieves a result it tags the source with an internal field describing where the URL came from. Ahrefs identified five such channels across 1.4 million prompts, and the gap between them is the single most actionable finding in the study.
| Retrieval channel | Citation rate | Data points |
|---|---|---|
| search | 88.46% | 25,563,589 |
| news | 12.01% | 3,940,537 |
| 1.93% | 16,182,976 | |
| youtube | 0.51% | 953,693 |
| academia | 0.40% | 185,337 |
Around 88% of everything ChatGPT cites comes straight from the general search index. That makes ordinary ranking the entry ticket rather than a legacy concern, which is the same conclusion the comparison of GEO and SEO as one workflow reaches from the other direction. YouTube and academic sources are pulled in at scale and almost never surface as attributed citations.
Two limits belong with this table. The search channel already contains Reddit and YouTube pages that come back through normal web results, so the separate channels appear to be additional feeds layered on top. And Ahrefs cautions that any study comparing cited against non-cited URLs without separating these channels risks reporting an artefact of its own dataset as a finding - a warning that applies to most of the citation research currently circulating.
How many retrieved pages actually get cited?
About half. Ahrefs found ChatGPT pulls roughly 16.57 cited and 16.58 non-cited URLs per prompt, then attributes only a handful of them.
The visible end of that funnel is narrower still. Profound counted around four unique citations in a turn that carries any, with 66% of cited turns showing between one and four sources, and roughly six unique citations across a whole conversation. So the contest is not for sole ownership of an answer. It is for one of about four slots, alongside sources the model has already decided to trust - frequently including direct competitors.
The wider market is unequal without being closed. Profound put the top ten domains at just 12% of all citations, with a Gini coefficient near 0.8. Wikipedia leads at roughly 5% of citations and appears in about 18% of cited conversations. Practically that means the realistic goal is not to displace the reference layer but to be the source that gets cited next to it.
Why is Reddit everywhere and cited almost never?
Because reading and crediting are different acts. Reddit supplies 67.8% of all non-cited URLs in the Ahrefs dataset while earning a 1.93% citation rate.
Over sixteen million Reddit data points entered the retrieval pipeline and almost none of them ended up as an attributed source. The plain reading is that ChatGPT uses Reddit heavily to understand a topic and gauge what people actually think, then hands the citation to a more institutional source. Profound's numbers point the same way from a different angle: Reddit is the second most cited domain there, at around 3% of citations, far behind its share of retrieval.
For anyone planning a distribution strategy this is a useful correction. Community presence still shapes what the model believes about your category, so it is not wasted effort. It is simply not a citation channel, and treating it as one confuses influence with attribution.
What gets a page selected once it is in the pool?
Semantic match between your title and the sub-question ChatGPT wrote for itself. Selection happens on retrieval metadata, before the page is opened at all.
Each retrieved result comes back with a title, a URL, sometimes a snippet, and an identifier. ChatGPT uses that payload to decide which pages are worth opening. There is a gatekeeping layer ahead of your content, and inside it the title is doing most of the work. Ahrefs measured how closely titles matched the query using cosine similarity: 0.602 for cited pages against the original prompt, 0.484 for non-cited ones, and 0.656 when cited titles were compared against the best-matching internal sub-question.
That last number is the important one, because it is higher than the score against the prompt itself. ChatGPT breaks a question into sub-questions and hunts facts for each, so the target is a query set you cannot see in any keyword tool. Two smaller findings sit alongside it: pages with natural-language URL slugs were cited 89.78% of the time against 81.11% for opaque ones, and the median cited page in the search channel was around 500 days old while non-cited pages skewed young. Freshness helps across the population and breaks ties in news, but within one retrieval set relevance decides.
Where does page structure actually matter?
After selection, not before. Structure decides whether a page is easy to quote once ChatGPT has already chosen to open it.
This is worth separating carefully, because the standard advice collapses two different jobs into one. Getting selected is a retrieval-metadata problem: title, URL, channel. Getting quoted is a body-copy problem: a self-contained answer under a question-shaped heading, no links inside the answer block, claims that stand on their own without the surrounding paragraphs. Both matter. They are not the same lever, and no amount of the second fixes a failure of the first.
The confusion shows up in the numbers people repeat. The advice circulating about answer capsules carries two different lengths - roughly 40 to 60 words in one telling, 120 to 150 characters in another - which are not compatible. Our own reading is that they describe two different objects: a page-level opening answer runs 40 to 60 words, while a section-level capsule under a heading sits nearer 20 to 25. Treat any single number quoted without saying which one it applies to as unreliable. The mechanics of both are in the guide to write answer-first content, and the habit that makes claims quotable is pairing them with named sources.
Which crawler does ChatGPT use?
OAI-SearchBot for search. OpenAI now documents four separate agents, each with its own job, published IP list and independent robots.txt setting.
OAI-SearchBot surfaces websites in ChatGPT's search features; opting out removes a site from search answers, though it can still appear as a navigational link. GPTBot crawls content that may train the foundation models. ChatGPT-User handles user-initiated actions in ChatGPT and Custom GPTs, and OpenAI states plainly that because a person initiates those fetches, robots.txt rules may not apply. OAI-AdsBot is the newest of the four and only visits pages submitted as ads.
Three operational details are easy to miss. If both OAI-SearchBot and GPTBot are allowed, OpenAI may reuse a single crawl for both purposes, so log volume will understate activity. A robots.txt change takes roughly 24 hours to reach the search systems. And OpenAI recommends allowing its published IP ranges as well as the user agent, which matters if a WAF or CDN is filtering ahead of your server - the request never reaches the rule you wrote. Structured markup that matches the visible text is the other half of machine access, covered in the schema ChatGPT reads.
How ChatGPT differs from the other engines
Each surface runs its own retrieval, so eligibility does not transfer between them. What is documented for one engine is not evidence about another.
ChatGPT eligibility is governed by OpenAI's own crawler tokens. Google AI Overviews is generated inside Google Search, where the documented requirement is ordinary indexing and snippet eligibility and nothing more - see what Google documents for that surface. Claude reaches the web through a different arrangement again, set out in how Claude picks sources, and Perplexity runs its own crawler and index, covered in the guide to earn Perplexity citations. The cross-engine version of the checklist is how to get cited by AI answer engines.
One question deserves an honest non-answer. OpenAI's crawler documentation does not state which search index backs ChatGPT's answers, and third-party accounts disagree - some describe a Bing-backed layer, others OpenAI's own index, others a mix. Treat any confident claim about the backbone as unverified. Keeping the site indexable in both Bing and Google covers every version of the answer at no extra cost.
Automating ChatGPT citations
Running this properly means keeping several loops alive at once: crawler access checked per agent and at the CDN layer, titles and slugs written against the sub-questions rather than the headline keyword, answer blocks kept quotable, structured data matched to visible text, and a prompt set re-run on a schedule so citation changes are observed rather than guessed from traffic. That workflow is what Citematic exists to automate ChatGPT citations with, reporting the result as Share of Model.
Key takeaways
- Roughly 18% of ChatGPT conversations trigger a web search at all, and citations concentrate in the opening turn - 12.6% of first turns carry one, against 3.0% by turn twenty.
- 88.46% of URLs arriving through the general search channel get cited, against 1.93% for Reddit and 0.51% for YouTube. Ranking is the entry ticket.
- Selection happens on title, URL and channel before your page is opened; cited titles matched ChatGPT's internal sub-questions at 0.656 cosine similarity against 0.484 for non-cited pages.
- A cited turn carries about four sources and 66% carry one to four, so the contest is for a slot beside competitors, not for sole ownership of the answer.
- Page structure decides whether you are quotable once selected. It cannot compensate for being in the wrong channel or carrying the wrong title.
Frequently asked questions
How do you get cited by ChatGPT?
Clear three gates in order. A web search has to fire at all, which Profound measured at about 18% of conversations. Your URL has to enter through the general search channel, which Ahrefs found carries an 88.46% citation rate against 1.93% for Reddit. And your title has to match the sub-question ChatGPT generated internally. Page structure matters after that, not before.
Which crawler does ChatGPT use?
OAI-SearchBot. OpenAI documents it as the crawler that surfaces websites in ChatGPT's search features and states that sites opted out will not be shown in ChatGPT search answers, though they can still appear as navigational links. Allow it in robots.txt and allow OpenAI's published IP ranges. Expect about 24 hours for a robots.txt change to take effect.
Does blocking GPTBot remove me from ChatGPT Search?
No. OpenAI states the settings are independent: GPTBot governs whether content may train the foundation models, OAI-SearchBot governs whether a page can appear in ChatGPT search answers. You can disallow GPTBot and stay eligible for citation. If both are allowed, OpenAI may reuse a single crawl for both purposes.
Why does ChatGPT rarely cite Reddit when it reads so much of it?
Ahrefs found Reddit accounts for 67.8% of all non-cited URLs while being cited at just 1.93%. Reddit enters through its own retrieval channel, apparently a dedicated feed on top of ordinary search results. ChatGPT appears to use it for context and consensus, then attribute the answer to a different kind of source.
Does fresh content get cited more by ChatGPT?
Both things are true at once. Across the whole population ChatGPT skews fresher than Google's organic results. But inside a single retrieval set, Ahrefs found the median cited page was around 500 days old and the non-cited pages skewed young. Relevance to the sub-question does the heavy lifting; freshness breaks ties, and mostly in news.
Does ChatGPT read JavaScript-rendered content?
Treat it as no. AI retrieval crawlers read the server-rendered HTML they receive and do not reliably execute JavaScript, so a client-side-rendered app can be retrieved and still be unreadable. Static generation or server-side rendering is the safe baseline, and it costs nothing to guarantee.
Sources
- Ahrefs, Why ChatGPT Cites One Page Over Another (Study of 1.4M Prompts). Retrieval channels, title similarity, URL slugs and page age. Published April 15, 2026, updated May 31, 2026. Checked August 23, 2026.
- Profound, How ChatGPT sources the web. About 730,000 conversations containing at least one citation, US English users, October to December 2025. Published February 3, 2026. Checked August 23, 2026.
- OpenAI, Overview of OpenAI Crawlers. User agents, robots.txt tags, published IP ranges and the 24-hour propagation note. Checked August 23, 2026.
Back to AI citation strategies | Answer-first content | GEO vs SEO