ChatGPT has no ranking. When it cannot answer from memory it runs a search, pulls a small set of pages from the index behind ChatGPT search, writes an answer and cites some of them. Being named depends on three things: OpenAI’s search crawler can read your page, your page answers the exact question in its opening lines, the third-party pages the model reads instead of yours already mention you.
The third one decides most cases. Almost nobody works on it, because it is the only one that is not on your own website.
ChatGPT is one surface. Answer visibility covers the rest of them, along with the vocabulary underneath the work and the five ways any of it gets measured.
Which crawler decides whether you appear?
OpenAI documents four. Two of them have nothing to do with search, which is where most of the confusion in this subject lives.
| Crawler | What it is for | What blocking it does |
|---|---|---|
OAI-SearchBot |
Surfacing sites in ChatGPT’s search features | Removes you from ChatGPT search answers |
GPTBot |
Crawling content to train foundation models | Nothing to your search visibility |
ChatGPT-User |
Fetching a page when a user or a custom GPT asks for it | Affects neither search nor training |
OAI-AdsBot |
Checking the safety of pages submitted as ads | Affects neither |
OpenAI’s own wording on the first one:
Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links.
And on the second:
Disallowing GPTBot indicates a site’s content should not be used in training generative AI foundation models.
Read those two sentences next to each other. Blocking GPTBot is a decision about training data. It costs nothing in ChatGPT search. The directive that costs you the answer is OAI-SearchBot, which is the one hardly anyone types deliberately.
Each crawler publishes its IP ranges, at openai.com/searchbot.json, openai.com/gptbot.json, openai.com/chatgpt-user.json and openai.com/adsbot.json. That is how you confirm a hit in your logs is real rather than something wearing the user agent.
Is your firewall answering for you?
Check the WAF before the robots file. A robots.txt is a request. A firewall rule is a wall. It never appears in any SEO audit. The person who set it usually does not work in marketing.
Cloudflare is the one to check first, because the setting there changed recently and the change was never going to surface inside anyone’s CMS. On 1 July 2026 it retired the single “block AI bots” switch and split AI traffic into three categories, defined on its own blog as Search (“any behavior that collects or indexes your content, so it can answer questions about it later”), Training (“a crawler taking your content to train or fine-tune a model”) and Agent (“automated behavior that is acting, usually in real time, on a person’s behalf”).
OAI-SearchBot sits in Search. From 15 September 2026 the defaults for new domains block Training and Agent “on the pages that display ads”, while Search stays allowed. So the current defaults are not the problem. The old switch is: anyone who flipped block-all before July 2026, or who wrote their own blanket rule against anything with “bot” and “AI” in the name, took themselves out of ChatGPT search and got no warning.
One more trap in the same paragraph of Cloudflare’s post. Blocking Training also blocks Googlebot, Applebot and Bingbot, because those crawlers serve several purposes at once. A decision about model training quietly becomes a decision about Google.
Does allowing the crawlers get you cited?
No. We measured it.
Method. The ten UK agencies holding the first page of Google for “generative engine optimisation agency” are the sample, chosen because every one of them sells AI visibility, so any of them getting this wrong is meaningful. On 10 August 2026 we fetched each robots.txt directly and pulled ChatGPT citation counts from Ahrefs’ AI citation index (site-explorer-ai-responses-count, mode=subdomains) the same day. Ten domains, one method, one date.
| Agency | robots.txt on 10 Aug 2026 | ChatGPT citations | Cited pages |
|---|---|---|---|
| Team Lewis | Default WordPress rules, no AI directives | 92 | 49 |
| MarGen | Explicit allow-list naming 22 AI crawlers | 30 | 4 |
| Impression | User-agent: * with an empty Disallow |
18 | 15 |
| Reboot Online | Allow: /, one Cloudflare path excluded |
17 | 9 |
| Catalyst | Default WordPress and WooCommerce rules | 14 | 4 |
| IDHL | User-agent: * with no rules under it |
10 | 15 |
| Push Group | Tracking parameters excluded, nothing else | 5 | 5 |
| Quirky Digital | Explicit allow-list naming 9 AI crawlers | 0 | 0 |
| Havas Market | Empty Disallow, plus Crawl-delay: 10 |
0 | 0 |
| Tilio | Four private paths excluded | 0 | 0 |
Ten out of ten allow every OpenAI crawler. Not one disallows OAI-SearchBot, GPTBot or ChatGPT-User. The ChatGPT citation counts across those same ten domains run from 0 to 92.
Access explains none of that spread. It is the entry ticket, not the race.
Two of the ten wrote explicit AI allow-lists, which is the tactic being sold hardest at the moment. MarGen names 22 crawlers under a heading reading “EXPLICITLY ALLOWED” and holds 30 ChatGPT citations. Quirky Digital names 9 and holds zero. Both allow-lists are redundant, because both sites already permit everything under User-agent: *. Writing Allow: / twice does not get you read twice.
Havas Market is the only site in the set carrying a directive that could plausibly slow a crawler, a Crawl-delay: 10. Google ignores crawl-delay. Whether OpenAI’s crawlers honour it is not documented, so nothing is claimed here either way.
Where do the citations concentrate?
Look at the last column rather than the middle one.
Team Lewis carries 92 ChatGPT citations spread across 49 pages, which is a large content estate collecting mentions across every subject it writes about. MarGen carries 30 across 4. Same platform, same index, same day, two entirely different mechanisms: one is breadth, one is a handful of pages that answer a specific question well.
The second is the only one available to a business that is not a global PR group. It also implies the useful working rule, which is that a page either answers a question a buyer asks or it does nothing at all.
One caveat, stated because the figure does not reconcile. IDHL returns 10 citations across 15 pages. It also returns 1 Gemini citation across 2 pages. Citations cannot be fewer than the pages carrying them, so the two columns are counted differently or over different windows. The counts are reported as returned. The per-page reading above is directional, not arithmetic.
What actually moves it
Five things, in the order they bite.
- Confirm
OAI-SearchBotcan reach you. Not the robots file. The server. Search your logs for the user agent, verify the hit against the published IP list, then check your WAF, your bot-management rules and any security plugin. If you cannot find a singleOAI-SearchBotrequest in ninety days, that is your answer and nothing further on this list matters. - Answer in the first two sentences. A model assembling a paragraph takes text that already reads like an answer. A page that spends four hundred words on context gives it nothing to lift.
- Write to the question, not the phrase. Google documents query fan-out, where the system generates a set of related queries at once and gathers results for all of them, which is why an AI answer can cite pages you never ranked for on the original wording. Whether ChatGPT does the same is not published. Our working rule either way: title the page with the sentence a buyer would say out loud, not with a noun.
- Get named on the pages that are already cited. Directories, comparison posts, trade titles, forum threads, the lists other people write about your market. This is the expensive one. It sits outside your CMS. It is the reason a content plan alone rarely changes a citation count.
- Measure it monthly against a named index with a date on it. Anything else is a story.
What none of this tells you
ChatGPT’s index is not published. Retrieval is per-query and probabilistic, so the same question asked twice returns different shortlists. Being named once is not being named. There is no impressions report, no position, no coverage view.
There is also no way to know, from the outside, whether the questions behind any citation count resemble the ones your buyers ask. A number can go up because a model got interested in a subject you do not sell.
So run the cheap version first. The afternoon method has you writing the buying questions yourself, running each three times and reading the citations underneath. Seventy-five rows of your own data beats a dashboard, because you chose the prompts. Then read what AI visibility means and how to measure it for the five measurement methods and what each one hides.
For the wider label covering all of this plus what Google says about it, see what generative engine optimisation actually is.
Data in this article: OpenAI crawler documentation and Cloudflare’s 1 July 2026 post, both read on 10 August 2026. robots.txt fetched directly from each domain on 10 August 2026. Citation counts from Ahrefs site-explorer-ai-responses-count, mode=subdomains, 10 August 2026. Citation counts for the same ten domains across all six platforms are in our UK agency table.