AI Crawlers and robots.txt: A Practical Guide for AI Search Visibility
AI crawler policy is not one universal switch. Different user agents can serve different purposes, and a site owner can make distinct choices about search surfacing and model training. That distinction is important when writing robots.txt rules.
OAI-SearchBot is the search control
OpenAI documents OAI-SearchBot as the crawler used to surface websites in ChatGPT search features. OpenAI also states that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, although they can still appear as navigational links. If ChatGPT search visibility is a goal, allowing this crawler is therefore the relevant robots.txt decision.
Example: allow search crawling
User-agent: OAI-SearchBot
Allow: /
User-agent: *
Allow: /
Allowing access is only permission to crawl. It does not guarantee that a page will be indexed, cited, recommended, or shown for a particular query.
GPTBot is a separate decision
OpenAI documents GPTBot separately from OAI-SearchBot and says the controls are independent. A publisher can allow OAI-SearchBot for search surfacing while disallowing GPTBot to signal that content should not be used for training OpenAI's generative AI foundation models.
Example: allow ChatGPT search while disallowing GPTBot
User-agent: OAI-SearchBot
Allow: /
User-agent: GPTBot
Disallow: /
User-agent: *
Allow: /
OpenAI also documents ChatGPT-User for certain user-triggered visits; because those requests are user initiated, robots.txt rules may not apply in the same way. Use OAI-SearchBot—not ChatGPT-User—as the control for ChatGPT search opt-outs.
Robots.txt is only one layer
A correct robots.txt file does not prove that a crawler can actually reach a page. Authentication, CDN or WAF bot rules, IP blocking, rate limits, JavaScript-only content, redirect loops, and server errors can still prevent access.
Practical verification sequence
- Confirm the final robots.txt served from the production domain.
- Check that important URLs return a normal public response without authentication.
- Review CDN or security rules for bot blocks that override the intent of robots.txt.
- Inspect server logs when available to confirm requests are not being rejected.
- Re-check after changes; OpenAI notes that robots.txt updates can take time to be reflected in its systems.
Do not use crawler rules as an optimization gimmick
Robots.txt controls access policy; it is not a ranking lever. Allowing OAI-SearchBot does not make weak content more citeable, and blocking GPTBot does not automatically remove a page from other search systems. Keep crawler policy separate from content, entity, authority, and technical-quality work.
The useful question is governance: which automated uses do you want to permit, and which do you want to decline? Express that policy clearly, then verify that infrastructure behavior matches the file.
Document changes
Keep a simple change log with the date, previous rule, new rule, reason for the change, and any related CDN or firewall adjustments. If visibility later changes, the log prevents teams from guessing whether crawler policy changed at the same time.
For OpenAI-specific rules, re-check the current official crawler documentation before making production changes because user-agent guidance and product behavior can evolve.
Related reading
Sources & verification
Product capabilities and pricing can change. These first-party pages were used to verify factual claims for this article.