Robots.txt: Syntax, Examples, and AI Crawler Rules
Robots.txt is a plain text file that tells cooperating web crawlers which URLs they may crawl on a website. Its main purpose is to manage crawler requests. A blocked web page can still appear in Google Search, so crawl restrictions and indexing controls serve different purposes. Google’s robots.txt introduction

What robots.txt controls
Robots.txt is advisory: it does not deny access at the server, and the paths it lists are public (RFC 9309).
Choose the control that matches the required outcome:
| Required outcome | Relevant control |
|---|---|
| Prevent a cooperating crawler from fetching selected URLs | Disallow rules in robots.txt. Google’s crawl guidance |
| Keep a public page out of Google Search | A noindex meta tag or X-Robots-Tag response header that Googlebot can fetch. Google’s indexing guidance |
| Prevent unauthorized access to private content | Authentication or other access restrictions enforced by the application or server. RFC 9309’s security guidance |
Where to find or put the file
The filename must be lowercase robots.txt, and the file belongs at the root of the host. Save it as UTF-8 plain text. Its rules apply to the same protocol, host, and port, including paths beneath that root. Subdomains and HTTP versions require their own corresponding files (Google’s file creation guide).
These locations illustrate that scope:
| Website origin | Robots.txt location |
|---|---|
https://example.com | https://example.com/robots.txt |
https://www.example.com | https://www.example.com/robots.txt |
https://shop.example.com | https://shop.example.com/robots.txt |
A file at /articles/robots.txt does not control the articles directory: crawlers look at the host’s root location. Hosting services may expose crawler settings instead of allowing direct file editing (Google’s location and hosting requirements).
A website does not need a robots.txt file merely to allow Google crawling. Google treats a missing file returning 404 as having no crawl restrictions (Search Console’s robots.txt documentation).
Basic syntax
A group starts with one or more User-agent lines, followed by rules for those agents. Write each directive on its own line (Google’s rule-writing guide).
| Directive | Meaning |
|---|---|
User-agent: * | Fallback group for crawlers without a matching named group. RFC 9309 |
Disallow: /search/ | Requests that the applicable crawler avoid URLs starting with /search/. Google’s rule definitions |
Allow: /search/help/ | Allows this more specific path despite a broader /search/ restriction. RFC 9309 |
Sitemap: https://example.com/sitemap.xml | Declares a sitemap using its absolute URL. Google’s sitemap guidance |
Paths are case-sensitive: /Search/ and /search/ differ. A # begins a comment (Google’s syntax reference).
Allow crawling of the whole site
An empty Disallow: blocks nothing. The following file allows crawling for the fallback group and declares a sitemap (Google’s rule definitions).
User-agent: *
Disallow:
Sitemap: https://example.com/sitemap.xml
Replace the sitemap address with the actual sitemap URL, or omit that line. Google accepts multiple Sitemap: lines, but submitting a sitemap remains a discovery hint and does not guarantee crawling (Google’s sitemap submission guidance).
Block a directory and allow an exception
Within the applicable group, the longest matching rule wins. The protocol favors Allow when equally specific allow and disallow rules conflict (RFC 9309’s path-matching rules).
User-agent: *
Disallow: /search/
Allow: /search/help/
/search/results is disallowed; /search/help/index.html and /articles/robots-guide are allowed under the most-specific-match rule and default allowance for unmatched paths (RFC 9309).
Block all crawling for a group
Disallow: / covers every path for the applicable crawler group. It still cannot enforce access restrictions or reliably remove web pages from Google’s results (Google’s robots.txt limitations).
User-agent: *
Disallow: /
Match exact paths and patterns
Google supports * for zero or more characters and $ for the end of a match. Otherwise, rules match prefixes (Google’s URL-matching reference).
| Rule | Match |
|---|---|
Disallow: /draft | Matches /draft, /drafts/, and /drafting. Google’s prefix rules |
Disallow: /draft$ | Matches /draft, but not /drafts/ or /draft?version=2. Google’s end-marker rules |
Disallow: /*.pdf$ | Matches URLs ending in .pdf; a URL ending in .pdf?download=1 does not match. Google’s wildcard rules |
Named groups replace the wildcard fallback
Google uses the most specific agent group, merging repeated matching groups. It does not merge named groups with User-agent: *; group order is irrelevant (Google’s group precedence rules).
Consequently, a separate Googlebot group must include every exclusion Googlebot should follow. This file repeats the /search/ block in both groups:
User-agent: *
Disallow: /search/
User-agent: Googlebot
Disallow: /search/
Disallow: /archive/
The named Googlebot group disallows both paths. Other crawlers using the fallback group are restricted only from /search/ (Google’s group precedence rules).
Robots.txt and noindex
Google can discover a blocked URL from links elsewhere and display its address in results without crawling the page. Blocking JavaScript or CSS needed to understand a page can also impair Google’s analysis of that page (Google’s robots.txt limitations).
For an HTML page that should remain publicly accessible but be excluded from Google Search, put this tag inside the page’s HTML head: Google’s noindex instructions
<meta name="robots" content="noindex">
For a PDF or another non-HTML resource, the server can send the equivalent response header: Google’s response-header guidance
X-Robots-Tag: noindex
Googlebot must be able to fetch the resource to see either instruction. Combining noindex with a robots.txt block can prevent Google from reading the indexing rule. Writing Noindex: inside robots.txt is unsupported by Google (Google’s noindex requirements).
Search and AI crawler controls
AI services document separate controls for different uses. OpenAI, for example, allows its search and potential training settings to be configured independently (OpenAI’s crawler documentation).
| Operator | Search-related token | Other documented control |
|---|---|---|
| OpenAI | OAI-SearchBot surfaces websites in ChatGPT search. | GPTBot crawls content that may be used for foundation-model training. OpenAI |
| Anthropic | Claude-SearchBot supports search results. | ClaudeBot collects potential training material; Claude-User retrieves pages for user requests. Anthropic says its bots honor robots.txt. Anthropic |
Googlebot crawling affects Google Search and its features. | Google-Extended controls specified Gemini training and grounding uses without affecting Search inclusion or ranking. Google | |
| Perplexity | PerplexityBot surfaces and links websites in Perplexity search. | Perplexity-User performs user-requested fetches, which generally ignore robots.txt. Neither agent is described as collecting foundation-model training data. Perplexity |
The following OpenAI groups allow search crawling and disallow GPTBot crawling. Add these groups to the file only when those are the intended settings (OpenAI’s independent bot controls).
User-agent: OAI-SearchBot
Allow: /
User-agent: GPTBot
Disallow: /
The search group allows every path. Repeat required exclusions there; it does not inherit wildcard rules (RFC 9309’s user-agent matching rules).
OpenAI says opting out of OAI-SearchBot excludes a site from ChatGPT search answers, although navigational links can still appear. It also says robots.txt may not apply to user-initiated ChatGPT-User actions (OpenAI’s crawler documentation).
Google-Extended has no separate HTTP request user-agent string; it is a product-control token. Google’s crawler reference Google’s AI Overviews and AI Mode use Googlebot crawl controls, while nosnippet, data-nosnippet, max-snippet, and noindex govern information shown from pages in Search (Google’s AI features guidance).
For Google’s AI features, allowing crawling does not guarantee placement. Google states that eligible pages are not guaranteed to be crawled, indexed, or served (Google’s AI eligibility guidance).
How to check and update robots.txt
Start from the file actually served by the website. Google’s update procedure is to download the current file, edit it as UTF-8 text, and upload the revision to the root. Hosting platforms may provide their own settings instead (Google’s update instructions).
Open the exact robots.txt URL in a browser, or retrieve it with curl, substituting the actual hostname: Google’s download instructions
curl https://example.com/robots.txt
Then check the served rules against the intended URLs and agents:
- Inspect Google’s fetched copy. Search Console’s robots.txt report shows fetch status, the last checked time, and parsing errors or warnings. Its displayed copy may differ from the live file after an edit (Search Console’s report guide).
- Test individual pages. Use Search Console’s URL Inspection tool to check whether a specific URL is blocked. Include a public page, a blocked path, and any allowed exception among the checks (Google’s URL-testing guidance).
- Verify crawler identities in logs. A user-agent string alone does not establish that a request came from Google. Google documents verification through reverse and forward DNS or comparison with its published IP ranges (Google’s request-verification guide).
- Check the updated version. Google normally refreshes its robots.txt cache every 24 hours. For a critical correction, its robots.txt report provides a recrawl request (Google’s cache-update guidance).
Missing files and server errors
RFC 9309 permits crawling after a client-error response. Server or network failures instead require an initial assumption of complete disallow. Returning an error therefore does not provide a dependable crawl block (RFC 9309’s access results).
Google documents its own error-handling sequence: most 4xx responses, including 403 and 404, mean no robots.txt restrictions; 429 is an exception. Server errors can stop crawling initially and cause Google to use a previously cached file afterward (Google’s HTTP-status guidance).
Restore a readable file with the intended rules, inspect the version Google has fetched, and test the affected page through URL Inspection. Search Console supports these checks and a recrawl request after a critical fix (Search Console’s troubleshooting guidance).