Robots.txt: Syntax, Examples, and AI Crawler Rules

Robots.txt is a plain text file that tells cooperating web crawlers which URLs they may crawl on a website. Its main purpose is to manage crawler requests. A blocked web page can still appear in Google Search, so crawl restrictions and indexing controls serve different purposes. Google’s robots.txt introduction

robots.txt policy: a tilted paper plane, small globe, and closed folder arranged left to right, open laptop, pen, stack of books, and closed calendar

What robots.txt controls

Robots.txt is advisory: it does not deny access at the server, and the paths it lists are public (RFC 9309).

Choose the control that matches the required outcome:

Required outcomeRelevant control
Prevent a cooperating crawler from fetching selected URLsDisallow rules in robots.txt. Google’s crawl guidance
Keep a public page out of Google SearchA noindex meta tag or X-Robots-Tag response header that Googlebot can fetch. Google’s indexing guidance
Prevent unauthorized access to private contentAuthentication or other access restrictions enforced by the application or server. RFC 9309’s security guidance

Where to find or put the file

The filename must be lowercase robots.txt, and the file belongs at the root of the host. Save it as UTF-8 plain text. Its rules apply to the same protocol, host, and port, including paths beneath that root. Subdomains and HTTP versions require their own corresponding files (Google’s file creation guide).

These locations illustrate that scope:

Website originRobots.txt location
https://example.comhttps://example.com/robots.txt
https://www.example.comhttps://www.example.com/robots.txt
https://shop.example.comhttps://shop.example.com/robots.txt

A file at /articles/robots.txt does not control the articles directory: crawlers look at the host’s root location. Hosting services may expose crawler settings instead of allowing direct file editing (Google’s location and hosting requirements).

A website does not need a robots.txt file merely to allow Google crawling. Google treats a missing file returning 404 as having no crawl restrictions (Search Console’s robots.txt documentation).

Basic syntax

A group starts with one or more User-agent lines, followed by rules for those agents. Write each directive on its own line (Google’s rule-writing guide).

DirectiveMeaning
User-agent: *Fallback group for crawlers without a matching named group. RFC 9309
Disallow: /search/Requests that the applicable crawler avoid URLs starting with /search/. Google’s rule definitions
Allow: /search/help/Allows this more specific path despite a broader /search/ restriction. RFC 9309
Sitemap: https://example.com/sitemap.xmlDeclares a sitemap using its absolute URL. Google’s sitemap guidance

Paths are case-sensitive: /Search/ and /search/ differ. A # begins a comment (Google’s syntax reference).

Allow crawling of the whole site

An empty Disallow: blocks nothing. The following file allows crawling for the fallback group and declares a sitemap (Google’s rule definitions).

User-agent: *
Disallow:

Sitemap: https://example.com/sitemap.xml

Replace the sitemap address with the actual sitemap URL, or omit that line. Google accepts multiple Sitemap: lines, but submitting a sitemap remains a discovery hint and does not guarantee crawling (Google’s sitemap submission guidance).

Block a directory and allow an exception

Within the applicable group, the longest matching rule wins. The protocol favors Allow when equally specific allow and disallow rules conflict (RFC 9309’s path-matching rules).

User-agent: *
Disallow: /search/
Allow: /search/help/

/search/results is disallowed; /search/help/index.html and /articles/robots-guide are allowed under the most-specific-match rule and default allowance for unmatched paths (RFC 9309).

Block all crawling for a group

Disallow: / covers every path for the applicable crawler group. It still cannot enforce access restrictions or reliably remove web pages from Google’s results (Google’s robots.txt limitations).

User-agent: *
Disallow: /

Match exact paths and patterns

Google supports * for zero or more characters and $ for the end of a match. Otherwise, rules match prefixes (Google’s URL-matching reference).

RuleMatch
Disallow: /draftMatches /draft, /drafts/, and /drafting. Google’s prefix rules
Disallow: /draft$Matches /draft, but not /drafts/ or /draft?version=2. Google’s end-marker rules
Disallow: /*.pdf$Matches URLs ending in .pdf; a URL ending in .pdf?download=1 does not match. Google’s wildcard rules

Named groups replace the wildcard fallback

Google uses the most specific agent group, merging repeated matching groups. It does not merge named groups with User-agent: *; group order is irrelevant (Google’s group precedence rules).

Consequently, a separate Googlebot group must include every exclusion Googlebot should follow. This file repeats the /search/ block in both groups:

User-agent: *
Disallow: /search/

User-agent: Googlebot
Disallow: /search/
Disallow: /archive/

The named Googlebot group disallows both paths. Other crawlers using the fallback group are restricted only from /search/ (Google’s group precedence rules).

Robots.txt and noindex

Google can discover a blocked URL from links elsewhere and display its address in results without crawling the page. Blocking JavaScript or CSS needed to understand a page can also impair Google’s analysis of that page (Google’s robots.txt limitations).

For an HTML page that should remain publicly accessible but be excluded from Google Search, put this tag inside the page’s HTML head: Google’s noindex instructions

<meta name="robots" content="noindex">

For a PDF or another non-HTML resource, the server can send the equivalent response header: Google’s response-header guidance

X-Robots-Tag: noindex

Googlebot must be able to fetch the resource to see either instruction. Combining noindex with a robots.txt block can prevent Google from reading the indexing rule. Writing Noindex: inside robots.txt is unsupported by Google (Google’s noindex requirements).

Search and AI crawler controls

AI services document separate controls for different uses. OpenAI, for example, allows its search and potential training settings to be configured independently (OpenAI’s crawler documentation).

OperatorSearch-related tokenOther documented control
OpenAIOAI-SearchBot surfaces websites in ChatGPT search.GPTBot crawls content that may be used for foundation-model training. OpenAI
AnthropicClaude-SearchBot supports search results.ClaudeBot collects potential training material; Claude-User retrieves pages for user requests. Anthropic says its bots honor robots.txt. Anthropic
GoogleGooglebot crawling affects Google Search and its features.Google-Extended controls specified Gemini training and grounding uses without affecting Search inclusion or ranking. Google
PerplexityPerplexityBot surfaces and links websites in Perplexity search.Perplexity-User performs user-requested fetches, which generally ignore robots.txt. Neither agent is described as collecting foundation-model training data. Perplexity

The following OpenAI groups allow search crawling and disallow GPTBot crawling. Add these groups to the file only when those are the intended settings (OpenAI’s independent bot controls).

User-agent: OAI-SearchBot
Allow: /

User-agent: GPTBot
Disallow: /

The search group allows every path. Repeat required exclusions there; it does not inherit wildcard rules (RFC 9309’s user-agent matching rules).

OpenAI says opting out of OAI-SearchBot excludes a site from ChatGPT search answers, although navigational links can still appear. It also says robots.txt may not apply to user-initiated ChatGPT-User actions (OpenAI’s crawler documentation).

Google-Extended has no separate HTTP request user-agent string; it is a product-control token. Google’s crawler reference Google’s AI Overviews and AI Mode use Googlebot crawl controls, while nosnippet, data-nosnippet, max-snippet, and noindex govern information shown from pages in Search (Google’s AI features guidance).

For Google’s AI features, allowing crawling does not guarantee placement. Google states that eligible pages are not guaranteed to be crawled, indexed, or served (Google’s AI eligibility guidance).

How to check and update robots.txt

Start from the file actually served by the website. Google’s update procedure is to download the current file, edit it as UTF-8 text, and upload the revision to the root. Hosting platforms may provide their own settings instead (Google’s update instructions).

Open the exact robots.txt URL in a browser, or retrieve it with curl, substituting the actual hostname: Google’s download instructions

curl https://example.com/robots.txt

Then check the served rules against the intended URLs and agents:

  1. Inspect Google’s fetched copy. Search Console’s robots.txt report shows fetch status, the last checked time, and parsing errors or warnings. Its displayed copy may differ from the live file after an edit (Search Console’s report guide).
  2. Test individual pages. Use Search Console’s URL Inspection tool to check whether a specific URL is blocked. Include a public page, a blocked path, and any allowed exception among the checks (Google’s URL-testing guidance).
  3. Verify crawler identities in logs. A user-agent string alone does not establish that a request came from Google. Google documents verification through reverse and forward DNS or comparison with its published IP ranges (Google’s request-verification guide).
  4. Check the updated version. Google normally refreshes its robots.txt cache every 24 hours. For a critical correction, its robots.txt report provides a recrawl request (Google’s cache-update guidance).

Missing files and server errors

RFC 9309 permits crawling after a client-error response. Server or network failures instead require an initial assumption of complete disallow. Returning an error therefore does not provide a dependable crawl block (RFC 9309’s access results).

Google documents its own error-handling sequence: most 4xx responses, including 403 and 404, mean no robots.txt restrictions; 429 is an exception. Server errors can stop crawling initially and cause Google to use a previously cached file afterward (Google’s HTTP-status guidance).

Restore a readable file with the intended rules, inspect the version Google has fetched, and test the affected page through URL Inspection. Search Console supports these checks and a recrawl request after a critical fix (Search Console’s troubleshooting guidance).

One person. A whole marketing team.

Invite only