robots.txt is a plain-text file at the root of a host that tells crawlers which paths they may fetch. Search engines, AI crawlers and most well-behaved automated clients read it before anything else, so one wrong line can remove a whole site from view. It carries 15 of the 100 points in Flowpane's site score.
What it is for
RFC 9309 standardised the Robots Exclusion Protocol. The file lives at /robots.txt and applies only to the protocol, host and port it is served from, so https://www.example.com/robots.txt says nothing about https://example.com.
The file is a list of groups. Each group starts with one or more User-agent lines followed by Allow and Disallow rules. Precedence works like this:
- A crawler looks for groups naming its product token, matched case-insensitively. If several groups name it, their rules are combined.
- If no group names it, the crawler follows the
User-agent: *group. If there is no such group, nothing is restricted. - Within the selected rules, the longest matching path wins. When an
Allowand aDisallowmatch equally, the RFC saysAllowshould win.
Sitemap: lines sit outside the group model and apply to the whole file. A crawler may cap how much of the file it parses, but the cap must be at least 500 KiB, so rules beyond that point may be ignored. The file is a request to cooperating crawlers, not access control, and it is public: every path it lists is advertised to anyone who reads it.
How it goes stale
Staging ships to production
Staging environments often serve User-agent: * with Disallow: / to stay out of search indexes. When robots.txt is a static file in the build rather than per-environment configuration, a routine deploy promotes the staging version. Nothing errors and every page still loads for people. Crawlers comply, and the loss shows up days later in search traffic.
AI crawler policy drifts
New AI crawler tokens appear regularly, and one operator often runs separate tokens for model training, search indexing and user-triggered retrieval. A policy written a year ago names some of them and silently leaves the rest to the wildcard group. Preference signals such as Content-Signal can then say one thing while the Disallow rules say another.
Declarations point at the past
Sitemap: lines keep pointing at files retired in a migration, or at an old protocol or host. Platforms that generate robots.txt can overwrite manual edits on the next update.
Empty is not the same as missing
A misconfigured host can return an empty 200, or an HTML error page, at /robots.txt. Both look like a healthy file to a monitor that only tests the status code.
What true looks like
User-agent: *
Content-Signal: search=yes, ai-train=no, ai-input=no
Disallow: /checkout/
Disallow: /search
User-agent: GPTBot
User-agent: Google-Extended
User-agent: CCBot
Disallow: /
Sitemap: https://www.example.com/sitemap.xml- A
User-agent: *group with explicit rules, and named groups only where policy genuinely differs. - Every
Sitemap:line is an absolute URL on the live protocol and host, pointing at a file that exists. - AI policy is explicit by crawler purpose, and
Content-Signalexpresses the same intent as theDisallowrules. - The production file is reviewed as release configuration, never copied from staging.
How Flowpane checks it
Flowpane parses the file against RFC 9309 and preserves every line, including comments and unknown directives. It resolves the effective group for each crawler, evaluates Allow and Disallow matching, validates Sitemap: directives as absolute URLs, checks file size against the 500 KiB parsing limit and records managed-platform evidence.
Blanket disallow. A wildcard Disallow: / is reported as a Warning, and Warnings are score-neutral, because only the publisher knows whether the restriction is intended. Its consequences are not neutral: the cross-file coherence engine reports sitemap URLs that robots.txt now blocks.
AI crawlers. Policy is assessed against a directory of named AI crawlers, including GPTBot, ChatGPT-User, OAI-SearchBot, Google-Extended, ClaudeBot, anthropic-ai, PerplexityBot, CCBot, Bytespider, Applebot-Extended, Meta-ExternalAgent, Amazonbot and cohere-ai. For each one, Flowpane works out the effective status: its own group first, then the wildcard group, then the default of allowed. The overall posture is summarised as open, selective, restrictive or no policy. Content-Signal directives (search, ai-train, ai-input) are read, and incomplete signals or signals that contradict the crawler rules are flagged.
Empty versus missing. A 404 is a missing result. Because robots.txt is always required, it counts as 0 in the site score, and Flowpane drafts a conservative starter that never invents AI policy, Crawl-delay or platform paths. An empty file is a present result that scores 0. An HTML page served at /robots.txt also scores 0.
Sitemap lines. Flowpane does not guess sitemap URLs. A draft adds a Sitemap: line only when same-site evidence confirms the file exists.
If a firewall challenge blocks the request, the result is Blocked, not Missing, and the site is shown as unscorable rather than given a partial score. Flowpane's own Deep Scan crawler obeys robots.txt under the product token Flowpane.
Common failures
| Symptom | What it means | Effect |
|---|---|---|
Disallow: / under User-agent: * | Staging file promoted, or a deliberate block | Warning, plus coherence findings |
No Sitemap: line | Crawlers must find the sitemap unaided | Recommendation |
| AI crawlers not addressed | The wildcard group decides AI access | Recommendation |
ai-train=yes while training crawlers are blocked | Signals and rules disagree | Issue |
| Empty 200 or HTML page | Host misconfiguration | Score 0 |