Skip to content
Flowpane
Level 3 / Method

robots.txt

robots.txt tells crawlers which paths they may fetch. One stray line can hide a whole site, so Flowpane checks it against RFC 9309 on every run.

Updated 29 September 2026·4 min read·Weight 15 of 100·Markdown

robots.txt is a plain-text file at the root of a host that tells crawlers which paths they may fetch. Search engines, AI crawlers and most well-behaved automated clients read it before anything else, so one wrong line can remove a whole site from view. It carries 15 of the 100 points in Flowpane's site score.

What it is for

RFC 9309 standardised the Robots Exclusion Protocol. The file lives at /robots.txt and applies only to the protocol, host and port it is served from, so https://www.example.com/robots.txt says nothing about https://example.com.

The file is a list of groups. Each group starts with one or more User-agent lines followed by Allow and Disallow rules. Precedence works like this:

  1. A crawler looks for groups naming its product token, matched case-insensitively. If several groups name it, their rules are combined.
  2. If no group names it, the crawler follows the User-agent: * group. If there is no such group, nothing is restricted.
  3. Within the selected rules, the longest matching path wins. When an Allow and a Disallow match equally, the RFC says Allow should win.

Sitemap: lines sit outside the group model and apply to the whole file. A crawler may cap how much of the file it parses, but the cap must be at least 500 KiB, so rules beyond that point may be ignored. The file is a request to cooperating crawlers, not access control, and it is public: every path it lists is advertised to anyone who reads it.

How it goes stale

Staging ships to production

Staging environments often serve User-agent: * with Disallow: / to stay out of search indexes. When robots.txt is a static file in the build rather than per-environment configuration, a routine deploy promotes the staging version. Nothing errors and every page still loads for people. Crawlers comply, and the loss shows up days later in search traffic.

AI crawler policy drifts

New AI crawler tokens appear regularly, and one operator often runs separate tokens for model training, search indexing and user-triggered retrieval. A policy written a year ago names some of them and silently leaves the rest to the wildcard group. Preference signals such as Content-Signal can then say one thing while the Disallow rules say another.

Declarations point at the past

Sitemap: lines keep pointing at files retired in a migration, or at an old protocol or host. Platforms that generate robots.txt can overwrite manual edits on the next update.

Empty is not the same as missing

A misconfigured host can return an empty 200, or an HTML error page, at /robots.txt. Both look like a healthy file to a monitor that only tests the status code.

What true looks like

User-agent: *
Content-Signal: search=yes, ai-train=no, ai-input=no
Disallow: /checkout/
Disallow: /search

User-agent: GPTBot
User-agent: Google-Extended
User-agent: CCBot
Disallow: /

Sitemap: https://www.example.com/sitemap.xml
  • A User-agent: * group with explicit rules, and named groups only where policy genuinely differs.
  • Every Sitemap: line is an absolute URL on the live protocol and host, pointing at a file that exists.
  • AI policy is explicit by crawler purpose, and Content-Signal expresses the same intent as the Disallow rules.
  • The production file is reviewed as release configuration, never copied from staging.

How Flowpane checks it

Flowpane parses the file against RFC 9309 and preserves every line, including comments and unknown directives. It resolves the effective group for each crawler, evaluates Allow and Disallow matching, validates Sitemap: directives as absolute URLs, checks file size against the 500 KiB parsing limit and records managed-platform evidence.

Blanket disallow. A wildcard Disallow: / is reported as a Warning, and Warnings are score-neutral, because only the publisher knows whether the restriction is intended. Its consequences are not neutral: the cross-file coherence engine reports sitemap URLs that robots.txt now blocks.

AI crawlers. Policy is assessed against a directory of named AI crawlers, including GPTBot, ChatGPT-User, OAI-SearchBot, Google-Extended, ClaudeBot, anthropic-ai, PerplexityBot, CCBot, Bytespider, Applebot-Extended, Meta-ExternalAgent, Amazonbot and cohere-ai. For each one, Flowpane works out the effective status: its own group first, then the wildcard group, then the default of allowed. The overall posture is summarised as open, selective, restrictive or no policy. Content-Signal directives (search, ai-train, ai-input) are read, and incomplete signals or signals that contradict the crawler rules are flagged.

Empty versus missing. A 404 is a missing result. Because robots.txt is always required, it counts as 0 in the site score, and Flowpane drafts a conservative starter that never invents AI policy, Crawl-delay or platform paths. An empty file is a present result that scores 0. An HTML page served at /robots.txt also scores 0.

Sitemap lines. Flowpane does not guess sitemap URLs. A draft adds a Sitemap: line only when same-site evidence confirms the file exists.

If a firewall challenge blocks the request, the result is Blocked, not Missing, and the site is shown as unscorable rather than given a partial score. Flowpane's own Deep Scan crawler obeys robots.txt under the product token Flowpane.

Common failures

SymptomWhat it meansEffect
Disallow: / under User-agent: *Staging file promoted, or a deliberate blockWarning, plus coherence findings
No Sitemap: lineCrawlers must find the sitemap unaidedRecommendation
AI crawlers not addressedThe wildcard group decides AI accessRecommendation
ai-train=yes while training crawlers are blockedSignals and rules disagreeIssue
Empty 200 or HTML pageHost misconfigurationScore 0