# robots.txt

> robots.txt tells crawlers which paths they may fetch. One stray line can hide a whole site, so Flowpane checks it against RFC 9309 on every run.

robots.txt is a plain-text file at the root of a host that tells crawlers which paths they may fetch. Search engines, AI crawlers and most well-behaved automated clients read it before anything else, so one wrong line can remove a whole site from view. It carries 15 of the 100 points in Flowpane's site score.

## What it is for

[RFC 9309](https://www.rfc-editor.org/rfc/rfc9309) standardised the Robots Exclusion Protocol. The file lives at `/robots.txt` and applies only to the protocol, host and port it is served from, so `https://www.example.com/robots.txt` says nothing about `https://example.com`.

The file is a list of groups. Each group starts with one or more `User-agent` lines followed by `Allow` and `Disallow` rules. Precedence works like this:

1. A crawler looks for groups naming its product token, matched case-insensitively. If several groups name it, their rules are combined.
2. If no group names it, the crawler follows the `User-agent: *` group. If there is no such group, nothing is restricted.
3. Within the selected rules, the longest matching path wins. When an `Allow` and a `Disallow` match equally, the RFC says `Allow` should win.

`Sitemap:` lines sit outside the group model and apply to the whole file. A crawler may cap how much of the file it parses, but the cap must be at least 500 KiB, so rules beyond that point may be ignored. The file is a request to cooperating crawlers, not access control, and it is public: every path it lists is advertised to anyone who reads it.

## How it goes stale

### Staging ships to production

Staging environments often serve `User-agent: *` with `Disallow: /` to stay out of search indexes. When robots.txt is a static file in the build rather than per-environment configuration, a routine deploy promotes the staging version. Nothing errors and every page still loads for people. Crawlers comply, and the loss shows up days later in search traffic.

### AI crawler policy drifts

New AI crawler tokens appear regularly, and one operator often runs separate tokens for model training, search indexing and user-triggered retrieval. A policy written a year ago names some of them and silently leaves the rest to the wildcard group. Preference signals such as `Content-Signal` can then say one thing while the `Disallow` rules say another.

### Declarations point at the past

`Sitemap:` lines keep pointing at files retired in a migration, or at an old protocol or host. Platforms that generate robots.txt can overwrite manual edits on the next update.

### Empty is not the same as missing

A misconfigured host can return an empty 200, or an HTML error page, at `/robots.txt`. Both look like a healthy file to a monitor that only tests the status code.

## What true looks like

```text
User-agent: *
Content-Signal: search=yes, ai-train=no, ai-input=no
Disallow: /checkout/
Disallow: /search

User-agent: GPTBot
User-agent: Google-Extended
User-agent: CCBot
Disallow: /

Sitemap: https://www.example.com/sitemap.xml
```

- A `User-agent: *` group with explicit rules, and named groups only where policy genuinely differs.
- Every `Sitemap:` line is an absolute URL on the live protocol and host, pointing at a file that exists.
- AI policy is explicit by crawler purpose, and `Content-Signal` expresses the same intent as the `Disallow` rules.
- The production file is reviewed as release configuration, never copied from staging.

## How Flowpane checks it

Flowpane parses the file against RFC 9309 and preserves every line, including comments and unknown directives. It resolves the effective group for each crawler, evaluates `Allow` and `Disallow` matching, validates `Sitemap:` directives as absolute URLs, checks file size against the 500 KiB parsing limit and records managed-platform evidence.

**Blanket disallow.** A wildcard `Disallow: /` is reported as a Warning, and Warnings are score-neutral, because only the publisher knows whether the restriction is intended. Its consequences are not neutral: the [cross-file coherence](/learn/cross-file-coherence) engine reports sitemap URLs that robots.txt now blocks.

**AI crawlers.** Policy is assessed against a directory of named AI crawlers, including GPTBot, ChatGPT-User, OAI-SearchBot, Google-Extended, ClaudeBot, anthropic-ai, PerplexityBot, CCBot, Bytespider, Applebot-Extended, Meta-ExternalAgent, Amazonbot and cohere-ai. For each one, Flowpane works out the effective status: its own group first, then the wildcard group, then the default of allowed. The overall posture is summarised as open, selective, restrictive or no policy. `Content-Signal` directives (`search`, `ai-train`, `ai-input`) are read, and incomplete signals or signals that contradict the crawler rules are flagged.

**Empty versus missing.** A 404 is a missing result. Because robots.txt is always required, it counts as 0 in the site score, and Flowpane drafts a conservative starter that never invents AI policy, Crawl-delay or platform paths. An empty file is a present result that scores 0. An HTML page served at `/robots.txt` also scores 0.

**Sitemap lines.** Flowpane does not guess sitemap URLs. A draft adds a `Sitemap:` line only when same-site evidence confirms the file exists.

If a firewall challenge blocks the request, the result is Blocked, not Missing, and the site is shown as unscorable rather than given a partial score. Flowpane's own Deep Scan crawler obeys robots.txt under the product token `Flowpane`.

## Common failures

| Symptom | What it means | Effect |
| --- | --- | --- |
| `Disallow: /` under `User-agent: *` | Staging file promoted, or a deliberate block | Warning, plus coherence findings |
| No `Sitemap:` line | Crawlers must find the sitemap unaided | Recommendation |
| AI crawlers not addressed | The wildcard group decides AI access | Recommendation |
| `ai-train=yes` while training crawlers are blocked | Signals and rules disagree | Issue |
| Empty 200 or HTML page | Host misconfiguration | Score 0 |

## Related

- [Sitemaps](/learn/sitemaps)
- [ai.txt](/learn/ai-txt)
- [Cross-file coherence](/learn/cross-file-coherence)
- [AI crawlers](/use-cases/ai-crawlers)
- [How Flowpane fetches](/learn/how-flowpane-fetches)
