# How Flowpane fetches

> Flowpane reads public HTTPS only, never signs in, and signs requests with Web Bot Auth so firewalls can verify them. Blocked is never reported as Missing.

Flowpane reads your governance files the way any member of the public can: over HTTPS, from your canonical origin, without signing in. Its requests identify themselves with a stable user agent and carry a cryptographic signature, so a firewall can verify they really came from Flowpane. When a firewall stands between Flowpane and a file, the result is Blocked, never Missing.

## Public HTTPS only

Governance evidence is only useful if it matches what the public receives. A logged-in view, a staging server or an editor's preview can all differ from the live file, so Flowpane resolves the site's canonical origin, for example `https://www.example.com` after redirects, and fetches from there.

When it checks your site, Flowpane never signs in, submits forms or sends credentials, and checking never writes anything to your origin. If a file is visible only after login, it is not public, and Flowpane reports what the public receives.

## Identity

Governance checks and Deep Scan crawls send this user agent:

```text
Flowpane/1.0 (+https://crawler.flowpane.com)
```

The robots.txt product token is `Flowpane`. A Mozilla-compatible form of the same identity is also defined for firewalls that reject non-browser strings: `Mozilla/5.0 (compatible; Flowpane/1.0; +https://crawler.flowpane.com)`. Rules that match on the `Flowpane` token cover both.

One deliberate exception applies. On a manually started check, after the normal signed fetch of a sitemap has succeeded, the sitemap reliability test requests it three more times: once as Flowpane, and once each, unsigned, with a generic browser user agent and a search-crawler user agent. It exists to detect a firewall or cache that serves different clients different content.

## Web Bot Auth

A user agent string is a claim that anyone can copy. Flowpane backs it with proof: requests are signed with HTTP Message Signatures ([RFC 9421](https://www.rfc-editor.org/rfc/rfc9421)) using Ed25519 keys, following the Web Bot Auth approach. Each signed request carries three headers:

- `Signature-Agent` points to crawler.flowpane.com, where the public keys are published.
- `Signature-Input` lists the signed components, the creation and expiry times, the key ID, a nonce and the `web-bot-auth` tag.
- `Signature` holds the Ed25519 signature itself.

The key directory is at `https://crawler.flowpane.com/.well-known/http-message-signatures-directory`. A firewall or CDN that supports Web Bot Auth fetches the directory, selects the key by its ID and verifies the signature. Each signature is created per request and is valid for five minutes, which limits the value of a captured header.

```text
# Illustrative only. Every value below is a placeholder, not a real key or signature.
GET /robots.txt HTTP/1.1
Host: www.example.com
User-Agent: Flowpane/1.0 (+https://crawler.flowpane.com)
Signature-Agent: https://crawler.flowpane.com/.well-known/http-message-signatures-directory
Signature-Input: sig1=("@authority");created=1790000000;keyid="<key-id>";alg="ed25519";expires=1790000300;nonce="<nonce>";tag="web-bot-auth"
Signature: sig1=:<base64-signature>:
```

Where your firewall supports it, verify the signature rather than allow-listing the user agent string. If you must allow-list by string, keep the exception narrow: the governance file paths and the verification path, not the whole site.

## Allowing Flowpane in robots.txt

```text
User-agent: Flowpane
Allow: /
Crawl-delay: 5

User-agent: *
Disallow: /admin/
```

Under [RFC 9309](https://www.rfc-editor.org/rfc/rfc9309), a crawler follows only the most specific group that matches it. Once a `Flowpane` group exists, Flowpane ignores the `*` group entirely, so copy across any Disallow lines you still want applied.

## Blocked is not Missing

Every fetch ends in one of four states, and the difference between them is what keeps the evidence honest.

| State | What happened | Effect |
| --- | --- | --- |
| Present | The file was served and assessed | Scored |
| Missing | The server said the file does not exist, such as 404 | Scores 0 |
| Blocked | A challenge, block page or refusal stood in the way | Not scored |
| Error | Timeout, server error or transport failure | Last good result stands |

Blocked means Flowpane did not see the file, so it cannot say the file is absent. Reporting Missing would send a security team to publish a security.txt that already exists while the real cause, their own firewall, goes unexamined. If a required artefact is blocked, the site is shown as unscorable rather than given a partial score. The presence of a CDN or WAF is never treated as a challenge on its own: proxy headers on an ordinary response are attribution, not evidence. Blocked needs challenge, block or rate-limit evidence.

## Ownership verification

Monitoring starts only after you prove control of the site, by one of three methods:

- A file at `/.well-known/flowpane-verify-<id>.txt`.
- A DNS TXT record at `_flowpane.<host>`. No HTTP request is involved, so this works even when a firewall blocks Flowpane.
- A `flowpane-site-verification` meta tag in the homepage `<head>`.

## Deep Scan crawls

On Starter and above, Deep Scan crawls the site breadth-first from the homepage, following internal links within the plan's depth and page limits.

- It reads robots.txt first and obeys the group that matches the `Flowpane` token, or `*` when there is none.
- It honours `Crawl-delay` for that group, capped at 10 seconds, and follows redirects one hop at a time, checking robots.txt for each.
- A 404 for robots.txt means no restrictions, as RFC 9309 specifies. A robots.txt that is blocked, times out or errors stops the crawl on that origin: Flowpane does not guess the rules.
- It is never scheduled automatically. At most one completed crawl per site per hour.

## Search Console

With read-only access (`webmasters.readonly`), Flowpane imports your submitted sitemaps and their counts and runs bounded URL Inspection. Flowpane reads from Search Console. It never submits sitemaps.

That gives three evidence layers: the public pull (what is on the wire), the crawl (what is reachable) and Search Console (what Google accepted). The value is in the difference between them.

## Related

- [robots.txt](/learn/robots-txt)
- [How the score works](/learn/scoring)
- [Sitemaps](/learn/sitemaps)
- [Flowpane crawler](https://crawler.flowpane.com)
- [Migrations](/use-cases/migrations)
