Skip to content
Flowpane
Level 3 / Method

Sitemaps

A sitemap is a public claim about your URL set. Flowpane tests that claim against the live file, a crawl of the site and what Search Console accepted.

Updated 29 September 2026·5 min read·Weight 35 of 100·Markdown

An XML sitemap is a public claim about a site's URL set: which pages exist, where they live and when they last changed. Search engines and other crawlers use it to decide what to fetch and how often. It carries 35 of the 100 points in Flowpane's site score, more than any other artefact, because a wrong sitemap misdirects every crawler that trusts it.

What it is for

The sitemap protocol defines two shapes. A urlset lists pages, each with a <loc> and an optional <lastmod>. A sitemapindex lists other sitemap files. Each file is limited to 50,000 entries and 50 MB uncompressed, so larger sites publish an index at the root that points to child urlsets split by section or date.

Crawlers treat the sitemap as a hint, not a command. It shapes discovery and recrawl priority: which URLs are fetched soon, and which are left alone. A crawler that keeps finding redirects, errors and meaningless timestamps in a sitemap learns to discount it.

How it goes stale

Sitemaps are usually generated, so teams assume they are correct. A generator is only as current as its configuration, and its output is rarely reviewed after launch.

The root that stopped being the root

A common pattern: /sitemap.xml is an index that still lists campaign children from several years ago. A platform migration moved the live inventory into new sitemap files, and those are declared only in robots.txt Sitemap: lines. Nobody removed the old index because nothing visibly broke. Crawlers and tools that start from the conventional path read an outdated claim. Those that read robots.txt read a different one. Both are "the sitemap".

lastmod that cannot be trusted

Build pipelines often stamp every URL with the deploy time, so every page appears to change on every release. Other failures include dates in the future, malformed dates, and timestamps that have not moved in years on a site that publishes weekly. Once lastmod moves without the content moving, it stops carrying information.

URLs that no longer resolve cleanly

Pages are deleted, redirected, marked noindex or canonicalised elsewhere while the sitemap keeps listing them. The reverse also happens: a new section launches and never reaches any sitemap.

What true looks like

<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <sitemap>
    <loc>https://www.example.com/sitemaps/pages.xml</loc>
    <lastmod>2026-09-24T09:12:00+00:00</lastmod>
  </sitemap>
  <sitemap>
    <loc>https://www.example.com/sitemaps/articles.xml</loc>
    <lastmod>2026-09-28T16:40:00+00:00</lastmod>
  </sitemap>
</sitemapindex>
  • One root, declared in robots.txt, and the same root the conventional path serves.
  • Every listed URL is canonical, indexable and served on the site's own protocol and host.
  • lastmod changes when the page's primary content changes, and at no other time.
  • Retired children are removed from the index, not left to return errors.

How Flowpane checks it

Discovery reads robots.txt Sitemap: lines, /sitemap.xml, /sitemap_index.xml and known platform paths, then expands each index up to 50 children. Because every source is read, a stale root and the live inventory appear side by side rather than one hiding the other.

Each sitemap is then compared across three evidence layers. The value is in the differences between them.

LayerQuestionSource
DeclaredWhat does the public file claim?Public pull of every discovered sitemap
ReachableWhat can a crawler actually reach?Deep Scan crawl, Starter and above
AcceptedWhat has Google accepted?Google Search Console, read-only

The crawl comparison sorts URLs into missing (crawled but not declared), orphan (declared but not reached), broken, redirected, canonical mismatch and noindex conflict. The crawl runs breadth-first from the homepage, respects robots.txt and Crawl-delay (capped at 10 seconds), runs at most once per site per hour and is never scheduled automatically. Search Console data, including submitted sitemaps, their counts and a bounded set of URL inspections, is shown beside Flowpane's own observation as a separate layer. Flowpane reads from Search Console. It never submits sitemaps.

The score is layered, and the weights change depending on whether crawl evidence exists.

LayerWith a crawlWithout a crawl
Structure1520
Truthfulness3035
Freshness2025
Reliability1520
Coverage20Excluded
  • Structure: XML parse, index or urlset shape, namespace, the 50,000 URL and 50 MB limits, reachable children, and a declaration in robots.txt.
  • Truthfulness: duplicates across children, protocol and host consistency, URLs that robots.txt blocks, and query or fragment URLs.
  • Freshness: lastmod coverage, bulk stamping (more than 80% of URLs sharing one value), future dates, invalid formats, and a newest date more than two years old.
  • Reliability: whether the file is delivered consistently, including compression integrity and response behaviour.
  • Coverage: the crawl comparison above.

Two rules keep the score honest. If the sitemap cannot be fetched, only structure is evaluated and the other layers contribute zero. A sitemap index can never score higher than the average of its children, so a tidy root cannot mask stale or broken files beneath it.

The Sitemap workbench includes a generator that outputs gzip files and splits above 45,000 URLs. You publish the output through your own host or CMS. Only the next public check changes the score.

Common failures

SymptomWhat it meansEffect
Root index lists years-old childrenRoot no longer tracks live inventoryFreshness falls
Most URLs share one lastmodTimestamps set at build timeFreshness falls
Declared URLs not reached by crawlOrphaned or deleted pagesCoverage falls
Sitemap URLs blocked by robots.txtTwo files disagreeCross-file coherence finding