Library  /  Crawl efficiency

I checked 216 sitemap URLs across 18 well known sites. Eight of them point at redirects, and they are not spread evenly.

The redirect problem in sitemaps is real but small, and it hides in a place nobody looks: it shows up on sites that are otherwise immaculate. I also found one site I had to throw out of the study, and that turned out to be the more useful finding.

Marek Ostrowski · Technical SEO lead  /  August 21, 2026  /  9 min read  /  3 sources
What this piece concludes
  • Across 216 randomly sampled sitemap URLs from 18 sites, 96% returned 200, 3.7% returned a redirect, and 0.5% returned an error.
  • The redirects were not spread evenly. Three sites accounted for all eight of them: 4 of 12 sampled on Stripe, 3 of 12 on Linear, 1 of 12 on Notion.
  • Fifteen of the eighteen sites had a perfect sample, which means a clean sitemap is normal and a dirty one is a specific defect rather than a general condition.
  • Two of twenty sites did not declare a sitemap in robots.txt at all, and both were findable only by guessing the conventional filename.
  • One site had to be excluded: it returns 403 to non browser traffic, including on its own homepage, so its sitemap could not be audited at all.

I keep reading that sitemaps are full of redirects. I have never seen the number.

So I measured it. Twenty well known product sites, their live sitemaps, twelve URLs sampled at random from each, every URL requested with redirects disabled so the first response is what gets recorded.

Here is what came back.

The headline number is boring, and that is the finding

Of 216 URLs across 18 sites, 207 returned 200. That is 96%. Eight returned a redirect, one returned an error.

I expected worse. The received wisdom is that sitemaps rot quietly, and on this sample they mostly do not. A clean sitemap is the normal state.

Which makes the exceptions worth looking at, because they are not random.

All eight redirects came from three sites

SiteSampled200Redirect
Stripe1284
Linear1293
Notion12111
Other 15 sites1801790

Four of twelve on Stripe. If that ratio holds across their 108 sampled sitemap entries, roughly a third of the file points somewhere other than the final destination.

Fifteen sites had a perfect twelve out of twelve. Ghost, Vercel, Netlify, Figma, Asana, Monday, HubSpot, Mailchimp, Webflow, Squarespace, Wix, Miro, Airtable, Zapier and one more all came back clean.

That distribution matters more than the average. A 3.7% redirect rate across the sample sounds like background noise you cannot act on. A 33% rate on one file is a defect with an owner and a fix.

What a redirect in a sitemap actually costs

Nothing punitive. Google does not penalise it, and their own documentation treats sitemaps as a hint rather than a directive.

The cost is arithmetic. A URL that redirects consumes at least two fetches instead of one: the crawler requests the listed address, receives a 301 or 308, then requests the target. On a file with 100 entries that is a rounding error. On the sites in this sample with thousands of entries, it stops being one.

Webflow’s sitemap listed 13,800 URLs. Vercel’s, 6,299. At those sizes a 30% redirect rate would mean thousands of wasted requests against a budget nobody publishes to you.

The site I had to throw out

Twenty sites went in. Nineteen produced usable data. One did not, and the reason is the part of this study I would keep if I had to discard the rest.

Canva returned an error on all twelve sampled URLs. Twelve out of twelve is not a content problem, it is a pattern, so I checked before writing anything down.

Their robots.txt returns 200. Their sitemap returns 403. Their homepage returns 403.

That is bot protection responding to a non browser user agent, not a broken site. Canva loads perfectly in a browser. Had I published “12 of 12 Canva sitemap URLs are broken” it would have been false, and the falsehood would have been entirely my own doing.

I mention it because the same trap sits in every automated audit, including the ones sold as products. A tool that cannot distinguish “this URL is broken” from “this site declined to talk to me” will hand you a report full of confident errors.

The tell is uniformity. Real defects are patchy. When every single URL fails identically, suspect the door, not the rooms.

Two sites do not declare their sitemap at all

Separately from the redirect question: of the twenty sites, eighteen declared a sitemap in robots.txt. Two did not.

HubSpot and Shopify both have working sitemaps at conventional locations. Neither points to them from robots.txt.

For a site of that size and authority it changes little, because the crawler already knows the property intimately. For a new domain it changes everything, and I have written about that separately.

The fix is one line and takes a minute, and it costs less thought than the arguments people have about link attributes:

Sitemap: https://example.com/sitemap.xml

How to check your own, properly

Four steps, about ten minutes.

One. Fetch your sitemap and pull the <loc> values. If it is a sitemap index, descend one level first, because the index itself only contains links to other files.

Two. Request a random sample with redirects disabled. This matters. Most HTTP clients follow redirects silently by default, which turns a 301 into a 200 in your results and hides the exact thing you are measuring.

Three. Before believing a run of failures, request the site’s own homepage the same way. If that fails too, you are measuring the bot policy.

Four. Twelve URLs per file is enough to catch a systemic problem. It caught all three here. It will not find a single stale entry among ten thousand, and it is not meant to.

What I would not conclude from this

The sample is twenty sites, all product companies, all large, all English language. It says nothing about e-commerce catalogues, where the failure modes are different and usually worse: expired product pages, faceted URLs, seasonal categories.

It also measures one moment. A sitemap that is clean today can break with the next release, which is the real argument for checking on a schedule rather than once.

And the redirects I found may be deliberate. A team consolidating URLs might leave the old address in the sitemap during a transition on purpose. I did not ask them, so I am reporting the response codes, not the intent behind them.

Questions people ask about this

Does a redirect in a sitemap actually hurt?

It does not carry a penalty. It wastes crawl requests: each redirected URL costs at least two fetches instead of one, and on a large site that adds up against a budget you do not control.

How many URLs should I sample to check my own sitemap?

Twelve per sitemap file caught the problem on all three affected sites here. If you find zero issues in twelve, sample another thirty before concluding the file is clean.

Why exclude a site that returns 403?

Because the measurement would be about their bot policy, not about their sitemap. Reporting 12 failures as 12 broken URLs would have been wrong, and it was the first thing I checked.

Read next
Indexing
Six new sites, one launch day, 106 URLs. Six days later Google had indexed 94 of them.
Internal linking
We counted our own internal links. Two thirds were furniture.
Before you touch the next page

Run the free check on the page you care about. It reads the live HTML and reports what a crawler sees: status, noindex, canonical, headings, links and structured data.

Run the free check → Open the glossary