SEO glossary

XML sitemap

Definition

An XML sitemap is a file listing the URLs you want search engines to discover, with an optional lastmod date for each. One file holds up to 50,000 URLs or 50MB uncompressed; beyond that you split it and add a sitemap index. It aids discovery. It never guarantees indexing.

Updated 26 July 2026 · Written by the Last Agency team · See what SEO actually costs

The short version

  • 50,000 URLs or 50MB uncompressed per file. Past that, split and use a sitemap index.
  • Google ignores priority and changefreq entirely. It uses lastmod — but only if you've never lied to it.
  • Split sitemaps by page type. The submitted-versus-indexed ratio per file is the cheapest content diagnostic you own.
  • A sitemap is a suggestion. If a page isn't indexed, listing it again won't change Google's mind.

What it's for, and the limits

A sitemap tells search engines a URL exists and roughly when it last changed. That's the whole job. It helps most where discovery is genuinely hard — a large site, a new site with few external links, pages buried deep in pagination.

It helps least on a small, well-linked site. If your 60 pages are all two clicks from the homepage, Google was going to find them anyway. The sitemap still earns its place, but as an instrument rather than a discovery tool.

The constraints are fixed: 50,000 URLs or 50MB uncompressed per file, whichever you hit first. Exceed either and you split into multiple files and publish a sitemap index — a sitemap of sitemaps, itself capped at 50,000 files. Gzip is allowed; the 50MB ceiling applies uncompressed.

Two tags in the spec are dead weight. Google has said plainly it ignores <priority> and <changefreq>. Setting every page to priority 1.0 tells a reviewer only that whoever configured your plugin didn't read the docs.

Split it by page type and it becomes a diagnostic

This is the part most sites skip, and the part that pays. One sitemap.xml with 500 URLs tells you one number. Eight sitemaps, one per page type, tell you which of your templates Google actually respects.

Search Console lets you filter the Page Indexing report by sitemap. Publish sitemap-services.xml, sitemap-glossary.xml, sitemap-locations.xml, and you can read the submitted-versus-indexed ratio for each page type independently. A site-wide 78% indexed rate is noise. Service pages at 100% and location pages at 31% is a decision.

Here's how to read it.

What a low indexed ratio in a single sitemap is telling you.
SitemapA low ratio here usually means
Core service or product pagesSomething structural — a canonical pointing elsewhere, or a noindex left on the template. Fix now; this is your money.
Blog or article pagesQuality or duplication. Google crawled and passed. More posts won't help; consolidating will.
Location or city pagesThe classic template-variant verdict. If the only difference is the city name, Google has correctly identified them as one page.
Category or collection pagesThin category copy, or near-identical product lists across categories.
Tag or archive pagesIgnore — these mostly shouldn't be in the sitemap at all.

lastmod, and the cost of lying

lastmod is the one optional field Google genuinely uses — as a hint about whether a page is worth re-crawling. Which makes it the one worth being disciplined about.

The failure mode is a CMS that stamps today's date on every URL every night. Google checks a handful, finds the content identical to last time, and learns your dates are meaningless. From then on it ignores lastmod across your whole site — including the day you genuinely update something important.

So set it when the substantive content changes. Not when a comment lands, not on a nightly rebuild. If your build can't tell content changes from deploys, omit lastmod entirely. An absent field is honest; a wrong one is a signal you've spent.

Which URLs to leave out

  • Anything with a noindex tag. Listing it says "index this" and "don't index this" in one breath.
  • Anything that redirects. A sitemap holds final URLs returning 200.
  • Non-canonical variants. Every URL should be the canonical version of itself.
  • Paginated archives, tag pages, internal search results, filter combinations.
  • Anything blocked in robots.txt — you're asking Google to fetch what you told it not to.

Submitting it, and what the report actually says

Two steps. Add a Sitemap: line to your robots.txt with the absolute URL — that's how crawlers with no access to your Search Console account find it. Then submit it under Sitemaps in Search Console, which is how you get the reporting.

The Sitemaps report itself is thin: last read date, status, discovered URL count. Don't stop there. "Discovered" means Google parsed the file, not that it indexed anything. The number you want is in the Page Indexing report, filtered by sitemap.

One honest limit: if a page isn't indexed because Google judged it thin or duplicated, resubmitting changes nothing. The sitemap solved discovery. What's failing is the indexing decision — a content problem wearing a technical costume.

Related questions.

Does an XML sitemap improve rankings?

No. It affects discovery, not position. A sitemap can get a page found faster, which matters on large or poorly-linked sites, but it has no influence on where that page ranks once it's indexed.

How many URLs can one sitemap contain?

50,000 URLs or 50MB uncompressed, whichever comes first. Beyond either limit you split into multiple files and reference them from a sitemap index file, which can itself list up to 50,000 sitemaps.

Should I include noindexed pages in my sitemap?

No — it's a direct contradiction. The sitemap says "here's a page I want indexed" and the tag says the opposite. Search Console flags the conflict, and it's one of the quickest signals that nobody's checked the setup.

Do I need an HTML sitemap as well?

Rarely, for search. An HTML sitemap is a user-facing page of links, and if you need one to make pages reachable, your navigation and internal linking are the real problem. Fix those instead.

How often should a sitemap update?

Automatically, whenever you publish or substantively edit a page. Nearly every CMS and framework does this — WordPress SEO plugins generate it, and `app/sitemap.ts` handles it in Next.js. Hand-maintained sitemaps go stale within a month, without exception.

Why does Search Console say discovered URLs but not indexed?

Because those are different stages. Discovered means Google read the file and knows the URLs exist. Indexed means it chose to store them. To see the second number, open the Page Indexing report and filter by that sitemap.

Last slot's open

Make this the last growth call you book.

Grab the free strategy call and walk away with a 90-day growth plan — hired or not. Or just text us. Either way, you'll know exactly how we'd win.

Guaranteed or it's free · No lock-in · Free strategy call