SEO glossary

Orphan page

Definition

An orphan page is a live page with no internal links pointing at it. Nothing on your site leads there, so a crawler following links never arrives and the page inherits no internal authority. Finding them requires comparing a crawl against your sitemap and Search Console — a crawler alone can't do it.

Updated 26 July 2026 · Written by the Last Agency team · See what SEO actually costs

The short version

  • A crawler can never find orphan pages by itself. It works by following links, and an orphan has none — the absence is the definition.
  • The method is a three-list diff: crawl, sitemap, Search Console URLs. Anything in lists two or three but missing from list one is an orphan.
  • Being in the sitemap means Google can find it. It doesn't mean the page has any internal context or authority, which is why orphans rank badly even when indexed.
  • Three outcomes only: link it, redirect it, or remove it. Leaving it is the fourth option and it's how you get here.

Why a crawler alone will never find them

This trips up experienced people, so it's worth stating plainly. Every crawler works the same way: start at a seed URL, fetch it, extract the links, fetch those, repeat. It's a graph traversal.

An orphan page isn't in the graph. Nothing points at it, so the crawler never reaches it, and it can't appear in a report of pages the crawler reached. Running a crawl and finding no orphans proves nothing — it's structurally guaranteed.

So you bring in URLs from somewhere other than your link graph, then check which of them the crawl never touched.

The three-list method

Build three lists of URLs from three independent sources, then compare them. Everything you need is free.

  • Orphans are (B ∪ C ∪ D) − A. Paste every list into one column, dedupe, then flag anything absent from the crawl.
  • Screaming Frog does the diff for you: connect the Search Console and GA4 APIs, enable Crawl Linked XML Sitemaps, then run Crawl Analysis and open the Orphan URLs report. Filter out 404s, redirects and parameter variants first — you want live, indexable pages.
  1. List A — the crawl. Every page reachable by following internal links from your homepage.
  2. List B — the sitemap. Fetch /sitemap.xml and its children. Your CMS generates this from the database, not from the link graph, which is exactly why it's useful.
  3. List C — Search Console. Export the Performance report's Pages tab for 16 months, plus the Indexing → Pages report. URLs Google has actually seen, whatever your links say.
  4. List D — analytics and backlinks. GA4 landing pages catch URLs that get email and ad traffic; a backlink tool catches URLs other sites link to. Both surface old campaign pages nobody remembers.

How pages get orphaned in the first place

Nobody orphans a page deliberately. It's always a side effect, and the causes are the same across almost every site we audit.

  • Campaign landing pages. A Diwali sale page linked only from an ad and a WhatsApp broadcast. The campaign ended, the page stayed.
  • Deleted category pages. A collection is removed and its products survive at their own URLs, reachable from nowhere. The single largest source on ecommerce sites.
  • Blog pagination. Posts reachable only through /blog/page/12. Google crawls deep pagination reluctantly, so old posts orphan themselves by ageing.
  • Template and taxonomy changes. Tag archives get noindexed, then tag links get pulled from the post template, and every post that was only linked from a tag loses its route in.
  • Migrations. A replatform maps the top 500 URLs and leaves the tail live, in the sitemap, and unlinked.
  • Programmatic pages. Generate 4,000 location pages, link 40 from the footer, and you've orphaned 3,960. See programmatic SEO for the version that works.

Related questions.

Are orphan pages bad for SEO?

They're wasted, more than harmful. An orphan can still be indexed via your sitemap, but it receives no internal link equity and no contextual anchor text, so it rarely ranks for anything competitive. At volume, a few hundred orphans also make your site look thinner and less organised than it is.

How do I find orphan pages for free?

Export your sitemap URLs, export Search Console's Pages report, then run a crawl with the free Screaming Frog tier (500 URLs) or any crawler you have. Paste all three into a spreadsheet and flag anything present in the exports but absent from the crawl.

Will adding an orphan page to my sitemap fix it?

It fixes discovery, not authority. Google will find and probably index the page, but it still arrives with no internal links describing it or passing signals to it. Sitemap inclusion is the minimum, not the fix — the fix is a real link from a relevant page.

Should I delete orphan pages?

Only the ones with no impressions, no backlinks and no purpose. Check Search Console for the last 16 months before deleting anything — orphans that still get impressions are usually pages worth linking properly rather than removing.

Do noindexed pages count as orphans?

Technically yes, practically no. If a page is deliberately kept out of the index — a thank-you page, an internal search result — the fact that nothing links to it isn't a problem. Filter your orphan list by indexable status before you start making decisions.

How often should I check for orphan pages?

Quarterly for ecommerce with a changing catalogue, annually for a stable brochure site, and immediately after any migration, replatform or navigation redesign. Those three events create more orphans in a week than normal operations create in a year.

Last slot's open

Make this the last growth call you book.

Grab the free strategy call and walk away with a 90-day growth plan — hired or not. Or just text us. Either way, you'll know exactly how we'd win.

Guaranteed or it's free · No lock-in · Free strategy call