Crawling Indexing

Orphan page

A page with no internal links pointing to it, reachable only via sitemap, external link or direct URL.

In full

Orphans are found by comparing a crawl of the site's link graph against a full URL inventory from sitemaps, analytics, logs and the CMS database. They are crawled rarely, receive no internal link equity, and are often either forgotten legacy pages worth deleting or valuable pages accidentally cut off by a navigation change. On large sites, orphan detection is a routine part of migration QA.

Example

A crawl finds 340 landing pages that exist in the sitemap and in analytics but are linked from nowhere after a nav redesign.

Related terms

Internal linking

Links between pages on the same site, which control crawl paths, distribute link equity, and communicate topical relationships and page…

Log file analysis

Examining raw server access logs to see exactly which URLs crawlers fetched, how often, with what status codes and response times.

Crawl budget

The number of URLs a search engine is willing and able to crawl on a site in a given period, set by crawl capacity limit and crawl demand.

XML sitemap

An XML file listing URLs a site wants crawled, optionally with lastmod dates, used as a discovery aid by Google, Bing and other engines.

Site migration

Any change to a site's domain, URL structure, platform, template or content organisation that materially affects how search engines see it.