Orphan Page Finder
Run a small same-host crawl of at most 30 pages and list HTML URLs that the crawl found but the sitemap does not contain.
From the form
Where Orphan Page Finder fits a share
- The input for Orphan Page Finder was XML or a public sitemap URL. (Orphan Page Finder, check 1.)
- A truncated note was read before the list was treated as complete. (Orphan Page Finder, check 2.)
- HTML or a challenge page was not accepted as a sitemap. (Orphan Page Finder, check 3.)
- The copied result matches crawl URLs that are absent from the sitemap, not a different sitemap report. (Orphan Page Finder, check 4.)
- Generated files were published by you, not by this page. (Orphan Page Finder, check 5.)
Detail
A Orphan Page Finder preview
Input
Paste a urlset that contains https://example.com/ once and the same address a second time. Orphan Page Finder keeps one stored page and records the duplicate. Fetching https://example.com/sitemap.xml instead of pasting uses the same parser after the download. (Orphan Page Finder, example 1.)
What you should see
If https://example.com/sitemap.xml is a sitemap index and https://example.com/hidden-landing is the first child, a URL fetch can read that child until the depth and child caps. A paste of the index lists https://example.com/hidden-landing and does not download it until you fetch the index URL. (Orphan Page Finder, example 2.)
Note
Tags and frames from Orphan Page Finder
- 01
Orphan Page Finder shows crawl URLs that are absent from the sitemap from one shared parse, so the XML rules match the other sitemap tools. (Orphan Page Finder, result 1.)
- 02
A paste and a fetched URL return the same kind of result, including errors, duplicates, and counts. (Orphan Page Finder, result 2.)
- 03
Private hosts, entity declarations, and oversized responses are refused before they can be used as a proxy. (Orphan Page Finder, result 3.)
Note
The share card Orphan Page Finder is judging
Orphan Page Finder produces crawl URLs that are absent from the sitemap. The finder crawls the same registrable host as the sitemap, at most 30 HTML pages, depth 2, two pages at a time. A URL found in that sample crawl and missing from the sitemap is listed as a possible orphan. You can paste XML into the box, or you can enter a public sitemap URL and let the server fetch it. A pasted urlset is parsed on the server without a second copy of the parser in the browser. The form stays near the top of the page, under the single title. (Orphan Page Finder, guide 1.)
Use https://example.com/sitemap.xml as the file you are checking and https://example.com/ as a page address you expect to see or not see, depending on the job. The note on the result says this is a sample crawl. A page that is three clicks from the homepage can be absent from both lists. Absence from this crawl is not proof the page does not exist. If the file is a sitemap index, child documents such as https://example.com/hidden-landing are fetched only when you supplied the index URL, only while the depth is within 2, and only until 20 child sitemap requests have been made. (Orphan Page Finder, guide 2.)
Orphan Page Finder shares one fetcher with the other sitemap tools. The fetcher allows http and https only. It rejects file, gopher, and data URLs, then blocks localhost, loopback, link-local, private ranges, CGNAT space, and the cloud metadata address before it opens a socket. DNS answers are checked before the connection, and a redirect is checked again. A redirect loop is reported instead of being followed forever. Three hops is the maximum. (Orphan Page Finder, guide 3.)
Large files are cut at 5 MB, and the stored page list stops at 5,000 URLs. Gzip responses and addresses that end in .gz are decompressed before parsing. A DOCTYPE or an entity declaration is rejected so the parser cannot expand an external entity. An HTML document or a Cloudflare challenge body is reported as not a sitemap. The raw response body is not written to the log. (Orphan Page Finder, guide 4.)
When a URL must be fetched, the page sends the job to the server and shows a short progress line while it polls. Small pastes return immediately with the same result shape: source URL, index or urlset, child sitemaps, page rows, errors, duplicates, redirect hops, status, truncated flag, and counts. Repeat fetches of the same normalized URL can reuse a finished parse for 10 minutes. Jobs sit in a file folder so more than one server process can see them, and they expire after 30 minutes. (Orphan Page Finder, guide 5.)
Calling the list a complete orphan report. Thirty pages cannot describe a large site. After you read that distinction, look at the severity of any finding. An error means the document or the fetch failed the check this page is responsible for. A warning, such as a duplicate loc or a truncated file, means the rest of the result is still usable. A note explains a cap, a paste that did not download child files, or a sample crawl. (Orphan Page Finder, guide 6.)
A concrete pass through Orphan Page Finder starts with https://example.com/sitemap.xml. If that address is an index whose first child is https://example.com/hidden-landing, the file list and the page list answer different questions. The page address https://example.com/ belongs in the page list only when a urlset contains it. Keep the two lists apart when you copy them into a spreadsheet or a ticket. (Orphan Page Finder, guide 7.)
The result panel uses the same dark block as the other tools on this site. Copy is there when a text result is useful. A download button appears only for a file this page actually builds, such as CSV, XLSX, XML, or llms.txt. Orphan Page Finder does not add a second title to the page, and it does not restyle the tools around it. (Orphan Page Finder, guide 8.)
Take Orphan Page Finder as one step. The next page in this set is Sitemap vs Crawl, which answers a different question about the same sitemap. Sitemap URL Extractor is the other close check. Finish the reading on this page before you switch, so you do not mix crawl URLs that are absent from the sitemap with a different report. (Orphan Page Finder, guide 9.)
Where Orphan Page Finder fits
Orphan Page Finder is the iSkills page for crawl URLs that are absent from the sitemap. The finder crawls the same registrable host as the sitemap, at most 30 HTML pages, depth 2, two pages at a time. A URL found in that sample crawl and missing from the sitemap is listed as a possible orphan. Paste XML when you already have the file, or enter a public http or https sitemap URL when you want the server to fetch it within the published caps.
The caps are part of the result, not a hidden failure. Orphan Page Finder stores at most 5,000 page URLs, fetches at most 20 child sitemaps, and stops nested indexes at depth 2. A truncated flag and a note tell you when the list is incomplete.
Orphan Page Finder in order
Paste sitemap XML or enter the sitemap URL. For Orphan Page Finder, the address to try first is https://example.com/sitemap.xml. (Orphan Page Finder, step 1.)
Run the check. A short progress line stays on screen while a fetch is in a server job. (Orphan Page Finder, step 2.)
Read crawl URLs that are absent from the sitemap. Confirm whether the result says the list was truncated. (Orphan Page Finder, step 3.)
Copy the text you need, or download the file when this page offers one. Publish any generated XML yourself. (Orphan Page Finder, step 4.)
Easy to misread Orphan Page Finder
Orphan Page Finder: Mixing this result with another sitemap tool
Calling the list a complete orphan report. Thirty pages cannot describe a large site. (Orphan Page Finder, mistake 1.)
Orphan Page Finder: Ignoring a truncated list
Orphan Page Finder stops at the published caps. A truncated note means you do not have the full file. (Orphan Page Finder, mistake 2.)
Orphan Page Finder: Sending a private URL
Localhost, link-local, private ranges, and the metadata address are blocked before a socket is opened. (Orphan Page Finder, mistake 3.)
Orphan Page Finder: Treating HTML as XML
An HTML page or a challenge interstitial is reported as not a sitemap. Fix the URL or the firewall rule, then run the check again. (Orphan Page Finder, mistake 4.)
Terms used by Orphan Page Finder
- urlset
- The sitemap root that lists page URLs. Orphan Page Finder reads it with the shared parser. (Orphan Page Finder, urlset.)
- sitemap index
- A sitemap that lists other sitemap files. Nested indexes are fetched only to depth 2. (Orphan Page Finder, sitemap index.)
- loc
- The address element. Page locs and sitemap locs are kept in different lists. (Orphan Page Finder, loc.)
- truncated
- The flag that says a byte, URL, child, depth, or time cap stopped the result early. (Orphan Page Finder, truncated.)
- crawl URLs that are absent from the sitemap
- The specific report Orphan Page Finder is built to show from the shared parse. (Orphan Page Finder, crawl URLs that are absent from the sitemap.)
From the form
Sharing questions for Orphan Page Finder
What does Orphan Page Finder return?
Orphan Page Finder returns crawl URLs that are absent from the sitemap. The note on the result says this is a sample crawl. A page that is three clicks from the homepage can be absent from both lists. Absence from this crawl is not proof the page does not exist. (Orphan Page Finder, What does Orphan Page Finder return?.)
Does Orphan Page Finder fetch a pasted file?
A paste is parsed without a download. A sitemap URL is fetched on the server. Child files inside a pasted index are listed and are not downloaded until you fetch the index URL. (Orphan Page Finder, Does Orphan Page Finder fetch a pasted file?.)
What will Orphan Page Finder refuse?
It refuses non-http schemes, private and metadata addresses, DOCTYPE and entity declarations, responses over 5 MB, and HTML that is not a sitemap. (Orphan Page Finder, What will Orphan Page Finder refuse?.)
How large a file can Orphan Page Finder store?
The stored page list stops at 5,000 URLs. Child sitemap fetches stop at 20, nested indexes stop at depth 2, and redirects stop at 3. The screen says when the result is truncated. (Orphan Page Finder, How large a file can Orphan Page Finder store?.)
Field note
Ship the Orphan Page Finder tags
After Orphan Page Finder looks right, open Sitemap vs Crawl if you still need a different tag or preview.