New

Sitemap vs Crawl

Compare sitemap URLs with a 30-page same-host sample crawl and split the result into sitemap only, crawl only, and in both.

Sitemap input

This is a sample crawl of at most 30 HTML pages. It is not a full crawl of the site.

Note

Where a Sitemap vs Crawl reading goes wrong

MistakeWhat happens
Sitemap vs Crawl: Mixing this result with another sitemap toolTreating sitemap only as an error. The crawl stops at 30 pages, so most sitemap URLs will be sitemap only on a large site. (Sitemap vs Crawl, mistake 1.)
Sitemap vs Crawl: Ignoring a truncated listSitemap vs Crawl stops at the published caps. A truncated note means you do not have the full file. (Sitemap vs Crawl, mistake 2.)
Sitemap vs Crawl: Sending a private URLLocalhost, link-local, private ranges, and the metadata address are blocked before a socket is opened. (Sitemap vs Crawl, mistake 3.)
Sitemap vs Crawl: Treating HTML as XMLAn HTML page or a challenge interstitial is reported as not a sitemap. Fix the URL or the firewall rule, then run the check again. (Sitemap vs Crawl, mistake 4.)

In practice

The two sides Sitemap vs Crawl is separating

Sitemap vs Crawl produces sitemap only, crawl only, and both. The same 30-page sample crawl is split three ways. Sitemap only means the address was stored from the sitemap and was not seen in the crawl. Crawl only means the crawl saw it and the sitemap did not. Both means the address was in each set. You can paste XML into the box, or you can enter a public sitemap URL and let the server fetch it. A pasted urlset is parsed on the server without a second copy of the parser in the browser. The form stays near the top of the page, under the single title. (Sitemap vs Crawl, guide 1.)

Sitemap vs Crawl shares one fetcher with the other sitemap tools. The fetcher allows http and https only. It rejects file, gopher, and data URLs, then blocks localhost, loopback, link-local, private ranges, CGNAT space, and the cloud metadata address before it opens a socket. DNS answers are checked before the connection, and a redirect is checked again. A redirect loop is reported instead of being followed forever. Three hops is the maximum. (Sitemap vs Crawl, guide 3.)

When a URL must be fetched, the page sends the job to the server and shows a short progress line while it polls. Small pastes return immediately with the same result shape: source URL, index or urlset, child sitemaps, page rows, errors, duplicates, redirect hops, status, truncated flag, and counts. Repeat fetches of the same normalized URL can reuse a finished parse for 10 minutes. Jobs sit in a file folder so more than one server process can see them, and they expire after 30 minutes. (Sitemap vs Crawl, guide 5.)

A concrete pass through Sitemap vs Crawl starts with https://example.com/sitemap.xml. If that address is an index whose first child is https://example.com/about, the file list and the page list answer different questions. The page address https://example.com/pricing belongs in the page list only when a urlset contains it. Keep the two lists apart when you copy them into a spreadsheet or a ticket. (Sitemap vs Crawl, guide 7.)

Take Sitemap vs Crawl as one step. The next page in this set is Orphan Page Finder, which answers a different question about the same sitemap. Sitemap Compare is the other close check. Finish the reading on this page before you switch, so you do not mix sitemap only, crawl only, and both with a different report. (Sitemap vs Crawl, guide 9.)

Use https://example.com/sitemap.xml as the file you are checking and https://example.com/pricing as a page address you expect to see or not see, depending on the job. Use sitemap only to spot pages the sample crawl did not reach. Use crawl only as a hint that a linked URL is missing from the sitemap. Use both as the overlap, not as a ranking. If the file is a sitemap index, child documents such as https://example.com/about are fetched only when you supplied the index URL, only while the depth is within 2, and only until 20 child sitemap requests have been made. (Sitemap vs Crawl, guide 2.)

Large files are cut at 5 MB, and the stored page list stops at 5,000 URLs. Gzip responses and addresses that end in .gz are decompressed before parsing. A DOCTYPE or an entity declaration is rejected so the parser cannot expand an external entity. An HTML document or a Cloudflare challenge body is reported as not a sitemap. The raw response body is not written to the log. (Sitemap vs Crawl, guide 4.)

Treating sitemap only as an error. The crawl stops at 30 pages, so most sitemap URLs will be sitemap only on a large site. After you read that distinction, look at the severity of any finding. An error means the document or the fetch failed the check this page is responsible for. A warning, such as a duplicate loc or a truncated file, means the rest of the result is still usable. A note explains a cap, a paste that did not download child files, or a sample crawl. (Sitemap vs Crawl, guide 6.)

The result panel uses the same dark block as the other tools on this site. Copy is there when a text result is useful. A download button appears only for a file this page actually builds, such as CSV, XLSX, XML, or llms.txt. Sitemap vs Crawl does not add a second title to the page, and it does not restyle the tools around it. (Sitemap vs Crawl, guide 8.)

Where Sitemap vs Crawl fits

Sitemap vs Crawl is the iSkills page for sitemap only, crawl only, and both. The same 30-page sample crawl is split three ways. Sitemap only means the address was stored from the sitemap and was not seen in the crawl. Crawl only means the crawl saw it and the sitemap did not. Both means the address was in each set. Paste XML when you already have the file, or enter a public http or https sitemap URL when you want the server to fetch it within the published caps.

The caps are part of the result, not a hidden failure. Sitemap vs Crawl stores at most 5,000 page URLs, fetches at most 20 child sitemaps, and stops nested indexes at depth 2. A truncated flag and a note tell you when the list is incomplete.

Sitemap vs Crawl in order

Paste sitemap XML or enter the sitemap URL. For Sitemap vs Crawl, the address to try first is https://example.com/sitemap.xml. (Sitemap vs Crawl, step 1.)

Run the check. A short progress line stays on screen while a fetch is in a server job. (Sitemap vs Crawl, step 2.)

Read sitemap only, crawl only, and both. Confirm whether the result says the list was truncated. (Sitemap vs Crawl, step 3.)

Copy the text you need, or download the file when this page offers one. Publish any generated XML yourself. (Sitemap vs Crawl, step 4.)

What Sitemap vs Crawl returns

  • Sitemap vs Crawl shows sitemap only, crawl only, and both from one shared parse, so the XML rules match the other sitemap tools. (Sitemap vs Crawl, result 1.)
  • A paste and a fetched URL return the same kind of result, including errors, duplicates, and counts. (Sitemap vs Crawl, result 2.)
  • Private hosts, entity declarations, and oversized responses are refused before they can be used as a proxy. (Sitemap vs Crawl, result 3.)

Before you trust Sitemap vs Crawl

  • The input for Sitemap vs Crawl was XML or a public sitemap URL. (Sitemap vs Crawl, check 1.)
  • A truncated note was read before the list was treated as complete. (Sitemap vs Crawl, check 2.)
  • HTML or a challenge page was not accepted as a sitemap. (Sitemap vs Crawl, check 3.)
  • The copied result matches sitemap only, crawl only, and both, not a different sitemap report. (Sitemap vs Crawl, check 4.)
  • Generated files were published by you, not by this page. (Sitemap vs Crawl, check 5.)

Terms used by Sitemap vs Crawl

urlset
The sitemap root that lists page URLs. Sitemap vs Crawl reads it with the shared parser. (Sitemap vs Crawl, urlset.)
sitemap index
A sitemap that lists other sitemap files. Nested indexes are fetched only to depth 2. (Sitemap vs Crawl, sitemap index.)
loc
The address element. Page locs and sitemap locs are kept in different lists. (Sitemap vs Crawl, loc.)
truncated
The flag that says a byte, URL, child, depth, or time cap stopped the result early. (Sitemap vs Crawl, truncated.)
sitemap only, crawl only, and both
The specific report Sitemap vs Crawl is built to show from the shared parse. (Sitemap vs Crawl, sitemap only, crawl only, and both.)

In practice

A small Sitemap vs Crawl diff

Input

Paste a urlset that contains https://example.com/pricing once and the same address a second time. Sitemap vs Crawl keeps one stored page and records the duplicate. Fetching https://example.com/sitemap.xml instead of pasting uses the same parser after the download. (Sitemap vs Crawl, example 1.)

What you should see

If https://example.com/sitemap.xml is a sitemap index and https://example.com/about is the first child, a URL fetch can read that child until the depth and child caps. A paste of the index lists https://example.com/about and does not download it until you fetch the index URL. (Sitemap vs Crawl, example 2.)

Note

Questions while comparing in Sitemap vs Crawl

What does Sitemap vs Crawl return?

Sitemap vs Crawl returns sitemap only, crawl only, and both. Use sitemap only to spot pages the sample crawl did not reach. Use crawl only as a hint that a linked URL is missing from the sitemap. Use both as the overlap, not as a ranking. (Sitemap vs Crawl, What does Sitemap vs Crawl return?.)

Does Sitemap vs Crawl fetch a pasted file?

A paste is parsed without a download. A sitemap URL is fetched on the server. Child files inside a pasted index are listed and are not downloaded until you fetch the index URL. (Sitemap vs Crawl, Does Sitemap vs Crawl fetch a pasted file?.)

What will Sitemap vs Crawl refuse?

It refuses non-http schemes, private and metadata addresses, DOCTYPE and entity declarations, responses over 5 MB, and HTML that is not a sitemap. (Sitemap vs Crawl, What will Sitemap vs Crawl refuse?.)

How large a file can Sitemap vs Crawl store?

The stored page list stops at 5,000 URLs. Child sitemap fetches stop at 20, nested indexes stop at depth 2, and redirects stop at 3. The screen says when the result is truncated. (Sitemap vs Crawl, How large a file can Sitemap vs Crawl store?.)

From the form

Leave Sitemap vs Crawl with

Once Sitemap vs Crawl has split the two sides, open Orphan Page Finder for a closer read.

WhatsApp Advisor
Enroll Now