Sitemap Checker: Mixing this result with another sitemap tool
Treating HTTP 200 HTML as a valid sitemap. A challenge page or an HTML document is reported as not a sitemap even when the status is 200. (Sitemap Checker, mistake 1.)
Check whether a sitemap URL is reachable, then read the HTTP status, content type, URL count, and whether it is an index or a urlset.
On this tool
Sitemap Checker refuses a paste or download larger than 5 MB.
Sitemap Checker keeps at most 5,000 page URLs and marks a longer list truncated.
Sitemap Checker fetches at most 20 child sitemaps from an index.
Sitemap Checker does not follow a nested index past depth 2.
Sitemap Checker follows at most 3 redirects.
Note
Input
Paste a urlset that contains https://example.com/ once and the same address a second time. Sitemap Checker keeps one stored page and records the duplicate. Fetching https://example.com/sitemap.xml instead of pasting uses the same parser after the download. (Sitemap Checker, example 1.)
What you should see
If https://example.com/sitemap.xml is a sitemap index and https://example.com/sitemaps/pages.xml is the first child, a URL fetch can read that child until the depth and child caps. A paste of the index lists https://example.com/sitemaps/pages.xml and does not download it until you fetch the index URL. (Sitemap Checker, example 2.)
On this tool
Sitemap Checker produces reachability, status, content type, URL count, and index or urlset. The checker answers whether the URL responded, which HTTP status came back, which content type the host sent, how many page URLs were stored, and whether the root element was a sitemap index or a urlset. You can paste XML into the box, or you can enter a public sitemap URL and let the server fetch it. A pasted urlset is parsed on the server without a second copy of the parser in the browser. The form stays near the top of the page, under the single title. (Sitemap Checker, guide 1.)
Use https://example.com/sitemap.xml as the file you are checking and https://example.com/ as a page address you expect to see or not see, depending on the job. A paste is marked as not requested, because no host was contacted. A fetched URL shows the status from the final response after at most three redirects. If the file is a sitemap index, child documents such as https://example.com/sitemaps/pages.xml are fetched only when you supplied the index URL, only while the depth is within 2, and only until 20 child sitemap requests have been made. (Sitemap Checker, guide 2.)
Sitemap Checker shares one fetcher with the other sitemap tools. The fetcher allows http and https only. It rejects file, gopher, and data URLs, then blocks localhost, loopback, link-local, private ranges, CGNAT space, and the cloud metadata address before it opens a socket. DNS answers are checked before the connection, and a redirect is checked again. A redirect loop is reported instead of being followed forever. Three hops is the maximum. (Sitemap Checker, guide 3.)
Large files are cut at 5 MB, and the stored page list stops at 5,000 URLs. Gzip responses and addresses that end in .gz are decompressed before parsing. A DOCTYPE or an entity declaration is rejected so the parser cannot expand an external entity. An HTML document or a Cloudflare challenge body is reported as not a sitemap. The raw response body is not written to the log. (Sitemap Checker, guide 4.)
When a URL must be fetched, the page sends the job to the server and shows a short progress line while it polls. Small pastes return immediately with the same result shape: source URL, index or urlset, child sitemaps, page rows, errors, duplicates, redirect hops, status, truncated flag, and counts. Repeat fetches of the same normalized URL can reuse a finished parse for 10 minutes. Jobs sit in a file folder so more than one server process can see them, and they expire after 30 minutes. (Sitemap Checker, guide 5.)
Treating HTTP 200 HTML as a valid sitemap. A challenge page or an HTML document is reported as not a sitemap even when the status is 200. After you read that distinction, look at the severity of any finding. An error means the document or the fetch failed the check this page is responsible for. A warning, such as a duplicate loc or a truncated file, means the rest of the result is still usable. A note explains a cap, a paste that did not download child files, or a sample crawl. (Sitemap Checker, guide 6.)
A concrete pass through Sitemap Checker starts with https://example.com/sitemap.xml. If that address is an index whose first child is https://example.com/sitemaps/pages.xml, the file list and the page list answer different questions. The page address https://example.com/ belongs in the page list only when a urlset contains it. Keep the two lists apart when you copy them into a spreadsheet or a ticket. (Sitemap Checker, guide 7.)
The result panel uses the same dark block as the other tools on this site. Copy is there when a text result is useful. A download button appears only for a file this page actually builds, such as CSV, XLSX, XML, or llms.txt. Sitemap Checker does not add a second title to the page, and it does not restyle the tools around it. (Sitemap Checker, guide 8.)
Take Sitemap Checker as one step. The next page in this set is Sitemap Validator, which answers a different question about the same sitemap. Sitemap Audit is the other close check. Finish the reading on this page before you switch, so you do not mix reachability, status, content type, URL count, and index or urlset with a different report. (Sitemap Checker, guide 9.)
Sitemap Checker is the iSkills page for reachability, status, content type, URL count, and index or urlset. The checker answers whether the URL responded, which HTTP status came back, which content type the host sent, how many page URLs were stored, and whether the root element was a sitemap index or a urlset. Paste XML when you already have the file, or enter a public http or https sitemap URL when you want the server to fetch it within the published caps.
The caps are part of the result, not a hidden failure. Sitemap Checker stores at most 5,000 page URLs, fetches at most 20 child sitemaps, and stops nested indexes at depth 2. A truncated flag and a note tell you when the list is incomplete.
Paste sitemap XML or enter the sitemap URL. For Sitemap Checker, the address to try first is https://example.com/sitemap.xml. (Sitemap Checker, step 1.)
Run the check. A short progress line stays on screen while a fetch is in a server job. (Sitemap Checker, step 2.)
Read reachability, status, content type, URL count, and index or urlset. Confirm whether the result says the list was truncated. (Sitemap Checker, step 3.)
Copy the text you need, or download the file when this page offers one. Publish any generated XML yourself. (Sitemap Checker, step 4.)
On this tool
Treating HTTP 200 HTML as a valid sitemap. A challenge page or an HTML document is reported as not a sitemap even when the status is 200. (Sitemap Checker, mistake 1.)
Sitemap Checker stops at the published caps. A truncated note means you do not have the full file. (Sitemap Checker, mistake 2.)
Localhost, link-local, private ranges, and the metadata address are blocked before a socket is opened. (Sitemap Checker, mistake 3.)
An HTML page or a challenge interstitial is reported as not a sitemap. Fix the URL or the firewall rule, then run the check again. (Sitemap Checker, mistake 4.)
Field note
Sitemap Checker returns reachability, status, content type, URL count, and index or urlset. A paste is marked as not requested, because no host was contacted. A fetched URL shows the status from the final response after at most three redirects. (Sitemap Checker, What does Sitemap Checker return?.)
A paste is parsed without a download. A sitemap URL is fetched on the server. Child files inside a pasted index are listed and are not downloaded until you fetch the index URL. (Sitemap Checker, Does Sitemap Checker fetch a pasted file?.)
It refuses non-http schemes, private and metadata addresses, DOCTYPE and entity declarations, responses over 5 MB, and HTML that is not a sitemap. (Sitemap Checker, What will Sitemap Checker refuse?.)
The stored page list stops at 5,000 URLs. Child sitemap fetches stop at 20, nested indexes stop at depth 2, and redirects stop at 3. The screen says when the result is truncated. (Sitemap Checker, How large a file can Sitemap Checker store?.)
On this tool
Take what Sitemap Checker returned into Sitemap Validator before you mix it with a different report.