Sitemap URL Extractor
Fetch a sitemap or paste XML and get a flat list of page URLs, with duplicates removed and a note when the list is capped.
Note
Input and output for Sitemap URL Extractor
Input
Paste a urlset that contains https://example.com/pricing once and the same address a second time. Sitemap URL Extractor keeps one stored page and records the duplicate. Fetching https://example.com/sitemap.xml instead of pasting uses the same parser after the download. (Sitemap URL Extractor, example 1.)
What you should see
If https://example.com/sitemap.xml is a sitemap index and https://example.com/sitemaps/pages.xml is the first child, a URL fetch can read that child until the depth and child caps. A paste of the index lists https://example.com/sitemaps/pages.xml and does not download it until you fetch the index URL. (Sitemap URL Extractor, example 2.)
From the form
What Sitemap URL Extractor keeps and what it drops
Sitemap URL Extractor produces a flat list of page URLs. The URL extractor stores page loc values and drops a repeated loc after the first copy. Image and video tags stay attached to the page row, but the flat list you copy is the page address. You can paste XML into the box, or you can enter a public sitemap URL and let the server fetch it. A pasted urlset is parsed on the server without a second copy of the parser in the browser. The form stays near the top of the page, under the single title. (Sitemap URL Extractor, guide 1.)
Sitemap URL Extractor shares one fetcher with the other sitemap tools. The fetcher allows http and https only. It rejects file, gopher, and data URLs, then blocks localhost, loopback, link-local, private ranges, CGNAT space, and the cloud metadata address before it opens a socket. DNS answers are checked before the connection, and a redirect is checked again. A redirect loop is reported instead of being followed forever. Three hops is the maximum. (Sitemap URL Extractor, guide 3.)
When a URL must be fetched, the page sends the job to the server and shows a short progress line while it polls. Small pastes return immediately with the same result shape: source URL, index or urlset, child sitemaps, page rows, errors, duplicates, redirect hops, status, truncated flag, and counts. Repeat fetches of the same normalized URL can reuse a finished parse for 10 minutes. Jobs sit in a file folder so more than one server process can see them, and they expire after 30 minutes. (Sitemap URL Extractor, guide 5.)
A concrete pass through Sitemap URL Extractor starts with https://example.com/sitemap.xml. If that address is an index whose first child is https://example.com/sitemaps/pages.xml, the file list and the page list answer different questions. The page address https://example.com/pricing belongs in the page list only when a urlset contains it. Keep the two lists apart when you copy them into a spreadsheet or a ticket. (Sitemap URL Extractor, guide 7.)
Take Sitemap URL Extractor as one step. The next page in this set is Sitemap Extractor, which answers a different question about the same sitemap. Sitemap to CSV is the other close check. Finish the reading on this page before you switch, so you do not mix a flat list of page URLs with a different report. (Sitemap URL Extractor, guide 9.)
Use https://example.com/sitemap.xml as the file you are checking and https://example.com/pricing as a page address you expect to see or not see, depending on the job. The count is stored URLs, not the raw number of loc tags. A duplicate is listed under Duplicates and is not repeated in the page list. If the file is a sitemap index, child documents such as https://example.com/sitemaps/pages.xml are fetched only when you supplied the index URL, only while the depth is within 2, and only until 20 child sitemap requests have been made. (Sitemap URL Extractor, guide 2.)
Large files are cut at 5 MB, and the stored page list stops at 5,000 URLs. Gzip responses and addresses that end in .gz are decompressed before parsing. A DOCTYPE or an entity declaration is rejected so the parser cannot expand an external entity. An HTML document or a Cloudflare challenge body is reported as not a sitemap. The raw response body is not written to the log. (Sitemap URL Extractor, guide 4.)
Expecting child sitemap files to appear as pages. Those files are fetched only so their page loc values can be collected. After you read that distinction, look at the severity of any finding. An error means the document or the fetch failed the check this page is responsible for. A warning, such as a duplicate loc or a truncated file, means the rest of the result is still usable. A note explains a cap, a paste that did not download child files, or a sample crawl. (Sitemap URL Extractor, guide 6.)
The result panel uses the same dark block as the other tools on this site. Copy is there when a text result is useful. A download button appears only for a file this page actually builds, such as CSV, XLSX, XML, or llms.txt. Sitemap URL Extractor does not add a second title to the page, and it does not restyle the tools around it. (Sitemap URL Extractor, guide 8.)
Where Sitemap URL Extractor fits
Sitemap URL Extractor is the iSkills page for a flat list of page URLs. The URL extractor stores page loc values and drops a repeated loc after the first copy. Image and video tags stay attached to the page row, but the flat list you copy is the page address. Paste XML when you already have the file, or enter a public http or https sitemap URL when you want the server to fetch it within the published caps.
The caps are part of the result, not a hidden failure. Sitemap URL Extractor stores at most 5,000 page URLs, fetches at most 20 child sitemaps, and stops nested indexes at depth 2. A truncated flag and a note tell you when the list is incomplete.
Before you trust Sitemap URL Extractor
- The input for Sitemap URL Extractor was XML or a public sitemap URL. (Sitemap URL Extractor, check 1.)
- A truncated note was read before the list was treated as complete. (Sitemap URL Extractor, check 2.)
- HTML or a challenge page was not accepted as a sitemap. (Sitemap URL Extractor, check 3.)
- The copied result matches a flat list of page URLs, not a different sitemap report. (Sitemap URL Extractor, check 4.)
- Generated files were published by you, not by this page. (Sitemap URL Extractor, check 5.)
Easy to misread Sitemap URL Extractor
Sitemap URL Extractor: Mixing this result with another sitemap tool
Expecting child sitemap files to appear as pages. Those files are fetched only so their page loc values can be collected. (Sitemap URL Extractor, mistake 1.)
Sitemap URL Extractor: Ignoring a truncated list
Sitemap URL Extractor stops at the published caps. A truncated note means you do not have the full file. (Sitemap URL Extractor, mistake 2.)
Sitemap URL Extractor: Sending a private URL
Localhost, link-local, private ranges, and the metadata address are blocked before a socket is opened. (Sitemap URL Extractor, mistake 3.)
Sitemap URL Extractor: Treating HTML as XML
An HTML page or a challenge interstitial is reported as not a sitemap. Fix the URL or the firewall rule, then run the check again. (Sitemap URL Extractor, mistake 4.)
Terms used by Sitemap URL Extractor
- urlset
- The sitemap root that lists page URLs. Sitemap URL Extractor reads it with the shared parser. (Sitemap URL Extractor, urlset.)
- sitemap index
- A sitemap that lists other sitemap files. Nested indexes are fetched only to depth 2. (Sitemap URL Extractor, sitemap index.)
- loc
- The address element. Page locs and sitemap locs are kept in different lists. (Sitemap URL Extractor, loc.)
- truncated
- The flag that says a byte, URL, child, depth, or time cap stopped the result early. (Sitemap URL Extractor, truncated.)
- a flat list of page URLs
- The specific report Sitemap URL Extractor is built to show from the shared parse. (Sitemap URL Extractor, a flat list of page URLs.)
Detail
Pull a list with Sitemap URL Extractor
- 01
Paste sitemap XML or enter the sitemap URL. For Sitemap URL Extractor, the address to try first is https://example.com/sitemap.xml. (Sitemap URL Extractor, step 1.)
- 02
Run the check. A short progress line stays on screen while a fetch is in a server job. (Sitemap URL Extractor, step 2.)
- 03
Read a flat list of page URLs. Confirm whether the result says the list was truncated. (Sitemap URL Extractor, step 3.)
- 04
Copy the text you need, or download the file when this page offers one. Publish any generated XML yourself. (Sitemap URL Extractor, step 4.)
Field note
What comes back from Sitemap URL Extractor
- 01
Sitemap URL Extractor shows a flat list of page URLs from one shared parse, so the XML rules match the other sitemap tools. (Sitemap URL Extractor, result 1.)
- 02
A paste and a fetched URL return the same kind of result, including errors, duplicates, and counts. (Sitemap URL Extractor, result 2.)
- 03
Private hosts, entity declarations, and oversized responses are refused before they can be used as a proxy. (Sitemap URL Extractor, result 3.)
Read this
Straight answers for Sitemap URL Extractor
What does Sitemap URL Extractor return?
Sitemap URL Extractor returns a flat list of page URLs. The count is stored URLs, not the raw number of loc tags. A duplicate is listed under Duplicates and is not repeated in the page list. (Sitemap URL Extractor, What does Sitemap URL Extractor return?.)
Does Sitemap URL Extractor fetch a pasted file?
A paste is parsed without a download. A sitemap URL is fetched on the server. Child files inside a pasted index are listed and are not downloaded until you fetch the index URL. (Sitemap URL Extractor, Does Sitemap URL Extractor fetch a pasted file?.)
What will Sitemap URL Extractor refuse?
It refuses non-http schemes, private and metadata addresses, DOCTYPE and entity declarations, responses over 5 MB, and HTML that is not a sitemap. (Sitemap URL Extractor, What will Sitemap URL Extractor refuse?.)
How large a file can Sitemap URL Extractor store?
The stored page list stops at 5,000 URLs. Child sitemap fetches stop at 20, nested indexes stop at depth 2, and redirects stop at 3. The screen says when the result is truncated. (Sitemap URL Extractor, How large a file can Sitemap URL Extractor store?.)