New

Sitemap Cannibalization Checker

Group sitemap URLs whose slugs use the same words, including singular and plural forms and query variants. Titles are optional.

Sitemap input

Slugs match when they use the same words, including a trailing plural. Title tags are fetched only if you ask, and only up to 20.

Field note

Read a Sitemap Cannibalization Checker card

  1. 01

    Paste sitemap XML or enter the sitemap URL. For Sitemap Cannibalization Checker, the address to try first is https://example.com/seo-course. (Sitemap Cannibalization Checker, step 1.)

  2. 02

    Run the check. A short progress line stays on screen while a fetch is in a server job. (Sitemap Cannibalization Checker, step 2.)

  3. 03

    Read groups of near-duplicate slugs. Confirm whether the result says the list was truncated. (Sitemap Cannibalization Checker, step 3.)

  4. 04

    Copy the text you need, or download the file when this page offers one. Publish any generated XML yourself. (Sitemap Cannibalization Checker, step 4.)

Detail

Labels Sitemap Cannibalization Checker prints

TermMeaning
urlsetThe sitemap root that lists page URLs. Sitemap Cannibalization Checker reads it with the shared parser. (Sitemap Cannibalization Checker, urlset.)
sitemap indexA sitemap that lists other sitemap files. Nested indexes are fetched only to depth 2. (Sitemap Cannibalization Checker, sitemap index.)
locThe address element. Page locs and sitemap locs are kept in different lists. (Sitemap Cannibalization Checker, loc.)
truncatedThe flag that says a byte, URL, child, depth, or time cap stopped the result early. (Sitemap Cannibalization Checker, truncated.)
groups of near-duplicate slugsThe specific report Sitemap Cannibalization Checker is built to show from the shared parse. (Sitemap Cannibalization Checker, groups of near-duplicate slugs.)

Read this

A Sitemap Cannibalization Checker walkthrough

Input

Paste a urlset that contains https://example.com/seo-course?ref=1 once and the same address a second time. Sitemap Cannibalization Checker keeps one stored page and records the duplicate. Fetching https://example.com/seo-course instead of pasting uses the same parser after the download. (Sitemap Cannibalization Checker, example 1.)

What you should see

If https://example.com/seo-course is a sitemap index and https://example.com/seo-courses is the first child, a URL fetch can read that child until the depth and child caps. A paste of the index lists https://example.com/seo-courses and does not download it until you fetch the index URL. (Sitemap Cannibalization Checker, example 2.)

On this tool

Under the Sitemap Cannibalization Checker labels

Sitemap Cannibalization Checker produces groups of near-duplicate slugs. The last path segment is split into words. A trailing plural s is folded, and the words are sorted, so seo-courses and course-seo land in one group. The same path with a different query string is marked as query variants. You can paste XML into the box, or you can enter a public sitemap URL and let the server fetch it. A pasted urlset is parsed on the server without a second copy of the parser in the browser. The form stays near the top of the page, under the single title. (Sitemap Cannibalization Checker, guide 1.)

Use https://example.com/seo-course as the file you are checking and https://example.com/seo-course?ref=1 as a page address you expect to see or not see, depending on the job. A group of one is not shown. Titles appear only when you paste them, or when you ask for up to 20 title tags. The tool does not fetch a title for every URL. If the file is a sitemap index, child documents such as https://example.com/seo-courses are fetched only when you supplied the index URL, only while the depth is within 2, and only until 20 child sitemap requests have been made. (Sitemap Cannibalization Checker, guide 2.)

Sitemap Cannibalization Checker shares one fetcher with the other sitemap tools. The fetcher allows http and https only. It rejects file, gopher, and data URLs, then blocks localhost, loopback, link-local, private ranges, CGNAT space, and the cloud metadata address before it opens a socket. DNS answers are checked before the connection, and a redirect is checked again. A redirect loop is reported instead of being followed forever. Three hops is the maximum. (Sitemap Cannibalization Checker, guide 3.)

Large files are cut at 5 MB, and the stored page list stops at 5,000 URLs. Gzip responses and addresses that end in .gz are decompressed before parsing. A DOCTYPE or an entity declaration is rejected so the parser cannot expand an external entity. An HTML document or a Cloudflare challenge body is reported as not a sitemap. The raw response body is not written to the log. (Sitemap Cannibalization Checker, guide 4.)

When a URL must be fetched, the page sends the job to the server and shows a short progress line while it polls. Small pastes return immediately with the same result shape: source URL, index or urlset, child sitemaps, page rows, errors, duplicates, redirect hops, status, truncated flag, and counts. Repeat fetches of the same normalized URL can reuse a finished parse for 10 minutes. Jobs sit in a file folder so more than one server process can see them, and they expire after 30 minutes. (Sitemap Cannibalization Checker, guide 5.)

Reading a slug group as proof that two pages compete. The same words are a hint. The titles and the page intent still have to be checked by hand. After you read that distinction, look at the severity of any finding. An error means the document or the fetch failed the check this page is responsible for. A warning, such as a duplicate loc or a truncated file, means the rest of the result is still usable. A note explains a cap, a paste that did not download child files, or a sample crawl. (Sitemap Cannibalization Checker, guide 6.)

A concrete pass through Sitemap Cannibalization Checker starts with https://example.com/seo-course. If that address is an index whose first child is https://example.com/seo-courses, the file list and the page list answer different questions. The page address https://example.com/seo-course?ref=1 belongs in the page list only when a urlset contains it. Keep the two lists apart when you copy them into a spreadsheet or a ticket. (Sitemap Cannibalization Checker, guide 7.)

The result panel uses the same dark block as the other tools on this site. Copy is there when a text result is useful. A download button appears only for a file this page actually builds, such as CSV, XLSX, XML, or llms.txt. Sitemap Cannibalization Checker does not add a second title to the page, and it does not restyle the tools around it. (Sitemap Cannibalization Checker, guide 8.)

Take Sitemap Cannibalization Checker as one step. The next page in this set is Sitemap Topical Map, which answers a different question about the same sitemap. Sitemap Compare is the other close check. Finish the reading on this page before you switch, so you do not mix groups of near-duplicate slugs with a different report. (Sitemap Cannibalization Checker, guide 9.)

Where Sitemap Cannibalization Checker fits

Sitemap Cannibalization Checker is the iSkills page for groups of near-duplicate slugs. The last path segment is split into words. A trailing plural s is folded, and the words are sorted, so seo-courses and course-seo land in one group. The same path with a different query string is marked as query variants. Paste XML when you already have the file, or enter a public http or https sitemap URL when you want the server to fetch it within the published caps.

The caps are part of the result, not a hidden failure. Sitemap Cannibalization Checker stores at most 5,000 page URLs, fetches at most 20 child sitemaps, and stops nested indexes at depth 2. A truncated flag and a note tell you when the list is incomplete.

What Sitemap Cannibalization Checker returns

  • Sitemap Cannibalization Checker shows groups of near-duplicate slugs from one shared parse, so the XML rules match the other sitemap tools. (Sitemap Cannibalization Checker, result 1.)
  • A paste and a fetched URL return the same kind of result, including errors, duplicates, and counts. (Sitemap Cannibalization Checker, result 2.)
  • Private hosts, entity declarations, and oversized responses are refused before they can be used as a proxy. (Sitemap Cannibalization Checker, result 3.)

Before you trust Sitemap Cannibalization Checker

  • The input for Sitemap Cannibalization Checker was XML or a public sitemap URL. (Sitemap Cannibalization Checker, check 1.)
  • A truncated note was read before the list was treated as complete. (Sitemap Cannibalization Checker, check 2.)
  • HTML or a challenge page was not accepted as a sitemap. (Sitemap Cannibalization Checker, check 3.)
  • The copied result matches groups of near-duplicate slugs, not a different sitemap report. (Sitemap Cannibalization Checker, check 4.)
  • Generated files were published by you, not by this page. (Sitemap Cannibalization Checker, check 5.)

Easy to misread Sitemap Cannibalization Checker

Sitemap Cannibalization Checker: Mixing this result with another sitemap tool

Reading a slug group as proof that two pages compete. The same words are a hint. The titles and the page intent still have to be checked by hand. (Sitemap Cannibalization Checker, mistake 1.)

Sitemap Cannibalization Checker: Ignoring a truncated list

Sitemap Cannibalization Checker stops at the published caps. A truncated note means you do not have the full file. (Sitemap Cannibalization Checker, mistake 2.)

Sitemap Cannibalization Checker: Sending a private URL

Localhost, link-local, private ranges, and the metadata address are blocked before a socket is opened. (Sitemap Cannibalization Checker, mistake 3.)

Sitemap Cannibalization Checker: Treating HTML as XML

An HTML page or a challenge interstitial is reported as not a sitemap. Fix the URL or the firewall rule, then run the check again. (Sitemap Cannibalization Checker, mistake 4.)

From the form

Sitemap Cannibalization Checker limits

What does Sitemap Cannibalization Checker return?

Sitemap Cannibalization Checker returns groups of near-duplicate slugs. A group of one is not shown. Titles appear only when you paste them, or when you ask for up to 20 title tags. The tool does not fetch a title for every URL. (Sitemap Cannibalization Checker, What does Sitemap Cannibalization Checker return?.)

Does Sitemap Cannibalization Checker fetch a pasted file?

A paste is parsed without a download. A sitemap URL is fetched on the server. Child files inside a pasted index are listed and are not downloaded until you fetch the index URL. (Sitemap Cannibalization Checker, Does Sitemap Cannibalization Checker fetch a pasted file?.)

What will Sitemap Cannibalization Checker refuse?

It refuses non-http schemes, private and metadata addresses, DOCTYPE and entity declarations, responses over 5 MB, and HTML that is not a sitemap. (Sitemap Cannibalization Checker, What will Sitemap Cannibalization Checker refuse?.)

How large a file can Sitemap Cannibalization Checker store?

The stored page list stops at 5,000 URLs. Child sitemap fetches stop at 20, nested indexes stop at depth 2, and redirects stop at 3. The screen says when the result is truncated. (Sitemap Cannibalization Checker, How large a file can Sitemap Cannibalization Checker store?.)

Field note

After Sitemap Cannibalization Checker

Write down what Sitemap Cannibalization Checker showed, then compare that note with Sitemap Topical Map.

WhatsApp Advisor
Enroll Now