AI Visibility Checker
Check which AI crawlers may fetch one public URL. Reads robots.txt, llms.txt, meta robots, canonical, and a Sitemap hint. No chatbot is queried.
In practice
Findings to confirm in AI Visibility Checker
- The URL was a public http or https address, and a private or metadata host was refused before a connection.
- robots.txt was either read, reported missing as a warning with the default allow, or reported unreadable when the body was HTML.
- Each of the eleven crawlers shows allowed or blocked at the path, with the rule that matched, or the list is omitted because the file could not be read.
- llms.txt either starts with a markdown H1 or is marked missing as a warning. An HTML body does not pass.
- A missing llms-full.txt is a note. The summary still says whether llms.txt exists and whether the page is noindex.
- The copied report matches the rows on screen and does not add a score, a rank, or a traffic number.
In practice
How a AI Visibility Checker finding reads
Input
Worked example. The fictional URL is http://harbornotes.example/guides/tides. Its robots.txt is: User-agent: GPTBot / Disallow: / / User-agent: Google-Extended / Allow: / / User-agent: * / Allow: / / Sitemap: http://harbornotes.example/sitemap.xml. The HTML title is Harbor Notes, the only H1 is Tides at the harbor, the meta description is a sentence about a harbor journal, JSON-LD is present, and there is no robots meta. llms.txt begins with the line # Harbor Notes. llms-full.txt is not on the host.
What you should see
What the checklist should show: GPTBot is blocked because Disallow: / matches /guides/tides. Google-Extended is allowed because Allow: / matches that path. The other nine named crawlers follow the star group and are allowed. The summary reads 1 of 11 AI crawlers are blocked (10 of 11 allowed). llms.txt exists. The page is not noindex. The llms.txt row passes and prints the heading Harbor Notes. The llms-full.txt row is a note. The sitemap hint prints http://harbornotes.example/sitemap.xml and says that address was not fetched. The H1 row passes because there is exactly one heading.
From the form
Language in the AI Visibility Checker report
| Term | Meaning |
|---|---|
| User-agent group | The block of Allow and Disallow lines that follows one or more User-agent lines in robots.txt. |
| Default allow | What a crawler does when robots.txt is missing, or when no rule matches the path. |
| Markdown H1 | A first line that starts with a single hash and a space, which is the opening llms.txt asks for. |
| X-Robots-Tag | An HTTP header that can carry noindex, nosnippet, noai, or noimageai for the response. |
| Sitemap hint | A Sitemap URL copied from robots.txt and shown without being fetched. |
On this tool
Severity, scope, and what AI Visibility Checker skipped
Publishers keep asking whether a page is visible to the new crawlers, and the honest answer is a set of files you can open yourself. AI Visibility Checker fetches the URL you typed, then requests robots.txt, llms.txt, and llms-full.txt on that same origin. This check does not query a chatbot. Nothing in the result is a reply from a model, a citation rank, or a guess about how often a bot will visit. If a row is red, the evidence is a header, a meta tag, or a line from a file, and you can compare that line with the file on the server.
The crawler list is fixed so the math stays visible. The eleven names are GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, Google-Extended, PerplexityBot, Amazonbot, CCBot, Applebot-Extended, Bytespider, and Meta-ExternalAgent. For each name the parser selects the most specific User-agent group, falling back to the star group when that bot has no group of its own. A rule matches from the start of the path. An asterisk in a pattern matches across characters, and a dollar sign at the end means the path must end there. The longest matching pattern wins. When an Allow and a Disallow are the same length, Allow wins. If nothing matches, the default is allow. The summary then says how many of those eleven are blocked, for example one of eleven blocked and ten of eleven allowed, so you can see the fraction without a second number hiding it.
A missing robots.txt is not treated as a refusal. When the host answers 404 or 410, every named crawler is marked allowed and the row is a warning. The warning exists because you may have meant to publish a file and the server is serving a soft gap, not because the crawlers are blocked. An empty 200 is different: that file was published and it simply has no rules, so the default allow stands as a pass when nobody is blocked. If the response is HTML, a browser challenge, or a status that is not a successful text file, the row fails and says the file could not be read. In that case the summary does not pretend to count blocks. Gzip is decoded before any group is read, using the same decoder the sitemap tools use for a compressed robots.txt.
llms.txt is a convention, not a ranking factor this page invents. The check requests /llms.txt on the origin of the page that was actually fetched. It passes when the body, after a trim, starts with a markdown H1 such as a line that reads # Name, and the body is not an HTML document. The first heading is printed so you can see which title the file is offering. A 404 or 410 is a warning. Missing is not a fail, because many useful sites have not published the file yet and a missing guide is not the same as a block. A file that comes back as a full HTML page, including a challenge page that happens to use that path, fails the row. The failure is about the shape of the body, not about a model refusing the site.
llms-full.txt is the longer companion and it is optional here. The same H1 test is used, and the first heading is shown when the file matches. If the path 404s, the row is a note. A note is recorded so you can see that the optional file is absent, and it is not counted as a failure of the page. An HTML body at that path still fails, because a challenge page should not be mistaken for the longer markdown file. You can publish the short file without the long one. The checker will not scold you for that choice.
The page itself is read for the robots meta element and for the X-Robots-Tag header when the fetch exposes one. noindex fails that check, including the token none, which means the same thing for indexing. nosnippet, noai, and noimageai are warnings. They limit what a consumer may quote or reuse, but they are not the same decision as noindex, so they do not flip the summary into the noindex sentence. If both a hard noindex and a softer token are present, the row fails and the evidence still lists every token that was found. When neither the meta element nor the header is present, the row passes and says so. A page that never came back, because the status was not successful, fails this check with the status rather than claiming the page is indexable.
Extractable content is a bundle of signals taken from the HTML after the fetch. You see the title element, the exact text of the H1, the meta description, a rough word count of the main text after script and style elements are removed, and whether a JSON-LD script is present. The H1 row warns when the heading is missing and when more than one H1 is present, because a crawler that is building a title from the page then has to choose. The word count is a count of words left in the body, not a quality grade and not a promise that a model will quote them. JSON-LD is recorded as present or not found. It is not validated against a rich-result gallery. If you need a structural check of a script you already wrote, that work belongs on a schema tool, not inside this visibility list.
The canonical link is reported as present or absent, and the check says whether the absolute address points at another host. A same-host canonical passes. A missing canonical is a warning, because the page can still be fetched but you have not named the preferred URL. A canonical on a different host is also a warning, and the evidence prints that absolute URL so you can see the host yourself. Relative hrefs are resolved against the fetched page before the hosts are compared. The sitemap hint is smaller than a sitemap audit. If robots.txt contains one or more Sitemap lines, those URLs are printed and explicitly not fetched. This page does not walk child sitemaps, does not count the URLs inside them, and does not decide whether the page you typed appears in the file. When robots.txt is missing, the sitemap row is a note. When the robots file could not be read, the sitemap row says that too.
The summary line is the only total on the page, and it is a sentence rather than a badge out of one hundred. It states how many of the eleven crawlers are blocked, whether llms.txt exists, and whether the page is noindex. When robots.txt cannot be read, the sentence says the crawler rules were not counted, instead of printing a zero that would look like a clean allow. There is no AI score, no ChatGPT rank, and no traffic estimate. If you want a different cut of the same host, the Robots.txt Sitemap Checker compares Sitemap lines with a sitemap address, the LLMs.txt Checker + Validator reads markdown you paste, and the Sitemap Checker looks at one sitemap URL. None of those pages queries a chatbot either.
The fetch has the same walls as the other tools that download a public file on this site. The URL has to be http or https, with no username or password in it. Loopback, private ranges, link-local addresses, and the cloud metadata hosts are rejected by the existing guard, including 127.0.0.1 and 169.254.169.254, and a redirect that lands on one of those addresses is rejected again. Responses stop at five megabytes. Redirects stop after a few hops. A timeout ends the attempt. The HTML that is kept is the body that already fit under that cap. The checker does not open a second client of its own, so a rule you rely on for sitemap fetches is the same rule here. If the host answers with a challenge page, you will see that the file could not be read, which is the useful outcome: you learn that the check did not silently treat a block page as permission.
Where AI Visibility Checker fits
AI Visibility Checker reads one public page and the host files that tell a crawler whether that page is meant to be fetched. You get a checklist, not a rank. This check does not query a chatbot. It does not call ChatGPT, Claude, Gemini, or any other model, and it does not estimate traffic.
The page you enter has to be an http or https URL on a public host. Private, loopback, link-local, and metadata addresses are refused before a connection is opened. The same fetch client used by the sitemap tools caps the download, follows a short redirect chain, and checks each hop again.
AI Visibility Checker in order
Paste one public URL, including the path of the page you care about, not only the bare host.
Wait for the page, robots.txt, llms.txt, and llms-full.txt. A single check stays in one request.
Read the summary first: how many of the eleven named crawlers are blocked, whether llms.txt exists, and whether the page is noindex.
Open the crawler rows and the other checks, then copy the plain-text report if you need the evidence somewhere else.
What AI Visibility Checker returns
- Each named crawler is allowed or blocked at the URL path, with the Allow or Disallow line that decided it.
- A missing robots.txt is a warning. The report says the default is allow, because absence is not a block.
- llms.txt passes only when the body starts with a markdown H1. An HTML or challenge page does not pass.
- llms-full.txt is optional. A missing file is a note, so it does not fail the checklist by itself.
- Title, the exact H1 text, the meta description, a rough word count, and JSON-LD are shown as signals, not as a score.
Easy to misread AI Visibility Checker
Reading a missing robots.txt as a block
A 404 means the file is absent. The default is allow, so the row warns. It does not mean GPTBot was disallowed.
Treating the summary as a rank
One of eleven blocked is a count of the named crawlers. It is not a score out of one hundred and it is not a traffic number.
Expecting the sitemap to be crawled
A Sitemap line is printed. The checker does not download that file or the pages inside it.
Asking the page to query a chatbot
This check does not query a chatbot. It cannot tell you what a model would answer about the page.
On this tool
Questions after a AI Visibility Checker run
Does this page ask a chatbot if the URL is visible?
No. This check does not query a chatbot. It fetches the public page, robots.txt, llms.txt, and llms-full.txt, then prints what those files say. There is no call to OpenAI, Anthropic, Gemini, or any other model API.
What does 1 of 11 blocked mean?
It means one of the eleven named crawlers matched a Disallow for the path, and the other ten did not. The summary prints both the blocked count and the allowed count so the fraction is the whole result. It is not an AI score.
Why did llms.txt fail when the server returned a web page?
The convention is a markdown file that starts with an H1. An HTML body, including a challenge page, does not pass. A missing file is only a warning. Those two outcomes are different, and the evidence says which one happened.
Will a noindex page still show the crawler list?
Yes. noindex fails the robots meta row and the summary says the page is noindex. The robots.txt rows are still listed, because a crawler rule and a noindex token are separate decisions. nosnippet, noai, and noimageai warn instead of failing, unless noindex is also present.
Which address is refused before any download?
Any URL that is not http or https, any URL with a username or password, and any host the sitemap guard already blocks: loopback, private ranges, link-local, and metadata addresses such as 127.0.0.1 and 169.254.169.254. A redirect onto one of those addresses is refused as well.