Free tool

AI visibility checker

Your site can be perfectly open in robots.txt and still be invisible to AI. A firewall never reads robots.txt, it just answers 403. This check asks both questions separately, plus three more that decide whether a model can use your page at all.

One domain. We fetch your homepage once per crawler, so a run takes a few seconds. Results are cached for six hours; after a change to your site you can check again.

Why robots.txt is only half the answer

robots.txt is a policy. It tells well-behaved crawlers what they may fetch. A firewall, a security plugin or a bot-management layer sits in front of that and decides whether the request is answered at all. Those two can disagree, and when they do, the site looks open and behaves closed.

The most common cause is not a decision anyone made. Cloudflare and several managed hosts block AI crawlers by default, so a site can start refusing them after a platform change nobody on the team noticed. This check fetches your homepage once as each crawler and once as an ordinary browser, and compares the answers.

Why we never report a block on weak evidence

We ran this over 200 domains in September 2026 and first reported that a large share blocked AI crawlers. A good part of that was wrong: the crawl was fast enough to trigger rate limiting, and a 429 is not a block. The fixed rule is the one this tool uses: a browser control request runs alongside every check, and a crawler counts as blocked only when the control gets through and the crawler does not. A 429, a timeout or a site that is down for everyone all come back as unclear.

Common Crawl is the part you cannot buy

Common Crawl is an open crawl of the web, funded in part by the same companies that train large language models, and it is a major source of their training data. A page that is not in the index is unlikely to be in the next model. It is also the only AI-visibility measurement available without a paid tool, which is why it sits in this check.

Being absent is not fatal and being present is not a guarantee. Treat it as one signal: if your site is missing from the index and your server also turns CCBot away, those two facts are related and both are fixable.

Text that only exists after JavaScript runs

AI crawlers generally fetch HTML and do not run JavaScript. If your page builds its text in the browser, they receive an empty shell. We count the words in the raw HTML your server sends. A low number on a page that looks full in a browser is the signature of a client-rendered site, and it is the single most expensive AI-visibility problem to leave in place.

What to do with the result

If crawlers are blocked and you want them in: the setting is in your CDN or firewall, not in robots.txt. On Cloudflare it is the AI bot control under Security. If you want them out, do it deliberately in robots.txt rather than by accident at the firewall, so you can tell the difference later.

llms.txt is a young convention: a plain text file at the root that points language models to your important pages. Nothing consumes it reliably yet. It costs almost nothing to add and it does not replace any of the four checks above.

FAQ

Does blocking AI crawlers hurt my rankings in Google?

Not directly. Google-Extended controls training use and is separate from Googlebot, which handles search. Blocking GPTBot or ClaudeBot has no effect on Google Search at all. It affects whether ChatGPT, Claude and similar systems can read and cite you.

My robots.txt allows everything but the tool says blocked. What now?

That is the case this tool exists for. Something in front of your site is refusing the request: a CDN bot rule, a WAF, a security plugin or the host. Look there, not in robots.txt. One exception: some protection checks bots by their IP address. It refuses our request but lets the real crawler in - your firewall log shows which case it is.

Why does one crawler get through and another not?

Bot-management rules usually work from a list of known AI agents, and those lists differ between vendors and are updated at different times. A newer or smaller crawler is often simply not on the list yet.

What does "unclear" mean?

We could not prove a block. Either the crawler got a rate-limit answer, the request timed out, or the control request also failed. We would rather show unclear than report a block that is not there.

Should every site be open to AI crawlers?

No. A publisher who sells access to their archive has a good reason to close it. The point of this check is that the decision should be yours and visible, not an accident of a default setting.

Domain authority checker →

Invisible to AI and not sure why?

We audit the whole chain: CDN rules, robots.txt, rendering, structured data and whether your topic is even present in the corpora. Billed by the hour, never by the promise.