Free tool
AI visibility checker
Your site can be perfectly open in robots.txt and still be invisible to AI. A firewall never reads robots.txt, it just answers 403. This check asks both questions separately, plus three more that decide whether a model can use your page at all.
Access and policy: live requests with the crawlers’ own user agents plus a browser control request. We send only the crawler’s user agent, not its IP address: protection that verifies bots by IP can turn our request away and still let the real crawler in. Common Crawl: the latest public index. Content: the HTML your server sends, without running JavaScript.
Why robots.txt is only half the answer
robots.txt is a policy. It tells well-behaved crawlers what they may fetch. A firewall, a security plugin or a bot-management layer sits in front of that and decides whether the request is answered at all. Those two can disagree, and when they do, the site looks open and behaves closed.
The most common cause is not a decision anyone made. Cloudflare and several managed hosts block AI crawlers by default, so a site can start refusing them after a platform change nobody on the team noticed. This check fetches your homepage once as each crawler and once as an ordinary browser, and compares the answers.
Why we never report a block on weak evidence
We ran this over 200 domains in September 2026 and first reported that a large share blocked AI crawlers. A good part of that was wrong: the crawl was fast enough to trigger rate limiting, and a 429 is not a block. The fixed rule is the one this tool uses: a browser control request runs alongside every check, and a crawler counts as blocked only when the control gets through and the crawler does not. A 429, a timeout or a site that is down for everyone all come back as unclear.
Common Crawl is the part you cannot buy
Common Crawl is an open crawl of the web, funded in part by the same companies that train large language models, and it is a major source of their training data. A page that is not in the index is unlikely to be in the next model. It is also the only AI-visibility measurement available without a paid tool, which is why it sits in this check.
Being absent is not fatal and being present is not a guarantee. Treat it as one signal: if your site is missing from the index and your server also turns CCBot away, those two facts are related and both are fixable.
Text that only exists after JavaScript runs
AI crawlers generally fetch HTML and do not run JavaScript. If your page builds its text in the browser, they receive an empty shell. We count the words in the raw HTML your server sends. A low number on a page that looks full in a browser is the signature of a client-rendered site, and it is the single most expensive AI-visibility problem to leave in place.
What to do with the result
If crawlers are blocked and you want them in: the setting is in your CDN or firewall, not in robots.txt. On Cloudflare it is the AI bot control under Security. If you want them out, do it deliberately in robots.txt rather than by accident at the firewall, so you can tell the difference later.
llms.txt is a young convention: a plain text file at the root that points language models to your important pages. Nothing consumes it reliably yet. It costs almost nothing to add and it does not replace any of the four checks above.
FAQ
Does blocking AI crawlers hurt my rankings in Google?
Not directly. Google-Extended controls training use and is separate from Googlebot, which handles search. Blocking GPTBot or ClaudeBot has no effect on Google Search at all. It affects whether ChatGPT, Claude and similar systems can read and cite you.
My robots.txt allows everything but the tool says blocked. What now?
That is the case this tool exists for. Something in front of your site is refusing the request: a CDN bot rule, a WAF, a security plugin or the host. Look there, not in robots.txt. One exception: some protection checks bots by their IP address. It refuses our request but lets the real crawler in - your firewall log shows which case it is.
Why does one crawler get through and another not?
Bot-management rules usually work from a list of known AI agents, and those lists differ between vendors and are updated at different times. A newer or smaller crawler is often simply not on the list yet.
What does "unclear" mean?
We could not prove a block. Either the crawler got a rate-limit answer, the request timed out, or the control request also failed. We would rather show unclear than report a block that is not there.
Should every site be open to AI crawlers?
No. A publisher who sells access to their archive has a good reason to close it. The point of this check is that the decision should be yours and visible, not an accident of a default setting.
Invisible to AI and not sure why?
We audit the whole chain: CDN rules, robots.txt, rendering, structured data and whether your topic is even present in the corpora. Billed by the hour, never by the promise.