100% Free!

Free AI Crawler robots.txt Checker

Every AI provider now runs two crawlers: one that collects training data and one that fetches your pages to build the answer a person is reading right now. Blocking the wrong one removes you from AI answers entirely.

Sign up for a KeySearch account to see whether AI answers actually cite you.

Training Crawlers vs Answer Crawlers

This is the distinction that trips most sites up. Every major AI provider now operates two separate crawlers, and they do very different jobs:

  • Training crawlers collect pages to train future models. Blocking them is an editorial decision about whether your work feeds someone’s model. It costs you nothing today.
  • Answer crawlers fetch your page at the moment a person is asking a question, so the assistant can quote and link you. Blocking them removes you from those answers — permanently and invisibly.

The names are easy to confuse. GPTBot is OpenAI’s training crawler; OAI-SearchBot is the one that builds answers. ClaudeBot trains; Claude-SearchBot answers. A site that meant to opt out of training and pattern-matched its way to blocking everything with a familiar name has quietly opted out of being cited.

The Googlebot Trap

Google’s AI Overview is served by Googlebotitself. There is no separate agent to allow or block, and Google-Extended controls training and Gemini grounding, not AI Overview eligibility. If you block Googlebot you are out of AI Overviews and out of ordinary Google search at the same time.

Which Crawlers We Check

We look up the robots.txt rules that apply to the crawlers run by OpenAI, Anthropic, Google, Perplexity, Meta, Apple, Amazon, ByteDance, Cohere and Common Crawl, and tell you which of them are blocked and what blocking each one actually costs you. A rule set aimed at a specific crawler wins over the wildcard group, the same way Google evaluates it.

What a Good Configuration Looks Like

There is no single right answer, but the common intent is: allow the answer crawlers, and decide separately about training. If you want to stay out of model training while remaining citable, block GPTBot, ClaudeBot, Google-Extended and CCBot, and leave OAI-SearchBot, Claude-SearchBot, PerplexityBot and Googlebot alone.

Need to build the file itself? Our free robots.txt generator writes one for you.

Frequently Asked Questions

Does robots.txt guarantee a crawler stays out?

No. robots.txt is a request that well-behaved crawlers honour, not an access control. Treat it as a signal, not a lock.

Should I add an llms.txt file?

It is not a permission system and there is no good evidence that AI systems use it for retrieval. robots.txt is the half that actually changes what happens.

My robots.txt is clean — why am I still not cited?

Because being crawlable is necessary, not sufficient. Once you know nothing is blocking you, what gets quoted is the page itself: a direct answer near the top, sourced facts, and clear question headings.

Other Free Tools at Keysearch

It’s free to start. Sign up for a Keysearch account to track which of your pages AI answers actually cite.

Start growing your organic traffic today

Join thousands of successful websites using KeySearch for keyword research, competitor analysis, and content optimization.

KeySearch Research Page