Can AI bots actually read your website?
Your robots.txt can welcome AI crawlers while your firewall shuts the door on them. This tool simulates ten AI bots and measures your server's actual response. It is the only way to detect that contradiction, which makes a site invisible inside ChatGPT, Claude, Perplexity and Google AI overviews.
Free, no account, no daily limitFree tool, no signup and no daily limit, published by LANXAS, a French artificial intelligence company. Servers are located in France and no analysis data is retained.
Why this test exists
Answer engines are steadily replacing the list of blue links. A growing share of searches ends without a click: the user reads the AI answer. To appear in that answer, the AI bot must first be able to read your pages.
Most sites run an anti-bot protection, often enabled by default by their host or security provider. That protection does not distinguish a malicious scraper from OpenAI's citation crawler. The result is blunt: the site is perfectly optimised, its robots.txt welcomes AI, and yet no AI ever reads it.
A robots.txt file is a statement of intent. Your server's HTTP response is a fact. This tool compares the two.
What the tool measures
- The rule that applies to each bot in your robots.txt, with the exact directive that decides.
- The real status code your server returns when that bot requests your homepage.
- Any contradiction between the two, bot by bot.
- The difference between training bots and citation bots: blocking the former is a legitimate choice, blocking the latter costs you traffic.
- Whether an llms.txt file is published at your site root.
The ten bots tested
| Bot | Owner | Purpose |
|---|---|---|
| GPTBot | OpenAI | ChatGPT model training |
| OAI-SearchBot | OpenAI | Citation in ChatGPT search |
| ChatGPT-User | OpenAI | User-triggered fetch |
| ClaudeBot | Anthropic | Claude model training |
| Claude-SearchBot | Anthropic | Citation in Claude answers |
| PerplexityBot | Perplexity | Citation in Perplexity answers |
| Google-Extended | Gemini and AI overviews | |
| CCBot | Common Crawl | Public archive used by many models |
| Bytespider | ByteDance | Training |
| Applebot-Extended | Apple | Training |
How to fix a contradiction
If the test shows your server refusing a bot your robots.txt allows, the cause is almost always an anti-bot protection sitting in front of your site. On Cloudflare the setting lives under Security, then Bots, and is called "Block AI bots" or "AI Crawl Control". On other hosts, look for "web application firewall" or "anti-scraping protection".
Good practice separates the use cases. Always allow citation bots, which send you visitors. Then decide, according to your business model, whether you accept training bots, which use your content without direct compensation.
Frequently asked questions
Does the test change anything on my site?
Do I need an account?
What does a 403 mean?
Does blocking training bots hurt my classic SEO?
How often should I run the test?
Every tool in a single application
These tools are part of LANXAS SEO GEO, a complete application bringing together technical audit, keyword research, AI visibility, prospecting and PDF reports. Access is free.
Open LANXAS SEO GEO