How to Check If Bots Can Crawl Your Website
A bot crawl check compares robots rules, HTTP responses, indexing directives, and challenge pages. A request with a crawler User-Agent can reveal header-based differences. It does not impersonate a verified crawler IP or prove access from the real service. Confirm important findings with server logs and the search engine’s inspection tools.
A rule intended for abusive traffic can also affect a crawler. Separate access failures from indexing exclusions before changing the configuration. A failed test from one address does not prove that all organic traffic or AI citations disappear.
Why Crawler Purpose Matters
A crawler needs access to retrieve current page content. However, a blocked URL can still appear in search results from links elsewhere. Crawl access, indexing eligibility, and ranking are separate checks. Google explains the limits of robots.txt.
OpenAI documents separate purposes for OAI-SearchBot, GPTBot, and ChatGPT-User. Search access is not the same as training access. Blocking GPTBot does not by itself block ChatGPT search. Review the OpenAI crawler reference before changing a policy.
Social preview fetchers also need the intended page and metadata. Test those separately from search and training crawlers.
The 5 Blocking Layers You Need to Check
When diagnosing bot access issues, test each of these 5 independent layers. The checks have different effects. An HTTP denial blocks retrieval; a readable noindex directive addresses indexing.
Layer 1: robots.txt Rules
The robots.txt file is the first thing a crawler checks before accessing any URL. Fetch yours and look for bot-specific rules:
curl https://your-site.com/robots.txt
A common mistake: allowing all bots with User-agent: * / Allow: / while simultaneously blocking specific bots with their own directives:
User-agent: * Allow: / User-agent: GPTBot Disallow: / User-agent: CCBot Disallow: /
This example blocks the named crawlers. It does not express a complete policy for every AI search or user-requested fetcher. Review each operator’s documented agents.
Layer 2: HTTP Response Codes
Send requests with the User-Agent values you want to compare. These requests still use your own IP address. They reveal behavior for those test requests, not the real crawler:
# Test as Googlebot curl -A "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" \ -I https://your-site.com # Test as BingBot curl -A "Mozilla/5.0 (compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm)" \ -I https://your-site.com # Test as GPTBot curl -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.0; +https://openai.com/gptbot)" \ -I https://your-site.com
| Status Code | Meaning | Impact |
|---|---|---|
| 200 | OK | Inspect the body; it can still be an error or challenge page |
| 301/302 | Redirect | Follow the destination and check for loops or changed access |
| 403 | Forbidden | The test request was denied; inspect the reason |
| 405 | Method Not Allowed | Server rejects the request method |
| 429 | Too Many Requests | Bot is being rate-limited |
| 503 | Service Unavailable | May indicate a Cloudflare JS challenge |
Layer 3: X-Robots-Tag HTTP Headers
Some servers send indexing directives via HTTP response headers instead of HTML meta tags. Check for noindex and nofollow in the X-Robots-Tag header:
curl -I https://your-site.com | grep -i "x-robots-tag"
Example header that blocks indexing: X-Robots-Tag: noindex, nofollow
This is particularly dangerous because it's invisible in the HTML source. You'd only find it by inspecting HTTP response headers.
Layer 4: Meta Robots Tags
Check the HTML for <meta name="robots" content="..."> tags. A noindex directive here prevents the page from appearing in search results even if the bot can crawl it:
curl -s https://your-site.com | grep -i "noindex"
Common mistake: setting noindex during development and forgetting to remove it before going to production.
Layer 5: WAF Challenge Detection
Cloudflare and other WAFs can serve JavaScript challenges, CAPTCHAs, and interstitial pages that bots cannot solve. Look for these patterns in the response body:
curl -s -A "Mozilla/5.0 (compatible; Googlebot/2.1)" https://your-site.com | grep -i "challenge\|captcha\|just a moment"
Cloudflare challenges include challenge-platform markers and "Just a moment" challenge pages.
Step-by-Step: Running a Full Bot Crawl Audit
Step 1: Prepare Your Bot User-Agent List
Here are the most important User-Agent strings to test:
# Search Engines GOOGLEBOT="Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" BINGBOT="Mozilla/5.0 (compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm)" YANDEXBOT="Mozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots)" # AI Crawlers GPTBOT="Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.0; +https://openai.com/gptbot)" CLAUDEBOT="ClaudeBot/1.0; +https://www.anthropic.com/claude-bot" PERPLEXITYBOT="PerplexityBot/1.0" # Social Media TWITTERBOT="Twitterbot/1.0" FACEBOOKBOT="facebookexternalhit/1.1"
Step 2: Test Each Bot Against Your URL
A HEAD request gives an initial header check. Also send a GET request because a server can treat the methods differently:
URL="https://your-site.com"
for BOT in "$GOOGLEBOT" "$BINGBOT" "$GPTBOT" "$CLAUDEBOT" "$TWITTERBOT"; do
echo "--- $BOT ---"
curl -s -o /dev/null -w "%{http_code}" -A "$BOT" -I "$URL"
echo ""
doneStep 3: Check for X-Robots-Tag and Meta noindex
For any bot that returns HTTP 200, verify the response doesn't contain noindex directives:
# Check headers curl -s -A "$GOOGLEBOT" -I "$URL" | grep -i "x-robots" # Check HTML meta tags curl -s -A "$GOOGLEBOT" "$URL" | grep -i "noindex"
Step 4: Verify robots.txt
curl -s "$URL/robots.txt"
Look for Disallow: / under any bot-specific User-agent sections.
Step 5: Review Search Console and Webmaster Tools
- Google Search Console: Use URL Inspection and the page indexing report for the exact URL
- Bing Webmaster Tools: URL Inspection for "Discovered but not crawled" or "Blocked"
All Important Bots to Test
Search Engine Crawlers
| Bot | Operator | Purpose |
|---|---|---|
| Googlebot | Main Google Search indexing | |
| BingBot | Microsoft | Bing Search + Copilot indexing |
| YandexBot | Yandex | Russian search engine |
| Baiduspider | Baidu | Chinese search engine |
| DuckDuckBot | DuckDuckGo | Privacy-focused search |
| Applebot | Apple | Siri and Spotlight suggestions |
| SeznamBot | Seznam | Czech search engine |
| Yeti | Naver | Korean search engine |
AI Crawlers
| Bot | Operator | Purpose |
|---|---|---|
| OAI-SearchBot | OpenAI | Search indexing for ChatGPT search |
| GPTBot | OpenAI | Content that may be used for model training |
| ChatGPT-User | OpenAI | Real-time web browsing in ChatGPT |
| ClaudeBot | Anthropic | Claude AI content access |
| PerplexityBot | Perplexity | AI-powered search answers |
| Google-Extended | Robots product token for certain Gemini uses; not a separate HTTP crawler | |
| CCBot | Common Crawl | Open web crawl dataset for AI |
| cohere-ai | Cohere | Enterprise AI models |
Social Media Bots
| Bot | Operator | Purpose |
|---|---|---|
| Twitterbot | X (Twitter) | Link preview cards |
| Facebookbot | Meta | Link preview cards |
| LinkedInBot | Link preview cards |
How to Fix Common Bot Blocking Issues
Fix: Blocked by robots.txt
Edit your robots.txt file to add explicit Allow rules for the blocked bot:
User-agent: GPTBot Allow: / User-agent: ClaudeBot Allow: /
Use the FindUtils Robots.txt Generator to build a properly formatted robots.txt file.
Fix: HTTP Access or Challenge Failure
Inspect the security event for the exact test request. Check the applied product and rule before creating an exception. Some bot products cannot be bypassed through a general WAF skip rule.
Do not allow traffic solely because its User-Agent says Googlebot. Verify the source using the operator's documented method. Google documents IP and DNS verification.
Limit any change to the intended crawler and public routes. Retest ordinary visitors, denied traffic, and the crawler path. Do not lower site-wide protections just to make one spoofed request return 200.
Fix: Meta noindex Tag
Search your HTML templates for <meta name="robots" content="noindex"> and remove or conditionally render it. This tag is sometimes added during development and forgotten in production.
Fix: X-Robots-Tag Header
Check your server configuration (nginx, Apache) or CDN settings for X-Robots-Tag headers. Use the FindUtils Security Headers Analyzer to inspect all HTTP response headers, or run curl -I against your URLs.
Manual Testing vs Automated Tools
| Feature | Manual curl Testing | Automated Crawl Checkers |
|---|---|---|
| Price | Free | Free to paid |
| Bots tested | One at a time | Multiple simultaneously |
| robots.txt parsing | Manual interpretation | Automatic per-bot analysis |
| WAF detection | Must read response body | Automatic challenge detection |
| Meta robots check | Must grep HTML source | Automatic HTML parsing |
| Report format | Raw terminal output | Structured report |
| Timing | Depends on the URLs and checks | Depends on the tool and URLs |
| Technical skill | Requires CLI knowledge | Varies by tool |
Real-World Scenarios
Scenario 1: Cloudflare WAF Blocking All Crawlers
A developer launches a new site on Cloudflare and enables a custom WAF rule that challenges all non-browser traffic. Three months later, they notice zero organic search traffic. Testing with curl -A "Googlebot" reveals all bots receive 403 or Cloudflare challenges. Fix: add a WAF exception for verified bots using (cf.client.bot).
Scenario 2: Confusing Search and Training Access
A publisher blocks GPTBot for training preferences. That rule does not automatically block OAI-SearchBot or another provider's search crawler. Review the actual User-agent groups before attributing a search-access problem to a training policy.
Scenario 3: Meta noindex Left from Development
A staging site configuration includes <meta name="robots" content="noindex, nofollow"> which gets deployed to production. Google Search Console shows "noindex" warnings, but the developer doesn't check for 6 months. A quick curl -s | grep noindex catches this instantly.
Related Tools
- Robots.txt Generator — Create properly formatted robots.txt files
- Security Headers Analyzer — Inspect HTTP response headers including X-Robots-Tag
- DNS Security Scanner — Check DNS configuration and security records
- SSL Certificate Checker — Verify SSL/TLS certificate validity
FAQ
Q1: What is a bot crawl checker? A: It evaluates selected rules and responses using test requests. A User-Agent simulation is a diagnostic, not proof of verified crawler access. Confirm critical results with logs and the operator’s inspection tools.
Q2: Why is my site blocked by Googlebot? A: Common causes include a restrictive robots.txt file, Cloudflare Bot Fight Mode or aggressive WAF rules, a meta robots tag set to "noindex", X-Robots-Tag HTTP headers, or the server returning 403/429/503 errors to bot user agents.
Q3: Should I block AI crawlers like GPTBot? A: Choose a policy by crawler purpose. Training, search indexing, and user-requested retrieval are distinct. Blocking one named crawler does not automatically control the others.
Q4: How often should I check bot accessibility? A: Check after any infrastructure change: updating CDN settings, modifying robots.txt, deploying new server configurations, or changing WAF rules. Also check quarterly as a routine audit. Bot blocking can happen silently without any visible symptoms.
Q5: Can Cloudflare Bot Fight Mode block legitimate crawlers? A: Check the actual security event and current product rules. Do not assume a general WAF skip rule bypasses every bot product. Use verified crawler identity rather than trusting the User-Agent alone.
Next Steps
- Generate a robots.txt — Use the Robots.txt Generator to create properly formatted directives
- Audit your security headers — Run the Security Headers Analyzer to check for X-Robots-Tag and other header issues
- Check your DNS security — Use the DNS Security Scanner to verify your DNS configuration
- Verify your SSL certificate — Run the SSL Certificate Checker to ensure HTTPS is properly configured