How to Check If Bots Can Crawl Your Website

A bot crawl check compares robots rules, HTTP responses, indexing directives, and challenge pages. A request with a crawler User-Agent can reveal header-based differences. It does not impersonate a verified crawler IP or prove access from the real service. Confirm important findings with server logs and the search engine’s inspection tools.

A rule intended for abusive traffic can also affect a crawler. Separate access failures from indexing exclusions before changing the configuration. A failed test from one address does not prove that all organic traffic or AI citations disappear.

Why Crawler Purpose Matters

A crawler needs access to retrieve current page content. However, a blocked URL can still appear in search results from links elsewhere. Crawl access, indexing eligibility, and ranking are separate checks. Google explains the limits of robots.txt.

OpenAI documents separate purposes for OAI-SearchBot, GPTBot, and ChatGPT-User. Search access is not the same as training access. Blocking GPTBot does not by itself block ChatGPT search. Review the OpenAI crawler reference before changing a policy.

Social preview fetchers also need the intended page and metadata. Test those separately from search and training crawlers.

The 5 Blocking Layers You Need to Check

When diagnosing bot access issues, test each of these 5 independent layers. The checks have different effects. An HTTP denial blocks retrieval; a readable noindex directive addresses indexing.

Layer 1: robots.txt Rules

The robots.txt file is the first thing a crawler checks before accessing any URL. Fetch yours and look for bot-specific rules:

curl https://your-site.com/robots.txt

A common mistake: allowing all bots with User-agent: * / Allow: / while simultaneously blocking specific bots with their own directives:

1
2
3
4
5
6
7
8
User-agent: *
Allow: /

User-agent: GPTBot
Disallow: /

User-agent: CCBot
Disallow: /

This example blocks the named crawlers. It does not express a complete policy for every AI search or user-requested fetcher. Review each operator’s documented agents.

Layer 2: HTTP Response Codes

Send requests with the User-Agent values you want to compare. These requests still use your own IP address. They reveal behavior for those test requests, not the real crawler:

1
2
3
4
5
6
7
8
9
10
11
# Test as Googlebot
curl -A "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" \
 -I https://your-site.com

# Test as BingBot
curl -A "Mozilla/5.0 (compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm)" \
 -I https://your-site.com

# Test as GPTBot
curl -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.0; +https://openai.com/gptbot)" \
 -I https://your-site.com
Status CodeMeaningImpact
200OKInspect the body; it can still be an error or challenge page
301/302RedirectFollow the destination and check for loops or changed access
403ForbiddenThe test request was denied; inspect the reason
405Method Not AllowedServer rejects the request method
429Too Many RequestsBot is being rate-limited
503Service UnavailableMay indicate a Cloudflare JS challenge

Layer 3: X-Robots-Tag HTTP Headers

Some servers send indexing directives via HTTP response headers instead of HTML meta tags. Check for noindex and nofollow in the X-Robots-Tag header:

curl -I https://your-site.com | grep -i "x-robots-tag"

Example header that blocks indexing: X-Robots-Tag: noindex, nofollow

This is particularly dangerous because it's invisible in the HTML source. You'd only find it by inspecting HTTP response headers.

Layer 4: Meta Robots Tags

Check the HTML for <meta name="robots" content="..."> tags. A noindex directive here prevents the page from appearing in search results even if the bot can crawl it:

curl -s https://your-site.com | grep -i "noindex"

Common mistake: setting noindex during development and forgetting to remove it before going to production.

Layer 5: WAF Challenge Detection

Cloudflare and other WAFs can serve JavaScript challenges, CAPTCHAs, and interstitial pages that bots cannot solve. Look for these patterns in the response body:

curl -s -A "Mozilla/5.0 (compatible; Googlebot/2.1)" https://your-site.com | grep -i "challenge\|captcha\|just a moment"

Cloudflare challenges include challenge-platform markers and "Just a moment" challenge pages.

Step-by-Step: Running a Full Bot Crawl Audit

Step 1: Prepare Your Bot User-Agent List

Here are the most important User-Agent strings to test:

1
2
3
4
5
6
7
8
9
10
11
12
13
# Search Engines
GOOGLEBOT="Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
BINGBOT="Mozilla/5.0 (compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm)"
YANDEXBOT="Mozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots)"

# AI Crawlers
GPTBOT="Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.0; +https://openai.com/gptbot)"
CLAUDEBOT="ClaudeBot/1.0; +https://www.anthropic.com/claude-bot"
PERPLEXITYBOT="PerplexityBot/1.0"

# Social Media
TWITTERBOT="Twitterbot/1.0"
FACEBOOKBOT="facebookexternalhit/1.1"

Step 2: Test Each Bot Against Your URL

A HEAD request gives an initial header check. Also send a GET request because a server can treat the methods differently:

1
2
3
4
5
6
URL="https://your-site.com"
for BOT in "$GOOGLEBOT" "$BINGBOT" "$GPTBOT" "$CLAUDEBOT" "$TWITTERBOT"; do
 echo "--- $BOT ---"
 curl -s -o /dev/null -w "%{http_code}" -A "$BOT" -I "$URL"
 echo ""
done

Step 3: Check for X-Robots-Tag and Meta noindex

For any bot that returns HTTP 200, verify the response doesn't contain noindex directives:

1
2
3
4
5
# Check headers
curl -s -A "$GOOGLEBOT" -I "$URL" | grep -i "x-robots"

# Check HTML meta tags
curl -s -A "$GOOGLEBOT" "$URL" | grep -i "noindex"

Step 4: Verify robots.txt

curl -s "$URL/robots.txt"

Look for Disallow: / under any bot-specific User-agent sections.

Step 5: Review Search Console and Webmaster Tools

  • Google Search Console: Use URL Inspection and the page indexing report for the exact URL
  • Bing Webmaster Tools: URL Inspection for "Discovered but not crawled" or "Blocked"

All Important Bots to Test

Search Engine Crawlers

BotOperatorPurpose
GooglebotGoogleMain Google Search indexing
BingBotMicrosoftBing Search + Copilot indexing
YandexBotYandexRussian search engine
BaiduspiderBaiduChinese search engine
DuckDuckBotDuckDuckGoPrivacy-focused search
ApplebotAppleSiri and Spotlight suggestions
SeznamBotSeznamCzech search engine
YetiNaverKorean search engine

AI Crawlers

BotOperatorPurpose
OAI-SearchBotOpenAISearch indexing for ChatGPT search
GPTBotOpenAIContent that may be used for model training
ChatGPT-UserOpenAIReal-time web browsing in ChatGPT
ClaudeBotAnthropicClaude AI content access
PerplexityBotPerplexityAI-powered search answers
Google-ExtendedGoogleRobots product token for certain Gemini uses; not a separate HTTP crawler
CCBotCommon CrawlOpen web crawl dataset for AI
cohere-aiCohereEnterprise AI models

Social Media Bots

BotOperatorPurpose
TwitterbotX (Twitter)Link preview cards
FacebookbotMetaLink preview cards
LinkedInBotLinkedInLink preview cards

How to Fix Common Bot Blocking Issues

Fix: Blocked by robots.txt

Edit your robots.txt file to add explicit Allow rules for the blocked bot:

1
2
3
4
5
User-agent: GPTBot
Allow: /

User-agent: ClaudeBot
Allow: /

Use the FindUtils Robots.txt Generator to build a properly formatted robots.txt file.

Fix: HTTP Access or Challenge Failure

Inspect the security event for the exact test request. Check the applied product and rule before creating an exception. Some bot products cannot be bypassed through a general WAF skip rule.

Do not allow traffic solely because its User-Agent says Googlebot. Verify the source using the operator's documented method. Google documents IP and DNS verification.

Limit any change to the intended crawler and public routes. Retest ordinary visitors, denied traffic, and the crawler path. Do not lower site-wide protections just to make one spoofed request return 200.

Fix: Meta noindex Tag

Search your HTML templates for <meta name="robots" content="noindex"> and remove or conditionally render it. This tag is sometimes added during development and forgotten in production.

Fix: X-Robots-Tag Header

Check your server configuration (nginx, Apache) or CDN settings for X-Robots-Tag headers. Use the FindUtils Security Headers Analyzer to inspect all HTTP response headers, or run curl -I against your URLs.

Manual Testing vs Automated Tools

FeatureManual curl TestingAutomated Crawl Checkers
PriceFreeFree to paid
Bots testedOne at a timeMultiple simultaneously
robots.txt parsingManual interpretationAutomatic per-bot analysis
WAF detectionMust read response bodyAutomatic challenge detection
Meta robots checkMust grep HTML sourceAutomatic HTML parsing
Report formatRaw terminal outputStructured report
TimingDepends on the URLs and checksDepends on the tool and URLs
Technical skillRequires CLI knowledgeVaries by tool

Real-World Scenarios

Scenario 1: Cloudflare WAF Blocking All Crawlers

A developer launches a new site on Cloudflare and enables a custom WAF rule that challenges all non-browser traffic. Three months later, they notice zero organic search traffic. Testing with curl -A "Googlebot" reveals all bots receive 403 or Cloudflare challenges. Fix: add a WAF exception for verified bots using (cf.client.bot).

Scenario 2: Confusing Search and Training Access

A publisher blocks GPTBot for training preferences. That rule does not automatically block OAI-SearchBot or another provider's search crawler. Review the actual User-agent groups before attributing a search-access problem to a training policy.

Scenario 3: Meta noindex Left from Development

A staging site configuration includes <meta name="robots" content="noindex, nofollow"> which gets deployed to production. Google Search Console shows "noindex" warnings, but the developer doesn't check for 6 months. A quick curl -s | grep noindex catches this instantly.

FAQ

Q1: What is a bot crawl checker? A: It evaluates selected rules and responses using test requests. A User-Agent simulation is a diagnostic, not proof of verified crawler access. Confirm critical results with logs and the operator’s inspection tools.

Q2: Why is my site blocked by Googlebot? A: Common causes include a restrictive robots.txt file, Cloudflare Bot Fight Mode or aggressive WAF rules, a meta robots tag set to "noindex", X-Robots-Tag HTTP headers, or the server returning 403/429/503 errors to bot user agents.

Q3: Should I block AI crawlers like GPTBot? A: Choose a policy by crawler purpose. Training, search indexing, and user-requested retrieval are distinct. Blocking one named crawler does not automatically control the others.

Q4: How often should I check bot accessibility? A: Check after any infrastructure change: updating CDN settings, modifying robots.txt, deploying new server configurations, or changing WAF rules. Also check quarterly as a routine audit. Bot blocking can happen silently without any visible symptoms.

Q5: Can Cloudflare Bot Fight Mode block legitimate crawlers? A: Check the actual security event and current product rules. Do not assume a general WAF skip rule bypasses every bot product. Use verified crawler identity rather than trusting the User-Agent alone.

Next Steps