A website can open in your browser but return an error, challenge, or indexing restriction to a crawler. Check the response and the indexing controls separately. A missing search result alone does not prove bot blocking. This guide gives a diagnostic workflow; it does not claim a measured FindUtils incident.
Five access and indexing checks
Each check answers a different question. Record the requested URL, final URL, request method, status, and response body before changing a setting.
1. Access challenges and security rules
A security rule can reject a request or return a challenge instead of the article. Review the request record to identify the rule that actually matched. Do not assume that every bot protection feature blocks verified search crawlers.
An access exception needs a narrow scope and a verified identity. A user-agent string alone is not proof that a request comes from Google or Bing. Keep authentication on private routes. Do not disable broad security controls because a crawler simulation receives a 403 response.
After a correction, test the public URL again. Confirm that the response contains the article, not merely a successful status with a challenge page.
2. robots.txt policy
A robots file tells compliant crawlers which paths they may fetch. For example:
User-agent: * Disallow: /api/ Disallow: /admin/ User-agent: GPTBot Disallow: / User-agent: CCBot Disallow: /
This example expresses different crawler policies. GPTBot relates to potential model training. It is not OpenAI's search crawler. Blocking GPTBot does not by itself block OAI-SearchBot. Check each provider's documented role before changing these rules.
Robots rules are not access control. Protect private content with authentication. A disallowed URL can still appear in search without a retrieved description. See Google's robots.txt introduction.
3. GET and HEAD responses
A method restriction can explain why a header-only diagnostic fails:
if (request.method !== 'GET') { return new Response('Method not allowed', { status: 405 }); }
This example returns 405 for HEAD. A curl -I check sends HEAD, so it does not show what a normal GET request receives. A 405 on HEAD alone does not prove that a search crawler cannot retrieve the page with GET.
Where you support HEAD, return the appropriate response headers without the body. Test both methods. Keep method restrictions on endpoints that require them.
4. A noindex meta tag
A page can return readable content while asking search engines not to index it:
<class="text-rose-400">meta name="robots" content="noindex, nofollow">
Check whether this is intentional. A staging site and a public article usually have different publication requirements. A crawler must retrieve the directive to act on it. Blocking the page in robots.txt can prevent that retrieval.
5. X-Robots-Tag headers
A response can carry X-Robots-Tag: noindex in its headers. That directive does not need to appear in the HTML source. Inspect headers as well as the body.
Use Google's robots meta and header reference for supported directives and their scope. Check the final response after redirects.
How to diagnose crawler access
Compare request methods and user-agent strings
These commands send HEAD requests with selected user-agent headers. They help identify response differences. They do not reproduce a verified crawler's IP address, request history, or full behavior.
# Header-only Googlebot simulation curl -A "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" \ -I https://your-site.com # Header-only Bingbot simulation curl -A "Mozilla/5.0 (compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm)" \ -I https://your-site.com # Header-only GPTBot simulation: training crawler, not search curl -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.0; +https://openai.com/gptbot)" \ -I https://your-site.com
Then compare a GET request. This command prints the headers and discards the body:
curl -sS -L -D - -o /dev/null https://your-site.com
Fetch the body separately when you need to check whether it contains the article:
curl -sS -L https://your-site.com
Interpret the evidence carefully:
| Result | What it establishes | Next check |
|---|---|---|
| 200 | This request succeeded at the HTTP level | Check the body and indexing directives |
| 403 | This request was refused | Identify the access rule or application cause |
| 405 on HEAD | HEAD is not accepted for this request | Compare GET and the intended method contract |
| 503 | The service could not complete this request | Check the body and service record for the cause |
| Redirect | The request points to another URL | Check the final URL and response |
noindex | The response asks supported engines not to index it | Confirm the intended publication policy |
Inspect the robots file and page directives
Fetch the robots file:
curl https://your-site.com/robots.txt
Check the matching user-agent group and path. Do not read one Disallow line without its group. More specific groups can change which rules apply.
For an initial text search of the HTML:
curl -s https://your-site.com | grep -i "noindex"
This command can also match comments or script text. Inspect the actual meta element. Review HTTP headers separately. A missing text match does not prove the final rendered page contains no directive.
Check real crawler evidence
Use the page inspection features in Google Search Console and Bing Webmaster Tools. Compare their crawl information with the request records for that URL. A “discovered” status is not itself a diagnosis of blocking.
When identity matters, follow the search provider's verification procedure. See Google's crawler verification instructions.
Choose the crawler roles you intend to permit
Search discovery
Googlebot and Bingbot support their search services. Check the relevant policy when search visibility is a publication goal. Access supports retrieval; it does not guarantee indexing or traffic.
For OpenAI, OAI-SearchBot supports search discovery. GPTBot covers potential training use. ChatGPT-User handles user-requested visits. These controls are not interchangeable. Use OpenAI's bot reference.
Perplexity also distinguishes its crawler from user-requested fetches. Check its crawler documentation before writing rules.
Training and other uses
Choose a training policy separately from a search policy where the provider exposes separate controls. Read the exact scope. Do not assume that blocking one named bot prevents every possible reuse of public content.
Social previews
Social services retrieve page metadata and images to form previews. A missing preview can result from access failures, bad metadata, an unavailable image, or a cached result. Check these causes before treating the problem as proof of a blocked bot.
Correct the specific failure
- Test one affected public URL with GET and HEAD.
- Record the final status, headers, and body.
- Review the robots policy for the intended crawler.
- Check meta and header indexing directives.
- Identify the matching access rule if a request is refused.
- Make a narrow correction that preserves private-route protection.
- Repeat the request and the provider's inspection check.
Use the Robots.txt Generator for syntax and the Bot Crawl Checker for initial comparisons. Keep the distinction between a simulation and a real crawler visit.
Repeat these checks after changes to access rules, redirects, publication settings, or response handling. Keep a dated record so a later regression has a known working reference.
Related Tools
- Robots.txt Generator — Create properly formatted robots.txt files
- Security Headers Analyzer — Inspect HTTP response headers for X-Robots-Tag
- SSL Certificate Checker — Verify HTTPS configuration
Frequently asked questions
Does a 200 response prove that my page can rank?
No. Check the body, canonical URL, indexing directives, and the search provider's index status. A successful fetch is only one requirement.
Does a bot user-agent string verify crawler identity?
No. Clients can choose that string. Use the provider's documented IP or reverse-DNS verification procedure where identity matters.
Does HEAD 405 prove that GET is blocked?
No. The methods can follow different handlers. Test GET before diagnosing a page retrieval failure.
Should I remove every crawler block?
No. Choose a policy for search, training, and private content. Correct unintended restrictions while retaining deliberate access controls.
Why does a social link have no image?
Check metadata, image accessibility, and the service's cached preview. Bot blocking is one possible cause, not the only cause.