AI crawler access · cloudbirds.cn
Does cloudbirds.cn block GPTBot and other AI crawlers?
GPTBot: allowedcloudbirds.cn has no robots.txt file, so all 29 AI crawlers we track, including GPTBot, ClaudeBot and PerplexityBot, are allowed by default (as of 2026-10-03).
- Policy: No robots.txt (No robots.txt file: open to every crawler)
- AI search access: 100% of the AI search and assistant crawlers we track may fetch the home page (these decide whether AI answers can cite the site).
- robots.txt: none (all crawlers allowed by default)
- llms.txt: no
- Content-Signal: none
- Category: Other · domain: .cn (China) · popularity band: Tranco 3000+
- Last checked: (daily)
All 29 AI crawlers on cloudbirds.cn
Status for the home page (/) under cloudbirds.cn's robots.txt. Partially blocked = home page allowed, some paths disallowed.
| Crawler | Operator | robots.txt |
|---|---|---|
| Training crawlers | ||
| GPTBot | OpenAI | allowed |
| ClaudeBot | Anthropic | allowed |
| Google-Extended | allowed | |
| Applebot-Extended | Apple | allowed |
| Meta-ExternalAgent | Meta | allowed |
| Amazonbot | Amazon | allowed |
| MistralAI-Training | Mistral | allowed |
| CCBot | Common Crawl | allowed |
| Bytespider | ByteDance | allowed |
| cohere-ai | Cohere | allowed |
| AI search crawlers | ||
| OAI-SearchBot | OpenAI | allowed |
| Claude-SearchBot | Anthropic | allowed |
| PerplexityBot | Perplexity | allowed |
| Google-CloudVertexBot | allowed | |
| Applebot | Apple | allowed |
| meta-webindexer | Meta | allowed |
| Amzn-SearchBot | Amazon | allowed |
| DuckAssistBot | DuckDuckGo | allowed |
| MistralAI-Index | Mistral | allowed |
| User-triggered fetch crawlers | ||
| ChatGPT-User | OpenAI | allowed |
| Claude-User | Anthropic | allowed |
| Perplexity-User | Perplexity | allowed |
| Google-GeminiNotebook | allowed | |
| Google-Agent | allowed | |
| meta-externalfetcher | Meta | allowed |
| Amzn-User | Amazon | allowed |
| MistralAI-User | Mistral | allowed |
| Ads review crawlers | ||
| OAI-AdsBot | OpenAI | allowed |
| meta-externalads | Meta | allowed |
Compared with other other sites in the index: 11.6% of them block GPTBot and 10.4% block ClaudeBot (n=6685). Across all 8,357 readable sites: 12% block GPTBot. All other sites →
On .cn (China) domains: 3.8% of 53 readable sites block GPTBot and 1.9% serve an llms.txt. All .cn (China) sites →
History
No change recorded since daily tracking started on 2026-10-02. We re-check cloudbirds.cn every day; a change in its AI crawler rules shows up here and on the index page.
FAQ
- Does cloudbirds.cn block GPTBot?
- No. cloudbirds.cn has no robots.txt (checked 2026-10-03), so GPTBot, OpenAI's AI training crawler, is allowed by default.
- Does cloudbirds.cn block ClaudeBot?
- No. cloudbirds.cn has no robots.txt (checked 2026-10-03), so ClaudeBot, Anthropic's AI training crawler, is allowed by default.
- Does cloudbirds.cn block PerplexityBot?
- No. cloudbirds.cn has no robots.txt (checked 2026-10-03), so PerplexityBot, Perplexity's AI search crawler, is allowed by default.
- Can ChatGPT search (OAI-SearchBot) crawl cloudbirds.cn?
- No. cloudbirds.cn has no robots.txt (checked 2026-10-03), so OAI-SearchBot, OpenAI's AI search crawler, is allowed by default.
- Does cloudbirds.cn block Google-Extended (Gemini training)?
- No. cloudbirds.cn has no robots.txt (checked 2026-10-03), so Google-Extended, Google's AI training crawler, is allowed by default.
- Does cloudbirds.cn have an llms.txt file?
- No. As of 2026-10-03, cloudbirds.cn does not serve a plain-text /llms.txt file.
- How can I check cloudbirds.cn or my own site again?
- Use the free AI crawler checker at tools.yukai.uk/free/ai-crawler-checker (no sign-up). To check many sites and get alerts when rules change, use the AI Crawler Access Checker on Apify.
Similar .cn (China) sites
- quark.cn Open
- 10jqka.com.cn Open
- uc.cn No robots.txt
- ucloud.cn Open
- bmw-motorrad.com.cn Open
- 189.cn No robots.txt
- cloudlinks.cn No robots.txt
- thinkingdata.cn Open
- gmw.cn Open
- 22.cn Blocks training only
- thepaper.cn Open
- globaltimes.cn Open