TidyToolsAI Crawler Index · data as of · updated daily

AI crawler access · theguardian.com

Does theguardian.com block GPTBot and other AI crawlers?

GPTBot: allowedAs of 2026-10-01, the robots.txt of theguardian.com blocks 12 of 29 AI crawlers on the home page: GPTBot is allowed, ClaudeBot is blocked and PerplexityBot is blocked.

All 29 AI crawlers on theguardian.com

Status for the home page (/) under theguardian.com's robots.txt. Partially blocked = home page allowed, some paths disallowed.

CrawlerOperatorrobots.txt
Training crawlers
GPTBotOpenAIallowed
ClaudeBotAnthropicblocked
Google-ExtendedGoogleallowed
Applebot-ExtendedAppleblocked
Meta-ExternalAgentMetablocked
AmazonbotAmazonblocked
MistralAI-TrainingMistralallowed
CCBotCommon Crawlblocked
BytespiderByteDanceblocked
cohere-aiCohereallowed
AI search crawlers
OAI-SearchBotOpenAIallowed
Claude-SearchBotAnthropicblocked
PerplexityBotPerplexityblocked
Google-CloudVertexBotGoogleblocked
ApplebotAppleallowed
meta-webindexerMetaallowed
Amzn-SearchBotAmazonblocked
DuckAssistBotDuckDuckGoblocked
MistralAI-IndexMistralallowed
User-triggered fetch crawlers
ChatGPT-UserOpenAIallowed
Claude-UserAnthropicblocked
Perplexity-UserPerplexityallowed
Google-GeminiNotebookGoogleallowed
Google-AgentGoogleallowed
meta-externalfetcherMetaallowed
Amzn-UserAmazonallowed
MistralAI-UserMistralallowed
Ads review crawlers
OAI-AdsBotOpenAIallowed
meta-externaladsMetaallowed

Compared with other news & media sites in the index: 43.7% of them block GPTBot and 49.2% block ClaudeBot (n=126). Across all 853 readable sites: 17% block GPTBot. All news & media sites →

History

No change in the latest daily run. We re-check theguardian.com every day; changes show up here and on the index page.

FAQ

Does theguardian.com block GPTBot?
No. As of 2026-10-01, https://theguardian.com/robots.txt allows GPTBot, OpenAI's AI training crawler, on the home page.
Does theguardian.com block ClaudeBot?
Yes. As of 2026-10-01, https://theguardian.com/robots.txt disallows the home page for ClaudeBot, Anthropic's AI training crawler.
Does theguardian.com block PerplexityBot?
Yes. As of 2026-10-01, https://theguardian.com/robots.txt disallows the home page for PerplexityBot, Perplexity's AI search crawler.
Can ChatGPT search (OAI-SearchBot) crawl theguardian.com?
No. As of 2026-10-01, https://theguardian.com/robots.txt allows OAI-SearchBot, OpenAI's AI search crawler, on the home page.
Does theguardian.com block Google-Extended (Gemini training)?
No. As of 2026-10-01, https://theguardian.com/robots.txt allows Google-Extended, Google's AI training crawler, on the home page.
Does theguardian.com have an llms.txt file?
No. As of 2026-10-01, theguardian.com does not serve a plain-text /llms.txt file.
How can I check theguardian.com or my own site again?
Use the free AI crawler checker at tools.yukai.uk/free/ai-crawler-checker (no sign-up). To check many sites and get alerts when rules change, use the AI Crawler Access Checker on Apify.

Similar news & media sites

All 1,005 sites · All 29 AI crawlers · AI Crawler Index