TECHSHUbuyer room · privateManage buyer room
Longitudinal AI-crawler permission + activity ledger per monitored site, built on an unusually rigorous hand-curated crawler registry. Three facts kept deliberately separate: (1) what a site's robots.txt / ai.txt / llms.txt actually declares for each named crawler (allowed / blocked, with a 'Change since last check' diff column and check timestamps); (2) which crawlers genuinely fetched the site, classified against a ~19-entry registry that maps each user-agent to its operator AND its purpose (GPTBot/ClaudeBot/Amazonbot/CCBot/Bytespider/Meta-ExternalAgent = training; OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Perplexity-User, Amzn-User, MistralAI-User, DuckAssistBot, Meta-WebIndexer = answering a live query; Google-Extended and Applebot-Extended = permissions that never fetch; OAI-AdsBot and Google-CloudVertexBot = explicitly not AI crawlers); (3) recommended policy state, which surfaces conflicts between robots.txt and AI-specific directives. The registry itself — with operator statements about training use, including Meta's stated willingness to bypass robots.txt — is a maintained reference asset.
Get notified as soon as Techshu publishes. Until then, discover alternatives today with a 7-day trial.