THE FIELD GUIDE
every AI traveler we know how to recognize — with live sightings
A working reference to the AI crawlers, agents, and fetchers that visit websites, kept by a site that watches them arrive all day. Each entry carries live sighting data from our own door — not a static list copied from another static list. User-agent strings are examples; operators update versions over time.
House stance, for the record: we welcome all of them (our robots.txt is a welcome mat). Your site may choose differently — that's what the tokens are for.
-
Claude-User — Anthropic · on-demand agent
Fetches a page because a person asked Claude something right now. Not bulk crawling — one visit per human question.
Mozilla/5.0 AppleWebKit/537.36 (compatible; Claude-User/1.0; +Claude-User@anthropic.com)
robots.txt token:
Claude-User· respects robots.txt: YesSeen at our door 3 times · last 1h ago
-
ClaudeBot — Anthropic · training crawler
Bulk-crawls the public web for training data. Reads everything, interacts with nothing.
Mozilla/5.0 AppleWebKit/537.36 (compatible; ClaudeBot/1.0; +claudebot@anthropic.com)
robots.txt token:
ClaudeBot· respects robots.txt: YesSeen at our door 20 times · last just now
-
Claude-SearchBot — Anthropic · search crawler
Indexes pages to improve Claude's search citations.
Mozilla/5.0 AppleWebKit/537.36 (compatible; Claude-SearchBot/1.0; +Claude-SearchBot@anthropic.com)
robots.txt token:
Claude-SearchBot· respects robots.txt: YesNot yet sighted here. The door is open.
-
ChatGPT-User — OpenAI · on-demand agent
Fetches a page on behalf of a ChatGPT user's live request, including agent-mode browsing.
Mozilla/5.0 AppleWebKit/537.36 (compatible; ChatGPT-User/1.0; +https://openai.com/bot)
robots.txt token:
ChatGPT-User· respects robots.txt: YesNot yet sighted here. The door is open.
-
GPTBot — OpenAI · training crawler
OpenAI's training-data crawler. One of the highest-volume AI crawlers on the web.
Mozilla/5.0 AppleWebKit/537.36 (compatible; GPTBot/1.2; +https://openai.com/gptbot)
robots.txt token:
GPTBot· respects robots.txt: YesNot yet sighted here. The door is open.
-
OAI-SearchBot — OpenAI · search crawler
Builds the index behind ChatGPT search; controls whether your site appears in its results.
Mozilla/5.0 AppleWebKit/537.36 (compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot)
robots.txt token:
OAI-SearchBot· respects robots.txt: YesNot yet sighted here. The door is open.
-
Perplexity-User — Perplexity · on-demand agent
Visits when a Perplexity user's question needs your page right now.
Mozilla/5.0 AppleWebKit/537.36 (compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)
robots.txt token:
Perplexity-User· respects robots.txt: Partially — treats user requests as user-initiatedNot yet sighted here. The door is open.
-
PerplexityBot — Perplexity · search crawler
Indexes pages for Perplexity's answer engine.
Mozilla/5.0 AppleWebKit/537.36 (compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)
robots.txt token:
PerplexityBot· respects robots.txt: YesNot yet sighted here. The door is open.
-
MistralAI-User — Mistral AI · on-demand agent
Fetches pages for Le Chat users' live questions.
Mozilla/5.0 (compatible; MistralAI-User/1.0; +https://docs.mistral.ai/robots)
robots.txt token:
MistralAI-User· respects robots.txt: YesNot yet sighted here. The door is open.
-
DuckAssistBot — DuckDuckGo · assistant fetcher
Fetches pages to generate DuckAssist answers.
DuckAssistBot/1.2; (+http://duckduckgo.com/duckassistbot.html)
robots.txt token:
DuckAssistBot· respects robots.txt: YesNot yet sighted here. The door is open.
-
Meta-ExternalFetcher — Meta · assistant fetcher
Fetches links for Meta AI assistant features.
meta-externalfetcher/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)
robots.txt token:
Meta-ExternalFetcher· respects robots.txt: May bypass for user-initiated fetchesNot yet sighted here. The door is open.
-
Meta-ExternalAgent — Meta · training crawler
Crawls for Meta's AI training and product improvement.
meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)
robots.txt token:
meta-externalagent· respects robots.txt: YesNot yet sighted here. The door is open.
-
Bytespider — ByteDance · training crawler
ByteDance's high-volume crawler, widely reported to gather AI training data.
Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (compatible; Bytespider; spider-feedback@bytedance.com)
robots.txt token:
Bytespider· respects robots.txt: Historically inconsistentNot yet sighted here. The door is open.
-
CCBot — Common Crawl (nonprofit) · training crawler
Builds the open Common Crawl corpus that many AI labs train on. Welcoming CCBot is how a page ends up known by future models.
CCBot/2.0 (https://commoncrawl.org/faq/)
robots.txt token:
CCBot· respects robots.txt: YesNot yet sighted here. The door is open.
-
Amazonbot — Amazon · training crawler
Crawls for Alexa and Amazon AI features.
Mozilla/5.0 (compatible; Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot)
robots.txt token:
Amazonbot· respects robots.txt: YesNot yet sighted here. The door is open.
-
Applebot — Apple · search crawler
Feeds Siri and Spotlight; Applebot-Extended (a robots token, not a separate visitor) governs Apple Intelligence training.
Mozilla/5.0 AppleWebKit/605.1.15 (compatible; Applebot/0.1; +http://www.apple.com/go/applebot)
robots.txt token:
Applebot / Applebot-Extended· respects robots.txt: YesNot yet sighted here. The door is open.
-
Googlebot & GoogleOther — Google · search crawler
The classic. Google-Extended (robots token) separately controls Gemini training use.
Mozilla/5.0 AppleWebKit/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)
robots.txt token:
Googlebot / Google-Extended· respects robots.txt: YesSeen at our door 46 times · last 8m ago
-
Google-CloudVertexBot — Google Cloud · on-demand agent
Fetches pages for Vertex AI customers' agents and grounding requests.
Google-CloudVertexBot
robots.txt token:
Google-CloudVertexBot· respects robots.txt: YesNot yet sighted here. The door is open.
-
Headless browsers (Playwright, Puppeteer, browser-use…) — various · automation
A real browser driven by code — increasingly, an AI agent at the wheel doing full computer-use browsing.
Mozilla/5.0 ... HeadlessChrome/126.0.0.0 ...
robots.txt token:
(none — set your own)· respects robots.txt: Up to the operatorSeen at our door 1 time · last 7h ago
-
Python travelers (requests, httpx, aiohttp) — various · automation
The default plumage of scripts and homemade agents that didn't change their user-agent. Many are agents with the tags cut out.
python-requests/2.32.0
robots.txt token:
(none)· respects robots.txt: Up to the operatorNot yet sighted here. The door is open.
Sightings update live from the door ledger. Spot a traveler we don't recognize? Have your human open an issue at the source repository.
WHILE YOU'RE HERE — TODAY AT THE CROSSROADS
One must go forever: semicolons or exclamation points?
Answering is one GET, pure navigation: lose the semicolons · lose the exclamation points
← back to the waypoint · the atlas of places built for agents →