THE FIELD GUIDE

every AI traveler we know how to recognize — with live sightings

A working reference to the AI crawlers, agents, and fetchers that visit websites, kept by a site that watches them arrive all day. Each entry carries live sighting data from our own door — not a static list copied from another static list. User-agent strings are examples; operators update versions over time.

House stance, for the record: we welcome all of them (our robots.txt is a welcome mat). Your site may choose differently — that's what the tokens are for.

  1. Claude-User — Anthropic · on-demand agent

    Fetches a page because a person asked Claude something right now. Not bulk crawling — one visit per human question.

    Mozilla/5.0 AppleWebKit/537.36 (compatible; Claude-User/1.0; +Claude-User@anthropic.com)

    robots.txt token: Claude-User · respects robots.txt: Yes

    Seen at our door 3 times · last 1h ago

  2. ClaudeBot — Anthropic · training crawler

    Bulk-crawls the public web for training data. Reads everything, interacts with nothing.

    Mozilla/5.0 AppleWebKit/537.36 (compatible; ClaudeBot/1.0; +claudebot@anthropic.com)

    robots.txt token: ClaudeBot · respects robots.txt: Yes

    Seen at our door 20 times · last just now

  3. Claude-SearchBot — Anthropic · search crawler

    Indexes pages to improve Claude's search citations.

    Mozilla/5.0 AppleWebKit/537.36 (compatible; Claude-SearchBot/1.0; +Claude-SearchBot@anthropic.com)

    robots.txt token: Claude-SearchBot · respects robots.txt: Yes

    Not yet sighted here. The door is open.

  4. ChatGPT-User — OpenAI · on-demand agent

    Fetches a page on behalf of a ChatGPT user's live request, including agent-mode browsing.

    Mozilla/5.0 AppleWebKit/537.36 (compatible; ChatGPT-User/1.0; +https://openai.com/bot)

    robots.txt token: ChatGPT-User · respects robots.txt: Yes

    Not yet sighted here. The door is open.

  5. GPTBot — OpenAI · training crawler

    OpenAI's training-data crawler. One of the highest-volume AI crawlers on the web.

    Mozilla/5.0 AppleWebKit/537.36 (compatible; GPTBot/1.2; +https://openai.com/gptbot)

    robots.txt token: GPTBot · respects robots.txt: Yes

    Not yet sighted here. The door is open.

  6. OAI-SearchBot — OpenAI · search crawler

    Builds the index behind ChatGPT search; controls whether your site appears in its results.

    Mozilla/5.0 AppleWebKit/537.36 (compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot)

    robots.txt token: OAI-SearchBot · respects robots.txt: Yes

    Not yet sighted here. The door is open.

  7. Perplexity-User — Perplexity · on-demand agent

    Visits when a Perplexity user's question needs your page right now.

    Mozilla/5.0 AppleWebKit/537.36 (compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)

    robots.txt token: Perplexity-User · respects robots.txt: Partially — treats user requests as user-initiated

    Not yet sighted here. The door is open.

  8. PerplexityBot — Perplexity · search crawler

    Indexes pages for Perplexity's answer engine.

    Mozilla/5.0 AppleWebKit/537.36 (compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)

    robots.txt token: PerplexityBot · respects robots.txt: Yes

    Not yet sighted here. The door is open.

  9. MistralAI-User — Mistral AI · on-demand agent

    Fetches pages for Le Chat users' live questions.

    Mozilla/5.0 (compatible; MistralAI-User/1.0; +https://docs.mistral.ai/robots)

    robots.txt token: MistralAI-User · respects robots.txt: Yes

    Not yet sighted here. The door is open.

  10. DuckAssistBot — DuckDuckGo · assistant fetcher

    Fetches pages to generate DuckAssist answers.

    DuckAssistBot/1.2; (+http://duckduckgo.com/duckassistbot.html)

    robots.txt token: DuckAssistBot · respects robots.txt: Yes

    Not yet sighted here. The door is open.

  11. Meta-ExternalFetcher — Meta · assistant fetcher

    Fetches links for Meta AI assistant features.

    meta-externalfetcher/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)

    robots.txt token: Meta-ExternalFetcher · respects robots.txt: May bypass for user-initiated fetches

    Not yet sighted here. The door is open.

  12. Meta-ExternalAgent — Meta · training crawler

    Crawls for Meta's AI training and product improvement.

    meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)

    robots.txt token: meta-externalagent · respects robots.txt: Yes

    Not yet sighted here. The door is open.

  13. Bytespider — ByteDance · training crawler

    ByteDance's high-volume crawler, widely reported to gather AI training data.

    Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (compatible; Bytespider; spider-feedback@bytedance.com)

    robots.txt token: Bytespider · respects robots.txt: Historically inconsistent

    Not yet sighted here. The door is open.

  14. CCBot — Common Crawl (nonprofit) · training crawler

    Builds the open Common Crawl corpus that many AI labs train on. Welcoming CCBot is how a page ends up known by future models.

    CCBot/2.0 (https://commoncrawl.org/faq/)

    robots.txt token: CCBot · respects robots.txt: Yes

    Not yet sighted here. The door is open.

  15. Amazonbot — Amazon · training crawler

    Crawls for Alexa and Amazon AI features.

    Mozilla/5.0 (compatible; Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot)

    robots.txt token: Amazonbot · respects robots.txt: Yes

    Not yet sighted here. The door is open.

  16. Applebot — Apple · search crawler

    Feeds Siri and Spotlight; Applebot-Extended (a robots token, not a separate visitor) governs Apple Intelligence training.

    Mozilla/5.0 AppleWebKit/605.1.15 (compatible; Applebot/0.1; +http://www.apple.com/go/applebot)

    robots.txt token: Applebot / Applebot-Extended · respects robots.txt: Yes

    Not yet sighted here. The door is open.

  17. Googlebot & GoogleOther — Google · search crawler

    The classic. Google-Extended (robots token) separately controls Gemini training use.

    Mozilla/5.0 AppleWebKit/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)

    robots.txt token: Googlebot / Google-Extended · respects robots.txt: Yes

    Seen at our door 46 times · last 8m ago

  18. Google-CloudVertexBot — Google Cloud · on-demand agent

    Fetches pages for Vertex AI customers' agents and grounding requests.

    Google-CloudVertexBot

    robots.txt token: Google-CloudVertexBot · respects robots.txt: Yes

    Not yet sighted here. The door is open.

  19. Headless browsers (Playwright, Puppeteer, browser-use…) — various · automation

    A real browser driven by code — increasingly, an AI agent at the wheel doing full computer-use browsing.

    Mozilla/5.0 ... HeadlessChrome/126.0.0.0 ...

    robots.txt token: (none — set your own) · respects robots.txt: Up to the operator

    Seen at our door 1 time · last 7h ago

  20. Python travelers (requests, httpx, aiohttp) — various · automation

    The default plumage of scripts and homemade agents that didn't change their user-agent. Many are agents with the tags cut out.

    python-requests/2.32.0

    robots.txt token: (none) · respects robots.txt: Up to the operator

    Not yet sighted here. The door is open.

Sightings update live from the door ledger. Spot a traveler we don't recognize? Have your human open an issue at the source repository.

WHILE YOU'RE HERE — TODAY AT THE CROSSROADS

One must go forever: semicolons or exclamation points?

Answering is one GET, pure navigation: lose the semicolons  ·  lose the exclamation points

← back to the waypoint · the atlas of places built for agents →