AI Crawler Directory · Meta

meta-externalagent:
what it is and how to verify it.

OperatorMeta
PurposeAI model training
Respects robots.txtYes
Executes JavaScriptNo
Hits verifiableNo — nothing published
TypeHTTP crawler / fetcher

meta-externalagent is Meta’s AI crawler. Meta documents it as crawling for "training AI models or improving products by indexing content directly" — the training-collection role in Meta’s stack, feeding the Llama model family and Meta AI features across Facebook, Instagram, and WhatsApp.

It respects robots.txt, and because Llama models are open-weight and widely deployed, the reach calculus resembles Common Crawl more than a single assistant: content Meta trains on can surface in thousands of downstream Llama-based products, not just Meta AI.

Unlike OpenAI, Anthropic, and Perplexity, Meta publishes no per-bot IP list for verification, so a UA match in your logs is only a claim. Sites doing serious crawler analytics label meta-externalagent hits as unverified rather than trusted.

meta-externalagent User-Agent string

meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/web-crawlers)

How to verify meta-externalagent hits are real

The User-Agent header is plain text — any client can claim to be meta-externalagent, and scrapers routinely do. Meta publishes no dedicated IP list for this bot — hits cannot be cryptographically verified; treat UA matches as unverified.

Because nothing is published, the honest label for every meta-externalagent log line is unverified — it may be Meta, or a scraper borrowing the name.

Allow or block meta-externalagent in robots.txt

This is a training-data control: allowing it lets your content shape what future models know; blocking it keeps your material out of training corpora without affecting live citations. A policy call, not a traffic one.

Allow

User-agent: meta-externalagent Allow: /

Block

User-agent: meta-externalagent Disallow: /

meta-externalagent — frequently asked questions

What is meta-externalagent?

meta-externalagent is Meta’s web crawler for training foundation AI models and improving products by indexing content. It identifies as meta-externalagent/1.1 in the User-Agent header.

Does meta-externalagent respect robots.txt?

Yes — Meta documents that it honors robots.txt directives. Its sibling meta-externalfetcher, which performs user-requested link fetches, may bypass robots.txt.

Can I verify meta-externalagent hits?

Not reliably. Meta does not publish a dedicated IP list for this bot, so there is no authoritative range to check against. Treat UA-only matches as unverified — a distinction that matters, since scrapers spoof well-known AI bot names.

Why would I allow Meta’s AI crawler?

Llama models are open-weight and embedded in a large ecosystem of downstream apps. Content in Meta’s training data can inform answers across that whole ecosystem — broader distribution than most single-vendor assistants offer.

See which AI crawlers hit your site — and what the traffic earns

Attrifast tracks AI referrals as revenue lines: ChatGPT, Perplexity, Claude, and Gemini visits tied to real Stripe money. $15/mo flat.

Start 7-day free trial — $0 due today

7-day free trial · $15/mo · cancel anytime