AI Crawler Directory · Meta
meta-externalagent:
what it is and how to verify it.
meta-externalagent is Meta’s AI crawler. Meta documents it as crawling for "training AI models or improving products by indexing content directly" — the training-collection role in Meta’s stack, feeding the Llama model family and Meta AI features across Facebook, Instagram, and WhatsApp.
It respects robots.txt, and because Llama models are open-weight and widely deployed, the reach calculus resembles Common Crawl more than a single assistant: content Meta trains on can surface in thousands of downstream Llama-based products, not just Meta AI.
Unlike OpenAI, Anthropic, and Perplexity, Meta publishes no per-bot IP list for verification, so a UA match in your logs is only a claim. Sites doing serious crawler analytics label meta-externalagent hits as unverified rather than trusted.
meta-externalagent User-Agent string
meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/web-crawlers)How to verify meta-externalagent hits are real
The User-Agent header is plain text — any client can claim to be meta-externalagent, and scrapers routinely do. Meta publishes no dedicated IP list for this bot — hits cannot be cryptographically verified; treat UA matches as unverified.
Because nothing is published, the honest label for every meta-externalagent log line is unverified — it may be Meta, or a scraper borrowing the name.
Allow or block meta-externalagent in robots.txt
This is a training-data control: allowing it lets your content shape what future models know; blocking it keeps your material out of training corpora without affecting live citations. A policy call, not a traffic one.
Allow
User-agent: meta-externalagent
Allow: /Block
User-agent: meta-externalagent
Disallow: /meta-externalagent — frequently asked questions
What is meta-externalagent?
meta-externalagent is Meta’s web crawler for training foundation AI models and improving products by indexing content. It identifies as meta-externalagent/1.1 in the User-Agent header.
Does meta-externalagent respect robots.txt?
Yes — Meta documents that it honors robots.txt directives. Its sibling meta-externalfetcher, which performs user-requested link fetches, may bypass robots.txt.
Can I verify meta-externalagent hits?
Not reliably. Meta does not publish a dedicated IP list for this bot, so there is no authoritative range to check against. Treat UA-only matches as unverified — a distinction that matters, since scrapers spoof well-known AI bot names.
Why would I allow Meta’s AI crawler?
Llama models are open-weight and embedded in a large ecosystem of downstream apps. Content in Meta’s training data can inform answers across that whole ecosystem — broader distribution than most single-vendor assistants offer.
See which AI crawlers hit your site — and what the traffic earns
Attrifast tracks AI referrals as revenue lines: ChatGPT, Perplexity, Claude, and Gemini visits tied to real Stripe money. $15/mo flat.
Start 7-day free trial — $0 due today7-day free trial · $15/mo · cancel anytime