Attrifast
ProductAI visibilityPricingDocsBlog
Log inStart free trial
ProductAI visibilityPricingDocsBlogLog in
  1. Home
  2. /AI Crawlers
  3. /CCBot
Product
  • Track Website Traffic
  • Attribution Software
  • Website Visitor Tracking
  • SEO Dashboard
  • Analytics for SaaS
  • Revenue by Source
  • Traffic Source Tracking
  • Revenue Channel Attribution
  • Revenue Attribution
  • Privacy-First Analytics
  • Cookieless Analytics
  • UTM to Revenue
  • AI Visibility Score
  • Share of Voice (AI)
  • Prompt Tracking
  • AI Citation Tracking
  • ChatGPT Rank Tracker
  • AI Revenue Attribution
  • Pricing
Track AI Traffic
  • Track ChatGPT Traffic
  • Track Perplexity Traffic
  • Track Claude Traffic
  • Track Gemini Traffic
  • Track AI Overviews
  • Track Copilot Traffic
  • ChatGPT Revenue Attribution
  • Perplexity Revenue Attribution
  • Claude Revenue Attribution
  • Gemini Revenue Attribution
  • AI Visibility to Revenue
Use Cases
  • Stripe Analytics
  • Shopify Analytics
  • Stripe Attribution
  • For Bootstrapped SaaS
  • Affordable Attribution
Compare
  • vs Profound
  • vs Loamly
  • vs Peec AI
  • vs Otterly
  • vs Cometly
  • vs Segment
  • vs Google Analytics
  • vs Plausible
  • vs Fathom
  • vs Simple Analytics
  • vs PostHog
  • vs Matomo
  • vs Umami
  • vs Pirsch
  • vs Mixpanel
  • vs Amplitude
  • vs Heap
  • vs Hyros
  • vs AnyTrack
  • vs DataFast
  • vs Similarweb
  • All comparisons
Resources
  • GEO Hub
  • AEO Hub
  • AI Search Hub
  • Research
  • Best Conversion Tracking Software
  • ChatGPT vs Google Traffic
  • Mixpanel Alternative
  • Track Channel Revenue
  • First vs Last Touch
  • Cookieless Conversion Tracking
  • GA4 Attribution Limits
  • CAC by Channel
  • Stripe Conversion Tracking
  • Stripe Revenue Tracking
  • AEO vs SEO 2026
  • AI Traffic Benchmark
  • How to Rank in ChatGPT
  • Best AEO Tools 2026
  • Measure GEO ROI
  • Schema for AI Search
  • Dark AI Traffic in GA4
  • What Is Referral Traffic?
  • What Is Direct Traffic?
  • What Is Cookieless Analytics?
  • What Is Conversion Attribution?
  • AI Share of Voice
  • Documentation
  • View all posts
  • Multi-Touch Attribution
  • Free Tools
  • UTM Builder
  • UTM Checker
  • ROI Calculator
  • AI Readiness Checker
  • AI Crawler Directory
  • SEO + GEO Workflow
Company
  • About
  • Contact
  • Return Delay Penalty
  • Backlink RPV Scoring
  • AI Instructions
  • Live Demo
  • FAQ
  • Log in
© 2026 Attrifast · built by Vincent Ruan & Jessica Huang
AboutContactTermsPrivacy

AI Crawler Directory · Common Crawl

CCBot:
what it is and how to verify it.

OperatorCommon Crawl
PurposeOpen web corpus (training feedstock)
Respects robots.txtYes
Executes JavaScriptNo
Hits verifiableYes — published IPs
TypeHTTP crawler / fetcher

CCBot is the crawler of Common Crawl, a non-profit that maintains an open repository of web crawl data available to anyone. It is not owned by an AI company — but its dataset is the single most important training feedstock in the industry: most large language models trained before per-vendor crawlers existed learned the web substantially through Common Crawl snapshots.

That gives CCBot outsized leverage. Blocking it removes your content from a public dataset used by researchers, startups, and every AI lab that licenses or downloads Common Crawl — including many that run no crawler of their own. Allowing it is the widest possible distribution of your content into future models, one crawl for the whole ecosystem.

CCBot respects robots.txt and crawls from dedicated IP ranges with reverse-DNS verification, published as JSON. It fetches periodically to build monthly-cadence archive snapshots rather than continuously.

CCBot User-Agent string

CCBot/2.0 (https://commoncrawl.org/faq/)

How to verify CCBot hits are real

The User-Agent header is plain text — any client can claim to be CCBot, and scrapers routinely do. Common Crawl crawls from dedicated ranges published at index.commoncrawl.org/ccbot.json, with reverse-DNS verification.

https://index.commoncrawl.org/ccbot.json

Fetch the list and check the hit’s source IP against it — for a quick manual look:

curl -s https://index.commoncrawl.org/ccbot.json

Allow or block CCBot in robots.txt

This is a training-data control: allowing it lets your content shape what future models know; blocking it keeps your material out of training corpora without affecting live citations. A policy call, not a traffic one.

Allow

User-agent: CCBot Allow: /

Block

User-agent: CCBot Disallow: /

CCBot — frequently asked questions

What is CCBot?

CCBot is Common Crawl’s web crawler. Common Crawl is a non-profit that publishes an open repository of web crawl data used by researchers and AI developers worldwide. The bot identifies itself as CCBot/2.0.

Which AI models train on Common Crawl data?

Common Crawl does not publish a client list, but its archives are a documented major component of training corpora across the industry — historically including the datasets behind GPT-class and many open-source models. Blocking CCBot affects all downstream users at once.

Does CCBot respect robots.txt?

Yes. Adding "User-agent: CCBot / Disallow: /" to robots.txt keeps your site out of future Common Crawl snapshots. Existing archived snapshots are not retroactively deleted by a robots change.

How do I verify CCBot hits?

Common Crawl publishes its crawl IP ranges (IPv4 and IPv6) at index.commoncrawl.org/ccbot.json and supports reverse-DNS verification. A "CCBot" UA from outside those ranges is a spoof.

Should a small business block CCBot?

Usually not. For most companies, presence in the open corpus means AI assistants of every brand can learn your product facts. The main reasons to block are licensing-sensitive content or strict data-control policies.

Related crawlers

GPTBotClaudeBotGoogle-Extendedmeta-externalagentBytespiderAll AI crawlers →

Related reading

AI readiness checkerAI revenue attributionAI Readiness Checker

See which AI crawlers hit your site — and what the traffic earns

Attrifast tracks AI referrals as revenue lines: ChatGPT, Perplexity, Claude, and Gemini visits tied to real Stripe money. $9.99/mo flat.

  • ✓One script tag and a Stripe key — live in minutes
  • ✓Cookieless, so no consent banner for analytics
  • ✓Every AI referral matched to the payment it produced
Start 7-day free trial — $0 due today

7-day free trial · $0 due today · then $9.99/mo · cancel anytime

Attrifast dashboard: prompt-level AI visibility with estimated value, revenue split by channel across ChatGPT, Google, Perplexity, Claude and Direct, competitor position tracking, and per-engine scan settings.