AI Crawler Activity in Website Logs: Cloudflare AI Bots, Visibility Tools, Google AI Overviews Crawler, and Bing Copilot Crawler Analytics (2026)
Which bot in your logs actually feeds ChatGPT, AI Overviews and Copilot, what Cloudflare shows on every plan, and where a visibility tool has to take over.
Most teams open their logs looking for an "AI Overviews crawler" and a "Copilot crawler." Neither exists as a separate user agent. That one fact changes how you read AI crawler activity, which tools you actually need, and which robots.txt lines are safe to touch. This guide maps each AI surface to the bot that really shows up in your logs, walks through what Cloudflare reports out of the box, and then looks at where a visibility tool has to pick up the trail.
Which bot feeds which surface
Start with the mapping. Nearly every bad decision about AI crawlers comes from getting it wrong.
| AI surface | What appears in your logs | What the vendor says |
|---|---|---|
| Google AI Overviews and AI Mode | Googlebot | Googlebot crawl preferences apply to Google Search "including Discover and all Google Search features." No separate AI Overviews user agent exists. |
| Gemini Apps and Vertex AI (training and grounding) | Nothing new | Google-Extended is a robots.txt token with no HTTP user agent of its own. You will never see it in a log. |
| No specific product | GoogleOther | A generic Google crawler that Google says doesn't affect any specific product. |
| Microsoft Copilot | Bingbot | Bing says Copilot relies on the same crawling, indexing and ranking foundation as traditional search. |
| ChatGPT | GPTBot, OAI-SearchBot, ChatGPT-User | OpenAI documents GPTBot (training) and OAI-SearchBot (search) as different robots. ChatGPT-User is user-initiated fetching. |
Two consequences follow. If you block Googlebot to keep content out of AI Overviews, you also take it out of Search. Google-Extended won't help with AI Overviews either, since Google states it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal." And a "Copilot crawler" spike in a dashboard is Bingbot. Bing lists blocking Bingbot in robots.txt as a mistake, because doing it cuts off Copilot's source as well.
Bing's real controls for Copilot are page-level directives. robots.txt controls crawling, not indexing. NOINDEX keeps a URL out of Bing, Copilot and grounding results. NOARCHIVE prevents content from being used in Copilot responses and grounding results. NOCACHE limits Copilot to the URL, title and snippet. If legal wants a page out of Copilot answers, those are the levers. A Bingbot block is not.
One caveat before you trust any of these names: user agents can be spoofed. Google tells you to verify Googlebot with a reverse DNS lookup (the host should resolve to googlebot.com) or against its published IP ranges. Microsoft offers a Verify Bingbot tool for the same check. A log line that says Googlebot is a claim until it passes.
What Cloudflare shows natively
If your site runs through Cloudflare, AI Crawl Control (it used to be called AI Audit) is the first place to look, and it's on every plan. The Overview, Crawlers and Metrics tabs break requests down by crawler and by operator, covering OpenAI, Microsoft, Google, ByteDance, Anthropic and Meta. You get allowed versus unsuccessful requests, robots.txt violations, status codes, paths and hosts. You can allow or block each crawler individually. Pay per crawl exists but is in private beta.
The free plan has real limits. Detection is by user-agent string only, so the spoofing problem above stays yours. The Metrics tab covers the past 24 hours. Referral counts are reserved for paid plans. Enterprise customers with Bot Management get detection IDs and configurable timeframes, and the GraphQL Analytics API exposes the data if you want it somewhere else.
For a quick read on whether OpenAI's bots are hitting you and getting 200s back, that's plenty. What Cloudflare doesn't do is connect any request to a prompt someone asked, a citation in an answer, or a visit that converted. It's a traffic control panel. It has no idea what happened after the bot left.
The questions logs alone can't answer
Suppose the Metrics tab shows OAI-SearchBot fetching your pricing page all week. Did ChatGPT then cite that page? For which prompts? Did anyone click through from the answer, and did they sign up? Or take the page that returned errors to PerplexityBot on Tuesday. Was that the page Perplexity would have cited for your most valuable buyer prompt?
Answering any of that means joining the crawl, the answer and the visit. A CDN dashboard holds the first. A prompt tracker holds the second. Your analytics holds part of the third, if it separates AI referrers at all. Joining them by hand in a spreadsheet works for about a week, and then nobody keeps it up.
Where a visibility tool takes over
This is the job Promptwatch built Agent Analytics for, and per its founders it shipped AI crawler logs before most of the GEO category had anything comparable. Agent Analytics ingests real-time crawler logs (the product lists ChatGPTBot, ClaudeBot, PerplexityBot, GoogleOther and the Meta AI crawler), shows the crawl-to-citation path, and tracks errors. The same workspace runs prompt tracking and citation analytics, so a crawl of your pricing page sits next to the prompts where that page did or didn't get cited. Visitor analytics, installed with a lightweight script or a GTM template, adds the AI-referred visits and the conversions behind them.
The Cloudflare connection is simple. Enterprise accounts use Logpush. On other plans Promptwatch auto-deploys a Worker, and the DNS record has to be orange-cloud proxied so traffic passes through it. Not on Cloudflare? AWS CloudFront, Fastly, Vercel, Netlify, Akamai, Google Cloud CDN and a custom HTTP option are supported. Our Cloudflare setup walkthrough has the steps.
Check the plan before you buy. Essential at $95/mo includes visitor analytics but lists no crawler-log allowance. Crawler logs start on Professional at $245/mo with 25M logs, and Business at $579/mo raises that to 100M. Agencies get them from the first self-serve tier: Kick-off at $199/mo includes 10M, Growth at $399/mo includes 25M and Scale at $799/mo includes 100M.
Leave Cloudflare's AI Crawl Control switched on alongside it. That's where you allow and block, and it costs nothing. Promptwatch covers what Cloudflare can't see, which is whether a crawl became a citation and whether that citation sent you a customer.
A short checklist for this quarter
- Confirm Googlebot and Bingbot are allowed everywhere you want to appear. They feed AI Overviews, AI Mode and Copilot.
- Treat Google-Extended and GPTBot as separate training and grounding decisions. Neither one changes your Google Search inclusion.
- Verify any bot you plan to act on with reverse DNS or the vendor's own tool.
- Open AI Crawl Control, filter for unsuccessful requests, and fix the status codes on pages you want cited.
- Put crawl, citation and visit data in one place. For us that's Promptwatch Professional. If you want to compare the alternatives first, read our ranking of GEO tools with AI crawler log tracking.