---
title: "How to Track AI Bots in Cloudflare and Act on the Data"
description: "Learn where Cloudflare shows AI bot requests, paths, status codes, and referrals so you can decide what to allow, block, or charge for with confidence."
canonical: "https://zerply.ai/resources/blog/how-to-track-ai-bots-in-cloudflare"
author: "Anshul Motwani"
date: "2026-09-20T18:30:00+00:00"
updated: "2026-10-07T11:01:07+00:00"
category: "AI Search Visibility"
image: "https://storage.zerply.ai/teams/92/blogs/346/b4625a6b0954c4ff-1791364576323-zerply-blog-banner-track-ai-bots-cloudflare-v2-2026-10-07.png"
---
# How to Track AI Bots in Cloudflare and Act on the Data

Cloudflare makes AI blocking easy. That is very useful, but it also makes it easy to act before you know what you are shutting off.

There is a way to approach this correctly. First, see which crawlers reach your site, which paths they request, how your site responds, and whether any AI platform sends visits back. Then decide what to allow, what to block, and whether any crawler belongs in a paid access path.

Cloudflare docs also draw an important line between controls. Behavior-based AI bot policies can target Search, Agent, and Training traffic. The older legacy Block AI bots setting was training-focused, so you should not treat every Cloudflare AI block as if it means the same thing.

## Where do you find AI bot data in Cloudflare?

If you need to track AI bots in Cloudflare, open **AI Crawl Control** for the zone you want to review.

The data is split across three tabs: Overview, Crawlers, and Metrics.

Use this path when you want to find the data fast:

1. Log in to Cloudflare.
2. Select the right account and domain.
3. Open **AI Crawl Control**.
4. Start in **Overview**, then move to **Crawlers**, then **Metrics**.

Overview tells you whether anything has changed. Crawlers tells you who has changed. Metrics tells you where the requests landed and how your site responded.

If a name looks unfamiliar, check whether it appears as a **verified bot** in the [bot reference](https://developers.cloudflare.com/ai-crawl-control/reference/bots/) and the [Cloudflare Radar Bots Directory](https://radar.cloudflare.com/bots/directory). A **verified bot** is one whose identity Cloudflare has confirmed against the operator's published signals, rather than one that merely claims a name. That check separates a known crawler from an imitator before you act on the traffic.

## What does AI Crawl Control show you?

AI Crawl Control shows request volume, response patterns, crawler-level activity, and referral data.

In the [Overview tab](https://developers.cloudflare.com/ai-crawl-control/features/analyze-ai-traffic/), Cloudflare gives you a snapshot of total request volume, volume change, the most common status code, the most popular path, the status of **managed robots.txt**, and trend views for total requests, allowed requests, unsuccessful requests, and total referrals.

The **Crawlers** tab inspects individual crawlers by name, operator, data transfer, requests, and current action. The crawler table can also show robots.txt violations and makes clear that unsuccessful requests can come from any rule or response error, not only from a manual block in AI Crawl Control.

The **Metrics** tab is where the interpretation work happens. Cloudflare documents views for all requests, allowed requests, data transfer, status code distribution, most popular paths, and URI pattern grouping such as `/blog/*` or `/docs/*`. That lets you move from a crawler name to the actual site areas it touches.

Keep one boundary clear while you read the data. Cloudflare shows requests, status codes, paths, and referrals. Referrals are the only metric here that connects crawler traffic to visits. Blocking a bot in Cloudflare happens at the **edge**, through WAF. 

The **edge** is Cloudflare's network, so the request is stopped before it reaches your origin server. Robots.txt, including Cloudflare's [managed robots.txt setting](https://developers.cloudflare.com/bots/additional-configurations/managed-robots-txt/), is only a request.

Several of these views are paid-plan only. Cloudflare marks referrals in Overview widgets as paid-plan only, and crawler groupings show referrals only on paid plans.

[Top referrers](https://developers.cloudflare.com/ai-crawl-control/features/analyze-ai-traffic/) is also a paid-plan feature, and [custom block responses](https://developers.cloudflare.com/ai-crawl-control/features/manage-ai-crawlers/) are limited to paid plans. AI Crawl Control itself is available on all plans, but free-plan identification relies on user-agent strings, while deeper detection depends on Bot Management detection IDs.

## How do you read the data before acting?

You read the data before acting by moving from trend, to response quality, to path-level impact, and only then to policy.

Start with the broad trend. If one operator or crawler suddenly jumps, do not block it on sight. Check whether the spike matches a new content section, a sitemap change, a rollout, or a media push that made a page worth fetching.

Next, check response quality. Cloudflare's status code distribution and unsuccessful-request counts tell you whether the problem is the crawler or your own delivery path. A high share of 4xx or 5xx responses may point to a rule conflict, a broken path, or an origin issue before it points to a policy decision.

Then look at **Most popular paths** and **Patterns**. A crawler that keeps hitting `/blog/*` and never touches product or docs pages tells a different story from one that concentrates on `/pricing`, `/api/*`, or a gated library. The path pattern is often more useful than the crawler name by itself.

Look at referrals last. Referral traffic is the only place where crawler activity connects to visits, so pair it with your [AI referral traffic](https://zerply.ai/resources/blog/how-to-measure-and-act-on-ai-referral-traffic) review once you understand the request and path patterns. That does not prove citation value by itself, but it does tell you that access may be tied to traffic you can observe.

![](https://storage.zerply.ai/teams/92/blogs/346/ba831b21caeb6932-1790585278042-image.png)

Before you change a setting, ask five plain questions:

- Did the volume change come from one crawler, one operator, or one site section?
- Are unsuccessful requests mostly 4xx, mostly 5xx, or spread across several codes?
- Are the requests landing on pages you want AI systems to read, or on pages you would rather limit?
- Is the crawler a confirmed Cloudflare bot, or only a claimed identity?
- If referrals exist, do the pages receiving them justify keeping access open?

That sequence keeps you from turning a messy dashboard into a blunt policy.

## What should you do about each signal?

You should map each signal to a check, an action, and an owner instead of jumping straight to a block.

| Signal in the data                      | What it likely means                                                                           | What to check                                                                                                         | Action                                                                                                             | Who owns it               |
| --------------------------------------- | ---------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ | ------------------------- |
| Traffic spike from one crawler          | A crawler found a fresh cluster of pages, a sitemap change, or a path pattern worth revisiting | Compare the date with deployments, new content, changed internal links, and sitemap updates                           | Keep access open while you inspect the affected paths, then decide if those pages are worth continued crawl access | SEO lead + web team       |
| High share of unsuccessful requests     | The crawler is hitting blocked, broken, redirected, or unstable paths                          | Review status code distribution, WAF rules, redirect chains, and origin errors                                        | Fix the path or rule conflict first, then reassess whether a block is still needed                                 | Web team + infra owner    |
| A bot you did not expect                | A new operator or category has started requesting content                                      | Verify the bot in Cloudflare's bot reference and Radar directory, then inspect its top paths                          | Allow temporarily if the paths are safe, or limit access after verification if the pattern conflicts with policy   | SEO lead + security owner |
| An unverified bot claiming a known name | Spoofing or weak identification, especially on free-plan user-agent matching                   | Compare the claimed name with Cloudflare's verified listings and, if available, stronger detection signals            | Do not trust the label alone. Investigate and block through security rules if needed                               | Security owner            |
| Referrals from an AI platform           | Access may be leading to actual visits from an AI surface                                      | Check which paths receive the referrals, which platform sent them, and whether those pages support your business goal | Protect the pages that earn useful visits before changing a crawler policy that could reduce them                  | SEO lead + content owner  |

That table does two things. It slows down reactive blocking, and it assigns real responsibility. A crawler policy is rarely just an SEO call or just a security call.

## How do you block, allow, or charge AI crawlers in Cloudflare?

![](https://storage.zerply.ai/teams/92/blogs/346/77886f01a94fa093-1790585698797-image.png)

You block, allow, or charge AI crawlers in Cloudflare by using per-crawler actions in AI Crawl Control and broader behavior policies in Security Settings.

For a specific crawler, start in AI Crawl Control and follow the [Manage AI crawlers](https://developers.cloudflare.com/ai-crawl-control/features/manage-ai-crawlers/) path:

1. Open the **Security** tab.
2. Find the crawler in the table.
3. Use the **Actions** column to choose **Allow** or **Block**.

If your account is in Cloudflare's **pay-per-crawl** beta, a control that charges the crawler operator per successful request instead of allowing or blocking it, you can also choose **Charge**.

If you block a crawler and want to change the response, the path for paid plans is **AI Crawl Control** > **Settings** tab > **Block response** > **Edit**. That is where you can set a `403 Forbidden` or `402 Payment Required` response and add a plain-text body.

For broader policy, use [Block AI Bots](https://developers.cloudflare.com/bots/additional-configurations/block-ai-bots/) carefully. You can block AI traffic by behavior across **Search**, **Agent**, and **Training**. Mixed-purpose crawlers that combine Search and Training are blocked by all configurations that aim to block AI training. That means a category-wide choice can remove search or live assistant access, not only training access.

Cloudflare also separates that behavior-based policy from the older legacy **Block AI bots** setting. The legacy option was training-focused and excluded mixed-purpose Search plus Training bots. If your team still talks about a one-click AI block, pause and name the exact control before you copy the policy somewhere else.

When you want to express a preference rather than enforce one, the path for **managed robots.txt** is **Security Settings** > Filter by **Bot traffic** > **Set your preference to block training in robots.txt**. That setting is useful for signaling intent. It is not the same as edge enforcement.

## What can Cloudflare data not tell you?

Cloudflare data cannot tell you whether a crawler visit turned into model training, a visible citation, or a customer touchpoint.

A request log is still a request log. It tells you that a crawler reached a path and how your site answered. It does not tell you whether that page later shaped an AI answer, whether the answer [named your brand](https://zerply.ai/resources/blog/chatgpt-citations), or whether a human ever saw the output.

That is why Cloudflare is best used as the access and enforcement layer. It tells you who reached the site, where they went, and how the site responded. It does not replace the separate work of [measuring AI visibility](https://zerply.ai/resources/blog/measure-brand-visibility-in-llm-search) over time.

## What does a week of AI Crawl Control data look like row by row?

A proper reading pattern is to move row by row from observed signal to next action.

**Illustrative example:**

| Day or row | Cloudflare signal                                                                        | What it suggests                                                                | What to check next                                                                                                  | Decision                                                                          |
| ---------- | ---------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------- |
| Monday     | `OAI-SearchBot` requests jump from 180 to 1,050, mostly on `/compare/*`                  | Search-focused AI fetching now centers on comparison pages                      | Confirm whether those pages were refreshed, linked from nav, or added to sitemap                                    | Keep access open and review whether those comparison pages deserve a fresh update |
| Tuesday    | `GPTBot` sends 620 requests, with 78 percent landing on `/docs/*`                        | Training crawl interest is heavy on documentation, not on core commercial pages | Check whether the docs section contains content you are comfortable exposing for training use                       | Decide on a docs-specific policy, not a domain-wide reaction                      |
| Wednesday  | `ChatGPT-User` activity appears on `/pricing` and `/integrations/*`                      | Live assistant fetches may be tied to in-session user questions                 | Check for clear pricing text, stable page load, and current integration details                                     | Keep access open if the pages are accurate and meant to answer buyer questions    |
| Thursday   | A crawler label resembles a known bot, but it is not confirmed as verified               | Claimed identity may not match a real Cloudflare listing                        | Compare the name with bot reference and Radar, then inspect request pattern and response codes                      | Treat it as suspicious until verified and hand it to security if needed           |
| Friday     | `chatgpt.com` and `perplexity.ai` appear in referrals for `/blog/ai-crawl-control-guide` | AI platforms are sending visits back to a specific content page                 | Check whether the page is still current, whether the CTA fits, and whether any planned block would reduce that flow | Preserve access while you improve the page that is already drawing visits         |

Read the example as a decision chain, and not a scoreboard. 

Monday says review the page set before touching policy. Tuesday says training access may need a path-level discussion. Wednesday says user-triggered fetches can land on bottom-funnel pages, which changes what a block would cost you. Thursday says identity verification comes before policy. Friday says any referral page deserves protection before a broad rule changes the traffic.

## Conclusion

Tracking AI bots in Cloudflare is straightforward. Acting on the data is where teams either stay precise or get sloppy.

The best rule is always simple. Start with **AI Crawl Control**, read the trends, inspect the paths, confirm the bot identity, and look for referral evidence before you touch a block. Then apply the smallest control that fits the signal, whether that is Allow, Block, Charge where the beta is available, or a request through managed robots.txt.

Cloudflare tells you a crawler reached the page. It stops there. Zerply's [AI Traffic Analytics](https://zerply.ai/platform/ai-traffic-analytics/) reads the same server-layer data with no page script on your site, so the crawl request and the answer that page later showed up in live in the same platform. [Signup for free](https://app.zerply.ai/signup) to see it in action.

## Frequently asked questions

### Where does Cloudflare show AI bot activity?

Cloudflare shows AI bot activity inside AI Crawl Control. The main views are Overview, Crawlers, and Metrics, where you can inspect request volume, status codes, paths, and, on paid plans, referral data.

### Does blocking a crawler in Cloudflare work the same way as robots.txt?

No. A block in AI Crawl Control is enforced at the edge through Cloudflare WAF. Robots.txt, including Cloudflare managed robots.txt, is a request that compliant crawlers may follow voluntarily.

### Are AI referrals available on every Cloudflare plan?

No. Cloudflare marks referrals in AI Crawl Control as a paid-plan feature. That includes referral totals in Overview and the Top referrers section.

### Can Cloudflare tell me whether an AI crawler trained on my content or cited it in an answer?

No. Cloudflare can show that a crawler requested a path and how your site responded, but it does not show whether the content was trained on, cited, or seen by a customer.

```json
{"@context":"https://schema.org","@type":"Article","headline":"How to Track AI Bots in Cloudflare and Act on the Data","description":"Learn where Cloudflare shows AI bot requests, paths, status codes, and referrals so you can decide what to allow, block, or charge for with confidence.","url":"https://zerply.ai/resources/blog/how-to-track-ai-bots-in-cloudflare","image":"https://storage.zerply.ai/teams/92/blogs/346/b4625a6b0954c4ff-1791364576323-zerply-blog-banner-track-ai-bots-cloudflare-v2-2026-10-07.png","datePublished":"2026-09-20T18:30:00+00:00","dateModified":"2026-10-07T11:01:07+00:00","author":{"@type":"Person","name":"Anshul Motwani"}}
{"@context":"https://schema.org","@type":"FAQPage","mainEntity":[{"@type":"Question","name":"Where does Cloudflare show AI bot activity?","acceptedAnswer":{"@type":"Answer","text":"Cloudflare shows AI bot activity inside AI Crawl Control. The main views are Overview, Crawlers, and Metrics, where you can inspect request volume, status codes, paths, and, on paid plans, referral data."}},{"@type":"Question","name":"Does blocking a crawler in Cloudflare work the same way as robots.txt?","acceptedAnswer":{"@type":"Answer","text":"No. A block in AI Crawl Control is enforced at the edge through Cloudflare WAF. Robots.txt, including Cloudflare managed robots.txt, is a request that compliant crawlers may follow voluntarily."}},{"@type":"Question","name":"Are AI referrals available on every Cloudflare plan?","acceptedAnswer":{"@type":"Answer","text":"No. Cloudflare marks referrals in AI Crawl Control as a paid-plan feature. That includes referral totals in Overview and the Top referrers section."}},{"@type":"Question","name":"Can Cloudflare tell me whether an AI crawler trained on my content or cited it in an answer?","acceptedAnswer":{"@type":"Answer","text":"No. Cloudflare can show that a crawler requested a path and how your site responded, but it does not show whether the content was trained on, cited, or seen by a customer."}}]}
```
