Measuring
Bot detection
Bot traffic is labelled and kept, never silently discarded. Human metrics always filter it out, and "which of my pages is GPTBot reading" stays answerable.
The distinction that matters
| AI crawler | AI referral | |
|---|---|---|
| What it is | GPTBot, ClaudeBot, PerplexityBot fetching your page | A real browser arriving from chatgpt.com |
| Is there a human? | No | Yes |
| Counted as a visitor? | Never | Yes |
| Where you see it | ai.crawler.* metrics | The ai channel |
Most tools conflate these. If they are shown together, "I am getting AI traffic" is both inflated and misread — you cannot tell whether a model read you or sent you someone.
Categories
| Kind | Examples |
|---|---|
ai-crawler | GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended, Bytespider, CCBot, Meta-ExternalAgent, Applebot-Extended, Amazonbot, Diffbot — reading in bulk, nobody waiting |
ai-agent | ChatGPT-User, Claude-User, Perplexity-User — one page fetched because a person asked a question just now. A human is waiting, but never lands on your site. |
search-crawler | Googlebot, bingbot, DuckDuckBot, YandexBot, Baiduspider, Applebot |
seo | AhrefsBot, SemrushBot, MJ12bot, DotBot, Screaming Frog, DataForSeoBot |
preview | facebookexternalhit, Twitterbot, Slackbot, Discordbot, LinkedInBot, WhatsApp |
monitor | UptimeRobot, Pingdom, StatusCake, Lighthouse, GTmetrix |
generic | HeadlessChrome, Puppeteer, Playwright, curl, python-requests, and an empty user agent |
An empty user agent counts as a bot. Real browsers always send one, so an empty value is either a script or a tool hiding itself — counting it as human would inflate your numbers silently.
What a JavaScript beacon can and cannot see
This is the limitation to understand before reading any crawler number, here or anywhere else. Our tracker is a script in your page. A crawler that renders the page runs it and appears in these reports; a crawler that fetches the HTML and leaves never does.
That split is not even: Googlebot renders, and so our own production data showed Google crawlers and essentially nothing else. GPTBot, ClaudeBot, CCBot and Bytespider do not execute JavaScript, so through the browser script alone they are invisible — to us and to every other client-side analytics tool, including the ones that advertise AI crawler reporting.
Saying otherwise would be exactly the kind of unverifiable claim this product exists to stop making, so: to see non-rendering crawlers you need server-side ingest, where your server forwards what it served. It is one POST per request and it is also the only path on which a signed agent can prove who it is.
POST https://app.vitrus.dev/api/collect/server
{ "site": "SITE_ID", "url": "/pricing", "ip": "203.0.113.9",
"headers": { "user-agent": "...", "signature-agent": "...",
"signature-input": "...", "signature": "..." } }
Agent sessions: the third kind of traffic
There is a case no bot table can catch. An agentic browser — ChatGPT Atlas, OpenAI Operator — is a real Chrome session driven by an agent on someone's behalf. It executes JavaScript, so it reaches the tracker like any visitor; it clicks, fills forms, and sometimes converts.
And it sends an ordinary Chrome user-agent. The string observed from ChatGPT's agent is exactly
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) … Chrome/138.0.0.0
Safari/537.36 — nothing in it can be matched on. What it does send is
Signature-Agent: "https://chatgpt.com", which is why agent sessions are
identified by signature rather than by any list.
What everyone does with it, and why we do not
| Tool | Behaviour | What it costs |
|---|---|---|
| Umami, Plausible | Discards the traffic | An agent that completes a checkout is revenue, and the revenue disappears. |
| Rybbit | Blocks it | Same, plus you cannot tell whether agents can use your site at all. |
| GA4 | Counts it as a person | Your funnel, bounce rate and intent data are all mixed with something that is not a customer. |
| Vitrus | Counted as its own class | Excluded from your visitor numbers, reported separately, with its own funnel. |
The number is a floor, not a total
An agent that browses without signing is indistinguishable from a person, and we do not guess. This page used to say heuristics were "not a trade available to us"; that was too strong, and the next section is the correction. Heuristics are computed — but they are kept out of this number, because a signed agent is a fact and a suspicious one is an opinion, and summing the two produces a figure that means neither. Agent sessions rise only as more operators start signing.
Automation signals: the second layer
A user-agent table catches clients that admit what they are. It catches nothing that lies — and that is the loudest complaint in this category.
Self-hosters report 200 real visitors showing up as 5,000. Plausible's cloud runs user-agent filtering plus roughly 32,000 datacentre IP ranges plus behavioural analysis; its Community Edition ships the user-agent filter alone. Umami users have measured about a third of their recorded visits as fake. Ours was one layer too.
So every request is also checked against what a browser making that call would actually look like. This is in the Apache-2.0 core, not held back for the paid tier.
| Rule | Weight | What it observes |
|---|---|---|
no-accept-language | 3 | Claims to be a browser, sent no Accept-Language. |
no-fetch-metadata | 3 | No Sec-Fetch-* headers. Chromium since 2020, Firefox since 90, Safari since 16.4 — hence 3, not 5. |
chromium-without-client-hints | 3 | Says Chrome, sent no Sec-CH-UA. Asked only of Chromium claims; these headers do not exist in Firefox or Safari. |
platform-mismatch | 5 | The user-agent names one operating system and Sec-CH-UA-Platform names another. Two different programs. |
no-accept | 2 | No Accept header at all. |
The threshold is 5, so no single rule can accuse anyone: one missing header is normal somewhere
on the internet — an old browser, a privacy extension, a corporate proxy that strips things. The rules
are applied only to clients claiming to be a browser; a request that honestly says
curl/8.4 is already named by the table above and is left alone.
Recorded, not enforced
This is the part that differs from every other tool. A request that scores above the threshold is still stored as a human visit and still counted in every number on your dashboard. What we store alongside it is the list of rules that fired, so the suspicion can be explained later.
bot_signals no-accept-language,no-fetch-metadata,chromium-without-client-hints
bot_score 9
The Traffic quality page shows how many visitors reached the threshold and which rules did it. Excluding them is a switch you flip, it is not remembered between visits, a banner stays on screen for as long as it is on, and the evidence panel shows the predicate doing the excluding:
SELECT COUNT(DISTINCT visitor_id) AS value
FROM events
WHERE site_id = ? AND ts >= ? AND ts < ?
AND bot_kind = '' AND agent_trust = 'human'
AND bot_score < 5
The reason for going to that trouble: a silent filter is indistinguishable, from your side, from a bug that loses traffic. "Rather undercount than guess" cuts both ways — we will not guess that someone is a robot either.
What this is not
It is not a browser fingerprint. No canvas, no font enumeration, no WebGL renderer, no plugin list — which is what a "client signals" layer usually means, and all of which would undercut the reason to run this instead of Google Analytics. Every rule reads a header the client sent anyway, and a test in the repository fails the build if a fingerprinting surface ever appears in that module.
Known blind spots. A headless browser driven by Playwright with stealth patches sends plausible headers and scores zero. So does a residential-proxy botnet running real Chrome. The rule table is versioned and will grow; it will never reach certainty, and anything claiming to is selling something.
Why bots are stored, not dropped
Umami and Plausible throw bot hits away. We keep them because the question "which pages are AI models actually reading" is the only real feedback loop for GEO/AEO work — you cannot optimise for citation if you cannot see what was fetched.
The safety is at the query layer, not at ingest: every human metric carries
bot_kind = '', and that constant is used rather than copy-pasted so it
cannot be forgotten in one place.
SELECT COUNT(DISTINCT visitor_id) AS value
FROM events
WHERE site_id = ? AND ts >= ? AND ts < ?
AND bot_kind = ''
A versioned table
Detection is a deterministic rule table with a version stamp, not a model. Each row records the table version that labelled it, so a reclassification is visible rather than retroactive. A golden-set eval runs in CI and has to stay at 100% — when a new assistant is added, the old rows are proven unbroken.
Quota
Crawler requests are stored but do not count against your plan's event quota. You did not ask GPTBot to visit, so you should not pay for it.