Bot protection
UsherStats can decide, for every request to your site, whether to serve it, ask the visitor to prove they are a person, or refuse it. It runs inside your own Cloudflare Worker (the Worker SDK) or in front of your origin (proxy mode), before your code, and it is built so a mistake of ours never takes your site down: anything we cannot answer quickly is treated as "serve the request".
Modes
Each site has one mode, set in the dashboard or the API:
| Mode | What happens |
|---|---|
off | Nothing is decided. Requests are still recorded for analytics. This is the default. |
log | Every request is served. Each request record says what the rules would have done (log: would block: ...). Start here. |
challenge | Requests the rules challenge are sent to a short check; requests they deny get a 403. |
block | As challenge, but where a challenge would be shown the request gets a 403 instead. |
Run a site in log for a few days and read the Crawlers and Bot protection views before switching to challenge.
What is decided, in order
robots.txtand/.well-known/security.txtare always served: a crawler that cannot read your rules cannot obey them.- Your rules, top to bottom; the first rule that matches decides.
- Verified search engines (Googlebot, Bingbot and the others Cloudflare verifies) are always served, unless one of your rules names them (see below).
- Impersonators: a request using the name of a crawler Cloudflare can verify, which Cloudflare has not verified, is refused. Measured on our first customer's sites, every unverified "Googlebot" was someone else.
- Other verified bots are served. Bots that say what they are but cannot be verified follow your
botssetting. - Browsers. A person who passed a check in the last day is served. Otherwise a browser is checked when:
it comes from a hosting network (cloud servers, where people rarely browse from); its client software is on the
shared list of software that never runs pages' scripts; or its behaviour over the last day shows at least two
independent signs of automation (pages read faster than a person reads, probing for files that are not there,
pages loaded with no sign the page was ever seen). The network alone never refuses anyone: in
blockmode a browser on a hosting network is served.
Rules
{
"mode": "challenge",
"bots": "allow",
"rules": [
{ "action": "allow", "label": "partner feed", "match": { "path": "/feeds/", "asn": 13335 } },
{ "action": "deny", "label": "no AI training", "match": { "category": "ai-training" } },
{ "action": "challenge", "label": "check browsers at checkout", "match": { "path": "/checkout", "kind": "browser" } },
{ "action": "deny", "match": { "path": ["/wp-login.php", "/xmlrpc.php"] } }
]
}A rule has an action (allow, challenge or deny), an optional label (it appears in your request records), and
a match. Every field in a match must hold:
| Field | Matches | Example |
|---|---|---|
bot | a bot's name, one or a list | "GPTBot" |
category | search-engine, ai-search, ai-assistant, ai-training, preview, seo, monitor, archiver, tool, other | "ai-training" |
ai | any AI crawler or AI fetcher | true |
kind | browser, bot or verified-bot | "bot" |
verified | whether Cloudflare verified the bot | false |
asn | network numbers | [16276, 24940] |
country | two-letter country codes | ["CN", "RU"] |
path | path prefixes | "/admin" |
Search engines are protected from broad rules. A rule that matches only by asn, country or path does not apply
to a verified search engine, so a rule written for scrapers cannot take your site out of Google by accident. A rule
that names them (bot, category, ai, kind or verified) does apply: {"category": "search-engine", "path":
"/drafts/"} keeps search engines out of /drafts/.
bots sets what happens to a bot no rule names: allow (the default), challenge, or deny. To let in only
the bots you name, set "bots": "deny" and add allow rules above it for the ones you want.
A passed check satisfies challenge rules, never deny rules.
import assert from 'node:assert/strict';
import { normalizeConfig, decide } from '@usherstats/bot';
const settings = normalizeConfig({
mode: 'challenge',
rules: [{ action: 'deny', label: 'no AI training', match: { category: 'ai-training' } }],
});
const gptbot = new Request('https://shop.example/', { headers: { 'user-agent': 'Mozilla/5.0 (compatible; GPTBot/1.1; +https://openai.com/gptbot)' } });
gptbot.cf = { verifiedBotCategory: 'AI Crawler', country: 'US', asn: 8075 };
assert.equal(decide(gptbot, { config: settings }).action, 'block');
assert.equal(decide(gptbot, { config: settings }).reason, 'rule: no AI training');
// A person's assistant opening a link for them is not training, and is served.
const assistant = new Request('https://shop.example/', { headers: { 'user-agent': 'Mozilla/5.0 (compatible; ChatGPT-User/1.0; +https://openai.com/bot)' } });
assistant.cf = { verifiedBotCategory: 'AI Assistant', country: 'US', asn: 8075 };
assert.equal(decide(assistant, { config: settings }).action, 'allow');Every decision is { action, reason, risk, signals }: risk is the 0-100 score for a browser (-1 for a bot) and
signals says which behaviours were seen.
The check
Visitors who are challenged are sent to challenge.usherstats.com, where a Cloudflare Turnstile check usually passes
in about a second, often without anything to click. They are then sent back to the page they asked for, through
/_us/pass on your own hostname, which sets one cookie, us_pass:
- it is first-party (your domain),
HttpOnly,Secure,SameSite=Lax, and lasts a day; - it is signed for your site and for that browser on that network, so copying it to another device or a script does nothing;
- it is set only after a visitor was challenged; no other visitor gets a cookie from UsherStats.
The check only ever returns visitors to hostnames registered for your site and verified: add every hostname your
site uses, publish the TXT record each one shows (_usherstats.<hostname>), and verify it
(POST /v1/sites/{id}/hostnames/{hostname}/verify). Proxy-mode hostnames are verified by Cloudflare's validation.
Without this, a challenge for an unverified hostname is refused, so anyone listing a hostname they do not own cannot
use the check page to send people there.
Setting it up in a Cloudflare Worker
import { withUsherStats } from '@usherstats/worker';
const site = async (request) => new Response('<!doctype html><html><head><title>Shop</title></head><body>Hello</body></html>', {
headers: { 'content-type': 'text/html; charset=utf-8' },
});
export default {
fetch: withUsherStats(site, { pageType: (path) => (path.startsWith('/blog/') ? 'article' : 'page') }),
};Put the site's server secret in the Worker secret USHERSTATS_SECRET and its public key in USHERSTATS_SITE_KEY
(both on the site's Install page). The wrapper:
- serves the analytics script and its reports from
/_uson your own hostname, and adds the script to HTML pages; - records every request (with Cloudflare's network and TLS details) and sends the records to UsherStats in batches, after the response, so no page waits for us;
- reads your bot protection settings at most once a minute and decides before your code runs;
- answers
/_us/passfor visitors returning from a check.
Options: pageType(path) labels pages for the dashboard; botProtection: false turns protection off in this Worker
whatever the mode; firstParty: '/_stats' moves the /_us routes if your site already uses that path;
inject: false stops the script being added (requests are still recorded).
If UsherStats is slow or down, your site is served as if protection were off: settings that cannot be read in a second, a visitor record that does not answer quickly, and a failed upload of records are all ignored. Errors thrown by your own code are passed on to Cloudflare exactly as without the wrapper.
Privacy
Bot protection keeps a short memory per visitor: how many pages, how fast, whether pages were seen. It is filed under a one-way code made from the visitor's network, browser and your site with a secret that changes every day, so it cannot be turned back into an address, cannot follow a person from one day to the next or from one site to another, and is deleted 24 hours after the visitor's last request. The shared list of client software that never runs pages' scripts holds only software fingerprints, never any site's traffic. It is built only from what UsherStats observed itself (proxy-mode requests and the connections of page-script reports), never from what a site's server reports, and a fingerprint goes on it only with evidence from several established customers (on a paid plan, or a month old with traffic on a week of days), no one of whom can outweigh the rest. Mainstream browsers' fingerprints are never put on it.