Clarvivo

Free tool by Clarvivo · no login

Find the traffic in GA4 that isn't human

Upload a GA4 acquisition export. We compare every segment against your own engagement baseline and flag the ones behaving like something that fetches pages rather than reads them.

GA4 acquisition export

A Session source / medium export finds more than a channel export, because a single bad referrer shows up as its own row. Engagement columns are required.

This tool answers: Which of my traffic segments look automated rather than human?

Enable JavaScript to run the check — the diagnosis itself is free and needs no account.

Why this happens

GA4 filters known bots automatically, and that filtering is genuinely good. What it cannot catch is the long tail: uptime monitors, security scanners, SEO crawlers, headless scrapers, link-preview fetchers, and click fraud on paid campaigns.

None of those announce themselves. They arrive as ordinary sessions, land on a page, and leave — and every average in your acquisition report moves slightly as a result. Engagement rate drops, conversion rate drops, and the channel carrying them looks like it is underperforming.

The reason it goes unnoticed is that nobody looks for it at segment level. Averaged across a whole property the effect is small; concentrated in one referral source or one campaign it can be most of that row.

But automated traffic is distinguishable, because its behaviour is extreme in a way real traffic never is. A segment with thousands of sessions and a 3% engagement rate is not a poorly performing channel. No human channel performs that badly at volume.

Next checks

These answer the questions this one tends to raise.

How this check works

So you can trust the result — and your developer can verify it.

  1. We build a baseline from your own data, then refine it

    Bot detection against an industry benchmark is meaningless — sites differ enormously. We compute a session-weighted engagement baseline from your export, then recompute it excluding the suspect rows. Without that second pass, a report that is mostly automated traffic normalises its own anomaly and we would find nothing exactly when there is most to find.

  2. Two independent signals

    Engagement rate far below your baseline, and average engagement time under two seconds. The second is the clearest tell there is: under two seconds the page was fetched, not read.

  3. Volume thresholds, so the result is actionable

    A row must have at least 40 sessions and half a percent of your traffic before we flag it. Flagging a twelve-session row as suspicious is noise, and noise is what makes people stop reading these reports.

  4. Each segment gets its own verdict

    You get one finding per suspect segment with its sessions, engagement rate and how far below baseline it sits — so you know exactly which row to exclude rather than being told a percentage and left to hunt.

Common causes and fixes

Uptime and performance monitors
Pingdom, StatusCake, internal health checks and similar hit the same page on a fixed interval forever. The regular timing is the giveaway — segment by hour and the pattern is unmistakable.
SEO crawlers and scrapers
Competitive-intelligence tools and content scrapers execute JavaScript well enough to register a session and then leave immediately. They usually cluster on a small number of high-value pages.
Referral spam
Sessions attributed to a domain you have never heard of, existing purely so you visit the referrer. Almost always near-zero engagement, and obvious once a source/medium export separates them out.
Click fraud on paid campaigns
The most expensive kind, because you are paying for it. Look for a paid segment whose engagement sits far below your other paid traffic — that is a different signature from a merely weak campaign.
Your own automation
Synthetic monitoring, end-to-end test suites and preview crawlers hitting production. Easy to fix once identified, and worth ruling out before assuming an external cause.

Frequently asked questions

Doesn't GA4 already filter bots?
It filters bots on the IAB known-bots list, which is a real and useful baseline. It does not catch the long tail — monitors, scanners, scrapers, preview fetchers and click fraud — because those are not on any published list. This tool looks at behaviour instead of identity.
How can you tell a bot from a genuinely bad traffic source?
You often cannot, with certainty, from aggregates alone — and we say so rather than pretending otherwise. What we can say is that engagement far below your own baseline at meaningful volume does not happen with human traffic. Whether the cause is automation or a landing page that fails to load, either way that segment is not what your reports think it is.
Which export finds the most?
A Session source / medium export, by a wide margin. On a channel export a single spam referrer is averaged into the entire Referral channel and disappears. On a source/medium export it is its own row.
What do I do once I find it?
Confirm the pattern first — segment by hour and by landing page, since automated traffic usually arrives on a schedule and concentrates on a few pages. Then exclude it with a GA4 filter or an internal traffic rule so it stops diluting your averages.
Why does bot traffic matter if it never converts?
Because it sits in the denominator. Every rate you calculate — engagement, conversion, revenue per session — is diluted by sessions that were never going to do anything, which makes real channels look worse than they are and can lead you to cut something that works.

Clean traffic data is the precondition for everything else.

Excluding automated traffic makes your rates real again. Clarvivo takes the next step and ties the remaining, real traffic to payments — so acquisition decisions rest on people who actually bought something.

Keep bad traffic out of acquisition decisions

What this tool reads, and what it keeps