Web analytics7 min read

What is clickstream data?

Every visitor leaves a trail of pages and clicks. Your analytics records the trail on your site, and a quieter industry collects trails across the whole web to sell in aggregate. Both are called clickstream data, and the difference matters.

By The Bigdelta team
What is clickstream data?

One visitor, one trail

Clickstream data is the sequence of pages and actions a visitor moves through, recorded in order: arrived from Google on the homepage at 9:14, clicked through to /pricing, opened the FAQ, left. The name is literal - a stream of clicks. Almost everything web analytics shows you is this trail after aggregation: sessions, paths, funnels and exit pages are all cuts of the same ordered record.

The term hides a split worth understanding before you trust any number built on it. There are two very different kinds of clickstream, owned by very different people.

First-party clickstream: the trail on your own site

When your analytics tool records a visitor's path through your pages, that's first-party clickstream. The pipeline behind it is the ordinary one - a script reports each pageview and click, the tool stitches them into sessions - and the data stays yours. It only covers your site: the trail starts when a visitor lands and goes dark the moment they leave.

This is the reliable kind. It is a real record of real visits, minus the visitors your script never sees, and it's what powers the reports you already use. A conversion funnel is clickstream filtered to the steps you care about. A user profile that lists someone's visits in order - Bigdelta's profiles work this way - is one person's clickstream read top to bottom.

Third-party clickstream: trails across the whole web

The second kind is collected by companies you've never installed anything from - or rather, that millions of other people have. Browser extensions, free apps and antivirus suites can observe the browsing of the people who run them. Aggregated across a large enough panel, those individual trails become a statistical picture of where the whole internet goes, and that picture is packaged and sold: to market researchers, to investors, and to the competitive-intelligence tools marketers use every day.

Neither the collection nor the sale is a secret. Semrush describes its panel openly: over 200 million anonymized users' worth of behavior, gathered through more than 100 apps and browser extensions. Similarweb lists its inputs as its own contributor network, direct measurement from sites that share analytics, public data, and purchased partnerships. When you look up a competitor's traffic, this is where the number comes from.

How a panel becomes a traffic estimate

A panel sees a sample of the internet, not the internet. If a site gets 2% of the panel's visits, the estimator models what 2% of the whole web's attention looks like and prints a visitor count. The bigger and better-balanced the panel, the better the model - and it is still a model, which both vendors say plainly. That's why estimates for the same site disagree with each other and with the site's own analytics, and why small sites, which a panel barely samples, often can't be estimated at all.

None of this makes the estimates useless. It makes them comparisons rather than counts - fine for ranking competitors against each other, wrong for reading anyone's absolute numbers.

What clickstream analysis is used for

On your own site, the order of events is the whole value. Path analysis shows the routes visitors actually take, which rarely match the navigation you designed. Funnel analysis finds the step where they quit. Ecommerce teams read the trail from product page to cart to checkout to see where buyers fall out. All of this is behavior analytics working on the sequence rather than the totals.

Third-party clickstream answers a different question - what happens outside your walls. Which competitor is growing, where audiences overlap, which sites send traffic you're not getting. Strategy questions, answered in estimates.

The privacy record

The panel model has an ugly history worth knowing. Avast, the antivirus maker, spent years collecting browsing trails from its security products and selling them through a subsidiary called Jumpshot - over 8 petabytes of browsing history, offered to more than 100 buyers. A joint press investigation in January 2020 found users hadn't meaningfully consented, and Avast shut Jumpshot down within days. In 2024 the FTC fined Avast $16.5 million and banned it from selling browser data.

The episode is why "anonymized panel" gets scrutiny. Aggregated trails are hard to tie back to a person. Raw trails are trivially identifying, because a browsing history is close to a diary. If you buy clickstream-based numbers, the vendor's transparency page is worth two minutes of your time. And your own first-party trail deserves the same respect you'd want from a panel - collect what you use, tell visitors honestly, and skip recording what you'd be uncomfortable explaining.

Clickstream vs click map

The names collide, the tools don't. A click map is a picture of where visitors click on one page - a heatmap, no sequence involved. Clickstream is the sequence: which pages, in what order, across a visit or across the web. One tells you what happens on a page, the other how visitors move between pages. Most analytics setups quietly give you both.

The practical takeaway

You already have clickstream data - any analytics tool collects the first-party kind from the moment it's installed, and the path and funnel reports are where it pays off. Treat the third-party kind as what it is: a modeled estimate of everyone else's trails, bought from panels of variable hygiene, good for direction and rankings and nothing after the decimal point. When a number about your own site and a number from a panel disagree, believe your own.