Essay · 19 January

The quiet cost of sampling

Sampling is not a moral failure. It is a budget. The failure is presenting a sampled funnel as if every rare journey had an equal chance to appear.

Someone typing on a laptop at a wooden table
Volume discounts in a tracker are not the same as a representative census.

Most App Analytics vendors will, at some volume, stop counting everyone. The documentation is accurate and unread. Product leads then compare a 2.1% checkout completion this month with 2.4% last month and commission a redesign. In the studio we have watched that gap vanish once the same question is asked of an unsampled warehouse export, and we have watched the opposite: a ‘stable’ funnel that was only stable because the rare failure path was too small to survive the sample.

Who is dropped is not random in the way people hope. High-frequency users fill the sample. Quiet users — the ones you are trying to retain — are the first to become rounding error. Accessibility journeys, older devices, and the one region with a flaky payment rail disappear from the highlight reel while remaining painfully present in support tickets.

How we label it

If a chart is sampled, the title says so, in words a director can see without hovering. We add the sample rate, the date it changed, and whether the warehouse copy disagrees. Experiment Readouts without Vanity spends a whole evening on this because A/B tools love to hide eligibility behind a pretty interval.

We also refuse to compare sampled and unsampled periods without a join in the margin. A vendor migration that ‘improved’ conversion by a fifth of a point is often a change in who was allowed into the count. That is not a product story. It is a metering story.

None of this requires you to buy a bigger plan on the day. It requires you to stop treating a cost-saving as a census. If finance will not fund full capture, say that in the pack. Honesty about the budget is part of the craft.

← Journal