Blue Camel
Back to all work A large social platform · 2023

Making the launch numbers trustworthy, for a global consumer hardware programme

The problem

A company was three months from a consumer hardware launch and had quietly stopped believing its own numbers.

The pattern is familiar. When a server goes down, somebody’s phone rings at two in the morning. When a number goes wrong, it sits on a dashboard until the quarterly review, and by then it has been wrong for a quarter. One top-line metric had been under-reporting for half a year before anyone caught it.

You cannot launch on numbers you do not trust. The people who most need to trust them are the ones who will be standing on a stage.

What we did

We treated the numbers like infrastructure.

First we ranked them. More than 200 production metrics, tiered by how much it would matter if each one were wrong. Most do not matter much. A handful decide whether the launch is going well. This ran as a 17-person cross-functional effort over four months.

Then we gave the ones that matter an on-call rotation, the same way a service has one. A documented bound, written in plain English with the reason for it. A runbook: what to check, what to roll back, who to wake up. Anomaly detection wired to page a person, not to fill a chart. A blameless review after every page.

The tooling is not the interesting part. The discipline is. When a number is going to wake somebody up if it drifts, the team argues about the bound before launch instead of after. The number that ships stops being the number the system happens to produce and becomes the number a team has agreed to be on the hook for. That is almost always a better number.

We also fixed what was already broken: the under-reporting, and a retention window that kept only thirty days of top-line history, which we extended past five years so that "is this normal" became an answerable question.

What happened

Data incidents stopped waiting for a quarterly review to be noticed. With bounds agreed before launch and anomalies paging a person, the team caught them as they happened and worked them the same way it worked a service outage.

Time to resolve came down 60%.

The org trusted its numbers on launch day. That was the whole assignment.

What it cost them to find out

Half a year of a top-line metric reading low, and every decision made against it in that window. Nobody writes that cost down, because the cost of a number being quietly wrong never appears as a line item. It appears as a plan that did not work and nobody could say why.

What we got wrong, and where this does not apply

This is overkill in two places. Do not pay for a rotation on a model whose drift produces no harm anyone can point at. And do not install it on a system that has not reached a steady state, because you will generate noise instead of signal.

The argument for treating metric drift as a paged severity, on Jeff’s personal site.

Contact

Recognise the problem? Write to jeff@bluecamelconsulting.com with two or three sentences on your version of it.