Measurement that survives the question
We put a number in front of legal, or a regulator, or a customer’s security review, and we cannot show our work.
You get a measured answer with its own error bars and a trail somebody can audit, built from a sample that people have checked rather than from a scanner’s raw count. It works on training corpora and on model behaviour, and on anything else where the real answer is a range rather than a point.