Blue Camel
AI Evaluation & Engineering

How you know a model does what you claim, with error bars a reviewer can check.

Six years designing and evaluating machine learning systems at a large social platform: privacy infrastructure for training data, the judge methodology that measures what is actually in a corpus, and the data science behind that platform's public creator-captioning launch in 2022.

A model is easy to demo and hard to trust. The work here is the second part: a judge that states its own error rate, an evaluation built around what breaks, and a number a team will stand behind on launch day.

What we take on

An LLM-as-judge you can calibrate

A model scores our content, or our risk, or our compliance. We cannot tell a regulator how often the score is wrong.

A judge that writes down a reason for every call, so a person can check it. Human raters score a slice, which tells you the judge’s precision and recall, and the raw count is then corrected for that error so the result ships as an interval rather than a bare number.

Evaluation you can hand to a skeptic

The demo passed. We still do not know what happens when it meets our real traffic.

Tests built around the failure rather than the happy path, so you find out what a process kill or a malformed input or a rate limit does to the thing before your customers do. The number that ships is one the team has agreed to be on the hook for.

Trust that survives the year after launch

The AI passed its evaluation and went live nine months ago. Nobody here can tell me what it is doing now.

Before launch it was measured against a standard: the outputs that count as correct, grounded in your own sources and inside the policy you wrote. We keep measuring against that same standard on live traffic, broken out by the kind of work, and we publish a range instead of one flattering figure. When a measurement goes out of date it gets pulled rather than redrawn, and only the affected slice needs re-labeling. This is middle management for an AI team, and it works on an assistant, an agent, a ranker or a classifier alike. See how we keep it valid after launch.

Buyer-side diligence when you are the one signing

A vendor’s accuracy line is internal benchmarks they cannot share. Our scorecard has nowhere to put that.

Scorecard weights fixed before anybody sees a demo, failure modes substantiated on a named dataset, and the word unscoreable written plainly wherever a claim cannot be tested at all. The structure is borrowed from a healthcare procurement that did it properly seventeen years ago.

The judge method, proven

This ran as a training-corpus measurement at a large social platform. The full case sits under Data Science, and the same judge method runs live as a prevalence audit you can drive.

Published and running

Evaluating a model, or a vendor’s claim about one? Start a conversation.

The other practices: Data Science & Experimentation Custom Build & Delivery Mission-Driven Systems