Model Evaluator

Model Evaluator is a discipline for tracking what happens after a classification model is deployed — not whether it scores well on a held-out benchmark, but what happens inside the systems it touches, to the people its outputs decide about, and to the costs those decisions push downstream.

What we track

Why it matters

Benchmark leaderboards measure model capability on curated evaluation sets. Deployment measures something different: which mistakes actually happen, to whom, and at what compounding cost. The gap between benchmark accuracy and deployment impact is where consequential harm accumulates — often silently, buffered by intermediate systems, until a downstream failure surfaces it far from the original classifier.

Model Evaluator names this gap as a distinct object of study, and treats it as the primary evaluation surface once a classifier is in production.

Three lanes

The site publishes in three lanes: methodology (disciplines for reading evaluation reports), consequence (case studies of deployed classifiers), and formation (analysis of the emerging evaluation market).

What's coming

A public tracker of classification-consequence patterns, methods for measuring them, and case studies of production classifier behavior in high-stakes decision systems — insurance, employment screening, benefits eligibility, content moderation, healthcare triage, credit decisioning.