Model Evaluator
Model Evaluator is a discipline for tracking what happens after a classification model is deployed — not whether it scores well on a held-out benchmark, but what happens inside the systems it touches, to the people its outputs decide about, and to the costs those decisions push downstream.
What we track
- Classification behavior in production, not on curated evaluation sets.
- Who bears the cost of each misclassification — the modeled subject, the operator, the third party downstream, or the taxpayer.
- How costs compound as classification outputs flow into automated decision, routing, and enforcement systems.
- Distributional impact across the populations a deployed classifier actually meets.
Why it matters
Benchmark leaderboards measure model capability on curated evaluation sets. Deployment measures something different: which mistakes actually happen, to whom, and at what compounding cost. The gap between benchmark accuracy and deployment impact is where consequential harm accumulates — often silently, buffered by intermediate systems, until a downstream failure surfaces it far from the original classifier.
Model Evaluator names this gap as a distinct object of study, and treats it as the primary evaluation surface once a classifier is in production.
Three lanes
The site publishes in three lanes: methodology (disciplines for reading evaluation reports), consequence (case studies of deployed classifiers), and formation (analysis of the emerging evaluation market).
What's coming
A public tracker of classification-consequence patterns, methods for measuring them, and case studies of production classifier behavior in high-stakes decision systems — insurance, employment screening, benefits eligibility, content moderation, healthcare triage, credit decisioning.