Clinical Evals

How Healthcare AI Benchmarks Work

6 guides · 18 benchmark profiles · updated August 17, 2026

Every week someone cites a healthcare AI score without knowing how it was produced. This site explains the machinery: how rubric grading works, why the graders are themselves models, what a leaderboard number can and cannot tell you, and how each benchmark that still tests frontier models is actually constructed. When you want the current standings rather than the mechanics, they live at clinicalbenchmarks.ai.

Guides

Benchmark mechanics

A profile of each benchmark with current frontier activity: how it was built, how it grades, and what its score means. Profiles explain mechanisms; current results live on the boards each profile links.

Rubric-graded

Composite indices

Safety

Agentic and workflow

Documentation and coding

Knowledge and exams

Writing

The Grader Problem Health AI benchmarks are graded by language models. How grader choice, length adjustment, and subset selection change published scores, and what to ask before believing one.