Clinical Evals

How Health Optimization Bench works

healthoptimizationbench.com · 89 tasks · updated August 17, 2026

Health Optimization Bench asks frontier models hard questions in preventive and optimization medicine, where the right answer depends on evidence that moved recently. The v1 release set is 89 tasks built around incretin therapeutics. Every task is written against a primary source, audited by model families that did not author it, and graded blind by a panel of independent families. Scores are rubric credit from 0 to 100.

publisherHealth Optimization Bench
released2026-08
size89 released tasks (v1 evidence suite)
scale0-100 rubric credit, higher better

Cross-family authoring

Every task is written against a primary source: a specific trial, guideline, or label rather than a general recollection. The family that drafts a task does not get the last word on it. Other model families audit the task and its rubric, checking that the stated answer follows from the cited source.

Splitting authoring from auditing keeps one family's blind spots from becoming the benchmark's blind spots. A question that only its author considers well posed does not survive to the release set.

Blind panel grading

Three model families grade each response, and the family that authored a task never grades it. Graders see the response and the rubric without knowing which model wrote the answer. Credit is combined across the panel, so no single grader's idiosyncrasy sets a score by itself.

Every published score carries a 95 percent bootstrap confidence interval. The interval is there so you can see when two models are separated by the benchmark and when they are separated by sampling. At 89 tasks those intervals are wide enough to matter, and the site shows them rather than leaving them implicit.

Freshness as the difficulty axis

Most medical benchmarks test knowledge that has been settled for years. This one targets questions where the evidence base changed, which is why the v1 suite is scoped to one fast-moving therapeutic area instead of spread across medicine. A model can hold the textbook well and still miss here, because the failure being measured is a stale prior rather than a gap in knowledge.

Limits

89 tasks is a small set, and confidence intervals of the width the site publishes are the honest consequence, so treat close rows as ties. The v1 suite covers one therapeutic area, which means a score describes handling of incretin evidence and not preventive medicine generally. Authoring and grading are both done by models with human-checked primary sources as the anchor, a design that controls for single-family bias but not for bias shared across families. Freshness-dependent tasks also age: a question that separates models today becomes ordinary once the evidence settles into training data, so the set has to be rebuilt to keep working.

Current scores and their confidence intervals are published at healthoptimizationbench.com, and the benchmark is indexed with the others at clinicalbenchmarks.ai.

The grading machinery behind pages like this one is covered in the rubric grading and model graders guides.