Clinical Evals

How MedScribe works

Vals AI · 100 SOAP-note cases · updated August 17, 2026

MedScribe scores the note a model writes after a clinical visit. Vals AI built it with Protege and released it in February 2026. The set is 100 cases, each graded against a rubric for documentation quality and compliance rather than compared against one reference note. That choice reflects the task, since two correct notes for the same visit can be worded very differently.

publisherVals AI (dataset with Protege)
released2026-02
size100 rubric-scored SOAP-note cases
scalepercentage accuracy 0-100, higher better

SOAP notes and rubric grading

The output format is a SOAP note: subjective, objective, assessment, plan. Rubric grading means the note is checked criterion by criterion, so credit depends on whether specific content is present and correct, not on similarity to a model answer.

The criteria cover two things at once. Documentation quality is whether the note captures the visit usefully. Compliance is whether it meets the requirements a note is expected to satisfy. A note can read well and still lose points on the second.

A small set, run in one place

100 cases is a small denominator. Each case carries one percent of the score, so a gap of a few points between two rows can come down to a handful of cases going differently. Vals AI runs every model itself, and the board carried 84 models as of its August 15, 2026 update. Self running keeps the configuration constant across rows, which is what makes comparison within the board meaningful.

Reading a score near the ceiling

Vals AI reports that top scores cluster near 90, so the leading group sits close to the maximum. When a board compresses at the top, ordering carries less information than the size of the gap between the top group and the rest. The site also does not always display exact decimals below first place, which limits fine comparisons among close rows. The paired coding board, /benchmarks/medcode, sits lower, and Vals AI notes that coding accuracy lags documentation quality; treat note writing and coding as separate skills.

Limits

A rubric encodes one set of documentation expectations, and note standards vary by specialty, institution, and jurisdiction, so compliance credit here is not a compliance guarantee anywhere specific. With 100 cases the resolution is limited, and near ties should be read as ties. Scores clustered near the ceiling also mean the benchmark has less room to separate strong models than it had at release. A note that scores well against a rubric is still not the same as a note a clinician would sign without editing, which is the outcome a deployment cares about.

Current MedScribe scores are tracked on the results index at clinicalbenchmarks.ai.

The grading machinery behind pages like this one is covered in the rubric grading and model graders guides.