MedHELM: Holistic Evaluation of Language Models for Medical Tasks

MedHELM is an open and extensible framework for evaluating large language models on tasks that reflect healthcare work. It was developed through a Stanford-led, multi-institutional collaboration and introduced in a 2026 Nature Medicine paper.

Rather than asking only whether an LLM can answer medical examination questions, MedHELM asks a more practical question: how well can a model perform specific tasks resembling those undertaken across healthcare?

What MedHELM evaluates?

Its clinician-validated taxonomy contains 121 tasks organised into 22 subcategories and five areas:

  • Clinical decision support
  • Clinical note generation
  • Patient communication and education
  • Medical research assistance
  • Administration and workflow

The original study operationalised this taxonomy through 37 benchmarks drawn from public, restricted-access and private datasets. The live MedHELM suite continues to evolve as models, datasets and evaluation methods change.

How it works?

MedHELM combines several forms of evaluation. For structured tasks, such as multiple-choice questions, calculations and classifications, it uses task-appropriate measures such as exact-match accuracy and F1 scores. For open-ended outputs, including clinical notes and patient communications, it uses an ensemble of three language models—the LLM-jury—to assess dimensions such as accuracy, completeness and clarity. The framework also supports comparison of performance and computational cost.

This produces a more detailed picture than a single overall accuracy score. A model may perform well in documentation or patient communication while remaining unreliable for calculations, administrative tasks or particular forms of clinical reasoning.

Why MedHELM is useful?

There is no universally “best” healthcare LLM. Suitability depends on the task, the population, the clinical context, the consequences of error and the controls surrounding its use. MedHELM is valuable because it:

  • Organises evaluation around recognisable healthcare functions.
  • Moves beyond examination-style medical knowledge tests.
  • Supports consistent comparison between models.
  • Includes structured and free-text tasks.
  • Incorporates real clinical data in parts of its benchmark suite.
  • Makes performance and cost trade-offs more visible.
  • Provides an extensible and reproducible evaluation infrastructure.

How MedHELM informs my review approach?

I draw on MedHELM to connect a proposed LLM use case with the relevant healthcare task and the evidence available for that task. This involves:

  1. Defining the intended use, users, workflow and consequences of error.
  2. Mapping the use case to relevant MedHELM categories and tasks.
  3. Reviewing applicable benchmark evidence and known limitations.
  4. Supplementing benchmark results with local data, policies and workflow-specific scenarios.
  5. Translating the findings into governance controls, decision criteria and monitoring requirements.

MedHELM therefore informs model evaluation, but it does not replace local clinical validation or sociotechnical assessment.

Governance value

MedHELM can support more defensible governance decisions by replacing general claims about model capability with task-specific evidence. Its findings can contribute to:

  • Model selection and vendor due diligence.
  • Definition of approval and exclusion criteria.
  • Identification of tasks requiring human review.
  • Documentation of the rationale for deployment decisions.
  • Establishment of pre-deployment performance baselines.
  • Re-evaluation following model, prompt or workflow changes.
  • Communication of limitations to clinicians, executives and oversight committees.

A benchmark score is evidence, not assurance. Assurance requires the evaluation evidence to be connected to accountable owners, decision thresholds, risk controls and ongoing monitoring.

Important limitations

MedHELM evaluates models on benchmark tasks; it does not reproduce the complete environment in which an AI system will operate. It cannot, on its own, demonstrate local workflow fit, usability, regulatory compliance, privacy protection, equitable performance, safe human–AI interaction or post-deployment effectiveness. Its LLM-jury also requires further validation across a wider range of tasks and clinical contexts.

For this reason, I position MedHELM as one component of a broader sociotechnical evaluation—not as a certification that an LLM is safe for clinical use.

Related resources

Last reviewed: August 2026