Stanford MARVL

MMBU
Challenge

Advance visual perception in biomedical multimodal models.

Massive Multimodal
Biomedical Understanding

Massive Multimodal
Biomedical Understanding

The MMBU Challenge aims to improve how multimodal models perceive biomedical images across modalities, scales, and anatomical regions.

Today’s multimodal large language models (MLLMs) can produce seemingly impressive biomedical answers. But beneath these capabilities lies a fundamental weakness: they may not even understand what they are looking at. In biomedical MLLMs, visual perception remains the primary bottleneck for reliable downstream reasoning. So we built a challenge to put biomedical visual understanding to the test.

Example MMBU open VQA tasks

Pick a track.
Submit one model.

A checkpoint can enter only one track. Models over 27B active parameters need an API key for evaluation.

Resources &
Prizes

Teams receive a public development set, weekly office hours with the organizers, special credits from Anthropic and GXL, and cash prizes. The top entry in each track will also be invited to contribute to the MMBU Challenge technical report.

Thanks to our partners

Official sponsors of the MMBU Challenge. Collaborator institutions are named in the FAQ, in text only.

Common
questions

MMBU evaluates open-ended visual question answering on a peer-reviewed biomedical benchmark. Models must recognize, localize, and characterize visual features — then identify imaging metadata in a separate pass, so they cannot shortcut the answer.