*Equal contribution
Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, from detecting lesions in chest X-rays to profiling cellular features in microscopy. Realizing this potential, however, requires robust and fine-grained visual perception. Models need to correctly interpret subtle features in images, and they must do so across diverse biomedical modalities, scales, and contexts. Nevertheless, current benchmarks remain limited. To address these gaps, we introduce the Massive Multimodal Biomedical Understanding (MMBU) benchmark. It is the largest biomedical vision and language benchmark to date, covering 35 submodalities with rich structured metadata. It includes both open and closed versions of ungrounded classification, grounded classification, and object detection, enabling systematic evaluation of model performance across biological scales, clinical settings, and imaging modalities. Evaluating open-weight and frontier VLMs, we find that while medical adaptation provides measurable gains for some models, the high accuracy often reported on established benchmarks can mask deficiencies in visual perception and domain generalization.
The MMBU public set is officially releasing on 10/1/2026 after a 4-month embargo following publication in conjunction with the MMBU Challenge. This should be used for internal benchmarking and reporting unverified results. MMBU authors maintain an MMBU private set that can be used for verified benchmarking. Hugging Face links can be found below.
MMBU Public set: https://huggingface.co/datasets/mmbu-challenge/mmbu-public
410 curated datasets spanning 11 modalities, 35 submodalities, 20 specimens, and 95 body-part regions — with 13 structured metadata fields per sample including provenance, modality, stain, specimen, and institution. Stats shown in Figure 2
Ungrounded classification, grounded classification (from segmentation masks and bounding boxes), and object detection — each in open- and closed-ended VQA formats with three context levels (no context, modality, full context).
Frontier and Open VLMs including Claude Opus 5.5, OpenAI GPT-6-Astra, Qwen3.8-27B, MedGemma, Lingshu, OctoMed, LLaVA-Med, InternVL3.5— plus their general-domain base counterparts where applicable.
Multidisciplinary expert taxonomy → dataset discovery across Zenodo, Kaggle, and Hugging Face → metadata standardization → metadata-driven question construction with human-in-the-loop validation.
MMBU evaluates vision-language models on ungrounded classification, grounded classification, and object detection, in open-ended and closed-ended VQA, at three context levels: no context, modality, and full context. Questions span medical domains, modalities, submodalities, specimens, and body regions, with structured metadata on each sample. Figure 2 summarizes that composition.
From MMBU, we curated a hard subset consisting of full-context open-ended VQA and split that into public and private sets. The public set is released for internal benchmarking and reporting unverified results, and the private set is held for verified results. The plots are the same UMAP of image embeddings: public versus private on the left, and imaging modality on the right. Hover a point to see its split and modality.
From the MMBU Hard set, we benchmarked the last frontier models as of September 2026 with mean accuracy shown below. Earlier models from MMBU get <10% on the MMBU Hard set. Each latest frontier model has its own color, held constant across both charts. Hover a bar to compare that model on the other split. More models benchmarked on MMBU are shown below.
~3k open-ended questions across all domains and modalities in MMBU released publicly on Hugging Face
~9k open-ended questions across all domains and modalities in MMBU held privately
Due to the fast pace of development, original models in MMBU are used for downstream analysis across all of MMBU, but frontier models are mainly displayed on the website due to their advanced performance. On the MMBU hard subset used for the public and private set, all original MMBU models score <10%. The results shown here include MMBU for all 3 context questions with both open-ended and closed-ended VQA.
We evaluated frontier and open state-of-the-art vision-language models (VLMs) across diverse biomedical imaging tasks and found that performance remains surprisingly limited. Although some models achieve strong results on individual benchmarks, no model consistently performs well across the full spectrum of biomedical perception tasks (Figure 3).
Across classification tasks, performance drops dramatically when models must generate answers freely rather than choose from predefined options (Figure 3). This large gap suggests that many current biomedical VLMs rely on answer-choice cues instead of robust visual understanding.
While several models can correctly classify findings when provided with a region of interest, they struggle to accurately localize those findings themselves (Figure 3). Detection performance remains near random across most models, highlighting significant limitations in spatial reasoning and visual grounding.
Different models excel on different tasks (Figure 3). GPT-based models perform strongly on some classification tasks, InternVL excels on certain grounded reasoning tasks, and Qwen-based models achieve the highest performance on segmentation-derived classification. However, no model consistently ranks first across all evaluation settings.
Models trained specifically on medical data not always outperfom their general-purpose counterparts (Figure 5A). The models that do (OctoMed and MedGemma), show modest gain. In many cases, medically adapted models perform similarly or worse to their base versions, suggesting that current adaptation strategies only partially address the challenges of biomedical visual understanding.
The success of medical adaptation appears to depend more on the amount of biomedical training data than on model size alone (Figure 5B). Models trained on larger medical datasets tend to show more consistent improvements across tasks.
Many specialized medical models achieve strong results on widely used benchmarks such as PathVQA, VQA-RAD, and SLAKE (Figure 6). However, these gains frequently fail to transfer to broader and more diverse biomedical settings, suggesting that current benchmarks may overestimate real-world capability.
The MMBU Challenge evaluates models through open-ended VQA on MMBU. Three tracks cover frontier models, medically adapted models, and efficient models under 4B active parameters. Teams receive a public development set, and final rankings use a private held-out set. Registration is open through September 30, 2026, and the competition runs through the end of 2026. Sponsors Anthropic, GXL, Stanford AI Lab, Biohub, and Highlanders are providing over $100k in compute and prizes.




Our results indicate that biomedical visual perception remains an open challenge. Despite rapid progress in multimodal AI, current models still struggle with robust visual understanding, grounded reasoning, and spatial localization across diverse biomedical domains. MMBU provides a comprehensive framework for measuring these limitations and guiding the development of more reliable biomedical AI systems.
@article{dcunha2026mmbu,
title = {MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models},
author = {D'Cunha, Ryan and Lozano, Alejandro and Sun, Xiaoxiao and Vela Jarquin, Daniel and Sun, Min Woo and Aklilu, Josiah and Burgess, James and Zhang, Yuhui and Nayebi, Ryan and Avila Robayo, Paola and Ye, Jin and Hu, Ming and Deng, Zhongying and He, Junjun and Chen, Xin and Yao, Yue and Tibshirani, Robert and Nirschl, Jeffrey J. and Yeung-Levy, Serena},
journal = {arXiv preprint arXiv:2606.06696},
year = {2026},
eprint = {2606.06696},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2606.06696}
}