Correspondence
Do findings at different scales describe the same underlying phenomenon?
ScaleReasoner-R1:
1 National University of Singapore, Singapore2 PuzzleLogic Pte Ltd, Singapore3 Department of Pathology, Fujian Medical University Cancer Hospital & Fujian Cancer Hospital, Fuzhou, China
01 / Overview
Pathological images are inherently multi-scale, requiring pathologists to integrate evidence from global tissue architecture at low magnification to cellular morphology at higher magnification for accurate diagnosis.
While existing pathological datasets for vision-language models (VLMs) include various scales, they often lack explicit cross-scale reasoning objectives. This limitation prevents VLMs from capturing essential cross-scale representations and learning evidence-based reasoning.
To bridge this gap, we introduce the first cross-scale training and evaluation paradigm that formulates pathology interpretation as multi-magnification reasoning:
02 / Scale-VQA
A benchmark designed to require images, with 4,685 multiple-choice questions across 15 organs and five clinically aligned reasoning dimensions.
Do findings at different scales describe the same underlying phenomenon?
Does evidence at one scale confirm or contradict the interpretation at another?
Which tissue compartment does a finding belong to, based on cross-scale evidence?
What process explains the relationship between findings across scales?
What diagnosis is supported by integrating evidence across scales?
Leakage-aware curation
Questions should depend on the images. Our curation pipeline screens out questions that can be answered using linguistic or biomedical priors alone.
Expert annotations become per-scale evidence sets, with visual-grounding and scale-dependency constraints.
Gemini 3 Pro and Qwen3-Max attempt questions without images. If either answers correctly, constraints are tightened and questions are regenerated.
Senior pathologists review the final questions for visually grounded answers and clinically plausible distractors.
03 / ScaleReasoner-R1
Initialized from Patho-R1-7B, ScaleReasoner-R1 is fine-tuned on Scale-VQA using GRPO. The model takes multi-scale images, a question, and answer options, then produces a structured reasoning trace and a final answer.
04 / Evaluation
ScaleReasoner-R1 leads across all five reasoning dimensions among the compared models and improves overall accuracy on PathMMU.

| Model | Correspondence | Confirmation | Localization | Explanation | Diagnosis | Average |
|---|---|---|---|---|---|---|
| Qwen2.5-VL-7B | 41.79 | 47.26 | 71.14 | 44.28 | 63.68 | 53.63 |
| Gemini 3 Flash | 48.26 | 58.71 | 71.64 | 53.73 | 72.14 | 60.90 |
| GPT-5.2 | 47.76 | 59.70 | 74.13 | 45.77 | 65.17 | 58.51 |
| LLaVA-Med-7B | 25.87 | 16.92 | 28.36 | 20.90 | 24.88 | 23.38 |
| Quilt-LLaVA | 32.34 | 14.43 | 45.27 | 30.35 | 29.35 | 30.35 |
| CLOVER | 37.31 | 61.69 | 73.13 | 46.27 | 65.77 | 56.82 |
| Patho-R1 | 31.84 | 40.30 | 59.70 | 56.22 | 68.66 | 51.34 |
| ScaleReasoner-R1 | 80.60 | 89.05 | 84.58 | 76.12 | 84.08 | 82.89 |
| Model | Val overall | Test overall | PubMed | SocialPath | EduContent | Atlas | PathCLS | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Tiny | All | Tiny | All | Tiny | All | Tiny | All | Tiny | All | Tiny | All | ||
| Qwen2.5-VL-7B | 44.3 | 48.2 | 44.4 | 53.4 | 50.5 | 50.0 | 47.7 | 59.6 | 48.9 | 48.6 | 46.8 | 29.4 | 28.0 |
| Gemini Pro Vision† | 41.9 | 42.8 | 42.7 | 43.8 | 44.9 | 42.4 | 42.0 | 43.5 | 43.7 | 49.5 | 49.4 | 32.8 | 34.7 |
| GPT-4V† | 49.3 | 53.9 | 49.8 | 59.4 | 53.5 | 58.7 | 53.9 | 60.4 | 53.6 | 48.1 | 52.8 | 36.2 | 33.8 |
| LLaVA-Med-7B | 17.5 | 22.2 | 22.7 | 24.6 | 25.1 | 23.4 | 21.1 | 25.1 | 23.6 | 17.8 | 24.5 | 20.3 | 19.2 |
| HuatuoGPT-V-7B | 43.2 | 43.8 | 41.2 | 49.1 | 47.7 | 52.3 | 46.0 | 52.9 | 46.1 | 44.7 | 46.9 | 19.8 | 19.2 |
| Lingshu-7B | 51.9 | 56.0 | 53.7 | 60.5 | 57.6 | 61.6 | 58.1 | 69.4 | 58.7 | 58.1 | 62.5 | 30.4 | 31.6 |
| Quilt-LLaVA | 33.8 | 32.3 | 31.9 | 30.3 | 34.2 | 26.8 | 34.6 | 35.7 | 33.3 | 41.8 | 33.3 | 27.1 | 24.0 |
| CLOVER | 53.1 | 60.6 | 56.6 | 68.0 | 59.8 | 67.6 | 60.5 | 68.6 | 62.2 | 62.0 | 64.1 | 36.7 | 36.6 |
| Patho-R1 | 62.3 | 64.8 | 62.6 | 68.7 | 64.9 | 63.9 | 65.7 | 71.0 | 66.1 | 78.4 | 73.5 | 41.8 | 42.7 |
| ScaleReasoner-R1 | 62.4 | 66.2 | 63.8 | 71.2 | 67.3 | 67.6 | 67.6 | 74.9 | 69.2 | 79.3 | 76.2 | 37.9 | 38.5 |
Tiny and All denote Test-Tiny and the full test set; subset scores are shown for both splits. Bold marks the best model score in each column. † Results taken directly from the original PathMMU paper.
Qualitative evaluation
Explore multi-magnification evidence and model responses for correspondence, confirmation, localization, explanation, and diagnosis.
05 / Citation
If you find ScaleReasoner-R1 or Scale-VQA useful in your research, please cite our paper.
@article{phan2026enhancing,
title={Enhancing Pathological VLMs with Cross-scale Reasoning},
author={Phan, Chi and Zhang, Tianyi and Xue, Qiaochu and Wu, Yufeng and Hu, Dan and Liu, Zeyu and Wang, Sudong and Jin, Yueming},
journal={arXiv preprint arXiv:2606.17412},
year={2026}
}
This work was supported by the Ministry of Education, Singapore, under the Tier 1 grant (24-1250-P0001) and Tier 2 grant (T2EP20224-0028), and by PuzzleLogic Pte Ltd, Singapore.
We thank the contributors to verl, LLaMA-Factory, vLLM, and Patho-R1, along with the developers of the open-source models used in our experiments.