Embodied Reasoning
Across 16 multimodal perception and understanding benchmarks, UnifoLM-ER-1 leads open-source models on seven and delivers overall performance comparable to leading proprietary models. Built on Qwen3-VL-4B, UnifoLM-ER-1 is trained on more than 5 million samples spanning image point prediction, object detection, multi-image reasoning, 2D trajectory prediction, 3D object detection, and multi-image spatial question answering. These data are co-trained with general image–text data, preserving broad vision-language capabilities while substantially improving spatial understanding and reasoning in embodied environments.
Benchmark results
| Model | Open Source | Spatial Understanding | Multimodal Understanding | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| RoboVQA | Ego-Plan2 | RefSpatial- Bench | Where2Place | Pixmo-Point | BLINK | CV-Bench | EmbSpatial | RoboSpatial | SAT | VSI-Bench | VSR | ERQA | RealWorld QA | MME | MMMU_VAL | ||
| UnifoLM-ER-1-4B | Yes | 62.4 | 55.1 | 61.7 | 82.0 | 73.8 | 93.4† | 88.6 | 88.9 | 73.1 | 76.0 | 54.2 | 88.1 | 50.0 | 69.8 | 2223.3 | 54.7 |
| RoboBrain2.0-7B* | Yes | 30.0 | 33.23 | 42.2 | 63.6 | 54.7 | 83.9 | 85.7 | 76.3 | 54.2 | 75.3 | 36.1 | 84.0 | / | 69.4 | 2057.4 | 44.4 |
| Robix-7B* | No | 63.6 | / | / | 41.9 | 29.5 | 87.6 | 86.5 | 77.4 | / | 71.1 | 44.6 | 83.3 | 42.5 | 70.7 | 2332.8 | / |
| Pelican-7B* | Yes | 31.8 | 33.7 | 22.3 | 57.3 | 20.4 | / | 79.4 | 73.2 | 57.5 | 52.0 | 52.8 | 82.2 | 39.8 | 69.3 | 2141.9 | 51.1 |
| Cosmos-R1-7B* | Yes | 38.8 | 26.0 | 5.6 | 2.9 | 8.2 | / | 76.7 | 68.9 | 42.4 | 82.7 | 25.4 | 82.4 | / | 67.6 | 2157.4 | 37.4 |
| Cosmos3-Super-64B* | Yes | / | / | 57.0 | 71.0 | / | 90.3† | 88.0 | / | 70.0 | / | 60.9 | / | 51.2 | / | / | / |
| Qwen3-VL-4B* | Yes | 47.7 | 40.7 | 46.6 | 63.0 | 48.3 | 85.0† | 85.1 | 79.6 | 61.7 | 68.7 | 59.3 | 81.6 | 41.3 | 71.0 | 2325.2 | 57.8 |
| Qwen3-VL-8B* | Yes | 43.3 | 49.7 | 54.2 | 61.9 | 51.0 | 73.8† | 86.2 | 78.5 | 66.9 | 67.3 | 59.4 | 83.2 | 45.8 | 70.6 | 2412.5 | 62.3 |
| Embodied-R1-3B* | Yes | 51.8 | 26.5 | 39.7 | 69.5 | 49.4 | 78.5† | 82.7 | 67.4 | 47.4 | 76.3 | 26.6 | / | 35.2 | / | / | / |
| Embodied-R1.5-8B* | Yes | 61.0 | 53.8 | 54.2 | 74.0 | 64.8 | 83.0† | 86.9 | 78.1 | 69.7 | 74.7 | 56.1 | / | 46.0 | / | / | / |
| Molmo2-ER-4B* | Yes | / | / | 52.5 | 54.0 | / | 85.7† | 87.8 | 78.8 | / | 78.0 | 74.5 | / | 46.8 | / | / | / |
| Hy-Embodied-VLM-1.0 30B-A3B* |
Yes | / | 49.6 | 53.4 | 65.0 | 64.6 | 87.3† | 89.7 | 82.7 | 69.4 | 78.0 | / | / | 60.8 | / | / | / |
| MiMo-Emb-7B* | Yes | 62.0 | 43.0 | 48.0 | 63.6 | 42.35 | 81.3 | 88.2 | 76.2 | 61.7 | 78.6 | 48.5 | 79.0 | 46.7 | 66.3 | 2320.8 | 26.4 |
| Thinker-4B* | Yes | 62.7 | 63.7 | 61.0 | 72.0 | 57.4 | 84.6 | 86.3 | 80.2 | 70.8 | 72.7 | 65.4 | 81.5 | / | 71.9 | 2323.4 | 46.2 |
| Wall-OSS-0.5-3B* | Yes | / | / | / | 15.0 | / | / | / | / | / | / | / | / | 33 | 44 | / | / |
| Lumo-1-Stage1-7B* | No | / | / | 51.0 | 69.1 | / | 82.4 | 86.4 | 75.6 | 62.6 | 74.7 | / | / | / | / | / | / |
| Gemini-ER 2‡ | No | / | / | 35.4 | / | / | 90.6† | 90.4 | 81.4 | 51.1 | / | / | / | 71.0 | / | / | / |
| Gemini-ER 1.5‡ | No | / | / | 41.8 | 48.0 | / | / | 83.6 | 73.4 | 57.7 | 62.0 | 39.9 | / | 47.0 | / | / | / |
| Gemini 2.5 Pro‡ | No | / | / | 33.6 | 37.0 | / | 88.6† | 85.9 | 78.0 | 71.3 | 74.7 | 51.1 | / | 56.0 | / | / | / |
| Gemini 2.5 Flash‡ | No | / | / | 41.2 | 48.0 | / | 80.3† | 85.5 | 76.2 | 73.4 | 73.3 | 45.3 | / | 47.5 | / | / | / |
| Gemini 3.1 Pro‡ | No | / | / | 70.0 | 61.0 | / | 86.1† | 88.6 | / | 65.1 | / | 47.5 | / | 65.2 | / | / | / |
| GPT-5.6-sol‡ | No | / | / | 58.3 | 51.1 | / | 85.6† | 85.2 | 80.7 | 66.8 | 21.3 | / | / | 64.8 | / | / | / |
| GPT-6-Astra‡ | No | / | / | 79.6 | 69.0 | / | 90.4† | 87.3 | 83.3 | 73.4 | 31.3 | / | / | 77.7 | / | / | / |
Results are sourced from the models' official technical reports or publicly available papers.
Results were obtained through tests using the models' official APIs.
BLINK scores are averaged over the Relative Depth and Spatial Relation subtasks only; all reported results were obtained in our own testing.

