Introducing MiMo-V2.6 series
Frontier intelligence, all the modalities, built in public.
Two natively omnimodal models — MiMo-V2.6-Pro and MiMo-V2.6-Flash — trained with reinforcement learning at scale and released with everything: code, environments, and technical report.
Scaling RL, Fully Open-Sourced
Today we're releasing and open-sourcing the MiMo-V2.6 series — a key step in our exploration of the RSI path.
This marks a key step in our exploration of the RSI path: scaling RL compute on verifiable, complex tasks, so the model can continuously expand its capability frontier through exploration and feedback.
MiMo-V2.6-Pro is our most capable model to date, while MiMo-V2.6-Flash strikes the best balance between intelligence, efficiency, and cost. We're also rolling out MiMo-V2.6-Pro-UltraSpeed, delivering up to 20× faster output speed at the same quality.
— strongest open-source model
coding score (Pro)
(Pro)
at same quality
Higher Intelligence, Same Price
MiMo-V2.6-Pro scores 46.32 on the Artificial Analysis Intelligence Index — surpassing Kimi K3 and Qwen3.8 Max to become the strongest open-source model to date.
Critically, the MiMo-V2.6 series keeps the API pricing of the V2.5 series. Higher intelligence at the same price pushes the Pareto frontier of intelligence versus cost outward once again. Pro sits on the Pareto line, inside the most attractive quadrant (source: Artificial Analysis).
MiMo-V2.6-Pro
Most capable model — frontier coding, knowledge work, visual tasks
MiMo-V2.6-Flash
Best balance of intelligence, efficiency, and cost
How MiMo-V2.6 Stacks Up
Across coding, general reasoning, and visual tasks — measured against the frontier.
Coding
General
Visual & Cyber
Six Days, 750K Trajectories, Streamed Live
We worked through the research and engineering problems of RL training — and streamed the production run live as it happened.
In under six days, MiMo-V2.6-Flash and MiMo-V2.6-Pro each completed 30 RL steps over roughly 750,000 trajectories. Average pass rate on training tasks rose by 25% and 12% in relative terms. On DeepSWE v1.1, a held-out long-horizon software engineering benchmark, scores rose by about 17 points (48.8 → 65.7) and about 14 points (58.4 → 72.6). RL proved sample-efficient, kept improving throughout the run, and generalized beyond the training distribution.
We scaled RL compute along three axes:
- Larger batches and higher throughput: large batches on a fully asynchronous architecture, 1,568 samples per update, training at up to 1M context length, 3.5–3.7B tokens per step.
- More tasks and richer environments: a multi-task training suite spanning coding, general agents, visual and cyber, mixed across harnesses so gains in one capability reinforce the others.
- More grader compute: relative comparison within each group gives long-horizon RL tasks more precise, diverse reward signals and steers the model toward shorter paths and fewer tokens per task.
Beyond Software Engineering
Given a goal, a harness and time, MiMo takes on real productive work — across six domains.
Game Development
Builds playable games from natural language descriptions — physics, rendering, and game logic in a single pass.
3D Modeling
Generates and manipulates 3D assets programmatically — meshes, materials, and scene composition from text.
Embodied Simulation
Controls virtual robots in physics-based environments, completing complex manipulation and navigation tasks.
Visual & Design
Frontend design, presentation design, and visual content — generated with production-level polish.
Video & Music
Video clip generation and music composition — omnimodal output spanning visual and audio domains.
Research
Co-scientist for materials research and formalizing mathematical proofs — pushing into scientific discovery.
Full Benchmark Results
All MiMo-V2.6 results from Xiaomi's published benchmarks; comparison figures from respective public leaderboards, September 2026.
| Benchmark | V2.6 Pro | V2.6 Flash | Opus 5 | Fable 5.1 | GPT 6 Astra | V2.5 Pro |
|---|---|---|---|---|---|---|
| Coding | ||||||
| DeepSWE v1.1 | 71.9 | 67.9 | 74.0 | — | 74.0 | 19.0 |
| ProgramBench | 26.5 | 26.0 | 37.0 | 33.0 | — | 12.5 |
| MiMo Code Bench in-house | 63.2 | 61.2 | 68.6 | — | 61.4 | 40.4 |
| General | ||||||
| GDPVal 2.1 (AA) | 1673 | — | 1708 | 1735 | 1542 | 1107 |
| AutomationBench | 53.1 | 52.3 | 50.3 | — | 52.0 | 16.0 |
| Agents' Last Exam | 31.6 | 27.6 | 31.6 | 25.7 | 34.2 | 13.2 |
| Terminal Bench 4.0 | 34.9 | 28.8 | 49.0 | 55.1 | 59.6 | 1.5 |
| Visual & Cyber | ||||||
| MiMo Visual Coding in-house | 72.3 | 71.5 | 70.0 | 74.4 | 82.2 | — |
| CyberGym | 94.0 | 95.1 | — | — | — | 40.0 |
| ExploitGym | 17.8 | 6.0 | 22.1 | 30.4 | 42.4 | 0.1 |
Terminal Bench 4.0 figures: GPT 6 Astra leads at 59.6; Claude Fable 5.1 at 55.1. All results as published by Xiaomi, September 22, 2026.
Get MiMo-V2.6 Today
Desktop app, API, open weights, and the full technical report — all available now.
Sources
- Xiaomi MiMo — Introducing MiMo-V2.6 series (September 22, 2026). Primary announcement page.
- MiMo Desktop. Try MiMo-V2.6 directly.
- MiMo API. Developer API access.
- HuggingFace — XiaomiMiMo/mimo-v26. Open-source model weights and code.
- MiMo V2.6 Technical Report. Full RL training methodology and results.
- Artificial Analysis. Intelligence Index v4.3, September 2026.