⚡ Open Source · Frontier RL · All Modalities SEPTEMBER 22, 2026

Introducing MiMo-V2.6 series

Frontier intelligence, all the modalities, built in public.

Two natively omnimodal models — MiMo-V2.6-Pro and MiMo-V2.6-Flash — trained with reinforcement learning at scale and released with everything: code, environments, and technical report.

Built in public · Made with ❤ by Xiaomi

Scaling RL, Fully Open-Sourced

Today we're releasing and open-sourcing the MiMo-V2.6 series — a key step in our exploration of the RSI path.

This marks a key step in our exploration of the RSI path: scaling RL compute on verifiable, complex tasks, so the model can continuously expand its capability frontier through exploration and feedback.

MiMo-V2.6-Pro is our most capable model to date, while MiMo-V2.6-Flash strikes the best balance between intelligence, efficiency, and cost. We're also rolling out MiMo-V2.6-Pro-UltraSpeed, delivering up to 20× faster output speed at the same quality.

46.32 AA Intelligence Index v4.3
— strongest open-source model
71.9 DeepSWE v1.1
coding score (Pro)
1,673 GDPval-AA Elo
(Pro)
20× UltraSpeed output
at same quality

Higher Intelligence, Same Price

MiMo-V2.6-Pro scores 46.32 on the Artificial Analysis Intelligence Index — surpassing Kimi K3 and Qwen3.8 Max to become the strongest open-source model to date.

Critically, the MiMo-V2.6 series keeps the API pricing of the V2.5 series. Higher intelligence at the same price pushes the Pareto frontier of intelligence versus cost outward once again. Pro sits on the Pareto line, inside the most attractive quadrant (source: Artificial Analysis).

🧠

MiMo-V2.6-Pro

Most capable model — frontier coding, knowledge work, visual tasks

Same as V2.5 pricing
Higher intelligence, unchanged cost

MiMo-V2.6-Flash

Best balance of intelligence, efficiency, and cost

Same as V2.5 pricing
Higher intelligence, unchanged cost

How MiMo-V2.6 Stacks Up

Across coding, general reasoning, and visual tasks — measured against the frontier.

Coding

DeepSWE v1.1
DeepSeek V4.1 Flash74.2
Claude Opus 574.0
GPT 6 Astra74.0
MiMo-V2.6-Pro71.9
Claude Fable 570.0
MiMo-V2.6-Flash67.9
ProgramBench
Claude Opus 537.0
Claude Fable 533.0
MiMo-V2.6-Pro26.5
MiMo-V2.6-Flash26.0
GPT 5.6 Sol25.0
MiMo Code Bench in-house
Claude Opus 568.6
MiMo-V2.6-Pro63.2
GPT 6 Astra61.4
MiMo-V2.6-Flash61.2

General

GDPVal 2.1 (AA)
Claude Fable 5.11735
Claude Opus 51708
MiMo-V2.6-Pro1673
DeepSeek V4.1 Flash1600
Automation Bench v1.0.6
DeepSeek V4.1 Flash54.8
MiMo-V2.6-Pro53.1
MiMo-V2.6-Flash52.3
GPT 6 Astra52.0
Agents' Last Exam
GPT 6 Astra34.2
DeepSeek V4.1 Flash31.8
MiMo-V2.6-Pro31.6
Claude Opus 531.6

Visual & Cyber

MiMo Visual Coding in-house
GPT 6 Astra82.2
Claude Fable 5.174.4
MiMo-V2.6-Pro72.3
MiMo-V2.6-Flash71.5
Claude Opus 570.0
CyberGym
MiMo-V2.6-Flash95.1
MiMo-V2.6-Pro94.0
DeepSeek V4.1 Flash88.1
GLM 5.384.5
ExploitGym
GPT 6 Astra42.4
Claude Fable 5.130.4
Claude Opus 522.1
MiMo-V2.6-Pro17.8

Six Days, 750K Trajectories, Streamed Live

We worked through the research and engineering problems of RL training — and streamed the production run live as it happened.

In under six days, MiMo-V2.6-Flash and MiMo-V2.6-Pro each completed 30 RL steps over roughly 750,000 trajectories. Average pass rate on training tasks rose by 25% and 12% in relative terms. On DeepSWE v1.1, a held-out long-horizon software engineering benchmark, scores rose by about 17 points (48.8 → 65.7) and about 14 points (58.4 → 72.6). RL proved sample-efficient, kept improving throughout the run, and generalized beyond the training distribution.

$0.85M Flash training cost
$2.62M Pro training cost
30 steps RL steps in under 6 days

We scaled RL compute along three axes:

  • Larger batches and higher throughput: large batches on a fully asynchronous architecture, 1,568 samples per update, training at up to 1M context length, 3.5–3.7B tokens per step.
  • More tasks and richer environments: a multi-task training suite spanning coding, general agents, visual and cyber, mixed across harnesses so gains in one capability reinforce the others.
  • More grader compute: relative comparison within each group gives long-horizon RL tasks more precise, diverse reward signals and steers the model toward shorter paths and fewer tokens per task.
Fully open-sourced: the full technical report, the training environments, and the RL code — so researchers can reproduce and verify these results and join us in exploring what scaled RL and model self-improvement can do next.

Beyond Software Engineering

Given a goal, a harness and time, MiMo takes on real productive work — across six domains.

🎮

Game Development

Builds playable games from natural language descriptions — physics, rendering, and game logic in a single pass.

🧊

3D Modeling

Generates and manipulates 3D assets programmatically — meshes, materials, and scene composition from text.

🤖

Embodied Simulation

Controls virtual robots in physics-based environments, completing complex manipulation and navigation tasks.

🎨

Visual & Design

Frontend design, presentation design, and visual content — generated with production-level polish.

🎬

Video & Music

Video clip generation and music composition — omnimodal output spanning visual and audio domains.

🔬

Research

Co-scientist for materials research and formalizing mathematical proofs — pushing into scientific discovery.

Full Benchmark Results

All MiMo-V2.6 results from Xiaomi's published benchmarks; comparison figures from respective public leaderboards, September 2026.

Benchmark V2.6 Pro V2.6 Flash Opus 5 Fable 5.1 GPT 6 Astra V2.5 Pro
Coding
DeepSWE v1.171.967.974.074.019.0
ProgramBench26.526.037.033.012.5
MiMo Code Bench in-house63.261.268.661.440.4
General
GDPVal 2.1 (AA)16731708173515421107
AutomationBench53.152.350.352.016.0
Agents' Last Exam31.627.631.625.734.213.2
Terminal Bench 4.034.928.849.055.159.61.5
Visual & Cyber
MiMo Visual Coding in-house72.371.570.074.482.2
CyberGym94.095.140.0
ExploitGym17.86.022.130.442.40.1

Terminal Bench 4.0 figures: GPT 6 Astra leads at 59.6; Claude Fable 5.1 at 55.1. All results as published by Xiaomi, September 22, 2026.

Get MiMo-V2.6 Today

Desktop app, API, open weights, and the full technical report — all available now.

Key takeaway: MiMo-V2.6 proves that open-source RL training at scale works. Frontier coding and knowledge-work performance, streamed live, fully reproducible — at the same price as the previous generation. The strongest open-source model to date, with the entire training pipeline public.

Sources

  1. Xiaomi MiMo — Introducing MiMo-V2.6 series (September 22, 2026). Primary announcement page.
  2. MiMo Desktop. Try MiMo-V2.6 directly.
  3. MiMo API. Developer API access.
  4. HuggingFace — XiaomiMiMo/mimo-v26. Open-source model weights and code.
  5. MiMo V2.6 Technical Report. Full RL training methodology and results.
  6. Artificial Analysis. Intelligence Index v4.3, September 2026.