The open coding model just closed a 45-point gap with rl, not parameters. jetbrains' mellum2.1 is a…
the open coding model just closed a 45-point gap with rl, not parameters. jetbrains' mellum2.1 is a 12b moe with 2.5b active that jumped from 2.0 to 47.0 on swe-bench verified after reinforcement learning in real repositories.
the score moved on data, not on size.
Context
JetBrains' blog post of October 8, 2026 says Mellum2.1 is the next version of the 12B mixture-of-experts model it open-sourced in June, with 2.5B active parameters, released under the Apache 2.0 license. It says the architecture has not changed since version 2 and that almost all the work went into post-training, mainly reinforcement learning (RL), with millions of sandboxed runs across thousands of environments.
MarkTechPost (October 8, 2026) reports from JetBrains' model card that SWE-bench Verified rose from 2.0 to 47.0, SWE-bench Pro from 0.0 to 28.0 and Terminal-Bench 2.1 from 0.6 to 17.4, with agentic runs on the open-source Pi v0.73.1 harness at a 114K-token context. It states all scores are self-reported by JetBrains, from one shared pipeline in thinking mode.
The JetBrains post says the model was compared with Mellum2, Qwen3.5-9B and Gemma 4 E4B under the same setup. It says Mellum2.1 serves almost twice as many tokens as Qwen3.5-9B under heavy load and is about 1.6 times faster on a single request with multi-token prediction. GGUF builds and the MTP head for vLLM are listed as coming soon.
The note's figures (12B mixture-of-experts, 2.5B active, 2.0 to 47.0 on SWE-bench Verified) match the sources read. The 45-point gap in the note is the difference of those two scores, a derived number. The scores are JetBrains' own and were not independently reproduced here.
The note says the score moved on data, not on size. The sources say the architecture and size are unchanged and that the gain came from RL training in real environments. That fits the note, though the sources do not isolate how much came from data versus the training method, since JetBrains describes changes to both.
The same MarkTechPost table lists Qwen3.5-9B at 50.0 on SWE-bench Verified, ahead of Mellum2.1's 47.0, and ahead on SWE-bench Pro (38.0 against 28.0) and Terminal-Bench 2.1 (21.7 against 17.4). So the model closed a gap with its own earlier version, and did not pass every comparable open model.
Related work
- Mellum2.1 Gets to Work: A Fast Open Model for Coding Agents (The JetBrains Blog) ↗Primary source for the release, the RL approach, the license and the speed claims.
- JetBrains Releases Mellum2.1: A 12B MoE Open Model for Coding Agents (MarkTechPost) ↗Source for the benchmark scores, the harness and the comparison with Qwen3.5-9B and Gemma 4 E4B.
- JetBrains/Mellum2-12B-A2.5B-Thinking (Hugging Face) ↗Model card page for the released weights; surfaced in search and not read in full here.
Watch next
- Open the Mellum2.1 model card for the exact evaluation settings. Look for independent reproductions of the SWE-bench Verified score.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 8 October 2026 at 23:17 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →