← All tool briefs

Tool brief · October 8, 2026

Mellum 2.1: a 12B open model aimed squarely at your agent's sub-agents

DeveloperFor Developer

The tool

JetBrains Mellum 2.1

Visit JetBrains Mellum 2.1 →

What it is

Mellum 2.1 is JetBrains' updated open-weights coding model. It's a 12B mixture-of-experts model with 2.5B active parameters, released under Apache 2.0, with the architecture unchanged from Mellum2 but trained further with reinforcement learning. JetBrains positions it as built for coding agents and fast sub-agents that run on your own hardware.

It is not a frontier model and does not try to be. It's a small, fast worker meant to sit inside an agent loop where a bigger model is doing the planning.

The next-work-session test

Concrete scenario: you have a LangGraph or custom agent where a frontier model (Claude, GPT-5, whatever) orchestrates, and you're paying per token for sub-tasks like "summarize this diff," "classify this stack trace," "pick which file to open next," or "rewrite this import block." Those calls are latency-sensitive and volume-heavy.

The next-session change is swapping those sub-agent calls to a self-hosted Mellum 2.1 endpoint behind an OpenAI-compatible interface, and running an eval harness over your existing traces to compare quality. You can run it locally with just two commands using Ollama, which means a prototype swap takes an afternoon, not a sprint.

That's the realistic test. If latency drops and your eval pass rate holds within tolerance on the narrow sub-tasks, keep it. If not, you've lost half a day.

Pricing

The model itself is free: Mellum is released under the Apache 2.0 license, with open weights available on Hugging Face. There is no per-token API fee from JetBrains for the open weights — your cost is whatever GPU you run it on, plus whatever inference server you wrap around it.

JetBrains does also sell AI Assistant and Junie subscriptions that use Mellum internally, but those are separate products. Pricing for a hosted Mellum 2.1 API from JetBrains: unverified — the public materials point you at self-hosting via Hugging Face or Ollama, not a metered endpoint.

What we'd actually use it for

Narrower than the pitch. We would use it for:

  • Routing and classification steps in an agent graph (which tool to call, which file matters).
  • Short code completions and rewrites inside a tight loop where a 2-second round-trip to a frontier API ruins the UX.
  • Local, offline dev inside a VPC where sending code to a hosted API is a non-starter.

We would not use it as the planning brain. JetBrains itself frames Mellum2 as efficient for routing, RAG, and sub-agents — that's the honest scope.

Limits

  • 12B MoE with 2.5B active is small. On ambiguous multi-file refactors or long-context reasoning, a frontier model will beat it. Don't put it in the planner slot.
  • Self-hosting means you own the ops: GPU capacity, autoscaling, observability, prompt-caching, the works. "Free weights" is not "free inference."
  • You will need your own eval set. The benchmarks a vendor publishes rarely match the shape of calls your actual agent makes. Treat any numbers in the launch post as the vendor's claim and re-measure on your traces.
  • The original Mellum was a 4B-parameter model focused on code completion, deployed as part of JetBrains' AI Assistant. The 2.x line is broader but still a specialist, not a generalist chat model.
  • RL-trained models can be uneven — strong on the task shapes they were rewarded for, weaker just outside that distribution. Probe the edges before trusting it in production.

For an outside write-up that's useful for context on the self-hosting angle, see this community breakdown of the Mellum2 release.

Try it if

  • You run a multi-model agent and your sub-agent token bill is embarrassing.
  • You need on-prem or air-gapped inference for a coding workflow.
  • You already have an eval harness and can measure quality regressions in a day.
  • You're comfortable operating a model server (vLLM, TGI, Ollama) in production.

Skip it if

  • You just want one API key and one model behind your IDE — stay on a hosted frontier API.
  • You have no evals. Swapping models blind to save cost is how agents silently get worse.
  • Your workload is long-context planning or deep multi-file reasoning. Wrong tool.
  • You don't have GPU budget or an ops person. The weights are free; the serving is not.

Source: blog.jetbrains.com

More for Developer professionals →

Get the next one in your inbox

One daily brief. Every story gets a hype verdict.

No spam. Unsubscribe anytime.

No sponsored verdicts · We have no paid relationship with featured vendors