Behind local Junie sits a string of engineering trade-offs: a quantized model, reasoning turned off, remembered context. The most interesting one is the model choice. The newer Qwen 3.8 was dropped because, with reasoning on, it churns out around five times more tokens.
- Context is remembered across tasks and sessions, so already-opened files don't get re-read.
- Reasoning is turned off on purpose: twice the execution speed for a small quality loss.
- Qwen 3.8 was dropped because it needs reasoning turned on and ends up around four times slower.
The newer model lost. JetBrains picked version 3.6 over the fresher 3.8, and said why.
There isn't one model behind local Junie. There's a list of trade-offs, made one at a time, until the agent fit on a laptop.
Why newer comes out worse here
Qwen 3.8 is the stronger model on paper. Only it needs reasoning turned on, and at medium effort produces around five times more tokens than 3.6. The result is roughly four times slower work. On a machine with limited memory, that isn't a trade-off. It's a brake.
The list is instructive beyond JetBrains, too. None of these moves is new on its own. The gain comes from stacking them. If someone promises you a laptop agent just because they swapped models, you know what's missing.