Nemotron 3.5 Lightning is an open mixture-of-experts model built to run locally: on an RTX computer, not in someone else's cloud. Alongside it ships NeMo Switchyard, an open router between models.
- On August 11 NVIDIA released Nemotron 3.5 Lightning: 30 billion parameters, mixture-of-experts architecture, open weights under an NVIDIA license.
- Per company data, tokens come out four times faster than the previous generation. Runs on RTX computers and Jetson, with integrations in Ollama, llama.cpp, vLLM and LM Studio.
- Alongside it ships NeMo Switchyard: an open router that sends each request to the right model.
Our workstation holds a card from the previous decade and its ceiling for local models is painfully familiar to me. Which is why I read this announcement twice.
Mixture of experts in plain terms: the model is big, but each request wakes up only a small part of it. That is how 30 billion parameters can breathe on a card where a dense model of the same size would not fit.
One clarification before the excitement: open weights and open source are not the same thing. The license is NVIDIA's, not Apache. For playing at home the difference is zero. Building a product on top of it, read the license before the first line of code.
The math is simple: tokens without rent. The electricity you are paying for anyway.