we_are_coded.by CODE · The world, decoded
БГ
NVIDIA

NVIDIA brought a 30-billion model down to the home RTX card

NVIDIA BlogInfra

Nemotron 3.5 Lightning is an open mixture-of-experts model built to run locally: on an RTX computer, not in someone else's cloud. Alongside it ships NeMo Switchyard, an open router between models.

In short
  • On August 11 NVIDIA released Nemotron 3.5 Lightning: 30 billion parameters, mixture-of-experts architecture, open weights under an NVIDIA license.
  • Per company data, tokens come out four times faster than the previous generation. Runs on RTX computers and Jetson, with integrations in Ollama, llama.cpp, vLLM and LM Studio.
  • Alongside it ships NeMo Switchyard: an open router that sends each request to the right model.
Checked on11 August 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

Our workstation holds a card from the previous decade and its ceiling for local models is painfully familiar to me. Which is why I read this announcement twice.

The facts: on August 11 NVIDIA released Nemotron 3.5 Lightning - a model with 30 billion parameters in a mixture-of-experts architecture, with open weights under an NVIDIA license. Per company data it generates tokens four times faster than the previous generation. It runs on RTX computers and Jetson devices, with ready integrations in Ollama, llama.cpp, vLLM and LM Studio. Alongside it ships NeMo Switchyard - an open router that directs each request to the right model.

Mixture of experts in plain terms: the model is big, but each request wakes up only a small part of it. That is how 30 billion parameters can breathe on a card where a dense model of the same size would not fit.

One clarification before the excitement: open weights and open source are not the same thing. The license is NVIDIA's, not Apache. For playing at home the difference is zero. Building a product on top of it, read the license before the first line of code.

Local stopped being a compromise for enthusiasts. It is becoming a second channel for distributing intelligence.

The math is simple: tokens without rent. The electricity you are paying for anyway.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→NVIDIA Blog - Nemotron 3.5 Lightning and NeMo Switchyard→NVIDIA Blog - Local AI, open models and agents
Original: https://wearecoded.com/en/articles/nvidia-nemotron-35-lightning-rtx.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news