we_are_coded.by CODE · The world, decoded
БГ
Concept

The model that runs on your machine

The BasicsUpdated on 25 August 2026we are coded

The difference between a model that lives in someone else's cloud and one you download once and run on your own machine. Worth knowing, because it decides whether your conversations leave your house or not.

Checked on25 August 2026
In short: usually you ask an AI model and your question travels to someone else's computer, somewhere in the world. A local model works the other way. You download it once to your own disk, and after that it runs right on your machine, sending nothing out. You pay with a stronger computer and a weaker model than the best ones in the cloud, instead of paying per question.

What the difference actually is

When you write to a model in the cloud, your words travel over the network, reach a server the company owns, a model with billions of numbers works through the question there, and sends an answer back to you. Convenient, but it depends entirely on someone else's machine, someone else's network, and a company that decides what happens to your conversation. A local model breaks that chain. You download the file with the weights, the numbers the model computes with, put it on your own disk, and from there everything happens on your own processor or graphics card.

People want it for a few clear reasons. Code and data don't travel to an outside server, which matters if you work with anything sensitive. There's no per-question bill, because once downloaded, the model is yours. And it works without internet, which sounds minor until the exact moment your connection drops.

A local model doesn't save you money right away. It saves you dependence.

When it makes sense, and when it's just for show

The cost is real. A local model wants a machine with a lot of memory and a strong graphics card, and that isn't cheap, and even then it falls behind the biggest cloud models, because those sit on thousands of chips at once. JetBrains showed exactly where that line sits in August 2026: they got their coding agent Junie running fully locally on a Mac with an M5 chip and 64 gigabytes of memory, on a model called Qwen 3.6-27B, about 20 gigabytes on disk, for free.

If you just want to try the idea, Ollama is the simplest way in, a program that downloads and runs local models with a few commands. It makes sense when your data genuinely shouldn't leave the machine, or when you do a lot of small work and the per-question bill stings. It's for show when you only do it because it sounds independent.

The visual is generated code art. No third-party images.
Official primary sources
→JetBrains: Junie goes local→Ollama