The difference between a model that lives in someone else's cloud and one you download once and run on your own machine. Worth knowing, because it decides whether your conversations leave your house or not.
What the difference actually is
When you write to a model in the cloud, your words travel over the network, reach a server the company owns, a model with billions of numbers works through the question there, and sends an answer back to you. Convenient, but it depends entirely on someone else's machine, someone else's network, and a company that decides what happens to your conversation. A local model breaks that chain. You download the file with the weights, the numbers the model computes with, put it on your own disk, and from there everything happens on your own processor or graphics card.
People want it for a few clear reasons. Code and data don't travel to an outside server, which matters if you work with anything sensitive. There's no per-question bill, because once downloaded, the model is yours. And it works without internet, which sounds minor until the exact moment your connection drops.
When it makes sense, and when it's just for show
The cost is real. A local model wants a machine with a lot of memory and a strong graphics card, and that isn't cheap, and even then it falls behind the biggest cloud models, because those sit on thousands of chips at once. JetBrains showed exactly where that line sits in August 2026: they got their coding agent Junie running fully locally on a Mac with an M5 chip and 64 gigabytes of memory, on a model called Qwen 3.6-27B, about 20 gigabytes on disk, for free.
If you just want to try the idea, Ollama is the simplest way in, a program that downloads and runs local models with a few commands. It makes sense when your data genuinely shouldn't leave the machine, or when you do a lot of small work and the per-question bill stings. It's for show when you only do it because it sounds independent.