Here is something I really believe in: a lot of AI is going to move out of the big cloud and onto the device, or into the application layer, running on open models. Cheaper, faster, and the data stays where it is. A couple of years ago that was a hobbyist thing. I think it is quietly becoming the default for a lot of work.
Let me try to lay out why, step by step.
How we got here
It started in February 2023, when Meta released LLaMA to researchers. Since then the pace has not really slowed:
1/ July 2023 - Llama 2, with a licence that allowed commercial use.
2/ September 2023 - Mistral 7B, under Apache 2.0, a small model that did a lot for its size.
3/ September 2024 - Llama 3.2 added 1B and 3B text models that Meta said fit on edge and mobile devices.
4/ January 2025 - DeepSeek-R1, a reasoning model released with its weights under the MIT licence, along with smaller versions distilled from it.
5/ August 2025 - OpenAI released gpt-oss (20B and 120B) under Apache 2.0, and described them as able to run locally on desktops and laptops.
6/ 2026 - Qwen3.5, Gemma 4 (Google’s first Gemma under Apache 2.0, with small edge sizes) and a DeepSeek-V4 preview.
Look at that list for a second. Even OpenAI is now shipping open weights you can run on a laptop. That tells me something about where this is going
How far behind are they?
Epoch AI tracks this. Between January 2023 and October 2025 the best open model trailed the best closed model by about three months on their capability index. Since January 2026 the gap is about four months. So it is not closing to zero, but it is small, and it has stayed small while the frontier moved fast.
The number that matters more for my argument is the next one. In August 2025 Epoch found that a single high-end gaming graphics card, under $2,500, could run models that matched the frontier from 6 to 12 months earlier. They also warn that small models are often tuned for the benchmarks, so the real-world lag may be longer. I take that caveat seriously
Why this matters for cost and speed
Epoch’s latest report says the cost of reaching a given level of AI performance has fallen about 13 times a year since 2023. Cheap thinking keeps getting cheaper, and more and more of it is cheap enough to run on hardware we already own.
This is where I think the real change lands:
Cost. A model on the device has no per-token bill. For high-volume, boring work (classifying, extracting, summarising, short replies) that adds up.
Speed. No network round trip. For anything that should feel instant, that matters more than a few extra benchmark points.
Privacy. The data does not leave the machine. Apple’s on-device model for developers suggests the big platforms see it the same way.
I do not think it is either/or, though. The sensible design is to send each job to the cheapest model that can do it. Hard, rare, high-stakes reasoning goes to a big cloud model. The everyday volume stays local
What I am not sure about
Running locally is not free. You give up some quality, you take on setup and updates, and for the hardest tasks the frontier models are still clearly ahead. Licences differ too - some “open” models come with conditions, so read them before building on one.
But if the gap stays at a few months while the price of a given level of performance keeps falling, the local option gets better every quarter without anyone doing anything. That is a trend I would want to be on the right side of.
I will leave you with a question - what is one thing you pay a cloud model for today that a small model on your own machine could do just as well?





