Liquid AI released LFM2.5-2.6B - a model built from the ground up for agents on phones and laptops: it follows instructions, handles tools, delivers 220 tokens per second on an Apple M5 Max CPU, and the weights are free on Hugging Face. According to the company, it leads its class on most agentic benchmarks.
- LFM2.5-2.6B: a 2.6-billion-parameter model from Liquid AI, built for on-device agents; weights free on Hugging Face (04.08).
- According to Liquid AI: first in instruction following among those tested, and leading in tool use on most benchmarks (except BFCLv4).
- 220 tokens per second on Apple M5 Max and 113 on AMD Ryzen - on CPU alone; support for llama.cpp, MLX, vLLM, ONNX.
For a long time the small model was a trimmed-down copy of the big one - the same ambitions, less brain. Here it's the opposite. A small model trained end to end for one single job: to be an agent.
Here's what survives my skeptical eye: they write down where they don't lead. You read a vendor benchmark with a grain of salt - they measured their own thing and chose what to show - but a table where you admit a lost test carries more weight than a table full of first places.
I find the training itself more interesting. The last stage is reinforcement learning inside real agent harnesses - the model isn't learning conversation, it's learning work: call a tool, read the result, keep going. This is where the category's maturity shows - the on-device model is now being built for the specific harness, with agentic work baked in from training onward.
I plan to try it myself. Your routine tasks - classification, extraction, a local document assistant - don't need a giant in the cloud. If 2.6 billion parameters on a laptop can carry that work reliably, the heavy model stays reserved for the heavy task. The check is easy, because the weights are public.