we_are_coded.by CODE · The world, decoded
БГ
Liquid AI

LFM2.5: 2.6 billion parameters trained for agent work right on the device

Liquid AI (Hugging Face blog)Frontier

Liquid AI released LFM2.5-2.6B - a model built from the ground up for agents on phones and laptops: it follows instructions, handles tools, delivers 220 tokens per second on an Apple M5 Max CPU, and the weights are free on Hugging Face. According to the company, it leads its class on most agentic benchmarks.

In short
  • LFM2.5-2.6B: a 2.6-billion-parameter model from Liquid AI, built for on-device agents; weights free on Hugging Face (04.08).
  • According to Liquid AI: first in instruction following among those tested, and leading in tool use on most benchmarks (except BFCLv4).
  • 220 tokens per second on Apple M5 Max and 113 on AMD Ryzen - on CPU alone; support for llama.cpp, MLX, vLLM, ONNX.
Checked on5 August 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

For a long time the small model was a trimmed-down copy of the big one - the same ambitions, less brain. Here it's the opposite. A small model trained end to end for one single job: to be an agent.

The facts: LFM2.5-2.6B is a 2.6-billion-parameter model, published by Liquid AI on 4 August together with a base variant, both free to download on Hugging Face. On the company's own benchmarks: first in instruction following among everything tested, and leading in tool use on most benchmarks, with the exception of BFCLv4. Speed: 220 tokens per second on Apple M5 Max and 113 on AMD Ryzen - on CPU alone, no GPU; on H100 at high concurrency, around 15,000 output tokens per second. Training runs through four stages, the last of which is agentic reinforcement learning inside real agent harnesses. Supported runtimes: llama.cpp, MLX, vLLM, ONNX. Source: Liquid AI's Hugging Face blog, 04.08.2026.

Here's what survives my skeptical eye: they write down where they don't lead. You read a vendor benchmark with a grain of salt - they measured their own thing and chose what to show - but a table where you admit a lost test carries more weight than a table full of first places.

I find the training itself more interesting. The last stage is reinforcement learning inside real agent harnesses - the model isn't learning conversation, it's learning work: call a tool, read the result, keep going. This is where the category's maturity shows - the on-device model is now being built for the specific harness, with agentic work baked in from training onward.

An agent with no monthly bill and no data leaving the machine.

I plan to try it myself. Your routine tasks - classification, extraction, a local document assistant - don't need a giant in the cloud. If 2.6 billion parameters on a laptop can carry that work reliably, the heavy model stays reserved for the heavy task. The check is easy, because the weights are public.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→Liquid AI (Hugging Face blog) - LFM2.5-2.6B, 04.08.2026
Original: https://wearecoded.com/en/articles/liquid-ai-lfm25-agenti-na-ustroistvoto.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news