{"id":75151,"date":"2026-08-06T22:42:01","date_gmt":"2026-08-07T02:42:01","guid":{"rendered":"https:\/\/overcentral.com\/en\/?p=75151"},"modified":"2026-08-06T22:42:01","modified_gmt":"2026-08-07T02:42:01","slug":"lfm2-5-2-6b-raspberry-pi","status":"publish","type":"post","link":"https:\/\/overcentral.com\/en\/lfm2-5-2-6b-raspberry-pi\/","title":{"rendered":"Liquid AI Brings LFM2.5-2.6B AI Agents to Raspberry Pi"},"content":{"rendered":"<article>\n<h1>Liquid AI\u2019s New LFM2.5-2.6B Model Brings Agentic AI to the Raspberry Pi<\/h1>\n<p>The idea that a useful artificial intelligence agent could run on a device smaller than a smartphone, with no internet connection, and costing less than a hundred dollars, has moved from theoretical possibility to commercial reality. Earlier this week, <a href=\"https:\/\/overcentral.com\/en\/liquid-ai-lfm25-encoders\/\" title=\"Liquid AI Slashes CPU Latency with LFM2.5 Bidirectional Encoders\" data-iacss-internal=\"1\">Liquid AI<\/a>, a startup founded in 2023 by former MIT computer scientists, released LFM2.5-2.6B, an open-weight language model designed specifically for agentic workloads that can operate entirely on local hardware. The model runs on everything from a smartphone to a Raspberry Pi, bypassing cloud inference entirely and opening the door for edge AI applications in regulated industries and environments where connectivity is limited or data privacy is paramount.<\/p>\n<h2>What LFM2.5-2.6B Is Designed to Do<\/h2>\n<p>LFM2.5-2.6B packs 2.6 billion parameters and supports a 128,000-token context window, with native tool-calling capabilities built directly into the architecture. The name itself encodes the model\u2019s identity: the \u201c2.5\u201d refers to the generation of the architecture, while \u201c2.6B\u201d denotes the parameter count. This is not a general-purpose chatbot meant to compete with frontier models like GPT-4 or Claude. Instead, Liquid AI positions it as a specialized engine for high-volume, well-defined agentic tasks that run locally: tool calling, document management, calendar and workflow automation, and always-on background routines. It is particularly suited for connectivity-limited environments like vehicles and robotics, as well as for enterprises that handle sensitive information and cannot afford to send data to the cloud.<\/p>\n<p>\u201cI do also believe that the best models will be in the cloud, and there\u2019s no problem with that,\u201d Maxime Labonne, Liquid AI\u2019s head of post-training, told VentureBeat. \u201cWe want to make models for another type of user, and the best way of describing it is: you should use [edge AI] when you can\u2019t use a cloud model.\u201d<\/p>\n<h2>Performance That Fits in Your Pocket<\/h2>\n<p>Labonne emphasized that the LFM2 architecture was explicitly designed around real-world CPU performance rather than GPU benchmarks. \u201cI think the best example is a Raspberry Pi,\u201d he said. \u201cWe have a lot of demos that show that actually, it works pretty fast on the Raspberry Pi.\u201d Company-reported measurements back up that claim: decoding throughput reaches approximately 220 tokens per second on an Apple M5 Max and 113 tokens per second on an <a href=\"https:\/\/www.amd.com\/en\/products\/processors\/desktops\/ryzen.html\" target=\"_blank\" rel=\"sponsored noopener noreferrer\" data-iacss-external=\"1\">AMD Ryzen<\/a> <a href=\"https:\/\/overcentral.com\/en\/google-ai-max\/\" title=\"Google AI Max Unlocks Billions of New Monetizable Searches\" data-iacss-internal=\"1\">AI Max<\/a>aaaa+ 395, while consuming less than 2.5 GB of memory. On a smartphone, the model achieves around 30 tokens per second, and users can test it directly through Apollo, Liquid AI\u2019s mobile app. At the other end of the deployment spectrum, when running on a single <a href=\"https:\/\/www.nvidia.com\/\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">Nvidia<\/a> H100 GPU under sustained concurrent load, Liquid reports nearly 15,000 output tokens per second \u2014 roughly 1.3 billion tokens per day on one card. These figures are vendor benchmarks and have not been independently verified.<\/p>\n<p>For Labonne, these numbers are not just performance bragging points but hard constraints that determine what can actually be deployed. \u201cWhat we want to show is that it\u2019s a really good trade-off, because you get the level of quality that you get with much bigger models, but in a tiny, tiny form factor,\u201d he said. \u201cYou can deploy it in target devices where you are not able to deploy the other ones at all.\u201d<\/p>\n<h2>Trained for Agents, Not Chatbots<\/h2>\n<p>The model\u2019s training regimen reflects a fundamental shift in how AI is consumed. \u201cModels are not consumed in chatbots anymore. They\u2019re really consumed through agentic harnesses, like OpenClaw, like Hermes Agent,\u201d Labonne said. \u201cWe wanted to make sure that this model is not just good at math or at code, but it\u2019s good at using tools.\u201d LFM2.5-2.6B was pretrained on approximately 34 trillion tokens, with a vocabulary doubled to 128K tokens to better support non-Latin scripts. A dedicated mid-training phase extended the context window to 128K tokens, enabling long-running agent workflows.<\/p>\n<p>The post-training pipeline is a four-stage process: supervised fine-tuning, teacher specialization (training separate expert models for instruction following, math, code, and tool use), multi-domain on-policy distillation (MOPD) to merge those experts\u2019 capabilities back into a single student model, and finally agentic reinforcement learning. During that final stage, the model was trained directly inside production agent harnesses \u2014 including Hermes Agent and OpenClaw \u2014 on realistic productivity tasks involving research, coding, document management, tool invocation, and workflow automation. This exposed the model to those harnesses\u2019 actual tools, system prompts, and interaction patterns.<\/p>\n<p>Labonne described the pipeline overhaul as yielding a \u201chappy accident\u201d: gains that extended well beyond the agentic targets. \u201cThrough these new training techniques, we also got a lot better at everything. We got better at math, at instruction following. We\u2019ve never been good at code, actually \u2014 and with this, we even got really good at code,\u201d he said.<\/p>\n<h2>Building the Model and the Harness Together<\/h2>\n<p>Rather than relying solely on existing frameworks, Liquid AI built its own agent harness and demonstrated the model running inside it on a phone, planning and calling tools entirely on-device. \u201cThis is a harness running on a phone, and I don\u2019t know if there\u2019s any other harness running on a phone,\u201d Labonne said. The company had two reasons for this: no phone-native harness existed, and Liquid wanted a different interaction model. Existing harnesses wait for a prompt; Liquid envisions assistants that act on their own. \u201cWe want proactive agents. We want agents that run in the background, check what you\u2019re doing, check your calendar, and based on this context, do tasks,\u201d he said. \u201cThat doesn\u2019t exist today, really.\u201d<\/p>\n<p>Co-designing the harness and model also lets the software compensate for the model\u2019s weak spots. \u201cEverything that the model is bad at, the harness should help the model with \u2014 provide as much assistance as possible to make it more reliable,\u201d Labonne said. \u201cEnd users don\u2019t care if it\u2019s the model or the harness. What they want is that the task is achieved at the end of the day.\u201d The model nevertheless works out of the box with established harnesses including Hermes Agent, OpenClaw, and Pi, served behind any OpenAI-compatible endpoint.<\/p>\n<h2>Swap the Harness, Not the Model<\/h2>\n<p>For enterprise deployment, Labonne argues the release marks a shift in what small models can be used for. Until now, local models made economic sense mainly as narrowly fine-tuned specialists \u2014 trained to do one thing at cloud-model quality, much faster and cheaper. Agentic capability changes that calculus, because the same model can be repurposed by changing the tools around it rather than the model itself. \u201cYou can have a calendar assistant, and you can reuse the same model and make a meeting assistant that will record what everybody said and summarize it \u2014 a bit like Granola, for example,\u201d he said. \u201cYou don\u2019t change the model; you just change the harness. You just change the tools around it. This gives much more generalizability, and it\u2019s a lot easier to do and a lot cheaper as well.\u201d<\/p>\n<p>Labonne still recommends fine-tuning for production deployments whenever feasible: \u201cIf you don\u2019t fine-tune it, you leave some quality on the table. If you fine-tune it well, it\u2019s going to match the performance of GPT and Claude \u2014 really, if your task is not the most complex task in the world.\u201d He added that the barrier to entry has collapsed: \u201cThe bar to be able to do fine-tuning now is super low. It\u2019s very accessible to everyone.\u201d<\/p>\n<h2>How LFM2.5-2.6B Stacks Up Against DeepSeek, Gemma, and Qwen<\/h2>\n<p>Liquid AI released its own benchmark comparisons pitting LFM2.5-2.6B against models most likely to compete for the same edge deployments: Google\u2019s Gemma 4 E2B (5.1B parameters) and E4B (8B), and <a href=\"https:\/\/overcentral.com\/en\/alibaba-qwen-3-8-leisure\/\" title=\"Alibaba frames Qwen 3.8 as AI replacing jobs for leisure\" data-iacss-internal=\"1\">Alibaba<\/a>\u2019s Qwen3.5-4B (4.7B) and Qwen3.5-9B (9.7B). A separate test by local AI client platform Atomic Chat found that LFM2.5-2.6B completed 35 tool calls to check weather and local time in six cities, convert one budget into six currencies, and check four hotels and book for a date \u2014 3.7 times faster than DeepSeek-V4-Flash (284B parameters), the model that has skyrocketed to the top of OpenRouter since its release last week.<\/p>\n<p>Gemma 4\u2019s small models are multimodal generalists, accepting image and audio input alongside text, using a Per-Layer Embeddings design that keeps only a fraction of their weights active per token \u2014 which is why <a href=\"https:\/\/www.google.com\/\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">Google<\/a> markets them by \u201ceffective\u201d size (2.3B and 4.4B) despite total footprints of 5.1B and 8B. Alibaba\u2019s Qwen3.5 small series, released in March, is natively multimodal from 4B up and leans on scaled reinforcement learning to chase frontier-style reasoning. Alibaba touts the 9B model as matching or beating OpenAI\u2019s far larger GPT-oss-120B on reasoning benchmarks.<\/p>\n<p>LFM2.5-2.6B takes a narrower path: it is text-only, dense, and specialized for agentic work, with Liquid AI shipping separate vision and audio variants of the LFM family rather than folding everything into one checkpoint. Where Qwen\u2019s post-training reinforcement learning targets reasoning, Liquid\u2019s targets tool use inside real agent harnesses. The result, per Liquid\u2019s published numbers, is that the smallest model in the comparison leads every instruction-following benchmark (IFBench, Multi-IF, IFStruct) and nearly every tool-use benchmark \u2014 77.83 on ToolSandbox versus 76.44 for Qwen3.5-9B, a model nearly four times its size \u2014 trailing only that 9B model on BFCLv4.<\/p>\n<p>On agentic evaluations, LFM2.5-2.6B beats both Gemma models across the board and essentially ties the Qwens: 26.89 on BrowseComp+ versus 27.23 for Qwen3.5-9B. It also posts the best score on AA Omniscience, a knowledge benchmark that penalizes hallucination. The Qwen models keep the edge where their training focus lies: math (Qwen3.5-9B leads AIME25) and coding, where larger models retain an advantage on LiveCodeBench \u2014 though Labonne noted the gap is smaller than the parameter counts would suggest. \u201cWith LiveCodeBench v6, we might not be the best among these models, but we\u2019re also by far the smallest. Showing that we\u2019re competitive with them is already quite a win for me,\u201d he said.<\/p>\n<h2>The Licensing Puzzle: Open Weights, Commercial Restrictions<\/h2>\n<p>One differentiator works against Liquid AI\u2019s adoption curve: licensing. Gemma 4 and Qwen3.5 ship under the permissive Apache 2.0 license \u2014 a change Google made specifically to court enterprises. DeepSeek-V4-Flash ships under a similarly permissive MIT License. Meanwhile, Liquid AI\u2019s LFM Open License v1.0 permits use, modification, and redistribution \u2014 including commercial use \u2014 for organizations with less than $10 million in annual revenue. Commercial use by larger companies requires a separate arrangement with Liquid AI; qualified nonprofits are exempt from the threshold for non-commercial and research purposes.<\/p>\n<p>Labonne framed the structure as a way to sustain model development: \u201cThe models are really the moats, so we need to be sensible in the way that we license them; otherwise, we cannot make money, so we can\u2019t make more models.\u201d He characterized the threshold as a light-touch mechanism in practice. Asked how the company would even know if a large enterprise quietly deployed the open weights, he was candid: \u201cI think this is a question for our legal team, but personally, I don\u2019t know. And even<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Liquid AI\u2019s New LFM2.5-2.6B Model Brings Agentic AI to the Raspberry Pi The idea that a useful artificial intelligence agent could run on a device smaller than a smartphone, with no internet connection, and costing less than a hundred dollars, has moved from theoretical possibility to commercial reality. Earlier this week, Liquid AI, a startup [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":83697,"comment_status":"closed","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/75151.png","fifu_image_alt":"Liquid AI Brings LFM2.5-2.6B AI Agents to Raspberry Pi","footnotes":""},"categories":[31],"tags":[],"class_list":["post-75151","post","type-post","status-publish","format-standard","has-post-thumbnail","category-technology"],"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/75151.png","fifu_image_alt":"Liquid AI Brings LFM2.5-2.6B AI Agents to Raspberry Pi","_links":{"self":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/75151","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/comments?post=75151"}],"version-history":[{"count":0,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/75151\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media\/83697"}],"wp:attachment":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media?parent=75151"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/categories?post=75151"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/tags?post=75151"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}