{"id":65177,"date":"2026-07-29T13:50:49","date_gmt":"2026-07-29T17:50:49","guid":{"rendered":"https:\/\/overcentral.com\/en\/?p=65177"},"modified":"2026-07-29T13:50:49","modified_gmt":"2026-07-29T17:50:49","slug":"liquid-ai-lfm25-encoders","status":"publish","type":"post","link":"https:\/\/overcentral.com\/en\/liquid-ai-lfm25-encoders\/","title":{"rendered":"Liquid AI Slashes CPU Latency with LFM2.5 Bidirectional Encoders"},"content":{"rendered":"<p><a href=\"https:\/\/www.liquid.ai\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">Liquid AI<\/a> has released two open-weight bidirectional encoders, the LFM2.5-Encoder-230M and <a href=\"https:\/\/huggingface.co\/LiquidAI\/LFM2.5-Encoder-350M\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">LFM2.5-Encoder-350M<\/a>, that directly address a persistent bottleneck in production NLP: the cost of running encoders over long inputs on CPU. Both models are masked language models built on the LFM2 hybrid backbone, carry an 8,192-token context window, and are designed for the kind of continuous, GPU-free inference that powers classifiers, intent routers, safety filters, and PII detectors. The claim is that the LFM2 architecture grows more slowly as inputs get longer, giving it a decisive latency advantage over existing models built on the BERT lineage.<\/p>\n<h2>How Liquid AI Converted a Decoder into a Bidirectional Encoder<\/h2>\n<p>The encoders are not trained from scratch. They are initialized from the LFM2.5-230M and LFM2.5-350M decoder backbones, then converted with three specific changes. First, the causal attention mask is replaced with a bidirectional one, so every token attends to both left and right context. Second, the LFM2 short convolutions are made non-causal using symmetric center padding, so each token&#8217;s convolution mixes in neighbors on both sides. Third, the model is trained with a masked language modeling objective at a 30 percent mask rate, which is denser than BERT&#8217;s 15 percent, following evidence Liquid AI cites that a higher mask rate helps at this scale.<\/p>\n<p>Training runs in two stages. Stage one establishes general language competence with a short-context MLM objective on a large web corpus at 1,024 tokens. Stage two extends context to 8,192 tokens on the full data mix, strengthening factual, legal, and multilingual competence. Architecturally, the backbone interleaves gated short-convolution blocks with grouped-query attention, the same design described in the LFM2 technical report. Both checkpoints use a hidden size of 1024 and a 65,536-token vocabulary, and support 15 languages. The license is the LFM Open License v1.0.<\/p>\n<h2>Benchmark Results: Where the Two Encoders Land<\/h2>\n<p>Liquid AI evaluated 14 models on 17 tasks pulled from GLUE, SuperGLUE, and multilingual classification. Every model is fully fine-tuned per task, and the reported score is that fine-tuned model&#8217;s result. The LFM2.5-Encoder-350M posts a 17-task mean of 81.02, with a standard deviation of 1.00, ranking fourth. The three models ahead of it are all larger: XLM-R XL at 3.5B parameters scoring 83.06, ModernBERT-large at 395M scoring 81.68, and XLM-R large at 560M scoring 81.34. The top model is nearly ten times its size.<\/p>\n<p>The LFM2.5-Encoder-230M posts 79.29 with a standard deviation of 1.02, ranking sixth. It beats ModernBERT-base at 78.19 and every EuroBERT model in the table, including EuroBERT-610M at 75.87 and EuroBERT-2.1B at 72.19. Both new encoders also score above Liquid AI&#8217;s own retrieval siblings, the LFM2.5-ColBERT-350M at 76.18 and the LFM2.5-Embedding-350M at 75.68. That gap is the stated reason Liquid AI built a general-purpose encoder instead of reusing the retrievers.<\/p>\n<p>The methodology is open-sourced under Apache-2.0. Every model is loaded with fp32 master weights and bf16 autocast, so the table compares models rather than number formats. Every model uses the same AdamW recipe, taken from the EuroBERT card. Learning rate is selected per model and task across 10 rates and 3 seeds. Scores are then reported as the mean over 5 fresh seeds that never touched selection. The transformers version is pinned to 4.56.2 so dependency drift is not an uncontrolled variable.<\/p>\n<h2>CPU Latency: The Defining Advantage<\/h2>\n<p>Liquid AI reports two measured points at full 8K context on CPU. The LFM2.5-Encoder-230M completes one forward pass in roughly 28 seconds. ModernBERT-base, at 149M parameters, takes over one minute and thirty seconds for the same workload. The quoted speed-up varies between 3.3x and 3.7x depending on the source, but the practical framing is consistent: a full contract or patient record can be processed in under 30 seconds on a machine with no GPU. On GPU, the margin is smaller. ModernBERT-base leads below roughly 1K tokens on the Apple GPU, and the LFM2.5 encoders take the lead from about 2K tokens, with a smaller margin than on CPU.<\/p>\n<h2>Use Cases and Deployment Environments<\/h2>\n<p>The release names three primary settings. Edge and embedded devices come first, such as a car&#8217;s onboard compute or an industrial controller that has no spare GPU and cannot afford a cloud round trip. Regulated and on-premise systems in finance, healthcare, and legal, where documents are long, sensitive, and cannot leave in-house infrastructure, form the second category. And high-volume cost-sensitive pipelines, where a small encoder acts as a cheap first pass in front of a larger model, are the third. Liquid AI also puts a useful number on the context window: 8,192 tokens is roughly 13 to 15 pages. One forward pass covers a full contract or a complete patient record.<\/p>\n<p>To show what a fine-tuned encoder looks like, the research team shipped five demos. Each runs in a CPU-only <a href=\"https:\/\/overcentral.com\/en\/autonomous-ai-hack-hugging-face\/\" title=\"OpenAI AI Autonomously Hacks Hugging Face During Test\" data-iacss-internal=\"1\">Hugging Face<\/a> Space. They cover zero-shot prompt routing, zero-shot policy linting, and spell checking. A PII detector handles 40 information types <a href=\"https:\/\/overcentral.com\/en\/qwen-audio-3-0-tts\/\" title=\"Qwen-Audio-3.0-TTS Launches in Flash and Plus Tiers Across 16 Languages\" data-iacss-internal=\"1\">across 16 languages<\/a>. A bonus masked-diffusion demo runs the encoder as a chatbot that generates by iteratively unmasking rather than left to right.<\/p>\n<h2>Getting the Encoders Running<\/h2>\n<p>Both encoders load through transformers. The body is exposed as Lfm2BidirectionalModel and masked-LM loading uses Lfm2BidirectionalForMaskedLM. Both are wired through auto_map, so trust_remote_code=True is required on every load call. A base encoder produces general-purpose representations, not task outputs, so fine-tuning is mandatory. Liquid AI&#8217;s fine-tuning tutorial walks through long legal documents at an 8K context configuration. The model selection guidance is straightforward: 350M when accuracy matters most, 230M for tighter hardware or higher throughput.<\/p>\n<p>For maximum efficiency on supported GPUs, the model cards recommend Flash Attention 2. No packaged browser build is documented in the release, though the 230M model card states that the model runs in the browser on WebGPU, calling its small footprint especially well suited to on-device and browser deployment. As of this writing, no inference provider is serving either encoder <a href=\"https:\/\/overcentral.com\/en\/openai-autonomous-hack-huggingface\/\" title=\"OpenAI Model Carries Out First Autonomous Hack on Hugging Face\" data-iacss-internal=\"1\">on Hugging Face<\/a> at launch, and the published speed-up multiplier differs between sources: the release blog says about 3.7x, the model cards say 3.3x, both describing the same 8K CPU comparison against ModernBERT-base.<\/p>\n<h2>Strategic Significance of the LFM2.5 Encoder Release<\/h2>\n<p>The release is significant because it extends the LFM2 architecture into the encoder family, which is the backbone of most production NLP pipelines. BERT established the class, and ModernBERT pushed its accuracy, speed, and context, but Liquid AI&#8217;s argument is that the LFM2 architecture continues that line because its cost grows more slowly as inputs get longer. For teams running classifiers, safety filters, or PII detectors on CPU at scale, the difference between 28 seconds and 90 seconds per forward pass at 8K tokens is not marginal. It changes what is feasible. The open-weight release under the LFM Open License v1.0, combined with the open-sourced evaluation harness, means that teams can verify the claims on their own hardware and data before committing to a pipeline change.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Liquid AI has released two open-weight bidirectional encoders, the LFM2.5-Encoder-230M and LFM2.5-Encoder-350M, that directly address a persistent bottleneck in production NLP: the cost of running encoders over long inputs on CPU. Both models are masked language models built on the LFM2 hybrid backbone, carry an 8,192-token context window, and are designed for the kind of [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":84042,"comment_status":"closed","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/65177.png","fifu_image_alt":"Liquid AI Slashes CPU Latency with LFM2.5 Bidirectional Encoders","footnotes":""},"categories":[349],"tags":[],"class_list":["post-65177","post","type-post","status-publish","format-standard","has-post-thumbnail","category-articles"],"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/65177.png","fifu_image_alt":"Liquid AI Slashes CPU Latency with LFM2.5 Bidirectional Encoders","fifu_redirection_url":"https:\/\/towardsdatascience.com\/gpu-resident-top-k-for-agentic-rag-i-built-a-cuda-kernel-so-my-retrieval-step-would-stop-bouncing-off-the-gpu\/","_links":{"self":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/65177","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/comments?post=65177"}],"version-history":[{"count":0,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/65177\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media\/84042"}],"wp:attachment":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media?parent=65177"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/categories?post=65177"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/tags?post=65177"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}