Liquid AI Launches LFM2.5 Retrieval Models for 11-Language Search

Liquid AI's new retrieval models deliver fast, accurate search across 11 languages, optimized for edge devices and enterprise deployments.

By Central
The new bidirectional models enable cross-lingual search without per-language indexes, available under the LFM Open License v1.0.
Highlights
  • Both models pack 350 million parameters and are purpose-built for short-context search like product catalogs and support documentation.
  • LFM2.5-Embedding-350M uses dense bi-encoder vectors for speed, while LFM2.5-ColBERT-350M uses late interaction for higher accuracy.
  • The models are available immediately on Hugging Face and can run on local hardware via llama.cpp GGUF builds.

Liquid AI has released two new retrieval models purpose-built for multilingual and cross-lingual search, marking the first bidirectional entries in its Liquid Foundation Model family. The pair — LFM2.5-ColBERT-350M and LFM2.5-Embedding-350M — each pack 350 million parameters and target fast, accurate search across 11 languages on hardware ranging from edge devices to enterprise GPU clusters. Both models are available immediately on Hugging Face under the LFM Open License v1.0.

How the Two Retrieval Models Differ

Although both models share the same backbone, they represent text in fundamentally different ways. LFM2.5-Embedding-350M is a dense bi-encoder that compresses each document into a single 1024-dimensional vector. This approach delivers the fastest search speeds and the smallest index footprint, making it the natural choice when storage economy and latency are the primary constraints.

LFM2.5-ColBERT-350M takes a late-interaction approach. Instead of one vector per document, it assigns a 128-dimensional vector to every token. This enables word-by-word matching between queries and documents through a MaxSim operation, yielding higher accuracy and better generalization at the cost of a larger index. The model caps query length at 32 tokens and can also serve as a reranker for first-stage retrieval pipelines without requiring a dedicated index.

Both models are designed for short-context search scenarios such as product catalogs, FAQ knowledge bases, and support documentation. Liquid AI positions them as drop-in replacements for existing retrieval-augmented generation pipelines.

What Changed Under the Hood: Causal to Bidirectional

Both models begin from LFM2.5-350M-Base, a general-purpose checkpoint that Liquid AI released in March. The critical architectural change is a small set of bidirectional patches applied to the LFM2 architecture, converting it from a causal decoder into a bidirectional encoder suited for retrieval tasks.

The team replaced the causal attention mask — which restricts each token to attend only to itself and previous tokens — with a bidirectional mask that allows every token to draw context from both left and right. The LFM2 short convolutions were also made non-causal, enabling them to mix local information symmetrically around each token rather than only from the past. The resulting models preserve the backbone’s computational efficiency while producing the full-context representations that retrieval demands. Each model contains 17 layers: 10 convolution layers, 6 attention layers, and 1 pooling or dense layer. The context window extends to 32,768 tokens, though documents are tuned to 512 tokens during training.

Training Recipe and Data Strategy

Both models follow an identical three-stage training pipeline. Stage one involves large-scale contrastive pretraining in English. Stage two uses multilingual and cross-lingual distillation from a strong teacher model across all 11 supported languages. Stage three applies final fine-tuning on hard-mined negatives.

The Embedding model receives slightly more cross-lingual data than ColBERT, reflecting the fact that cross-lingual retrieval emerges more naturally in the late-interaction setup. Training data combines curated internal datasets with open-source English retrieval collections, and LLMaa-based translation was used to expand multilingual and cross-lingual pairs.

Benchmark Performance Across 11 Languages

Liquid AI evaluated the models on two benchmarks: multilingual retrieval using NanoBEIR Multilingual Extended (measuring NDCG@10) and cross-lingual open-domain question answering using MKQA-11 (measuring Recall@20). Both benchmarks cover Arabic, German, English, Spanish, French, Italian, Japanese, Korean, Norwegian, Portuguese, and Swedish.

LFM2.5-ColBERT-350M leads on both averages with an NDCG@10 of 0.605 on NanoBEIR and a Recall@20 of 0.694 on MKQA-11. LFM2.5-Embedding-350M follows closely at 0.577 and 0.691 respectively. Both models outperform Qwen3-Embedding-0.6B, a larger 600-million-parameter model, and the new ColBERT model represents a significant improvement over its predecessor, LFM2-ColBERT-350M, which scored 0.540 on NanoBEIR. For reference, the English-language subset of NanoBEIR tracks the more expensive full BEIR benchmark with high correlation, scoring roughly 15 percent higher on average, which makes it a practical proxy during training.

Latency and Deployment Flexibility

Liquid AI has released GGUF variants for llama.cpp, enabling both models to run on CPUs, laptops, and edge devices. On a MacBook Pro M4 Max at FP16 with 32-token queries and 256-token documents, the Embedding model achieves a median query latency of 7.3 milliseconds when document embeddings are precomputed. The ColBERT model reaches 8.2 milliseconds under the same conditions. Encoding documents at query time increases ColBERT latency to 34.3 milliseconds.

For enterprise deployments, Liquid AI built an internal GPU stack that runs on H100 hardware at FP16, where query latency can fall as low as 1 millisecond. This makes the models viable for both on-device semantic search and high-throughput production retrieval systems.

Use Cases That Benefit from the Approach

The cross-lingual capability is particularly relevant for e-commerce platforms where a shopper might type a Korean query and expect to surface an English product listing without maintaining per-language indexes. FAQ and support knowledge bases benefit similarly: a French support question can reliably map to an English help article. For privacy-sensitive applications, the GGUF builds allow on-device semantic search across files, emails, and notes on consumer hardware with near-zero infrastructure cost. Enterprise knowledge assistants handling legal, financial, or technical documents across multiple languages can lean on the ColBERT variant when answer accuracy matters more than index size.

Getting Started with the Models

The Embedding model works through the sentence-transformers library. Users must pass asymmetric prompts — query: and document: — as omitting them silently degrades retrieval quality. The ColBERT model runs through PyLate, which provides a PLAID index implementation using FastPLAID for efficient similarity search. Both models can be fine-tuned on custom data using standard losses such as MultipleNegativesRankingLoss.

Who Should Try These Models Now

LFM2.5-ColBERT-350M and LFM2.5-Embedding-350M represent a practical advance for anyone building multilingual retrieval or RAG systems with limited compute budgets. The Embedding variant suits applications where index size and latency are the binding constraints, while ColBERT is the better choice when accuracy and cross-lingual generalization take priority. Both models are available for immediate download on Hugging Face, can run locally via llama.cpp GGUF builds, and integrate directly into existing pipelines through sentence-transformers and PyLate. For developers evaluating retrieval models for production, the open license and documented benchmark results make these strong candidates for side-by-side testing against current deployments.

Share This Article