{"id":57120,"date":"2026-06-18T12:41:22","date_gmt":"2026-06-18T16:41:22","guid":{"rendered":"https:\/\/overcentral.com\/en\/?p=57120"},"modified":"2026-06-18T12:41:22","modified_gmt":"2026-06-18T16:41:22","slug":"kv-cache-compression-methods","status":"publish","type":"post","link":"https:\/\/overcentral.com\/en\/kv-cache-compression-methods\/","title":{"rendered":"TurboQuant, OSCAR, EpiCache Vie for KV Cache Compression Lead"},"content":{"rendered":"<p>Long-context large language models (LLMs) have a memory problem that is not about the weights. During the decoding phase, transformers cache the key and value (KV) vectors for every token at every layer to avoid recomputing the attention mechanism. This cache grows linearly with sequence length and batch size, and at long contexts with high concurrency, it can eclipse the model\u2019s own footprint. For example, Llama-3.1-70B in BF16 requires roughly 0.31 MB per token for its KV cache. At a 128K token context, this amounts to around 40 GB; at 1M tokens, it exceeds 300 GB \u2014 more than the 140 GB of the model weights themselves. Every newly decoded token must stream the entire cache out of high-bandwidth memory (HBM), making decoding memory-bandwidth-bound rather than compute-bound. Shrinking the KV cache is therefore the most direct way to cut both cost and decode latency.<\/p>\n<p>Current approaches to this problem fall into roughly five families: token eviction (H2O, SnapKV), quantization (KIVI, GEAR), low-rank projection (Palu), merging (KVMerger), and architectural sharing (MLA). Recent 2026 work has pushed especially hard on the ultra-low-bit quantization frontier. Google and NYU\u2019s <a href=\"https:\/\/arxiv.org\/abs\/2406.12345\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">TurboQuant<\/a> (ICLR 2026) and Together AI\u2019s <a href=\"https:\/\/github.com\/togethercomputer\/OSCAR\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">OSCAR<\/a> attack the same problem from opposite directions, while <a href=\"https:\/\/overcentral.com\/en\/apple-siri-ai-eu-dma-dispute\/\" title=\"Apple withholds Siri AI from EU iPhones over DMA dispute\" data-iacss-internal=\"1\">Apple<\/a>\u2019s EpiCache tackles a problem neither one addresses.<\/p>\n<h2 id=\"h-the-outlier-channel-problem\">The Outlier Channel Problem<\/h2>\n<p>Most KV quantizers are fighting the same underlying enemy: outlier channels. These are a small number of channels with disproportionately large magnitudes that dominate the quantization range, squeezing the rest of the signal into just a few representable levels. This is why naive INT2 quantization (with only four levels) collapses to near-zero accuracy.<\/p>\n<p>KIVI established the standard baseline here. It showed that key vectors have fixed outlier channels across tokens while value vectors do not, so it quantizes keys per-channel and values per-token. That tuning-free 2-bit recipe cuts end-to-end peak memory (weights included) by about 2.6 times, and it is the reference point the newer methods build on.<\/p>\n<h2 id=\"h-turboquant-data-oblivious-and-theoretically-optimal\">TurboQuant: Data-Oblivious and Theoretically Optimal<\/h2>\n<p>TurboQuant handles outliers without ever looking at your data, in two stages. First, each vector is randomly rotated so its coordinates become nearly independent and approximately Gaussian, which lets an optimal precomputed scalar (Lloyd\u2013Max) quantizer be applied per coordinate. Second, a 1-bit Quantized Johnson\u2013Lindenstrauss (QJL) transform is applied to the residual, giving a provably unbiased estimate of attention logits with no normalization-constant overhead.<\/p>\n<p>The selling point is theoretical: TurboQuant\u2019s distortion is provably within a small constant factor (around 2.7 times) of the information-theoretic lower bound. In practice, it reaches essentially full-precision recall on Needle-in-a-Haystack at 4x compression. The paper reports absolute quality neutrality at 3.5 bits and only marginal degradation at 2.5 bits per channel. Because it needs no calibration, it works on any model untouched and doubles as a fast vector-database quantizer. One caveat: the widely repeated \u201c8x faster attention on H100\u201d figure comes from Google\u2019s blog post, not the paper, and refers to a narrow attention-logit microbenchmark. TurboQuant\u2019s documented sweet spot is the 3 to 4-bit near-lossless regime.<\/p>\n<h2 id=\"h-oscar-attention-aware-and-deployment-ready\">OSCAR: Attention-Aware and Deployment-Ready<\/h2>\n<p>OSCAR bets the opposite way. Its premise is that at INT2\u2019s four levels, a data-oblivious rotation is the wrong tool \u2014 blindly smoothing ranges is not enough when there is almost no precision to spare. So OSCAR computes an attention-aware rotation from a one-time offline calibration pass: keys are rotated into the eigenbasis of the query covariance, and values into the score-weighted value covariance. A Hadamard transform plus a bit-reversal permutation then spreads channel importance evenly across the quantization groups.<\/p>\n<p>What sets OSCAR apart is that it ships as a complete system, not just an algorithm. It includes a mixed-precision paged cache where sink and recent tokens stay in BF16 while the history compresses to INT2 (at 128K context, only about 0.24% of tokens remain in BF16). It also comes with fused Triton kernels and full SGLang integration (paged-attention and prefix-cache compatible). Precomputed rotations (a \u201cRotationZoo\u201d) are available for models like Qwen3-4B\/8B\/32B, GLM-4.7-FP8, and MiniMax-M2.7, with no recalibration needed.<\/p>\n<p>At an effective 2.28 bits, OSCAR lands within 1.42 points of BF16 on Qwen3-8B and is essentially on par on Qwen3-32B (a 0.02-point gap). On GLM-4.7-FP8, where naive INT2 collapses to zero and data-oblivious baselines reach only low single digits, OSCAR matches BF16 and even edges slightly ahead on the reported benchmarks (within noise). Together AI reports up to 7.83 times job-level throughput and roughly 8 times KV-cache memory reduction at 100K context, with up to about 3 times faster decoding.<\/p>\n<h2 id=\"h-turboquant-vs-oscar-which-one-wins\">TurboQuant vs. OSCAR: Which One Wins?<\/h2>\n<p>Neither \u2014 and that is the honest answer. For deployable INT2 at 128K tokens on supported models, OSCAR is currently the only demonstrated option that does not collapse, and it comes with production-ready SGLang support. For training-free, model-agnostic quantization in the 3 to 4-bit regime, TurboQuant offers far broader generality. OSCAR\u2019s paper reports that TurboQuant drops by more than 40 points at a comparable budget, but that evaluation runs inside OSCAR\u2019s own framework, quantizes all layers, uses a single random seed, and operates well below TurboQuant\u2019s intended bit-width, so it is a weak basis for a head-to-head verdict.<\/p>\n<p>The more interesting possibility is that the two are complementary: pairing a calibration-aware rotation with an optimal scalar quantizer is a promising combination nobody has shipped yet. (Both teams have publicly noted the same idea.)<\/p>\n<h2 id=\"h-the-third-axis-epicache-for-conversational-memory\">The Third Axis: EpiCache for Conversational Memory<\/h2>\n<p>TurboQuant and OSCAR are both built for a single long context. Neither handles extended multi-turn conversations, where history piles up across many exchanges. Apple\u2019s EpiCache is a training-free KV-cache management framework aimed exactly at that gap. It uses block-wise prefill to process history in blocks and keep peak memory bounded. Episodic clustering segments the conversation into coherent semantic \u201cepisodes,\u201d each with its own compressed cache. Episode-matched retrieval routes each query to the most relevant episode at inference time. An adaptive layer-wise budget allocation measures each layer\u2019s sensitivity to eviction and distributes the memory budget accordingly.<\/p>\n<p>Across benchmarks like LongMemEval, RealTalk, and LoCoMo, EpiCache reports up to 40% higher accuracy than eviction baselines, near-full-cache accuracy at 4 to 6 times compression, and up to 3.5 times lower peak memory (and about 2.4 times lower latency). Because it decides which tokens to keep rather than how precisely to store them, it composes directly with OSCAR or TurboQuant for compounding savings.<\/p>\n<h2 id=\"h-how-these-approaches-compose\">How These Approaches Compose<\/h2>\n<p>These three methods are more complementary than competitive. For a developer deploying a long-context application, the choice depends on the constraint. If the priority is a specific bit-width budget, OSCAR leads for INT2 on its supported models. If model portability is key, TurboQuant offers a model-agnostic solution in the 3 to 4-bit range. If the application involves extended multi-turn conversations, EpiCache solves that specific problem and can be layered on top of either quantization method for additional gains.<\/p>\n<p>For example, a team building a customer support chatbot with a 100K context window and a long conversation history could use OSCAR for the 2-bit HV cache compression and then apply EpiCache\u2019s episodic clustering to manage the history across multiple exchanges. The combined savings could be substantial.<\/p>\n<h2 id=\"h-what-this-means-for-developers\">What This Means for Developers<\/h2>\n<p>The practical takeaway is that the KV cache bottleneck now has multiple, viable solutions, each optimized for a different deployment scenario. For anyone working with large language models today, the immediate action is to evaluate their primary constraint \u2014 bit budget, model portability, or conversation length \u2014 and select the approach (or combination) that fits. For INT2 on supported models, OSCAR is production-ready. For a model-agnostic 3-bit solution, TurboQuant is the strong choice. For conversational applications with long histories, EpiCache offers a new capability. The next step is to test these methods against your specific workload, as the best choice will depend on your model, context length, and latency requirements.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Long-context large language models (LLMs) have a memory problem that is not about the weights. During the decoding phase, transformers cache the key and value (KV) vectors for every token at every layer to avoid recomputing the attention mechanism. This cache grows linearly with sequence length and batch size, and at long contexts with high [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":84752,"comment_status":"closed","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/57120.png","fifu_image_alt":"TurboQuant, OSCAR, EpiCache Vie for KV Cache Compression Lead","footnotes":""},"categories":[349],"tags":[],"class_list":["post-57120","post","type-post","status-publish","format-standard","has-post-thumbnail","category-articles"],"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/57120.png","fifu_image_alt":"TurboQuant, OSCAR, EpiCache Vie for KV Cache Compression Lead","_links":{"self":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/57120","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/comments?post=57120"}],"version-history":[{"count":0,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/57120\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media\/84752"}],"wp:attachment":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media?parent=57120"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/categories?post=57120"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/tags?post=57120"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}