Subquadratic Claims Breakthrough That Slashes LLM Costs and Power Use

A startup unveils a sparse attention method that could dramatically reduce the computational cost and energy consumption of large language models.

By Central
Subquadratic's sparse attention approach promises to slash LLM costs and energy use for many common tasks.
Highlights
  • Subquadratic's method replaces the dense attention mechanism in LLMs with a sparse alternative to cut calculations.
  • The quadratic scaling of dense attention makes long-context processing increasingly costly as context windows expand.
  • Subquadratic positions its sparse inference engine for specific use cases rather than a full replacement of existing pipelines.

Subquadratic, a startup specializing in efficient AI architectures, has unveiled an approach it claims can dramatically reduce the computational cost and energy consumption of large language models without sacrificing performance on many common tasks. The company’s method replaces the dense attention mechanism that powers virtually every modern LLM with a sparse alternative, potentially slashing the number of calculations required for long-context processing by orders of magnitude. While Subquadratic’s system is not positioned as a universal replacement for today’s top models, the company argues that for a broad class of use cases—summarization, document analysis, retrieval-augmented generation—the efficiency gains could fundamentally alter how LLMs are built. “We hope we’re kicking off a new age of efficiency,” says Justin Dangel, cofounder and CEO of Subquadratic. “We don’t think anybody will be building on transformers in a few years.”

Why Dense Attention Drives Up LLM Costs

To understand the significance of Subquadratic’s claim, it helps to look at how most large language models actually process text. The core mechanism inside an LLM is a type of neural network called a transformer, which runs a process known as dense attention. (The foundational paper of the LLM era, published by Google researchers in 2017, was titled “Attention Is All You Need.”) When a transformer processes a chunk of text, it first encodes each word—or part of a word, known as a token—as a numerical vector. To capture the relationships between every part of the input, it then multiplies each token’s representation with every other token’s representation. For a piece of text 10,000 words long, that operation triggers nearly 50 million individual multiplications. “If you want to summarize The Great Gatsby, you have to look at the first word and the last word together, and then you have to look at every other combination,” Dangel explains.

The Quadratic Scaling Problem

This brute-force pairwise computation creates a fundamental scalability bottleneck. As the length of the input text grows, the number of computations grows quadratically: each new token must be multiplied by every token that came before it. Double the number of words, and you roughly quadruple the number of operations. A simple mental model helps visualize the curve. Draw a circle and place dots around its edge; each dot represents a token. Connect every pair of dots with a line to represent the multiplication of those two tokens. Five dots produce 10 lines. Ten dots produce 45 lines. Twenty dots produce 190 lines. The pattern is relentless, and it is the primary reason LLMs are notorious power hogs. As context windows continue to expand—some models now support 100,000 tokens or more—the quadratic cost of dense attention becomes increasingly untenable for real-world deployment at scale.

Sparse Attention: How Subquadratic Cuts the Computation

Subquadratic’s solution is to abandon dense attention altogether in favor of what is known as sparse attention. Instead of multiplying every token against every other token, sparse attention selects only a subset of token pairs to compute. The insight is straightforward: not all relationships between words in a piece of text carry equal weight. Many word pairs contribute little to the overall meaning, especially in longer documents. By intelligently choosing which token pairs to compute and which to skip, Subquadratic’s method preserves the quality of the model’s understanding while eliminating a large fraction of the arithmetic. This is not the first attempt at sparsity in attention mechanisms, but Subquadratic claims its specific implementation achieves practical speed and power improvements at inference time that previous efforts could not deliver at scale.

What This Means for LLM Deployment

The practical implications are significant for any organization deploying LLMs for tasks involving long documents. Customer support systems processing lengthy conversation histories, legal teams reviewing contracts, and knowledge management tools indexing entire libraries could all benefit from substantially lower latency and reduced compute costs. Sparse attention does not eliminate the need for dense computation in every scenario—tasks that genuinely require global reasoning across every token pair may still demand a traditional transformer approach—but Subquadratic argues that a large and growing share of real-world LLM workloads can be served with sparse methods. The company is positioning its technology as an alternative inference engine that teams can drop in for specific use cases rather than a wholesale replacement of existing training pipelines.

Who Should Monitor This Development

Engineering teams building or serving LLM-based products should track Subquadratic’s progress closely, particularly if their workloads involve long-context processing where quadratic scaling creates cost or latency pain today. The company has not yet released public benchmarks or open-source code, so independent validation remains pending. However, the direction of the underlying research is consistent with a broader industry push toward more efficient attention mechanisms. If Subquadratic’s claims hold up under external scrutiny, the technology could offer a practical way to extend the reach of LLMs into cost-sensitive and latency-sensitive applications without waiting for the next generation of hardware. For now, teams should evaluate their own token-per-request profiles and ask whether a sparse attention path could serve a meaningful fraction of their inference load. That calculation is exactly the kind of due diligence that will determine whether Subquadratic’s “new age of efficiency” arrives quickly or incrementally.

Share This Article