{"id":61525,"date":"2026-06-30T09:29:00","date_gmt":"2026-06-30T13:29:00","guid":{"rendered":"https:\/\/overcentral.com\/en\/?p=61525"},"modified":"2026-06-30T09:29:00","modified_gmt":"2026-06-30T13:29:00","slug":"meituan-longcat-2-0-ai-model-chinese-chips","status":"publish","type":"post","link":"https:\/\/overcentral.com\/en\/meituan-longcat-2-0-ai-model-chinese-chips\/","title":{"rendered":"Meituan Releases LongCat-2.0, 1.6T AI Model Trained on Chinese Chips"},"content":{"rendered":"<p><a href=\"https:\/\/www.meituan.com\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">Meituan<\/a> has officially released <a href=\"https:\/\/longcat.chat\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">LongCat-2.0<\/a>, a 1.6-trillion-parameter Mixture-of-Experts (MoE) model, confirming what the developer community had suspected for weeks: this is the architecture behind &#8220;Owl Alpha,&#8221; the anonymous model that surged to the top of OpenRouter&#8217;s global charts over the past two months. The open-weight system, released under a permissive MIT license, brings a native 1-million-token context window and a commercial pricing model that includes free context-cache hits\u2014a structural shift in the economics of large-scale agentic AI deployment. Just as significantly, the model was trained entirely on a cluster of more than 50,000 domestic Chinese ASICs, demonstrating that frontier-class AI can be built without relying on Nvidia&#8217;s GPU infrastructure.<\/p>\n<h2>LongCat-2.0 Pricing: A Promotional Salvo in a Competitive Market<\/h2>\n<p>Meituan is introducing a two-tier commercial access model. Standard pay-as-you-go API pricing is set at $0.75 per million input tokens and $2.95 per million output tokens. A time-limited promotional discount, however, slashes these rates to $0.30 per million input tokens and $1.20 per million output tokens\u2014placing LongCat-2.0 squarely in the mid-range of current frontier model pricing but with a context window that far exceeds most competitors.<\/p>\n<p>What distinguishes this commercial framework is the zero-cost processing of context-cache hits. In agentic workflows where a model repeatedly reads and modifies the same large codebase, this eliminates the compounding cost penalty that typically makes long-context sessions prohibitively expensive. Alongside the standard API, Meituan has launched structured &#8220;Token Packs&#8221; sold through limited flash sales four times daily at 10:00, 16:00, 21:00, and 23:00 Beijing Time. These volumetric allocations are valid for 30 days and stack on top of existing API accounts.<\/p>\n<p>The pricing table below places LongCat-2.0 in context against current market offerings:<\/p>\n<table>\n<tbody>\n<tr>\n<td>\n<p><strong>Model<\/strong><\/p>\n<\/td>\n<td>\n<p><strong>Input ($\/1M)<\/strong><\/p>\n<\/td>\n<td>\n<p><strong>Output ($\/1M)<\/strong><\/p>\n<\/td>\n<td>\n<p><strong>Total ($\/1M)<\/strong><\/p>\n<\/td>\n<td>\n<p><strong>Source<\/strong><\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p>MiMo-V2.5 Flash<\/p>\n<\/td>\n<td>\n<p>$0.10<\/p>\n<\/td>\n<td>\n<p>$0.30<\/p>\n<\/td>\n<td>\n<p>$0.40<\/p>\n<\/td>\n<td>\n<p>Xiaomi<\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p>deepseek-v4-flash<\/p>\n<\/td>\n<td>\n<p>$0.14<\/p>\n<\/td>\n<td>\n<p>$0.28<\/p>\n<\/td>\n<td>\n<p>$0.42<\/p>\n<\/td>\n<td>\n<p>DeepSeek<\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p>deepseek-v4-pro<\/p>\n<\/td>\n<td>\n<p>$0.435<\/p>\n<\/td>\n<td>\n<p>$0.87<\/p>\n<\/td>\n<td>\n<p>$1.305<\/p>\n<\/td>\n<td>\n<p>DeepSeek<\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p>MiniMax-M3<\/p>\n<\/td>\n<td>\n<p>$0.30<\/p>\n<\/td>\n<td>\n<p>$1.20<\/p>\n<\/td>\n<td>\n<p>$1.50<\/p>\n<\/td>\n<td>\n<p>MiniMax<\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p><strong>LongCat-2.0 promo<\/strong><\/p>\n<\/td>\n<td>\n<p><strong>$0.30<\/strong><\/p>\n<\/td>\n<td>\n<p><strong>$1.20<\/strong><\/p>\n<\/td>\n<td>\n<p><strong>$1.50<\/strong><\/p>\n<\/td>\n<td>\n<p>LongCat<\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p>Gemini 3.1 Flash-Lite<\/p>\n<\/td>\n<td>\n<p>$0.25<\/p>\n<\/td>\n<td>\n<p>$1.50<\/p>\n<\/td>\n<td>\n<p>$1.75<\/p>\n<\/td>\n<td>\n<p>Google<\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p>Qwen3.7-Plus<\/p>\n<\/td>\n<td>\n<p>$0.40<\/p>\n<\/td>\n<td>\n<p>$1.60<\/p>\n<\/td>\n<td>\n<p>$2.00<\/p>\n<\/td>\n<td>\n<p>Alibaba Cloud<\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p>MiMo-V2.5<\/p>\n<\/td>\n<td>\n<p>$0.40<\/p>\n<\/td>\n<td>\n<p>$2.00<\/p>\n<\/td>\n<td>\n<p>$2.40<\/p>\n<\/td>\n<td>\n<p>Xiaomi<\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p><strong>LongCat-2.0 standard<\/strong><\/p>\n<\/td>\n<td>\n<p><strong>$0.75<\/strong><\/p>\n<\/td>\n<td>\n<p><strong>$2.95<\/strong><\/p>\n<\/td>\n<td>\n<p><strong>$3.70<\/strong><\/p>\n<\/td>\n<td>\n<p>LongCat<\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p>Grok 4.3 (low context)<\/p>\n<\/td>\n<td>\n<p>$1.25<\/p>\n<\/td>\n<td>\n<p>$2.50<\/p>\n<\/td>\n<td>\n<p>$3.75<\/p>\n<\/td>\n<td>\n<p>xAI<\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p>MiMo-V2.5 Pro (\u2264256K)<\/p>\n<\/td>\n<td>\n<p>$1.00<\/p>\n<\/td>\n<td>\n<p>$3.00<\/p>\n<\/td>\n<td>\n<p>$4.00<\/p>\n<\/td>\n<td>\n<p>Xiaomi<\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p>Kimi-K2.6<\/p>\n<\/td>\n<td>\n<p>$0.95<\/p>\n<\/td>\n<td>\n<p>$4.00<\/p>\n<\/td>\n<td>\n<p>$4.95<\/p>\n<\/td>\n<td>\n<p>Moonshot AI<\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p>GLM-5.2<\/p>\n<\/td>\n<td>\n<p>$1.40<\/p>\n<\/td>\n<td>\n<p>$4.40<\/p>\n<\/td>\n<td>\n<p>$5.80<\/p>\n<\/td>\n<td>\n<p>Z.ai<\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p><a href=\"https:\/\/overcentral.com\/en\/openai-gpt-5-6-delay-trump\/\" title=\"OpenAI Delays GPT-5.6 After Trump Administration Request\" data-iacss-internal=\"1\">GPT-5.6<\/a> Luna<\/p>\n<\/td>\n<td>\n<p>$1.00<\/p>\n<\/td>\n<td>\n<p>$6.00<\/p>\n<\/td>\n<td>\n<p>$7.00<\/p>\n<\/td>\n<td>\n<p>OpenAI<\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p>Grok 4.3 (high context)<\/p>\n<\/td>\n<td>\n<p>$2.50<\/p>\n<\/td>\n<td>\n<p>$5.00<\/p>\n<\/td>\n<td>\n<p>$7.50<\/p>\n<\/td>\n<td>\n<p>xAI<\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p>MiMo-V2.5 Pro (&gt;256K)<\/p>\n<\/td>\n<td>\n<p>$2.00<\/p>\n<\/td>\n<td>\n<p>$6.00<\/p>\n<\/td>\n<td>\n<p>$8.00<\/p>\n<\/td>\n<td>\n<p>Xiaomi<\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p>Qwen3.7-Max<\/p>\n<\/td>\n<td>\n<p>$2.50<\/p>\n<\/td>\n<td>\n<p>$7.50<\/p>\n<\/td>\n<td>\n<p>$10.00<\/p>\n<\/td>\n<td>\n<p>Alibaba Cloud<\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p>GPT-5.6 Sol<\/p>\n<\/td>\n<td>\n<p>$5.00<\/p>\n<\/td>\n<td>\n<p>$30.00<\/p>\n<\/td>\n<td>\n<p>$35.00<\/p>\n<\/td>\n<td>\n<p>OpenAI<\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p><a href=\"https:\/\/overcentral.com\/en\/anthropic-blocks-claude-fable-5\/\" title=\"Anthropic Blocks All Access to Claude Fable 5 After US Government Order\" data-iacss-internal=\"1\">Claude Fable 5<\/a> \/ Mythos 5<\/p>\n<\/td>\n<td>\n<p>$10.00<\/p>\n<\/td>\n<td>\n<p>$50.00<\/p>\n<\/td>\n<td>\n<p>$60.00<\/p>\n<\/td>\n<td>\n<p>Anthropic<\/p>\n<\/td>\n<\/tr>\n<\/tbody>\n<h2>Engineering the 1-Million-Token Sparse Context<\/h2>\n<p>LongCat-2.0 scales total parameters to 1.6 trillion while activating only 33 billion to 56 billion parameters per token, with an average of 48 billion. This aggressive MoE sparsity is achieved through a &#8220;Zero-Compute Experts&#8221; design where routine execution elements pass through lighter subnetworks, eliminating the idle computational overhead that penalizes dense architectures.<\/p>\n<p>To sustain a full 1-million-token context window, Meituan developed LongCat Sparse Attention (LSA), an evolution of DeepSeek Sparse Attention. LSA resolves the quadratic scoring costs and memory fragmentation typical of fine-grained sparse attention through three mechanisms:<\/p>\n<ul>\n<li>\n<p><strong>Streaming-aware Indexing (SI):<\/strong> Restructures token selection by blending hardware-aligned contiguous data reads with dynamic random selection, converting fragmented memory access into predictable sequential blocks for improved High Bandwidth Memory utilization.<\/p>\n<\/li>\n<li>\n<p><strong>Cross-Layer Indexing (CLI):<\/strong> Exploits the finding that attention saliency remains stable across adjacent hidden layers. A single indexing pass guides multiple consecutive layers during inference, reinforced by cross-layer distillation during training.<\/p>\n<\/li>\n<li>\n<p><strong>Hierarchical Indexing (HI):<\/strong> A coarse-to-fine two-stage scoring approach where the indexer performs rapid approximate block-level recall to filter candidates before running fine-grained token selection on the remaining population.<\/p>\n<\/li>\n<\/ul>\n<p>An N-gram Embedding module inherited from Meituan&#8217;s lighter model lines appends 135 billion parameters within a 5-gram token combination framework, expanding the embedding space roughly 100-fold. This captures dense local token relationships and accelerates large-batch inference by reducing memory I\/O bottlenecks.<\/p>\n<h2>Post-Training with MOPD: Specialized Expert Clusters<\/h2>\n<p>LongCat-2.0 is explicitly optimized for agentic software engineering tasks rather than general conversation. The model achieves a score of 59.5 on SWE-bench Pro, narrowly exceeding GPT-5.5&#8217;s 58.6, and posts a 70.8 on Terminal-Bench 2.1, 77.3 on SWE-bench Multilingual, and 73.2 on the FORTE corporate workflow simulator.<\/p>\n<p>This performance comes from a post-training architecture called Multi-Teacher Optimization via Mixture of Specialized Experts (MOPD). Rather than blending human feedback into a single reward function, MOPD segregates optimization into three independent expert clusters:<\/p>\n<ul>\n<li>\n<p><strong>Agent Experts:<\/strong> Fine-tuned strictly for structural execution\u2014precise tool invocation, multi-turn API parameter parsing, and self-correcting loop mechanisms.<\/p>\n<\/li>\n<li>\n<p><strong>Reasoning Experts:<\/strong> Optimized for multi-hop logic, chain-of-thought reasoning, mathematics, and high-level STEM problem-solving.<\/p>\n<\/li>\n<li>\n<p><strong>Interaction Experts:<\/strong> Focused on human alignment, instruction-following, factual grounding to suppress hallucinations, and safety guardrails.<\/p>\n<\/li>\n<\/ul>\n<p>A dynamic gate-routing mechanism fuses these behaviors at runtime, enabling the model to coordinate deep reasoning, stable tool execution, and safe interaction simultaneously. While LongCat-2.0 trails premium systems like Claude Opus 4.8 on broad general-agent benchmarks, it punches above its weight in software engineering\u2014proving highly competitive for complex coding tasks despite a leaner computational footprint.<\/p>\n<h2>Trained on Domestic Chinese ASICs: A Geopolitical Inflection Point<\/h2>\n<p>What makes this release a structural inflection point is its operational independence. LongCat-2.0 was trained entirely on a cluster of over 50,000 domestic Chinese ASICs, proving that near-frontier AI models can be scaled without Nvidia GPUs\u2014the hardware that has powered the vast majority of generative AI training to date.<\/p>\n<p>This successful deployment of alternative silicon arrives precisely as Washington pressures top-tier American labs to restrict access to their latest models. OpenAI was forced to limit access to its new GPT-5.6 models following a U.S. governmental request, and Anthropic was ordered to restrict access to Claude Fable 5 and Mythos 5, taking them entirely offline. Critics argue these defensive regulatory maneuvers have backfired: by locking down Western closed-source models and driving up API costs, the U.S. government has opened a wide operational window for global developers seeking affordable, high-performance alternatives from Chinese open-source ecosystems.<\/p>\n<p>The developer community&#8217;s response was emphatic even before Meituan claimed the architecture. During its unbranded residency on <a href=\"https:\/\/openrouter.ai\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">OpenRouter<\/a> as Owl Alpha, the model accounted for approximately 10.1 trillion monthly tokens\u2014averaging 559 billion tokens per day\u2014representing a 242% month-over-month explosion in volume that propelled it into the platform&#8217;s global top three. It secured the top ranking on the <a href=\"https:\/\/overcentral.com\/en\/nous-research-hermes-agent-profile-builder\/\" title=\"Nous Research Ships Hermes Agent Profile Builder with MCP Server Integration\" data-iacss-internal=\"1\">Hermes Agent<\/a> workspace, second place on Claude Code deployments, and third place across international OpenClaw environments.<\/p>\n<h2>MIT Licensing: Maximum Enterprise Flexibility<\/h2>\n<p>The MIT License allows near-unrestricted freedom for enterprise integration. Unlike copyleft paradigms such as the GPL, which legally obligate developers to open-source derivative frameworks, the MIT license permits deep modification, compilation, and hard-coding into closed-source commercial applications. Corporations can fork the repository, optimize the internal LSA mechanisms for private databases, and sell the resulting software stack without any obligation to disclose proprietary intellectual property.<\/p>\n<h2>Meituan&#8217;s Strategic Transformation<\/h2>\n<p>Founded in 2010 by Wang Xing as a daily deals platform, Meituan evolved into one of China&#8217;s dominant super apps after a 2015 merger with Dianping. The Beijing-based company claims over 770 million annual transacting users and supports a network of more than 14.5 million merchants. Faced with intense domestic competition and margin compression, Meituan publicly committed to investing billions into artificial intelligence and domestic chip capabilities. This strategic shift began materializing in late 2025 with the release of LongCat-Flash, a 560-billion-parameter MoE foundation model, followed by the advanced reasoning model LongCat-Flash-Thinking. LongCat-2.0 represents the most significant step yet in Meituan&#8217;s ambition to become a foundational player in global AI infrastructure rather than remaining strictly a regional e-commerce and delivery giant.<\/p>\n<h2>What This Means for Developers and Enterprises<\/h2>\n<p>For development teams, LongCat-2.0 is already available for direct use. The model can be accessed through Meituan&#8217;s API at longcat.chat with the limited-time promotional pricing, downloaded from GitHub or Hugging Face under the MIT license for self-hosting, or reached through OpenRouter where the Owl Alpha endpoint remains active. The combination of an open-weight architecture, a 1-million-token context window, and zero-cost cache hits fundamentally alters the operational cost economics of large-scale agentic software development. Engineering teams building autonomous codebase migration tools, continuous code review agents, or long-session repository analysis pipelines should evaluate LongCat-2.0&#8217;s pricing against its actual performance on their specific workflows\u2014particularly for tasks that benefit from repeated passes over large codebases where the cache-hit economics become decisive. The broader takeaway is that the compute infrastructure for frontier AI is diversifying rapidly, and developers who remain locked into a single hardware ecosystem or API provider may be missing significant cost and capability advantages emerging from alternative approaches.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Meituan has officially released LongCat-2.0, a 1.6-trillion-parameter Mixture-of-Experts (MoE) model, confirming what the developer community had suspected for weeks: this is the architecture behind &#8220;Owl Alpha,&#8221; the anonymous model that surged to the top of OpenRouter&#8217;s global charts over the past two months. The open-weight system, released under a permissive MIT license, brings a native [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":84227,"comment_status":"closed","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/61525.png","fifu_image_alt":"Meituan Releases LongCat-2.0, 1.6T AI Model Trained on Chinese Chips","footnotes":""},"categories":[349],"tags":[],"class_list":["post-61525","post","type-post","status-publish","format-standard","has-post-thumbnail","category-articles"],"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/61525.png","fifu_image_alt":"Meituan Releases LongCat-2.0, 1.6T AI Model Trained on Chinese Chips","_links":{"self":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/61525","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/comments?post=61525"}],"version-history":[{"count":0,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/61525\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media\/84227"}],"wp:attachment":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media?parent=61525"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/categories?post=61525"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/tags?post=61525"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}