{"id":62405,"date":"2026-07-07T12:23:24","date_gmt":"2026-07-07T16:23:24","guid":{"rendered":"https:\/\/overcentral.com\/en\/?p=62405"},"modified":"2026-07-07T12:23:24","modified_gmt":"2026-07-07T16:23:24","slug":"tencent-hy3-moe-model","status":"publish","type":"post","link":"https:\/\/overcentral.com\/en\/tencent-hy3-moe-model\/","title":{"rendered":"Tencent Drops Hy3: 295B MoE with 21B Active and 256K Context"},"content":{"rendered":"<p>Tencent&#8217;s Hy team has released Hy3, a 295-billion-parameter Mixture-of-Experts (MoE) model that activates only 21 billion parameters per token while supporting a 256,000-token context window. The weights are available under the Apache License 2.0, and the model is specifically designed for reasoning, agentic workflows, and long-context tasks. With 192 experts and top-8 routing, Hy3 keeps computational costs low while delivering competitive performance across coding, STEM, and agent benchmarks. An FP8 quantized checkpoint, Hy3-FP8, is also available for reduced memory footprint during serving.<\/p>\n<h2>What Is Hy3?<\/h2>\n<p>Hy3 is a sparse MoE model with 192 experts and top-8 routing \u2014 only eight experts fire per token, which keeps inference efficient. The architecture includes a Multi-Token Prediction (MTP) layer that predicts several tokens at once, enabling faster decoding through speculative decoding in both vLLM and SGLang. The model has 80 layers (excluding the MTP layer), 64 attention heads with Grouped Query Attention using 8 <a href=\"https:\/\/overcentral.com\/en\/kv-cache-compression-methods\/\" title=\"TurboQuant, OSCAR, EpiCache Vie for KV Cache Compression Lead\" data-iacss-internal=\"1\">KV<\/a> heads and a head dimension of 128, a hidden size of 4096, an intermediate size of 13312, and a vocabulary of 120,832 tokens. The context window spans 256K tokens, and the model supports BF16 precision.<\/p>\n<p>The MTP layer itself contains 3.8 billion parameters. Total model size is 295B parameters, with only 21B activated per token, which is the core efficiency advantage of the MoE design.<\/p>\n<h2>Benchmark and Performance<\/h2>\n<p>Tencent&#8217;s research team published scores across coding, agents, and STEM domains. On coding benchmarks, Hy3 achieves 78.0 on SWE-Bench Verified, 57.9 on SWE-Bench Pro, and 75.8 on SWE-Bench Multilingual. Terminal-Bench 2.1 lands at 71.7, and DeepSWE at 28.0. On STEM and reasoning, the model reports 90.4 on GPQA Diamond, 72.0 on USAMO 2026, 90.0 on IMOAnswerBench, and 53.2 on HLE (with tools).<\/p>\n<p>In a blind test with 270 experts collecting 312 valid comparisons on real workflows, Hy3 scored 2.67 out of 4, ahead of GLM-5.1 at 2.51. The largest advantage appeared in frontend development, CI\/CD, and data and storage tasks.<\/p>\n<h2>Reliability and Production Behavior<\/h2>\n<p>Tencent&#8217;s team focused three specific areas of production reliability. First, tool calling and output formatting received improvements that reduced invalid calls causing infinite loops. <a href=\"https:\/\/overcentral.com\/en\/meta-ai-brain2qwerty-v2-decoding\/\" title=\"Meta AI\u2019s Brain2Qwerty v2 Decodes Typed Sentences at 61% Accuracy\" data-iacss-internal=\"1\">Accuracy<\/a> variance across CodeBuddy, Cline, and KiloCode scaffoldings stays within 4% on SWE-Bench Verified. Second, anti-hallucination training reduced the hallucination rate from 12.5% to 5.4% and commonsense error rates from 25.4% to 12.7% in internal evaluations. Third, multi-turn intent tracking improved through joint supervised fine-tuning and reinforcement learning, dropping the internal issue rate from 17.4% to 7.9% and raising MRCR long-dialogue benchmark scores from 42.9% to 75.1%.<\/p>\n<h2>How to Call Hy3<\/h2>\n<p>Hy3 exposes an OpenAI-compatible API. You deploy it with vLLM or SGLang, then call the endpoint. A single flag, <strong>reasoning_effort<\/strong>, controls how much the model thinks before responding. Use <strong>no_think<\/strong> for direct answers, <strong>low<\/strong> for light reasoning, and <strong>high<\/strong> for deep chain-of-thought on math, coding, or multi-step tasks. The recommended settings are temperature 0.9 and top_p 1.0. You can also try Hy3 without local hardware \u2014 OpenRouter lists a tencent\/hy3:free route at no cost, with the free tier scheduled to end on July 21, 2026.<\/p>\n<h2>Hy3 vs GLM-5.2: Size Versus Raw Performance<\/h2>\n<p>Tencent&#8217;s research team benchmarked Hy3 against <a href=\"https:\/\/overcentral.com\/en\/glm-5-2-matches-mythos-cybersecurity\/\" title=\"Z.ai GLM-5.2 Matches Mythos on Cybersecurity\" data-iacss-internal=\"1\">GLM-5.2<\/a>, a roughly 744B MoE model with about 40B active parameters. GLM-5.2 leads on coding: 84.2 vs 78.0 on SWE-Bench Verified, 83.0 vs 75.8 on SWE-Bench Multilingual, 81.0 vs 71.7 on Terminal-Bench 2.1, and 46.2 vs 28.0 on DeepSWE. However, Hy3 operates at roughly half the total size and half the active parameter count. The tradeoff is clear: Hy3 trades some top-end coding accuracy for a dramatically smaller active footprint, which matters when self-hosting and paying for GPUs. Hy3 is also released under Apache 2.0, while GLM-5.2 uses open weights with different licensing terms.<\/p>\n<h2>Deployment Notes<\/h2>\n<p>Serving Hy3 requires significant memory. Tencent recommends 8 GPUs such as the H20-3e or cards with larger memory. Both vLLM and SGLang ship recipes with MTP enabled. A minimal vLLM launch uses tensor-parallel size 8, speculative config method mtp with 2 speculative tokens, and the hy_v3 tool-call parser and reasoning parser. For compression, Tencent provides the AngelSlim toolkit covering quantization, low-bit methods, and speculative sampling, along with a complete finetuning pipeline for Hy3.<\/p>\n<h2>What This Means for Developers<\/h2>\n<p>Hy3 gives developers an open, permissively licensed MoE model optimized for agentic and long-context workloads. Its 21B active parameter count makes it feasible to self-host on a cluster of 8 GPUs, and the 256K context window handles full repository scans, lengthy contracts, or multi-turn agent sessions. The <strong>reasoning_effort<\/strong> flag provides practical control over compute cost versus depth of thought. If you are building coding agents, document processing pipelines, or financial analysis tools that need grounded outputs and long-context support, Hy3 is worth testing now \u2014 either via the free OpenRouter route or by deploying the open weights on your own hardware.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Tencent&#8217;s Hy team has released Hy3, a 295-billion-parameter Mixture-of-Experts (MoE) model that activates only 21 billion parameters per token while supporting a 256,000-token context window. The weights are available under the Apache License 2.0, and the model is specifically designed for reasoning, agentic workflows, and long-context tasks. With 192 experts and top-8 routing, Hy3 keeps [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":84390,"comment_status":"closed","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/62405.png","fifu_image_alt":"Tencent Drops Hy3: 295B MoE with 21B Active and 256K Context","footnotes":""},"categories":[349],"tags":[],"class_list":["post-62405","post","type-post","status-publish","format-standard","has-post-thumbnail","category-articles"],"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/62405.png","fifu_image_alt":"Tencent Drops Hy3: 295B MoE with 21B Active and 256K Context","_links":{"self":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/62405","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/comments?post=62405"}],"version-history":[{"count":0,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/62405\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media\/84390"}],"wp:attachment":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media?parent=62405"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/categories?post=62405"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/tags?post=62405"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}