AI News

Tencent Hy3: The Budget Reasoning Model to Watch

By CoreAI · · 5 min read · 2 views
Tencent Hy3: The Budget Reasoning Model to Watch

Twenty cents per million input tokens. That is what Tencent Hy3 charges for a 295-billion-parameter reasoning model with a 262,144-token context window — released July 6, 2026 under an Apache 2.0 license with no regional restrictions. While the headlines that week chased trillion-parameter flagships, Tencent quietly shipped the most interesting budget-tier model of the summer: a mixture-of-experts design that activates just 21B parameters per token, supports configurable reasoning effort, and cut its own hallucination rate by more than half versus its predecessor. If you have been waiting for cheap reasoning that does not feel cheap, Hy3 is the one to test — and unlike most launches in this class, the specs and the license both hold up under scrutiny.

Key Takeaways

  • Tencent Hy3 is a 295B-parameter mixture-of-experts reasoning model (21B active, 192 experts with top-8 routing), released July 6, 2026.
  • It costs $0.20 per million input tokens and $0.80 per million output — firmly budget tier, with cached input at just $0.05 per million.
  • Context window is 262,144 tokens with up to 131,072 output tokens per response, and reasoning effort is configurable.
  • Tencent reports the hallucination rate dropped from 12.5% to 5.4% versus its predecessor, and says it rivals open models with 2-5x the parameters.
  • Weights are open under Apache 2.0, and the hosted model is live on CoreAI alongside 300+ others.

What Is Tencent Hy3?

Hy3 is the newest reasoning and agent model from Tencent's Hunyuan team, and the culmination of what the lab describes as a complete model-development loop in under six months — from a February 2026 architecture rebuild to the July release. The architecture is a textbook modern MoE: 295B total parameters spread across 192 experts, with top-8 routing activating about 21B per token. You get big-model knowledge with small-model serving costs, which is exactly how a 295B system ends up priced like a compact.

Two design choices stand out. First, configurable reasoning effort: you decide how long the model thinks before answering, so quick lookups stay fast while hard problems get real deliberation — the trade-off we unpacked in our guide to when thinking mode pays off. Second, the reliability push: Tencent reports the hallucination rate fell from 12.5% to 5.4% generation-over-generation. For a budget model — the class most prone to confident nonsense — that is arguably the most important number in the release notes.

How Cheap Is Hy3, Really?

Here is Hy3 against its budget-tier peers and one standard-tier reference point, using live per-token pricing:

ModelContextInput / 1MOutput / 1MTier
Tencent Hy3262,144$0.20$0.80Budget
Poolside Laguna XS 2.1262,144$0.06$0.12Budget
Nex-N2-Mini262,144$0.025$0.10Budget
Grok 4.5 (reference)500,000$2.00$6.00Standard

Hy3 is not the absolute cheapest row in that table — Laguna XS 2.1 and Nex-N2-Mini undercut it. But those are 33B and mini-class models; Hy3 brings 295B total parameters and Tencent's claim that it rivals flagship open models with two to five times its size. The honest framing: Hy3 is the premium end of the budget tier — a tenth of standard-tier pricing for reasoning quality that punches well above its bracket. It joins the value class we mapped in our budget AI models under a dollar roundup, and immediately becomes one of its strongest general-purpose entries.

CoreAI app — all AI models, one subscription

Why Does the Apache 2.0 Release Matter?

Tencent published full weights on Hugging Face and GitHub under Apache 2.0 with no regional restrictions — the most permissive mainstream license, covering commercial use, modification, and redistribution. For enterprises, that means Hy3 can be self-hosted, fine-tuned on private data, and audited, with the hosted version as the low-effort default.

It also cements a pattern: Chinese labs now dominate the open-weight value segment. DeepSeek, Qwen, GLM, and now Hunyuan are shipping permissively licensed models that Western labs keep behind APIs, a shift we traced in our analysis of Chinese AI models winning enterprise share. Hy3 arriving the same month as Kimi K3's record-setting open release makes July 2026 the strongest month open-weight AI has ever had.

What Should You Actually Use Hy3 For?

Our recommendation: make Hy3 your default for reasoning-heavy volume work — summarization with actual analysis, data extraction that requires judgment, agent pipelines that run thousands of times a day, drafting where you will edit anyway. The 262K context handles full documents and long sessions, the 131K output ceiling means it will not truncate a long report mid-table, and the improved hallucination numbers make it more trustworthy than the budget class average.

Where we would still pay up: client-facing writing, gnarly multi-step engineering, and anything where a wrong answer is expensive. That remains flagship territory. But the entire point of a model this cheap is that trying it costs nothing meaningful — and the gap between budget and standard tiers has narrowed enough in 2026 that your assumptions from even six months ago are probably stale.

How Can You Try Tencent Hy3 Today?

Self-hosting a 295B MoE requires serious hardware, and API access means keys, quotas, and per-token invoices. The shortcut: Hy3 is already live on CoreAI's model library on iOS, Android, and the web — one subscription covering 300+ models, plans from $57/month. The revealing experiment is to open Compare, run your standard workload through Hy3 and a standard-tier model like Grok 4.5 side by side, and see whether you can tell the difference on your tasks. If you cannot, you have just found a ten-fold cost improvement. If you can, you have learned exactly where your quality bar sits — either way, ten minutes well spent.

Try any AI model free on CoreAI

Frequently Asked Questions

What is Tencent Hy3?

Hy3 is a 295B-parameter mixture-of-experts reasoning and agent model from Tencent's Hunyuan team, released July 6, 2026 under Apache 2.0. It activates 21B parameters per token via 192 experts with top-8 routing, supports configurable reasoning effort, and has a 262,144-token context window.

How much does Tencent Hy3 cost?

$0.20 per million input tokens and $0.80 per million output, with cached input at $0.05 per million — budget-tier pricing, roughly a tenth of standard-tier flagships. On CoreAI it is included under one flat subscription with no per-token billing.

Is Hy3 good at reasoning for a budget model?

That is its calling card. Tencent reports it outperforms similar-size models and rivals open flagships with 2-5x its parameters, while the hallucination rate dropped from 12.5% to 5.4% versus its predecessor. Configurable reasoning effort lets you dial thinking depth up for hard problems and down for speed.

Is Tencent Hy3 open source?

Yes — full weights are published on Hugging Face, ModelScope, and GitHub under Apache 2.0 with no regional restrictions, permitting commercial use, fine-tuning, and redistribution. Most users will still find the hosted version far more practical than self-hosting a 295B MoE.

Where can I try Hy3 without an API key?

On CoreAI, across iOS, Android, and the web. It sits in the model library alongside 300+ others, and Compare mode lets you benchmark it against standard-tier models with your own prompts to see if the budget tier covers your needs.

Every AI model. One app.

Chat with Tencent Hy3, Claude, GPT, and 300+ other AI models, compare them side by side, and use 74 free AI tools — on iOS, Android, and the web.

Try CoreAI on the Web

Related Posts

Meta Muse Spark 1.1: The Everything-Input AI Model
AI NEWS

Meta Muse Spark 1.1: The Everything-Input AI Model

Meta's Muse Spark 1.1 takes text, images, video, audio, and PDF documents in a single model with a 1M-token context. Here is what the everything-input
5 min read
Kimi K3: Moonshot's 2.8T Multimodal Model Explained
AI NEWS

Kimi K3: Moonshot's 2.8T Multimodal Model Explained

Moonshot AI just shipped Kimi K3, a 2.8 trillion parameter open-weight multimodal model with a 1M-token context window. Here is what it actually does
6 min read
Gemini 3.5 Flash: Google's Everything Model, Explained
AI NEWS

Gemini 3.5 Flash: Google's Everything Model, Explained

Google's Gemini 3.5 Flash quietly became the default brain behind the Gemini app, AI Mode in Search, and half of Google's agent stack. It's fast, chea
5 min read