Tencentが2026年8月28日にオープンソース化した次世代フラグシップMoEモデル。総パラメータ770B(活性49B)、コンテキスト100万トークン、Apache 2.0ライセンス。78層のMoE(256ルーティング専門家+共有専門家、トークンあたりtop-8活性)を採用し、Hy3(295B/21B・256K)からサイズとコンテキストを大幅に拡張。Terminal-Bench 2.1で85.4、Toolathlon-Verified 74.1、APEX-Agents 37.1を記録し、ソフトウェア工学・オフィス分析・ゲーム開発・科学研究といった実生産性タスクに特化。キャッシュヒット$0.042/1Mと低価格で、OpenRouter・Tencent Cloud TokenHub経由で利用可能。
パラメータ
770B
コンテキスト長
1M
ライセンス
Apache 2.0
リリース日
2026-08-28
API料金
入力料金(1Mトークンあたり)
$0.834
出力料金(1Mトークンあたり)
$2.501
課金モード: standard
強み
- •総パラメータ770B/活性49BのMoEでオープン最高クラスの規模
- •100万トークン文脈で長文・長時間タスクに強い
- •Terminal-Bench 2.1 85.4でClaude Opus 5級のエージェント型コーディング
- •Apache 2.0で商用利用・微調整・自己ホストが自由
弱み
- •視覚・マルチモーダル非対応(テキストのみ)
- •複雑な課題で過剰に思考・自己検証し遅延・トークン増加
- •preview版で事前・事後学習とも改善余地あり
- •チェックポイントが極めて大きく自前展開のメモリ要件が高い
活用例
- •長期間・複数ファイルのソフトウェア工学エージェント
- •財務・保険などの複雑なオフィス文書分析
- •Unity・Unrealを用いた対話型ゲームプロトタイピング
- •分子動力学・基礎数学などの科学研究支援
深度分析
Total / Active Parameters
770B / 49B (MoE)
78 layers; 256 routed + 1 shared expert, top-8 active; 1M-token context
Context Window
1,048,576 tokens
4x the 256K of Hy3; native MTP speculative-decoding layer
License
Apache 2.0
Weights on Hugging Face and ModelScope; commercial use permitted
Input / Output Price
$0.834 / $2.501 per 1M
Cache hits $0.042/1M; a fraction of closed frontier cost
Terminal-Bench 2.1
85.4
14.6 pts above Hy3; Tencent claims parity with Claude Opus 5
SWE-Bench Pro
51.2
DeepSWE 64.3; Toolathlon-Verified 74.1; APEX-Agents 37.1
強み
- ・Open-weight flagship at 770B/49B that roughly doubles Hy3's capacity and context, landing in the top tier of permissively-licensed models.
- ・85.4 on Terminal-Bench 2.1 and 74.1 on Toolathlon-Verified put it at Opus 5-tier on agentic coding, per Tencent's internal blind test (2.99/4.00 vs GLM-5.3 2.92, Kimi K3 2.94).
- ・Apache 2.0 weights with official vLLM/SGLang recipes and API access via TokenHub and OpenRouter at $0.834/$2.501 per 1M tokens.
- ・Self-improving R&D loop: Tencent reports Hy4 helped optimize its own inference stack for a 31.8% end-to-end throughput gain.
弱み
- ・Text-only — no vision or other multimodal capability, unlike GLM-5.3 or DeepSeek V4-Flash-Vision-Exp.
- ・Tends to over-think and over-verify, which can raise latency and token consumption on complex tasks.
- ・Explicitly an early preview; Tencent states pre-training and post-training both have substantial headroom.
- ・The full checkpoint is very large, so self-hosting still demands serious memory and serving engineering.
競合比較
| Model | Arena | SWE | GPQA | Price |
|---|---|---|---|---|
| Tencent Hy4 preview | 85.4 (Terminal-Bench 2.1) | 51.2 (SWE-Bench Pro) | N/A | $0.834/$2.501 per 1M |
| GLM-5.3 | ~84.5 (Terminal-Bench 2.1) | N/A | N/A | MIT; OpenRouter |
| Kimi K3 | ~83 (Terminal-Bench 2.1) | N/A | N/A | Open weight |
| DeepSeek V4 Pro | ~85 (Terminal-Bench 2.1) | N/A | N/A | Closed API |
Tencent Hy4 preview is the company's largest permissively-licensed model to date — a 770B/49B MoE with a 1M-token context, open-sourced August 28, 2026 under Apache 2.0. It is built for real productivity work (coding, office analytics, game dev, science) and posts Opus 5-adjacent agentic benchmarks at a fraction of frontier cost.
出典
分析生成日: 2026-09-02