Tencent Hy reasoning model for coding, instruction following, and agent tasks
Hy3 by tencent is offered by 22 providers on this page. Public prices are shown for 19 of them. Its official list price is $0 per 1M tokens. The lowest paid channel is NanoGPT at $0.066 / $0.26 per 1M, about 2.7× below the list price. That channel is a third-party gateway, so confirm its availability and rate limits before depending on it. It also has 1 free ($0) channel and 2 covered by a paid subscription; free tiers usually carry rate limits, and subscription-covered access bills $0 per token only after the subscription fee.
Artificial Analysis rates it 25.3 on the Intelligence Index, with 58.8 for coding and 24.1 for agentic tasks. Against its lowest blended price of $0.115 per 1M, that is roughly 221 index points per dollar, which is the value ratio the leaderboards rank on. Median output speed is 88 tokens per second, with 3.07s to the first token. Running one task of the Artificial Analysis suite costs about $0.0718, which reflects how many tokens its reasoning consumes rather than the unit price alone.
The context window is 256,000 tokens at the reference host, but hosts report different limits, from 202,752 to 262,144, so the usable window depends on the provider you pick. It supports reasoning, tool use, and structured output. The weights are open, so it can also be self-hosted or served through a gateway of your choice. Providers report the capability flags inconsistently, so verify a specific feature against the host you plan to use.
Quality is independently evaluated by Artificial Analysis. Speed/latency are model-level medians.
| Provider | Tier | Input | Output | Cache read | Cache write | Context | Output limit | Status |
|---|---|---|---|---|---|---|---|---|
| Tencent TokenHubOfficial | First-party | Subscription | $0 | $0 | 256,000 | 128,000 | ||
| Kenari | Gateway | Free | — | — | 256,000 | 128,000 | ||
| Tencent Token Plan | Gateway | Subscription | $0 | $0 | 256,000 | 128,000 | ||
| NanoGPT | Gateway | $0.066 | $0.26 | $0.029 | — | 262,144 ⚠ | 128,000 | |
| Deep Infra tencent/Hy3 | Cloud | $0.13 | $0.53 | $0.033 | — | 262,144 ⚠ | 128,000 | |
| Kilo Gateway | Gateway | $0.13 | $0.53 | $0.033 | — | 262,144 ⚠ | 128,000 | |
| Eden AI deepinfra/tencent/Hy3 | Gateway | $0.13 | $0.53 | $0.033 | — | 262,144 ⚠ | 128,000 | |
| SiliconFlow tencent/Hy3 | Cloud | $0.13 | $0.53 | $0.033 | — | 262,144 ⚠ | 262,144 | |
| OpenRouter | Gateway | $0.13 | $0.53 | $0.033 | — | 262,144 ⚠ | 128,000 | |
| DevPass (LLM Gateway) | Gateway | $0.13 | $0.53 | $0.033 | — | 262,144 ⚠ | 128,000 | |
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.
Interleaved thinking (reasoning between tool calls) is declared by 2 of 22 providers.
1 / 22 providers expose no reasoning control (reasoning_options: []).
3 more channels offer $0 (Tencent TokenHub, Kenari, Tencent Token Plan); free tiers usually have rate limits and no SLA, excluded from ranking.
| Name | Conditions | Score | Metric | Source |
|---|---|---|---|---|
| SWE-Bench Verified | — | 78 | resolved | Source ↗ |
| SWE-Bench Multilingual | harness: SWE-agent · variant: highest reasoning effort | 75.8 | score | Source ↗ |
| SWE-Bench Pro | harness: SWE-agent · variant: highest reasoning effort | 57.9 | score | Source ↗ |
| Terminal-Bench | harness: Terminus 2 · variant: highest reasoning effort; 4h timeout; 500 episodes · v2.1 | 71.7 | score | Source ↗ |
| NL2Repo | harness: Claude Code · variant: highest reasoning effort; 250 turns; 12000s timeout | 45.6 | score | Source ↗ |
| DeepSWE | harness: mini-swe-agent · variant: highest reasoning effort; 2h timeout | 28 | score | Source ↗ |
| BrowseComp | harness: Tencent internal search harness · variant: highest reasoning effort | 84.2 | score | Source ↗ |
| WideSearch | harness: Tencent internal search harness · variant: highest reasoning effort | 76.4 | score | Source ↗ |
| DeepSearchQA | harness: Tencent internal search harness · variant: highest reasoning effort | 91 | score | Source ↗ |
| MCP Atlas | harness: Scale AI · variant: highest reasoning effort; April 2026; 100 tool calls · dataset: 500 public tasks | 79.1 | score | Source ↗ |
| Toolathlon | variant: highest reasoning effort | 48.5 | score | Source ↗ |
| APEX-Agents | variant: highest reasoning effort | 25.6 | pass@1 | Source ↗ |
| ClawEval | harness: Tencent internal harness · variant: highest reasoning effort · dataset: 105 queries · v20260325 | 68.5 | pass@3 | Source ↗ |
| WildClawBench | harness: OpenClaw · variant: highest reasoning effort · dataset: 35 text-only tasks | 53.6 | score | Source ↗ |
| SkillsBench | harness: Claude Code · variant: highest reasoning effort · dataset: 79 text-only tasks | 55.3 | average over 3 runs | Source ↗ |
| Humanity's Last Exam | variant: highest reasoning effort; with tools · dataset: text-only | 53.2 | score | Source ↗ |
| Humanity's Last Exam | variant: highest reasoning effort; without tools · dataset: text-only | 37 | score | Source ↗ |
| GPQA Diamond | variant: highest reasoning effort | 90.4 | score | Source ↗ |
| FrontierScience | variant: highest reasoning effort; research | 21.3 | score | Source ↗ |
| FrontierScience | variant: highest reasoning effort; olympiad | 74.8 | score | Source ↗ |
| USAMO | variant: highest reasoning effort · v2026 | 72 | score | Source ↗ |
| MathArena Apex | variant: highest reasoning effort | 38.7 | score | Source ↗ |
| ArxivMath | variant: highest reasoning effort | 52.2 | score | Source ↗ |
| HorizonMath | variant: highest reasoning effort | 7.1 | pass@12 | Source ↗ |
| PHYBench | variant: highest reasoning effort | 77.4 | score | Source ↗ |
| CMT Benchmark | variant: highest reasoning effort | 37.8 | score | Source ↗ |
| IMOAnswerBench | variant: highest reasoning effort | 90 | score | Source ↗ |
| SuperChem | variant: highest reasoning effort | 54.9 | score | Source ↗ |
| CL-bench | variant: highest reasoning effort | 23.8 | score | Source ↗ |
| CL-bench-life | variant: highest reasoning effort | 17 | score | Source ↗ |
| AA-LCR | variant: highest reasoning effort | 73.4 | score | Source ↗ |
The same benchmark scores very differently across harness / dataset, so the qualifying conditions must be shown together.
Individual evaluations run by Artificial Analysis, on the same reasoning tier as the intelligence score above. Each benchmark has its own task set and harness, so rows are not comparable with one another. The Intelligence Index above draws on a different, newer set of evaluations.