There is one model currently being heavily underestimated, namely Hunyuan Hy3 from Tencent. If you only look at leaderboards, Hy3 is very easy to overlook because its numbers do not stand out and it often loses to peers. Yet once you use it for daily work, especially bug fixing and refactoring, it feels far more mature and careful than a number of models that look far more capable on paper.
I initially overlooked Hy3 for the same reason, seeing its mid-tier Intelligence Index and a non-leading SWE-bench result as a sign of a middle-class model. Only after trying it directly for several days for refactors and bug fixes in a personal repository did I realize the scoreboard does not fully reflect day-to-day experience. Hy3 feels designed for long-horizon work, rarely making the careless mistakes that cost the most time when coding.
Hy3 Benchmark Numbers Are Not Outstanding
The official Hy3 released in July 2026 does not stand out on popular leaderboards. On the Artificial Analysis Intelligence Index it records a score of 42, trailing GLM-5.2 at 51 to 53, DeepSeek V4 Flash 0731 at 50, and Qwen3.8 27B at 52. On real-world agentic benchmarks like GDPval-AA it also does not sit at the very top and still trails those same peers.
On arena.ai (LMArena), the official Hy3 records 1457 on the text leaderboard and 1501 on the coding leaderboard, up from 1412 for the preview version. This places it in the upper-middle tier among hundreds of models, though still below GLM-5.2 and Qwen3.8 on the same arena. A similar picture appears on the coding side, where the official Hy3 scored 78.0 on SWE-bench Verified compared to GLM-5.2 at 84.2, as well as 71.7 on Terminal-Bench 2.1 compared to 81 and 28.0 on DeepSWE compared to 46.2. This coding data comes from Tencent’s technical report and is consistent with third-party testing.
In my view, the composite Intelligence Index covering nine evaluations naturally favors models that are strong across many domains at once, so Hy3’s middle position does not automatically mean it is weak for everyday coding. I also concluded it was middle-tier from the numbers alone, until I realized SWE-bench only measures success on a single isolated task, not comfort over hours of use. The arena climb from 1412 to 1457 also feels more relevant than the ranking itself, showing the official version’s improvement is felt in direct human evaluation.
Tencent’s Take on Benchmarks and Real-World Evaluation
What is interesting is that Tencent did not push a benchmark champion narrative. Instead of showing off leaderboards, they released an independent blind human study with 270 experts testing the model across 312 real-world scenarios. In my view, this kind of blind evaluation is much closer to what users actually experience in the field than automated scores that are often specifically optimized.
The official Hy3 scored 2.67 out of 4 and beat GLM-5.1 at 2.51, with the clearest advantages in frontend development, CI/CD, and data and storage work. Those areas demand caution and consistency rather than raw intelligence, and this is where Hy3 feels more reliable because it rarely changes things outside the instructions.
Tencent also shared internal numbers for the official version that are rarely published.
- The hallucination rate in internal evaluation was recorded at 5.4% for the official version.
- The commonsense error rate was recorded at 12.7% for the official version.
- The issue rate in multi-turn conversations was recorded at 7.9% for the official version.
In my view these numbers matter far more than a few points on the Intelligence Index. A 5.4% hallucination rate means the model far less often fabricates an answer when unsure, and a low multi-turn issue rate is strongly felt during long agent sessions without losing direction.
The post-training principle they instilled is simple, asking the model to answer when it has a clear basis, admit when unsure, and not fabricate. This explains why Hy3 feels more careful. A model like DeepSeek V4 is raw faster and more efficient in tool calling, yet more prone to being careless when writing code. From my experience comparing the two, DeepSeek often finishes steps faster but occasionally makes unrequested changes, while Hy3 almost never makes similar mistakes and feels safer for sensitive work.
Tencent’s focus is not on chasing leaderboard scores but on stabilizing what matters day to day, such as stable tool calling, consistent output formatting, and good recovery when an error occurs mid agentic process. Small things like these never show up on SWE-bench, yet they strongly determine whether a session lasting hours remains comfortable or becomes exhausting.
Consistency Across Coding Tools
One relevant technical detail is Hy3’s accuracy variance staying below 4% on SWE-bench Verified when run through different scaffolding agents such as CodeBuddy, Cline, and Kilo Code. This means the model is designed to stay stable no matter which tool wraps it and is not optimized only for one specific environment. This matches user experience in Kilo Code as well as opencode where performance does not suddenly drop when switching tools.
Personally, I greatly value this stability. With Hy3 the variance below 4% is truly felt and its behavior remains predictable without drastically adjusting instruction style every time I change tools.
Hy3 also comes with an adjustable reasoning effort mode.
- Default mode is a fast mode without long reasoning for standard tasks.
- Low and high modes can be activated specifically for heavy tasks such as complex refactors.
In my view this feature is very helpful for efficiency. For small fixes the default mode is more than enough and far more economical in tokens and time, while for multi-file refactors the high mode provides the depth truly needed without always paying the same overhead.
For me, the combination of cross-tool stability and control over reasoning effort is what makes Hy3 feel designed for real use, not for benchmark demonstrations.
Why Hy3 Is So Underrated
On pure coding benchmarks like SWE-bench, Terminal-Bench, and DeepSWE, the official Hy3 is indeed not yet the champion and still below GLM-5.2. Yet that is exactly where its uniqueness lies and the main reason its reputation does not yet reflect its true quality.
Most people judge coding models only by leaderboards, yet leaderboards like SWE-bench only measure ability on a single isolated task, not comfort over hours in a real session. In my view this habit is what makes Hy3 look underestimated, as many see 78.0 versus 84.2 and immediately consider it a decisive loss without considering reliability when used all day.
There is one assumption I personally find quite plausible. Perhaps Hy3 is simply a genuine model that does not engage in benchmaxxing at all, which is why its benchmark scores appear lower compared to models specifically optimized for leaderboards. If true, its seemingly average numbers are actually evidence that Tencent chose to optimize for real-world experience, which would also explain why real-world performance feels far better than benchmark numbers suggest.
It is on this dimension of real-world use that the official Hy3 quietly excels.
- The hallucination rate was recorded at 5.4% in internal evaluation for the official version.
- The issue rate in long conversations was recorded at 7.9% for the official version.
- Consistency across
codingtools is tightly maintained with variance below 4% on CodeBuddy, Cline, and Kilo Code.
All of the points above do not show up in SWE-bench numbers, yet they strongly determine whether users can trust a model’s work on large-scale refactors or bug fixing that risks breaking other code. I feel this directly when refactoring dozens of files, where low hallucination means the model does not invent non-existent APIs and cross-tool consistency keeps me productive without worrying about a drop when switching scaffolding.
In addition, Hy3 is open weight with an Apache 2.0 license and activates only 21 billion of its total 295 billion parameters. It can currently still be tried for free on Kilo Code and opencode. With far fewer active parameters, Hy3 is much more resource-efficient than models with twice the active count yet still reliable, so the barrier to trying it is very low.
This combination of free access, resource efficiency, and greater trustworthiness is what makes Hy3’s underrated status feel unfair. If your goal is to chase the highest ranking, Hy3 may not be the first choice. But if your goal is to get real work done more calmly, economically, and with fewer careless mistakes, the official Hy3 in my view is one of the most worth considering right now.





