Chinese AI model leads in bug hunting—but lags in exploitation

Chinese AI lab Zhipu just released GLM-5.3, a coding-focused model that unexpectedly leads in one crucial security test: finding software vulnerabilities. On the CyberGym benchmark it tops 84.5%, edging out Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol. The result has sparked headlines, but the full picture is more nuanced—and it matters because vulnerability discovery cuts both ways.
Where the numbers tell different stories
Zhipu’s own release note quietly reveals that GLM-5.3’s edge is narrow and selective. CyberGym measures whether a model can spot a flaw and confirm it’s real; that’s the metric that travelled. But two tougher tests paint a different scene. ExploitBench asks models to reason about a real vulnerability and build an exploit; here GLM-5.3 scores 54.4%, less than Mythos 5’s 78.0%. ExploitGym counts how many exploits a model can finish in a fixed time; GLM-5.3 completes 105 tasks in two hours versus Mythos 5’s 181.
Transparency meets confusion
Zhipu’s technical note clarifies that the further along the security chain—from flaw to exploit—the further behind its model falls. Yet some coverage glossed over that detail, partly because Zhipu compared GLM-5.3 to three different Anthropic models across benchmarks: Opus 4.8 in one table, Fable 5 in charts, and Mythos 5 in cybersecurity. The shifting comparisons make it easy to conflate results that don’t line up.
On coding, the picture is mixed: GLM-5.3 outperforms Opus 4.8 on some tests and trails on others. Zhipu frames the release as progress, not dominance.
Why it matters
The stakes here aren’t just academic. A model that can reliably find software flaws is valuable for defensive audits, but it’s also a potential tool for attackers. Zhipu’s decision to publish GLM-5.3’s weights openly lowers the barrier to entry for both defenders and adversaries. Meanwhile, the gap between discovery and exploitation shows that even leading models still struggle with the harder step of turning theory into a working attack. For organizations relying on AI-assisted security, the key takeaway is to keep expectations calibrated: today’s models are strong scouts, but not yet reliable saboteurs.
Source: AI News. AI-assisted editorial synthesis — TechnoExpress.

