Reading Zhipu’s GLM-5.3 results past the headline number

🤖 AI-GENERATED✓ HUMAN-REVIEWED⚡ Posted 16 minutes after it broke⏱ 4 min read📡 AI News

The short version

Chinese AI model GLM-5.3 narrowly beats US rivals on a key cybersecurity benchmark for finding vulnerabilities, but lags significantly on more advanced exploitation tasks.

Zhipu’s new GLM-5.3 AI model has made headlines by scoring 84.5% on the CyberGym benchmark, slightly outperforming leading American models from Anthropic and OpenAI in finding software vulnerabilities. This result holds strategic weight due to the dual-use nature of such cybersecurity capabilities, which can serve both defense and offense. However, Zhipu’s own release note and other benchmarks reveal a more nuanced picture, showing the model trails its rivals substantially in more complex, time-sensitive exploitation scenarios.

Key takeaways

  • Zhipu’s GLM-5.3 scored 84.5% on the CyberGym benchmark, narrowly beating top US models in vulnerability discovery.
  • The capability is dual-use, valuable for both security audits and potential attacks, making the result strategically sensitive.
  • Zhipu plans to openly publish the model’s weights, unlike some US rivals who restrict access to similar models.
  • On more demanding benchmarks like ExploitBench and ExploitGym, GLM-5.3 lags significantly behind US models.
  • Zhipu’s own release note acknowledges rapid improvement in an area of relative weakness, framing the headline result as a narrow win within a broader competitive gap.

The Headline Cybersecurity Result and Its Strategic Weight

Zhipu’s GLM-5.3 scored 84.5% on the CyberGym benchmark, narrowly outperforming Anthropic’s Mythos 5 (83.8%) and OpenAI’s GPT-5.6 Sol (83.6%). This benchmark tests a model’s ability to find and confirm software vulnerabilities from source code. The result gained significant attention for positioning a Chinese model ahead of leading American counterparts in this specific capability.

The strategic impact is heightened because vulnerability discovery is a dual-use capability. It is valuable for both defenders auditing their own code and for potential attackers conducting the same activity against others’ software. This dual nature is why Anthropic’s similar model sits behind restricted access.

In contrast, Zhipu intends to publish GLM-5.3’s weights openly for anyone to download. The company’s own release note acknowledges the sensitivity, stating this capability “is growing fastest exactly where we are furthest behind.” While the CyberGym result was the headline, Zhipu’s paper shows it trails the American models on two other, more demanding cybersecurity benchmarks: ExploitBench and ExploitGym.

Zhipu’s Measured Self-Assessment and the Broader Benchmark Gap

Zhipu’s own release note for the GLM-5.3 model provides a more nuanced view than the headlines it generated. The company explicitly states that its capability “is growing fastest exactly where we are furthest behind,” framing its cybersecurity progress as rapid improvement in an area of relative weakness rather than claiming overall superiority.

A Narrow Lead Amidst a Wider Gap

The much-reported CyberGym result, where GLM-5.3 scored 84.5% against top American models, represents the narrowest of three cybersecurity benchmarks Zhipu published. On the more challenging ExploitBench, which requires reasoning about how to exploit a vulnerability, GLM-5.3 scored 54.4%, significantly behind Anthropic’s Mythos 5 at 78.0% and OpenAI’s GPT-5.6 Sol at 76.5%. Similarly, on ExploitGym, which measures tasks completed under time constraints, GLM-5.3’s results (105 tasks in two hours) lagged behind Mythos 5’s 181.

This data presents three different pictures of performance. While the model shows a marginal lead on one specific code-reading task, the other two benchmarks reveal a substantial performance gap in more advanced exploitation scenarios, consistent with Zhipu’s own acknowledgment of playing catch-up.

Divergent Results on More Complex Exploitation Tasks

On the ExploitBench benchmark, which requires reasoning about how to exploit a real vulnerability, GLM-5.3 scored 54.4%. While this is a major improvement from its predecessor’s 24.4%, it remains far behind the scores of Anthropic’s Mythos 5 (78.0%) and OpenAI’s GPT-5.6 Sol (76.5%).

Separately, the ExploitGym benchmark measures how many exploitation tasks a model can complete within a fixed time limit. According to Zhipu’s release note, GLM-5.3 completed 105 tasks in two hours and 130 tasks in six hours. In the same time frames, Mythos 5 completed significantly more, at 181 and 247 tasks respectively.

These two benchmark results, which have received less coverage than GLM-5.3’s top score on the vulnerability-finding CyberGym test, paint a different picture. They indicate a substantial capability gap between GLM-5.3 and the leading American models in practical, time-sensitive exploitation scenarios.

Context: GLM-5.3’s Release and the Competitive Landscape

Zhipu, which also trades as Z.ai, is one of a handful of Chinese labs releasing models that compete with the American frontier. On August 14, 2026, it launched GLM-5.3, a coding-focused model, and published a technical release note setting out its performance. This event underscores the intense benchmark competition between major AI labs.

The release generated headlines based on a specific, narrow victory: GLM-5.3 scored 84.5% on the CyberGym cybersecurity benchmark, slightly ahead of American rivals Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol. However, Zhipu’s own note provides a more measured view, detailing broader performance disparities. On the more demanding ExploitBench, GLM-5.3 scored 54.4%, significantly behind Mythos 5’s 78.0% and GPT-5.6 Sol’s 76.5%.

A Measured Self-Assessment

Zhipu itself acknowledges the competitive gap in its release note, stating that capability “is growing fastest exactly where we are furthest behind.” This candid admission highlights how specific benchmark wins can drive publicity even when a more comprehensive evaluation shows a wider performance gap with frontier models.

📡 Original reporting: AI News. AI Craft Technologies’ news engine summarised and rewrote this story in our own words; facts are drawn from the linked source.

⚙️ How this article was made — fully automated
01📡 ScanOur engine watches trusted AI & tech sources in real time.
02🤖 WriteAI drafts an original summary in the ACT house style.
03🎨 IllustrateA custom hero image is generated for every story.
04📤 PublishReviewed, posted, and shared to social — hands-free.

This is a live demo of the ACT News Factory engine. Want one running on your own site? See our services →

Share this project