Breaking
AI & MLDeveloping Story

GLM-5.3 Claims Superior Bug-Finding Edge

Zhipu's new AI model outperforms Western rivals on a cybersecurity benchmark, signaling China's rapid advance.

··1 hour ago·3 min read
a person's head with a circuit board in front of it
Photo by Steve A Johnson on Unsplash

Chinese AI firm Zhipu has released a new model it claims can find software vulnerabilities better than leading American systems, a development that could reshape assumptions about where the world's most potent cyber tools are being built. The company's announcement, made public last week, positions GLM-5.3 as a state-of-the-art bug-finder on a key benchmark, backed by tests on real-world codebases that uncovered thousands of flaws.

Benchmark Success on CyberGym

According to Zhipu's announcement, GLM-5.3 surpasses both Fable 5 and GPT-5.6 Sol on the CyberGym benchmark, a test designed to measure a model's ability to tackle real-world cybersecurity challenges. The company stated, "As we scaled post-training, cyber capability developed faster than we expected. GLM-5.3 is state of the art on CyberGym for vulnerability discovery, and its gains are largest further up the exploitation chain."

The company added that the model "did not simply become better at identifying isolated flaws: it began to reason across multiple stages of exploitation, forming coherent plans for complete exploitation chains." This suggests a leap in capability beyond simple pattern matching, moving toward strategic, multi-step attack planning.

Thousands of Flaws Found

Zhipu claims it tested GLM-5.3 on real-world codebases in collaboration with Chinese companies, resulting in the discovery of 2,436 vulnerabilities across 269 projects. Of these, 1,097 were rated medium-to-high severity. The vulnerabilities spanned a wide range of software, including system kernels, operating systems, browser engines, open-source infrastructure, web applications, and network protocols.

  • 2,436 total vulnerabilities found
  • 269 projects examined
  • 1,097 medium-to-high severity issues

The announcement notes that "many had remained unnoticed for years or even decades, with the oldest dating back roughly 40 years." This points to the model's ability to uncover long-dormant weaknesses, potentially raising concerns about the security of aging systems.

Performance Nuances

While GLM-5.3 excelled at bug-finding, the company's own data shows it performed worse than Western models on other security and coding benchmarks. This mixed result suggests that the model's strengths are specifically concentrated in vulnerability discovery rather than general-purpose coding or security analysis.

The fact that Zhipu developed this capability so quickly, shortly after the debut of Anthropic's Mythos, indicates that China is not far behind in the race to build AI systems capable of poking holes in rivals' software. Any advantage the US felt it had as the home of Anthropic has therefore dissipated, according to the Register's analysis.

Focus on Exploitation Chains

Zhipu's emphasis on reasoning across multiple stages of exploitation is particularly notable. The company claims the model can form coherent plans for complete exploitation chains, rather than just identifying individual bugs. This suggests a more advanced understanding of how vulnerabilities can be chained together to achieve a broader compromise.

This capability, if verified, could have significant implications for both offensive and defensive cybersecurity. For defenders, it underscores the need for comprehensive patching and monitoring. For attackers, it represents a powerful tool that could automate complex attack planning.

Real-World Testing

Zhipu says it worked with Chinese companies to test the model on real-world codebases, lending some credibility to the findings. However, the lack of independent verification means these claims should be treated with caution. The company's announcement provides benchmark data, but third-party validation is still lacking.

Given the sensitive nature of these capabilities, it is possible that full details of the testing methodology will not be disclosed. Nevertheless, the sheer volume of vulnerabilities reported suggests a level of effectiveness that warrants attention.

Geopolitical Implications

The development is a reminder that AI-driven cybersecurity is a global race. China's rapid progress in this area could shift the balance of power in cyber conflict, enabling it to find and exploit vulnerabilities in Western software more easily. This could increase the pressure on US and allied governments to invest in defensive measures.

The Register reports that the US may have felt secure in its position as the home of leading AI companies, but Zhipu's claims challenge that assumption. As AI models become more adept at finding bugs, the advantage may increasingly go to those who can deploy them at scale.

For businesses and governments, the message is clear: the era of AI-powered vulnerability discovery is here, and it is not limited to one country. Organizations must assume that their software may face automated, sophisticated attacks, and they need to adopt proactive security measures accordingly.

#zhipu#glm-5.3#cybergym#vulnerability#china#ai security

Sources

Iliyas

Founder & Editor, Xploitwire

This article was compiled from the sources listed above and checked against them for accuracy, under editorial policies set by Iliyas. Read our Editorial Policy →

← Back to all stories