Open-Weight AI Models Are Reaching Advanced Cyber Capabilities
By Suad Seferi ·

Key takeaways
- NIST describes GLM-5.3 as the most cyber-capable open-weight AI model it has evaluated so far.
- Anthropic says simple techniques bypassed GLM-5.3's cyber safeguards in 64% to 100% of its simulated tests.
- The broader issue is that advanced cyber capabilities are moving from restricted frontier systems into downloadable models.
Advanced cybersecurity capabilities that only months ago were largely confined to tightly controlled frontier AI systems are beginning to appear in models that anyone can download. Anthropic said on September 29 that GLM-5.3, an open-weight model developed by Chinese AI company Z.ai, can autonomously build sophisticated end-to-end cyber exploits and has safeguards that the company found relatively easy to bypass. The findings add to growing evidence that the gap between restricted frontier AI systems and openly available models is narrowing in cybersecurity. GLM-5.3 was released by Z.ai in August, with its model weights subsequently made publicly available. Anthropic tested the model as part of its Frontier Red Team research and reported that simple techniques could bypass GLM-5.3's safeguards between 64% and 100% of the time in its simulated evaluations. By comparison, Anthropic said the same attacks were unsuccessful against the safeguarded Claude models included in its testing. The results do not mean GLM-5.3 is the world's most capable AI model for cybersecurity. An independent assessment published earlier by the U.S. National Institute of Standards and Technology's Center for AI Standards and Innovation found that GLM-5.3 still performs below the strongest U.S. frontier models. But NIST reached another important conclusion: GLM-5.3 is currently the most cyber-capable open-weight model it has evaluated. The access gap is getting smaller The distinction matters because some of the most powerful cybersecurity capabilities developed by U.S. AI companies are not freely available. Anthropic, for example, has previously restricted access to some advanced cyber capabilities through trusted-access programmes designed for vetted security organisations. Open-weight models work differently. Once their weights are publicly released, organisations and individuals can download them, modify them and run them on their own infrastructure. That makes conventional API restrictions much harder to enforce. According to Anthropic, GLM-5.3 demonstrated capabilities similar to those shown five months earlier by Claude Mythos Preview, a restricted Anthropic system capable of autonomously developing sophisticated exploits. NIST's benchmark results provide additional context. GLM-5.3 solved 40.4% of tasks in SEC-Bench Pro and achieved 61.1% on ExploitBench, while the strongest U.S. models evaluated by NIST performed considerably better. Across its combined cyber capability measure, NIST estimated that GLM-5.3 trails the current U.S. frontier by approximately four months. Four months may sound substantial in ordinary software development. In frontier AI, it is not. The more important change is that capabilities that previously required access to restricted systems are appearing in models whose weights can be downloaded publicly. Useful to defenders, too The same capabilities are not inherently malicious. Models capable of identifying vulnerabilities and developing exploits can help security teams discover weaknesses before attackers do. Anthropic explicitly acknowledged this dual use, arguing that defenders should also have access to highly capable AI systems as attackers adopt increasingly powerful tools. The challenge is that defensive and offensive cybersecurity often rely on many of the same technical skills. A system that can find a vulnerability for a security researcher may also help someone exploit it. That makes cybersecurity one of the clearest areas where the debate around open AI models becomes more complicated than a simple choice between open and closed development. Why this matters for the Balkans For organisations in the Western Balkans, the immediate issue is less about which country currently leads AI benchmarks and more about the falling cost of sophisticated cyber capability. Businesses, public institutions and civil society organisations across the region are adopting AI tools while often working with limited cybersecurity staff and budgets. If advanced vulnerability discovery and exploit development become available through downloadable models, the technical barrier for both attackers and defenders falls. That puts more pressure on ordinary security practices: patching known vulnerabilities, restricting system access, monitoring unusual activity, separating sensitive infrastructure and ensuring that AI agents do not receive unnecessary permissions. The development also complicates AI governance. Restrictions imposed by individual AI providers can reduce misuse through hosted services, but they cannot easily control models once weights have been released and distributed. GLM-5.3 is unlikely to be the last model to raise that problem. The question now is how quickly openly available models will close the remaining cybersecurity capability gap with the most advanced restricted systems.
Frequently asked questions
What is GLM-5.3?
GLM-5.3 is an artificial intelligence model developed by Chinese AI company Z.ai, formerly known as Zhipu AI. Its model weights have been publicly released.
Is GLM-5.3 more capable than U.S. frontier AI models?
No. NIST found that it remains below the strongest U.S. models on its cybersecurity benchmarks, estimating a gap of about four months on its combined capability measure.
Why is its open-weight release important?
Open-weight models can be downloaded, modified and operated independently. This makes provider-level restrictions and API safeguards harder to enforce once the model has been distributed.
Can these capabilities also help cybersecurity teams?
Yes. The same capabilities used to discover and exploit vulnerabilities can help defenders identify weaknesses, test systems and fix security problems.