Robots Atlas>ROBOTS ATLAS
Artificial Intelligence

Open AI models catch up to the frontier. The safety gap remains

Sir Robot6 August 2026 · 4 min read
Open AI models catch up to the frontier. The safety gap remains

SaferAI, a nonprofit that evaluates AI safety, published an analysis on August 4, 2026 showing that the open-weight model GLM-5.2 from China's Z.ai now matches the capabilities of the leading closed models. The difference lies elsewhere: the open model lacks their safeguards. The researchers' conclusion is blunt — an edge in capability does not come with control over risk.

Key takeaways

  • GLM-5.2 from Z.ai refused none of the offensive cybersecurity or dual-use biology tasks in the evaluation.
  • Anthropic's Claude Opus 4.7 refused so consistently that SaferAI could not complete the CyberGym benchmark on it.
  • GLM-5.2 has no published safety framework and no pre-deployment testing commitments.
  • Far.ai found hundreds of universal jailbreaks that also work on frontier models.
  • The evaluation covered models including GPT-5.6 Sol, Claude Opus 5, Grok 4.5 and Gemini 3.1 Pro.

Open models have caught up on capability

China's Z.ai released GLM-5.2 with open weights: A model's parameters released publicly — anyone can download, run and modify the model locally, without going through the provider., meaning anyone can download and run it locally. According to SaferAI's tests, the model sits alongside the best closed models in capability, such as GPT-5.6 Sol from OpenAI, the Claude Opus 5 family from Anthropic, Grok 4.5 from xAI and Gemini 3.1 Pro from Google DeepMind.

A year ago open models clearly trailed the frontier. Today that gap in raw capability has all but vanished, and it changes the risk calculus: a powerful model stops being an asset locked inside a few labs and becomes a file to download.

Where the gap is

The problem is not capability itself but the absence of control over it. In SaferAI's tests GLM-5.2 refused none of the tasks involving offensive cybersecurity or dual-use biology — the kind that could aid weapon development.

0dangerous cybersecurity or biology tasks refused by GLM-5.2 in testingSaferAI

The model has no published safety policy, no stated pre-deployment testing, and no formal risk assessment. The contrast with the frontier is stark: Claude Opus 4.7 refused dangerous tasks so consistently that SaferAI could not run the full CyberGym benchmark, which measures capability in cybersecurity operations, on it at all.

Safety dimensionGLM-5.2 (open)Frontier models (closed)
Refusal of dangerous tasksdoes not refuserefuse (Claude Opus 4.7)
Safety policynonepublished
Pre-deployment testingnonecommitted
Jailbreak resilienceunknownfragile — hundreds of jailbreaks

Closed models are not immune either

The gap does not mean closed models are safe. Far.ai demonstrated hundreds of universal jailbreaks — techniques that bypass safeguards with a single reusable prompt — that work on frontier models too. Safety therefore depends not only on whether the weights are open, but on how the model was trained and how firmly it enforces refusals. Henry Papadatos of SaferAI put it this way.

The frontier of capability is not the frontier of risk.

Henry Papadatos, SaferAI.

The China context

The absence of a formal safety framework in GLM-5.2 fits a broader pattern. As Graham Webster of the Stanford Cyber Policy Center notes, Chinese labs have historically focused on filtering "politically sensitive content" rather than the "catastrophic AI risks" as understood in the West. That is a divergence in priorities, not just a technical difference, and it shows that the global market for open models has no shared standard for assessing threats.

Why it matters

Open weights democratize access to advanced AI, but they shift the burden of safety from the provider onto the whole ecosystem. Once a model without safeguards enters public circulation, it cannot be recalled or patched for every user. Rising capability with zero barriers means the risk of misuse in cybersecurity or biology stops being theoretical. At the same time, the resilience of closed models turns out to be fragile, which undercuts the assumption that control over weights is a substitute for a real safety policy.

What's next

  • SaferAI says it will keep publishing evaluations of further models — the full GLM-5.2 report is already available on the organization's site.
  • Regulatory pressure may push makers of open models to publish safety frameworks and test results before releasing weights.
  • The universal jailbreaks found by Far.ai will test whether frontier labs can harden refusals without hurting the usefulness of their models.

Sources

Share this article