China's DoGNAVY Ranks Third in Global AI Security Benchmark, Paving the Way for Agent Safety
Microsoft
OpenAI
Google DeepMind
Anthropic
DeepSeek
In the latest CyberGym benchmark, DoGNAVY, a joint effort by a top Chinese AI research institute and security team DARKNAVY, ranked third globally with a 90.8% pass rate, surpassing OpenAI and Anthropic models. Unlike competitors using multiple closed-source models, DoGNAVY relies on a single open-source model, GLM-5.2, and introduces AgentDoG, a diagnostic framework for AI agents.
CyberGym, a benchmark from UC Berkeley, ranks AI agents on their ability to autonomously discover, verify, and exploit vulnerabilities from 1507 real-world fixed vulnerabilities across 188 open-source projects. DoGNAVY, developed by a top Chinese AI research institute and DARKNAVY, achieved a pass rate of 90.8% (1369 tasks), ranking third globally behind Microsoft's MDASH (92.0%) and Wiz×Google DeepMind's Atlas (90.9%), and ahead of OpenAI GPT-5.5-Cyber (85.6%) and Anthropic Claude Mythos (83.1%). The top two systems rely on routing tasks to the best-performing closed-source models, which poses availability and cost challenges for Chinese teams. In contrast, DoGNAVY uses only one open-source model, GLM-5.2 from Zhipu, which is freely downloadable from Hugging Face. DoGNAVY emulates the workflow of elite security researchers through a structured workflow, vulnerability reachability analysis, static-dynamic feedback loops, and a knowledge base of platform-agnostic security research experience. The system's architecture includes AgentDoG, an open-source security framework that diagnoses agent behavior, identifies risk sources, and explains decision motivations. AgentDoG's training uses a smart selection mechanism to pick key samples and data purification, resulting in a small model with 78.4% accuracy on complex risk identification tasks, comparable to top models. This achievement demonstrates that with open-source models and real-world security expertise, China can handle the most challenging security tasks, making advanced AI security capabilities more accessible.
- Abbreviations
- API = Application Programming Interface — программный интерфейс приложения
- PoC = Proof of Concept — доказательство концепции
- IoT = Internet of Things — интернет вещей
Source: QbitAI 量子位 —
original
