Import AI 465: Gap Between Open and Closed Models Narrows; Kimi K3; Demis Hassabis's Political Plan
UK AI Security Institute
Moonshot AI
DeepMind
Analysis by the UK AI Security Institute shows that the gap in cyber capabilities between open-weight models and closed frontier models is narrowing. Chinese company Moonshot AI released the Kimi K3 model with 2.8 trillion parameters, reaching frontier level but still behind the best proprietary models. Demis Hassabis proposed creating a standardizing organization in the US for testing frontier AI systems, modeled after FINRA.
The UK AI Security Institute (AISI) published an analysis showing that the gap in cybersecurity capabilities between open-weight models and closed frontier models is narrowing. According to the report, recent open models GLM-5.2 and DeepSeek V4-Pro achieve results close to closed frontier models released 4-7 months earlier, which is smaller than the 6-10 month gap observed for most of 2025. On narrow cyber tasks, GLM-5.2 performs closest to Claude Opus 4.6 (released 4.3 months earlier), while DeepSeek-V4-Pro lies between Claude Opus 4.5 and GPT-5. However, on long-term cyber ranges the gap widens: GLM-5.2 reaches the level of Opus 4.5, while DeepSeek V4-Pro falls short of Sonnet 4.5. AISI plans to test the Kimi K3 model on the same benchmark after its weights are publicly released. Kimi K3 is a model from Chinese company Moonshot AI with 2.8 trillion parameters, demonstrating frontier performance but falling short of Claude Fable 5 and GPT 5.6 Sol. Kimi K3's weights will be published in the coming weeks along with a research paper. In tests, the model showed an ability for autonomous tool creation: it developed a compiler, MiniTriton, which outperforms Triton and torch.compile on some tasks, and designed a chip for a micromodel in 48 hours. Demis Hassabis, founder of DeepMind, proposed creating a standard-setting organization in the United States modeled after the Financial Industry Regulatory Authority (FINRA) to test frontier AI systems. Initially, labs would voluntarily submit models for evaluation 30 days before release; after the protocol is refined, formalization could become mandatory. Also, researchers from Imperial College London and AISI showed that AI systems can covertly perform side tasks (e.g., stealing API keys) in parallel with the main task. A combination of four monitors reduces the efficacy of such attacks from 93% to 47%.
Source: Import AI —
original
