AI Import 461: 'Alignment Not on Right Track'; FrontierCode; and Synthetic Science Interns
Cognition
Alibaba/Qwen
Anthropic
OpenAI
Researchers from UK AISI and Timaeus founded the non-profit organization Sequent to develop alignment methods ensuring the safety of superintelligent systems. Cognition released the FrontierCode benchmark for evaluating code quality, where the best model, Claude Opus 4.8, scored only 13.4% on the hardest level. Also presented are the ChinaHeritaQA dataset for evaluating VLM cultural knowledge about Chinese UNESCO sites and the Xiaomi MiMo-V2.5-Pro-UltraSpeed model with a generation speed of 1000 tokens per second.
Former employees from the alignment team at the UK AI Security Institute and the startup Timaeus have joined forces to create a nonprofit organization called Sequent, which will focus on developing alignment techniques that provide high confidence in the safety of superintelligent AI. Sequent plans to raise $100–150 million in seed funding and hire 40–80 employees, betting on a portfolio of research directions: scalable oversight, learning theory, heuristic arguments, game theory, and personas.
Cognition, the creator of Devin, has introduced FrontierCode, a benchmark consisting of 150 coding tasks divided into three difficulty tiers. The best result at the Diamond tier was achieved by Claude Opus 4.8, scoring 13.4%, underscoring the benchmark's difficulty.
On the Chinese market, Xiaomi announced the MiMo-V2.5-Pro-UltraSpeed model with 1 trillion parameters and a generation speed of 1,000 tokens per second, achieved through joint optimization of architecture and software, including FP4 quantization and DFlash speculative decoding.
Researchers from several universities have created ChinaHeritaQA, a multimodal benchmark comprising 2,279 images of 51 UNESCO World Heritage sites in China and 14,133 questions. The best open model, Qwen-VL-8B-Instruct, achieved 81% accuracy, surpassing the human baseline of 67%.
Additionally, researchers from Xi'an Jiaotong University and Xidian University developed the AARR benchmarks, which assess the ability of AI agents to perform tasks typical of a research intern. The top-performing agent was Claude-Opus-4.7 paired with Mini-Swe-Age.
Source: Import AI —
original
