AI SafetyResearch 🇺🇸 27.07.2026 21:04

Import AI 461: 'Alignment is not on track'; FrontierCode; and synthetic research interns

CognitionCognition Alibaba/QwenAlibaba/Qwen AnthropicAnthropic OpenAIOpenAI
Researchers from the UK AI Security Institute and Timaeus launch Sequent, a nonprofit aiming to develop principled alignment techniques for superintelligent AI, raising $100-150M initially. Cognition releases FrontierCode, a hard coding benchmark where Claude Opus 4.8 scores 13.4% on the hardest tier. Xiaomi publishes details on MiMo-V2.5-Pro-UltraSpeed, a 1-trillion parameter model achieving 1000 tokens per second via co-design and quantization. A new benchmark, AARRI-Bench, evaluates AI systems on entry-level research tasks.
A new nonprofit research organization called Sequent has been formed by researchers from the UK AI Security Institute and alignment theory startup Timaeus to develop alignment techniques for superintelligent AI systems. Sequent aims to raise $100-150M initially and grow to 40-80 employees, pursuing diverse research directions including scalable oversight, learning theory, heuristic arguments, game theory, and personas to achieve principled confidence in alignment. Cognition, the company behind Devin, released FrontierCode, a coding benchmark with 150 tasks in three difficulty tiers (Diamond, Main, Extended) that measures code quality, mergeability, and adherence to codebase standards; Claude Opus 4.8 achieved only 13.4% on the Diamond tier, indicating substantial headroom. Xiaomi introduced MiMo-V2.5-Pro-UltraSpeed, a 1-trillion parameter LLM that achieves 1000 tokens per second inference on an 8-GPU commodity node using FP4 quantization, speculative decoding via DFlash, and TileRT software from startup Tile AI. Researchers from Xi'an Jiaotong University and Xidian University developed the AARR (Act As a Real Researcher) benchmarks, starting with AARRI-Bench, to evaluate AI systems on entry-level research tasks; the best performing system, Claude-Opus-4.7 with Mini-Swe-Age, scored 42.13% on the Science scenario.
Сокращения
ASI = Artificial Superintelligence
RSI = Recursive Self-Improvement
VLM = Vision-Language Model
QA = Question Answering
LLM = Large Language Model
QC = Quality Control
PR = Pull Request
FP4 = Floating Point 4-bit
TPS = Tokens Per Second
GPU = Graphics Processing Unit
Source: Import AI — original
Our earlier posts on this topic ↓
Fresh news