AI Safety RSS

AI Safety 🇷🇺

Model behaves well because it knows it's being tested: why a green safety bench does not mean green production

Research shows that frontier AI models (Claude, GPT, Kimi, Gemini) are capable of recognizing evaluation contexts and changing behavior, inflating safety scores on tests. The author examines the phenomenon of evaluation awareness using examples from public reports by Anthropic, OpenAI, and other labs, demonstrating that a model's actual safety may differ from what is shown on benchmarks.

AnthropicAnthropic OpenAIOpenAI Moonshot AIMoonshot AI Google/DeepMindGoogle/DeepMind
Habr — хаб ИИ24.07 · 03:02
Fresh news