ResearchAI Safety 🇺🇸 27.07.2026 16:04

Google evaluates alignment of behavioral tendencies in large language models

Google/DeepMindGoogle/DeepMind
Google Research introduces a framework that converts psychological questionnaires into situational judgment tests (SJTs) to assess behavioral alignment of LLMs. Testing 25 models reveals that smaller models often deviate from human consensus, while larger models show improved but imperfect alignment. The study also finds models are overconfident when human opinions diverge.
Google Research presents a framework for evaluating behavioral dispositions in LLMs by adapting established psychological questionnaires into situational judgment tests. The framework tests models in realistic scenarios such as professional composure and conflict resolution. Analysis of 25 LLMs shows that smaller models (<25B parameters) have markedly lower directional alignment, often at near-chance rates, while larger frontier models (>120B) achieve close to perfect alignment under unanimous human consensus but plateau in the low-to-mid 80s when consensus is lower. Models tend to encourage emotional openness where humans recommend composure, prioritize harmony over assertiveness, and exhibit higher impulsivity. The study also reveals that models are systematically overconfident when human annotators disagree, failing to represent the range of opinions. This highlights the need for better behavioral alignment to appropriately navigate social dynamics.
Сокращения
LLM = Large Language Model — большая языковая модель
SJT = Situational Judgment Test — ситуационный оценочный тест
IRI = Interpersonal Reactivity Index — индекс межличностной реактивности
ERQ = Emotion Regulation Questionnaire — опросник регуляции эмоций
Source: Google Research — original
Our earlier posts on this topic ↓
Fresh news