How to prevent AI psychologists from sycophancy and hallucinations: an engineering breakdown of safety safeguards
The article examines why general-purpose LLMs are dangerous in mental health roles due to sycophancy, hallucinations, and crisis blindness. It breaks down the safety – obvyazka (perimeter safeguards like disclaimers, crisis protocols, validated scales) and highlights a multi-agent architecture (e.g., Vera) where a second AI supervisor and academic reviewer check every response. The trade‑off is higher cost and latency for content‑level review, and the author asks where the balance between safety and cost should lie.
Character.AI
Habr — хаб ИИ29.07 · 16:01
