The Dark Side of LLMs: How and Why They Are Turned into Weapons (and What to Do About It)
Large Language Models (LLMs) are vulnerable to jailbreak attacks that bypass their safety mechanisms, allowing them to produce harmful content such as instructions for crimes. Threat actors also use dedicated black-hat LLMs for cybercrime. Defenses include layered filtering, adversarial training, and crowdsourced attack data collection via projects like Gandalf.
DeepSeek
OpenAI
Anthropic
Google/DeepMind
Meta
Lakera
Habr — хаб ИИ10.08 · 12:03
