Mistral unveils open model to control AI-generated content
Mistral
Mistral has launched Shieldstral, a 3-billion-parameter open-weight multimodal security model designed to detect risky content and enforce governance policies. It runs on a single Nvidia GPU with 16 GB of memory and is published under Apache 2.0. Developers can define control rules using natural language questions.
Amid the growing deployment of generative AI assistants in enterprises, content control is becoming a central issue. To address this, Mistral has launched Shieldstral, a 3-billion-parameter open-weight multimodal security model designed to identify risky content and apply appropriate governance policies. It supports 12 languages. Published under Apache 2.0, the model makes its weights available to developers for easy integration and adaptation, and it can run on a single Nvidia GPU with 16 GB of video memory. Shieldstral is available for download on Hugging Face. Unlike traditional filtering systems trained on predefined categories such as violence, hate speech, fraud, or explicit sexual content, Shieldstral uses a configurable approach where developers can define control rules as natural language questions, for example, 'Does this content incite physical violence?' The request can be supplemented with instructions specifying the context of analysis, expected severity, or the nature of the content (text or image). The model provides a binary safety evaluation, generating just one token for the final decision based on the probabilities of 'yes' and 'no' answers. It can be used at different stages of AI processing, either before sending a user request to a generative model or after generation to check a response before publication. According to Mistral, Shieldstral performs as well as or better than open guard models up to seven times larger across several benchmarks, particularly in textual safety, refusal detection, and multimodal evaluation, achieving an average global score of 84.9% on textual safety assessments and 83.8% on multimodal safety tests including image analysis.
- Abbreviations
- GPU = Graphics Processing Unit — графический процессор
Source: Le Monde Informatique — IA —
original
