Mistral Launches Shieldstral, Lightweight AI Safety Model

Mistral Launches Shieldstral, Lightweight AI Safety Model
Mistral AI has introduced Shieldstral, a 3-billion-parameter open-weight multimodal safety model that lets developers write moderation policies in natural language at runtime. The model returns a simple yes or no safety verdict, requires no retraining, and outperforms models up to seven times its size across safety benchmarks, scoring 84.9% on text and 83.8% on image safety. Shieldstral can distinguish between closely related but different policies, such as separating malware instructions from cybersecurity discussion. Lightweight enough to run on a single 16GB GPU, it can operate alongside larger models to guard inputs and outputs with minimal latency, even in edge environments.
Read the original article →