Mistral AI Releases Shieldstral 1.0 3B Moderation Model
Shieldstral lets services define moderation criteria in natural language, removing the need for retraining when safety rules change.
Reporting from 1 source: GIGAZINE.
Mistral AI released Shieldstral 1.0 3B on August 4, 2026, a 3-billion-parameter AI model that judges whether text or images meet safety standards specified in natural language. It supports 12 languages, runs on a single GPU with 16GB VRAM, and is available under Apache License 2.0.
Mistral AI's Shieldstral 1.0 3B takes a different approach to content moderation. Instead of relying on predefined categories of dangerous content, it accepts safety criteria as questions, such as whether a text recommends violence or whether an image is safe for minors. The model reads the target and outputs probabilities for yes and no, which services can use to hide content or route it to human review.
The model handles both text and images, and can also check whether an AI appropriately refused a dangerous request. In Mistral AI's evaluations, Shieldstral matched larger models on new criteria without additional training and scored highest in image moderation among compared models. It supports 12 languages including English, Japanese, Chinese, and French, and runs on a single GPU with 16GB of VRAM. The weights are on Hugging Face under the Apache License 2.0.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.