Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics
Mistral AI Releases Shieldstral, 3B-Parameter Open-Source Safety Classifier for AI Content Moderation
00

Mistral AI Releases Shieldstral, 3B-Parameter Open-Source Safety Classifier for AI Content Moderation

Aug 4, 2026

On August 4, 2026, Mistral AI released Shieldstral, a 3-billion-parameter open-weights multimodal safety classifier. Operating on a single 16GB GPU, Shieldstral allows developers to define custom moderation policies at inference time using plain-language questions rather than relying on fixed taxonomies. Trained on 54.1 million samples, the model achieved an 84.9% average F1 score on text safety benchmarks, matching models seven times its size, and scored 83.8% on multimodal safety benchmarks.

Shieldstral model release

  • ▪Shieldstral is designed to run on a single 16GB Nvidia GPU, enabling on-device or edge deployment.
  • ▪Shieldstral is released under the Apache 2.0 license and is available for download on Hugging Face.
  • ▪Mistral AI framed the release of Shieldstral as the inaugural release from the Open Secure AI Alliance, formed with Nvidia and other organizations.
  • ▪Mistral AI released Shieldstral, a 3-billion-parameter open-weights multimodal safety classifier, on August 4, 2026.

Policy-adaptive moderation design

  • ▪Shieldstral allows operators to define moderation policies at inference time using plain-language yes/no questions without retraining the model.
  • ▪Shieldstral is built on Mistral AI's Ministral-3-3B-Base-2512 backbone with a Pixtral vision encoder to handle both text and image inputs.
  • ▪Shieldstral processes requests using three fields: an instruction field for context, a query field for the yes/no question, and a content field for the text or image.
  • ▪Shieldstral returns a calibrated safety score between zero and one based on the softmax-normalized probabilities of yes and no tokens.

Contrastive training methodology

  • ▪Shieldstral was trained on approximately 54.1 million samples, including 45.2 million open-source text samples, 4.4 million synthetic contrastive text samples, and 4.5 million multimodal samples.
  • ▪The final Shieldstral checkpoint was created by merging three LoRA fine-tunes via spherical linear interpolation.
  • ▪Mistral AI trained Shieldstral using contrastive pairs generated by another language model that rewrote safe text into unsafe variants to teach fine-grained policy distinctions.

Benchmark performance results

  • ▪On a policy-adaptability evaluation, Shieldstral scored a 91.3% F1 score, trailing GPT-OSS-Safeguard-20B which scored 94.1%.
  • ▪Shieldstral achieved an average F1 score of 84.9% on text safety benchmarks, matching the 20-billion-parameter GPT-OSS-Safeguard-20B model.
  • ▪Mistral AI's technical report appendix indicates that Shieldstral trails several baseline models in Arabic and Indonesian language coverage.
  • ▪Shieldstral scored 83.8% on multimodal safety benchmarks, outperforming OmniGuard-7B at 77.6% and LlavaGuard-7B at 71.6%.

Mistral moderation product evolution

  • ▪Mistral AI launched its first content moderation API on November 7, 2024, as a hosted text classifier covering nine fixed categories across 11 languages.
  • ▪Shieldstral is Mistral AI's third moderation model and its first to be released as open weights, following two hosted API versions.

9 sources

Mistral
Introducing Shieldstral. | Mistral AI
View source article
Unite
Mistral’s Shieldstral Packs Policy-Adaptive Safety Screening Into 3B Parameters
View source article
Digg
Mistral AI Releases Shieldstral Safety Model · Digg
View source article
Aiweekly
Mistral open-sources Shieldstral, a 3B multimodal safety guard | AI Weekly
View source article
Siliconangle
Mistral introduces Shieldstral to provide lightweight policy-aware moderation for AI models - SiliconANGLE
View source article

Featured stories

View more in Multimodal models

Nvidia releases Open Agent Safety Platform to contain AI agents after security incidents

Sep 28, 2026 · 8 sources

YouTube and Meta reverse ad bans for Alex Gibney Musk documentary after backlash

Sep 26, 2026 · 2 sources

FTC opens investigation into OpenAI and Anthropic over consumer protection

Sep 30, 2026 · 7 sources

DeepSeek releases software tools for Huawei AI chips to challenge Nvidia

Sep 30, 2026 · 4 sources

Story comments

Loading comments…

Related entities

France

Related Projects

Mistral AI

Topics

Multimodal modelsAI content watermarking & provenanceOpen weights vs closed modelsAI misinformation & deepfakesOpen-source AIAI safety & social impact

Featured stories

View more in Multimodal models

Nvidia releases Open Agent Safety Platform to contain AI agents after security incidents

Sep 28, 2026 · 8 sources

YouTube and Meta reverse ad bans for Alex Gibney Musk documentary after backlash

Sep 26, 2026 · 2 sources

FTC opens investigation into OpenAI and Anthropic over consumer protection

Sep 30, 2026 · 7 sources

DeepSeek releases software tools for Huawei AI chips to challenge Nvidia

Sep 30, 2026 · 4 sources