On August 4, 2026, Mistral AI released Shieldstral, a 3-billion-parameter open-weights multimodal safety classifier. Operating on a single 16GB GPU, Shieldstral allows developers to define custom moderation policies at inference time using plain-language questions rather than relying on fixed taxonomies. Trained on 54.1 million samples, the model achieved an 84.9% average F1 score on text safety benchmarks, matching models seven times its size, and scored 83.8% on multimodal safety benchmarks.
Sep 28, 2026 · 8 sources
Sep 26, 2026 · 2 sources
Sep 30, 2026 · 7 sources
Sep 30, 2026 · 4 sources
Story comments
Loading comments…