On August 4, 2026, Mistral AI released Shieldstral, a 3-billion-parameter open-weights multimodal safety classifier. Operating on a single 16GB GPU, Shieldstral allows developers to define custom moderation policies at inference time using plain-language questions rather than relying on fixed taxonomies. Trained on 54.1 million samples, the model achieved an 84.9% average F1 score on text safety benchmarks, matching models seven times its size, and scored 83.8% on multimodal safety benchmarks.
Aug 10, 2026 · 8 sources
Aug 10, 2026 · 2 sources
Aug 10, 2026 · 3 sources
Aug 10, 2026 · 8 sources
Story comments
Loading comments…