Mistral released Shieldstral - a European, multimodal safety classifier with 3 billion parameters and open weights under Apache 2.0. The moderation policy goes in as plain text at the call itself, with no retraining. According to Mistral, the model holds its own against systems up to 7 times larger, and does it on a single card with 16GB of memory.
- Shieldstral: a multimodal safety classifier from Mistral with 3 billion parameters - prompts, responses, refusal, toxicity; text and images (04.08).
- The moderation policy is fed in as text at the call; a new policy needs no retraining. The output is a calibrated result from a single pass.
- Open weights on Hugging Face, Apache 2.0 license; according to Mistral it holds its own against systems up to 7 times larger and runs on a single 16GB card.
Every live AI product carries the same invisible problem: who guards the input and the output. Until now the answer was either someone else's API, or a guard trained once and for all on someone else's idea of 'bad'. Too static.
This one's good, and I'll say why. The guards we had until now were weights - the idea of what's allowed gets baked into the model during training, and after that you live with it as it is. Here the idea is text, your text. The rules for a kids' app, a medical assistant, and an internal company tool become three different files, not three different models.
And the other half: Apache 2.0 on a 16GB card means the guard reaches the small players too. Moderation with your own rules, on your own machine, with not a single line of client text leaving the building - until recently that was a luxury only companies with a whole ML team could afford.
What will stick around, I wonder. The format. The model itself will be overtaken - half a year is a whole lifetime in this league. But a guard that takes its policy on the fly gets swapped out without touching anything around it.