On 2 September Microsoft Research uploaded VibeVoice-ASR-Streaming to Hugging Face in two sizes, under the MIT licence. The model turns speech into text as it flows, separates the speakers and accepts a custom list of names and terms. There are ten languages, and Bulgarian is not among them.
- Two variants, 1.5B and 7B, both under MIT.
- Streaming transcription with speaker attribution and custom hotwords.
- Languages: Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian and Spanish.
Meeting minutes that write themselves while people talk, and know who said what.
There are plenty of speech recognition (ASR) models. What is interesting here is the combination: real-time streaming, speaker attribution and an open licence that allows commercial use.
Bulgarian is missing. Russian is there. For anyone who has run a foreign speech model on our language, that is no surprise. Open weights do not learn a new language on their own.
There is a practical upside, though. Custom hotwords do their job exactly where general models stumble: names of people, products, places. If you are building a tool for international meetings in English, there is something to take from here today.
For Bulgarian speech this is not the answer. The licence allows someone to fine-tune it on Bulgarian, and that is the next thing we will be looking for in the repositories.