we_are_coded.by CODE · The world, decoded
БГ
Microsoft

Microsoft released an open model that transcribes who says what in real time, but not in Bulgarian

Hugging Face · event date: 2 September 2026Audio

On 2 September Microsoft Research uploaded VibeVoice-ASR-Streaming to Hugging Face in two sizes, under the MIT licence. The model turns speech into text as it flows, separates the speakers and accepts a custom list of names and terms. There are ten languages, and Bulgarian is not among them.

In short
  • Two variants, 1.5B and 7B, both under MIT.
  • Streaming transcription with speaker attribution and custom hotwords.
  • Languages: Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian and Spanish.
Checked on1 October 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

Meeting minutes that write themselves while people talk, and know who said what.

There are plenty of speech recognition (ASR) models. What is interesting here is the combination: real-time streaming, speaker attribution and an open licence that allows commercial use.

The facts: on 2 September 2026 Microsoft Research published the models VibeVoice-ASR-Streaming-1.5B and VibeVoice-ASR-Streaming-7B on Hugging Face under the MIT licence. According to the team's description, it is a unified streaming speech recognition model that transcribes which speaker says what as speech arrives, and accepts custom hotwords, such as names and technical terms, for better recognition of domain-specific content. It supports ten languages: Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian and Spanish. The code is in the microsoft/VibeVoice repository on GitHub.

Bulgarian is missing. Russian is there. For anyone who has run a foreign speech model on our language, that is no surprise. Open weights do not learn a new language on their own.

Before you read the licence, read the list of languages.

There is a practical upside, though. Custom hotwords do their job exactly where general models stumble: names of people, products, places. If you are building a tool for international meetings in English, there is something to take from here today.

For Bulgarian speech this is not the answer. The licence allows someone to fine-tune it on Bulgarian, and that is the next thing we will be looking for in the repositories.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→Hugging Face - microsoft/VibeVoice-ASR-Streaming-7B (model card)→Hugging Face - microsoft/VibeVoice-ASR-Streaming-1.5B (model card)
Original: https://wearecoded.com/en/articles/microsoft-vibevoice-asr-streaming.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news