Audio
Audio
20 stories · page 1 of 3Suno launches Speech: spoken word and background music in one track, in beta for everyone
You type a text, describe the voice and the musical style, and get it read aloud with its own soundtrack. Suno calls it the first audio model that makes voice and music together as one piece. The beta has been open to the whole community since 1 October.
Read →ElevenLabs releases Eleven v4 - a voice that reads the tone of a sentence, not just the words
On 28 September ElevenLabs introduced Eleven v4, an entirely new text-to-speech architecture, and the fast Eleven v4 Turbo for voice agents. Per the company the model ranks first on the Artificial Analysis leaderboard, and in blind comparisons about 75 percent of listeners prefer it.
Read →Grok Voice Transcribe 2.0 transcribes twice as accurately as the first version in SpaceXAI tests, at the same price
The new speech-to-text model is built on the audio model behind Grok Voice, detects the language itself and follows a switch mid-recording. The price stays at $0.10 per hour of audio for batch and $0.20 for streaming.
Read →Gemini 3.8 Live talks and thinks at the same time, and it is now available to everyone
Google made two voice models for the Live API generally available. One is for fast conversation without pauses, the other reasons in the background while the conversation goes on.
Read →ElevenLabs put voice, music, image and video into one connector for your chat
The ElevenLabs MCP connector no longer just manages agents, it now makes content as well: speech, transcription, dubbing, music, sound effects, images and video. One install, sign in with your account, no server and no API key.
Read →ElevenLabs Music v2.5: the song you make is yours on every plan, unless it is built on someone else's
On 11 September ElevenLabs made Music v2.5 the default model in ElevenMusic and gave lossless downloads on every plan, including Free. Downloads are blocked, however, for songs built on other artists' songs.
Read →ElevenLabs signed its first major-label deal with Universal Music, and the fan platform doesn't exist yet
ElevenLabs and Universal Music Group entered a multi-year licensing agreement and strategic collaboration. The first step is a fan platform built on licensed music, with remixes, mashups and new interpretations from participating artists. The platform is in development.
Read →GPT-Live-1 listens and speaks at the same time, and it is now in the API, with the voice layer at 5 cents a minute
On 10 September OpenAI brought the voice model GPT-Live-1, so far in ChatGPT, to the API. It listens and speaks at the same moment, can leave the thinking and the tools to another model behind it, and costs $0.05 per minute for the voice layer.
Read →Microsoft released an open model that transcribes who says what in real time, but not in Bulgarian
On 2 September Microsoft Research uploaded VibeVoice-ASR-Streaming to Hugging Face in two sizes, under the MIT licence. The model turns speech into text as it flows, separates the speakers and accepts a custom list of names and terms. There are ten languages, and Bulgarian is not among them.
Read →