You type a text, describe the voice and the musical style, and get it read aloud with its own soundtrack. Suno calls it the first audio model that makes voice and music together as one piece. The beta has been open to the whole community since 1 October.
- The input is text: an idea, a poem or something you wrote. The model speaks it and composes the music under it in one track.
- Suno listed the weak spots itself: British accents sometimes wander off to Australia, and dramatic pauses are very dramatic.
- No languages and no price in the post. It was tested for a month with a small group before opening.
A voice note with an unnecessarily epic score under it. That is how Suno tested it, in its own words: turning friends' texts into wildly overproduced dramatic readings, making meditations, poems and bedtime stories for their kids. Out of that game, today, comes a product, open to everyone in beta.
I know this effect from live shows. The same text read over silence and read over the right music are two different texts, and the second one holds people to the end. Until now that took two steps and two tools: read in one, compose in the other, then glue them and argue with the levels. Here it is one text field, and the music arrives with the words.
For a podcast intro, a meditation, a birthday poem or a bedtime story this is more than enough. For an ad that will go on air, I do not see it yet, and Suno itself does not claim it. Beta really does mean beta, they write, and they prove it in the next paragraph.
The best part of the post is the paragraph on weaknesses. A company that writes, itself, that its accent drifts to Australia and back is telling you exactly where the edge of the product is. Give it something of your own and listen to the pauses. That is where you will hear how far it still has to go.