Song from a description and lyrics
A style note plus lyrics becomes a full song with vocals and arrangement.
A StepFun and ACE Studio model that turns a style description and lyrics into a full song. This page covers what it can do, how the paper compares it with Suno, and where the official demo lives. The studio button below runs Suno v5.5.
This opens the Suno v5.5 studio on sunov5-5.com and uses your account credits. It does not run StepAudio 3.
StepAudio 3 Music is a long-form music model from StepFun and ACE Studio. You give it a style description and lyrics. It returns a complete song, up to about 5 minutes 30 seconds, in 48 kHz stereo.
People who search the name usually want a sample or a finished song. Samples and generation for this model are on StepFun’s demo and Hugging Face Studio. Songs made on sunov5-5.com come from Suno v5.5, a different generator.
The technical report describes four modes. The official demo also says you can steer style, emotion, vocals, instruments, key, tempo, and song structure.
A style note plus lyrics becomes a full song with vocals and arrangement.
The same model can generate music with no vocals.
Cover-song synthesis is one of the trained tasks in the report.
An a cappella or dry vocal can be turned into a song with accompaniment.
Before the audio is synthesized, the model writes an arrangement plan in ABC notation. The paper calls this step ABC-CoT. The official demo spells the same planning pass ABC-COT: structure and arrangement first, then the recording.
StepFun’s model guide marks ABC notation input as coming soon. The ABC-CoT planning interface is not open. This site does not offer a StepAudio endpoint, and this page does not list request fields.
The weights are not published as an open model. On the Hugging Face Space, people have already asked when the weights will be released and whether an English interface is coming.
The figures below are the authors’ own results in the StepAudio 3 Music Technical Report (arXiv 2609.16034). sunov5-5.com did not re-run those tests.
The paper reports the highest AudioBox Content Enjoyment, Content Usefulness, and Production Quality scores in its comparison, and the highest MuQ-MuLan similarity.
On the preliminary Artificial Analysis Music Arena Vocals board, the paper gives StepAudio 3 Music a Quality Elo of 1105. That place is behind Suno V5.5 and Mureka, and ahead of Suno V5 and MiniMax.
| System | Place in the paper |
|---|---|
| Suno V5.5 | Ahead of StepAudio 3 Music |
| Mureka | Ahead of StepAudio 3 Music |
| StepAudio 3 Music | Quality Elo 1105 |
| Suno V5 | Behind StepAudio 3 Music |
| MiniMax | Behind StepAudio 3 Music |
This order is the paper’s statement about a preliminary board. To hear StepAudio 3 Music, use the official Studio. To make a song on this account, use Suno v5.5.
These pages belong to StepFun. None of them is the Suno v5.5 studio on sunov5-5.com.
Use this path when you want a track from sunov5-5.com today. It runs Suno v5.5, not StepAudio 3.
Write the genre, mood, and whether you want vocals. The form keeps that text on this site.
Continue sends you to the Suno v5.5 workspace. The prompt is passed the same way other guides on this site pass one.
Create the take in the studio, then download it from your history. Credits follow the pricing page.
Short answers for the name, the planning step, and where a song actually gets made.
It is a music model from StepFun and ACE Studio. Style text and lyrics go in. A full song comes out, up to about 5 minutes 30 seconds, at 48 kHz stereo.
No. sunov5-5.com does not run StepAudio 3 weights. The button on this page opens Suno v5.5. To try StepAudio 3 Music, use the official Studio.
ABC-CoT is the paper’s name for a planning pass that writes structure and arrangement in ABC notation before the audio is synthesized. The official demo spells it ABC-COT.
No. StepFun’s model guide marks ABC notation input as coming soon. This website does not publish a StepAudio API.
No. The weights are not released as an open model. People on the Hugging Face Space have asked when that will change, and whether an English interface is planned.
In the technical report, Music Arena Vocals quality Elo is 1105. The authors place that score behind Suno V5.5 and Mureka, and ahead of Suno V5 and MiniMax. Those are the paper’s figures, not a test run by sunov5-5.com.
On the official demo and the Hugging Face Studio linked on this page. The studio on sunov5-5.com is Suno v5.5.
StepFun and ACE Studio. The technical report also lists the Chinese University of Hong Kong and the University of California San Diego. This website is not a StepFun partner and is not the official product.
StepAudio 3 Music stays on StepFun’s demo and Studio. A song from this account is made with Suno v5.5.
Sources: Duration, 48 kHz output, the four modes, ABC-CoT, AudioBox, MuQ-MuLan, and the Music Arena Elo are taken from the StepAudio 3 Music Technical Report, arXiv 2609.16034. Stereo is stated on the official demo. ABC notation input is marked coming soon in StepFun’s model guide. This site did not re-run the benchmarks.