Seedance 2.5
Saiba maisByteDance's newest video model is here. Generate up to 30 seconds in one pass, direct it with 50 image, video, and audio references, and get dialogue, music, and effects locked to the timeline.
SeedMusic is ByteDance's unified music model family from the Doubao Seed team. It combines language-model and diffusion strengths to support controlled vocal music generation, score-to-song conversion, lyric editing, and zero-shot singing voice conversion.
SeedMusic is ByteDance's unified music model family from the Doubao Seed team, officially released in September 2024. It combines language-model and diffusion strengths to support both music generation and post-production editing.
Create vocal music from lyrics, style descriptors, audio references, musical scores, or voice prompts with fine-grained style control.
Edit lyrics, melodies, and timbres directly in generated audio for post-production workflows that need interpretable control.
Core Direction
The Seed team built a unified framework on audio tokens, symbolic music tokens, and vocoder latents. That lets SeedMusic support both expressive generation and editable music workflows for beginners and professional musicians.
Generate vocal music from style descriptions, references, scores, and voice prompts with stronger control over musical attributes.
Turn symbolic musical input into full vocal and instrumental mixes instead of starting from text alone.
Adjust lyrics and melodies in generated audio with note-level editing tools designed for real musician workflows.
What changed
A unified framework covering ten music creation and editing tasks.
Controllable vocal music generation with multi-modal conditioning.
Score-to-song conversion for ideation and arrangement workflows.
Lyrics and melody editing directly inside generated audio.
Zero-shot singing voice conversion from about 10 seconds of user audio.
Expressive vocals across multiple languages with fine-grained style control.
Multimodal Workflow
SeedMusic is designed for teams that want to move from prompt to song quickly, then refine lyrics, melody, or voice without rebuilding the entire track from scratch.
Combine explicit lyrics with genre, mood, structure, and style tags to shape the vocal and instrumental mix.
Steer style and performance with reference clips when you need a clearer sonic direction.
Use score-based guidance for arrangements that need stronger symbolic control and interpretability.
Integrate user voices into music creation with low-threshold singing voice conversion workflows.
How teams use it
Song demos and jingles for campaigns, social content, and product launches.
Lyrics-to-song exploration for writers, producers, and content teams.
Melody and lyric editing for iterative vocal production.
Voice-led music experiments with reference clips and short user recordings.
Recommended flow
Step 1
Start with lyrics, style descriptors, or a score outline that captures genre, mood, and structure.
Step 2
Use audio references or short voice prompts when you need tighter control over style, timbre, or performance.
Step 3
Produce the first mix, then edit lyrics or melody directly in the audio until the track feels ready to share.
FAQ
It excels at controlled vocal music generation plus post-production editing for lyrics, melody, and timbre.
Public release materials describe ten creation tasks across generation, score conversion, editing, and voice cloning.
Yes. SeedMusic is positioned for music novices exploring ideas and professional musicians who need editable control.
Yes. SeedMusic includes zero-shot singing voice conversion using roughly 10 seconds of singing or speech.
It combines auto-regressive language modeling and diffusion across audio tokens, symbolic tokens, and vocoder latents.
Provide lyrics plus descriptors for genre, mood, and structure. Add references or a score when you need stronger control.
Start Here
Use the generator above to iterate on song ideas, then refine lyrics, references, and style until the track feels ready for delivery.