SeedMusic AI 音乐生成器
SeedMusic 是字节跳动豆包 Seed 团队推出的统一音乐模型家族,结合语言模型与扩散模型能力,支持可控人声音乐生成、曲谱转歌曲、歌词编辑以及零样本歌声转换等创作任务。
Built for controlled music generation, score-to-song workflows, and editable vocal production
SeedMusic is ByteDance's unified music model family from the Doubao Seed team, officially released in September 2024. It combines language-model and diffusion strengths to support both music generation and post-production editing.
Generation mode
Create vocal music from lyrics, style descriptors, audio references, musical scores, or voice prompts with fine-grained style control.
Editing mode
Edit lyrics, melodies, and timbres directly in generated audio for post-production workflows that need interpretable control.
Core Direction
SeedMusic leans into musician-friendly creative control
The Seed team built a unified framework on audio tokens, symbolic music tokens, and vocoder latents. That lets SeedMusic support both expressive generation and editable music workflows for beginners and professional musicians.
Controllable generation
Generate vocal music from style descriptions, references, scores, and voice prompts with stronger control over musical attributes.
Score-to-song conversion
Turn symbolic musical input into full vocal and instrumental mixes instead of starting from text alone.
Editable vocal production
Adjust lyrics and melodies in generated audio with note-level editing tools designed for real musician workflows.
What changed
Highlights from SeedMusic
A unified framework covering ten music creation and editing tasks.
Controllable vocal music generation with multi-modal conditioning.
Score-to-song conversion for ideation and arrangement workflows.
Lyrics and melody editing directly inside generated audio.
Zero-shot singing voice conversion from about 10 seconds of user audio.
Expressive vocals across multiple languages with fine-grained style control.
Multimodal Workflow
Better control when music needs both generation and editing
SeedMusic is designed for teams that want to move from prompt to song quickly, then refine lyrics, melody, or voice without rebuilding the entire track from scratch.
Lyrics and descriptors
Combine explicit lyrics with genre, mood, structure, and style tags to shape the vocal and instrumental mix.
Audio references
Steer style and performance with reference clips when you need a clearer sonic direction.
Musical score input
Use score-based guidance for arrangements that need stronger symbolic control and interpretability.
Voice prompts and cloning
Integrate user voices into music creation with low-threshold singing voice conversion workflows.
How teams use it
Practical music production scenarios
Song demos and jingles for campaigns, social content, and product launches.
Lyrics-to-song exploration for writers, producers, and content teams.
Melody and lyric editing for iterative vocal production.
Voice-led music experiments with reference clips and short user recordings.
Recommended flow
A simple 3-step working pattern
Step 1
Define the musical intent
Start with lyrics, style descriptors, or a score outline that captures genre, mood, and structure.
Step 2
Add references or voice cues
Use audio references or short voice prompts when you need tighter control over style, timbre, or performance.
Step 3
Generate, edit, and refine
Produce the first mix, then edit lyrics or melody directly in the audio until the track feels ready to share.
FAQ
Questions teams usually ask before switching workflows
What is SeedMusic best at?
It excels at controlled vocal music generation plus post-production editing for lyrics, melody, and timbre.
What tasks does it cover?
Public release materials describe ten creation tasks across generation, score conversion, editing, and voice cloning.
Can beginners and pros both use it?
Yes. SeedMusic is positioned for music novices exploring ideas and professional musicians who need editable control.
Does it support voice cloning?
Yes. SeedMusic includes zero-shot singing voice conversion using roughly 10 seconds of singing or speech.
How does the unified framework work?
It combines auto-regressive language modeling and diffusion across audio tokens, symbolic tokens, and vocoder latents.
How should I prompt it?
Provide lyrics plus descriptors for genre, mood, and structure. Add references or a score when you need stronger control.
Start Here
Open the generator and test a SeedMusic-style workflow
Use the generator above to iterate on song ideas, then refine lyrics, references, and style until the track feels ready for delivery.
