Kubeez · Audio · Seed 1.0 Audio (ByteDance)
One prompt, a whole scene: voices, sound effects, and music together
ByteDance's Seed 1.0 Audio is prompt-driven: describe the scene and it generates the voices, sound effects, and background music together, not just a voice reading a line. Clone a voice from up to three short reference clips (tag them @Voice1-3), or guide the delivery with a single image. Great for audiobooks, video dubbing, games, and radio-drama skits. Billed per output second from the same Kubeez wallet as ElevenLabs, Gemini, music, and video. Section images below are Nano Banana 2 marketing stills; audio comes from the dialogue engine.
Showcase






























































































































































































































Describe the environment, the music, the action, and each voice, and Seed performs the whole scene at once.

From one free-form prompt Seed generates the spoken lines, the sound effects, and the background music together, so a scene lands finished instead of flat.

Upload up to three reference clips (30s each), tag them @Voice1, @Voice2, @Voice3 in your prompt, and Seed speaks each character in the cloned voice.

Attach a single image and Seed uses it to shape the voice and mood, with no separate voice picker needed.
Use MCP tool "generate_dialogue" with the Seed provider for automated prompt-driven audio and voice cloning, with the same workspace as Audio → Dialogue.
Turn a chapter into a performed scene with ambience and music, at a fraction of studio-recording cost, and revise by editing the prompt.

Clone a voice or describe one, then generate lines plus sound effects for dubs, trailers, and game scenes, all from one prompt.

Generate Seed audio from agents with the same generate_dialogue pipeline: select the Seed provider, write a prompt, and pass image or audio references.

Open Audio → Dialogue, pick Seed, and describe the audio you want: voices, effects, and music from one prompt; credits stay in one wallet with ElevenLabs, Gemini, music, and video.