Kubeez · Audio · Gemini 3.8 Flash TTS
Two voices, a delivery for every line, 70 voices to cast
Gemini 3.8 Flash TTS turns a script into a conversation. Write one or two speakers as an ordered list of lines, give each line its own delivery style, and drop in any of 40 delivery tags like <laugh>, <sigh>, <whispers> or <short pause> that the model performs instead of reading aloud. Cast from 70 voices grouped by use case, let the language come straight from your text, and render up to 5,000 characters per generation as WAV. Billed from the same Kubeez wallet as music and video. Section images below are illustrative marketing stills; every clip in the sample band is real output.
Showcase




































































































































































































































Every clip below is real output from Gemini 3.8 Flash TTS, not a mock-up. Play the two-voice scene first, then audition voices from different use cases.
One generation with two voices, Alnilam and Laomedeia, a delivery style on each line, inline delivery tags and filler words switched on.
Script as written (the model adds the filler words)
Write the lines, set the delivery on each one, and place tags exactly where the laugh, the breath or the pause belongs.

Every line carries its own delivery style, like 'tired night-shift radio host, slow and dry', so one speaker can change mood from line to line.

Tags in angle brackets such as <laugh>, <sigh>, <gasp> and <long pause> are performed, not read aloud, so the line lands with real expression.

Cast from 70 voices grouped by use case: tutor, commercial voiceover, podcast, storyteller, call center and more. The language is detected from your text.
Call MCP tool "generate_dialogue" with provider "gemini-3.8" for automated one and two speaker audio, with the same voices and limits as Audio → Dialogue.
Turn notes into a back-and-forth. Filler words add natural hesitations between the two speakers, and a rewritten line is one more render away.

Pick from the use case that fits, a commercial voiceover, a concierge or a tech support agent, and keep a whole script in one render of up to 5,000 characters.

Call generate_dialogue with provider 'gemini-3.8', a list of speakers and ordered dialogue turns: the same pipeline as the Dialogue studio.

Open Audio → Dialogue, choose the 3.8 Flash model, and cast your speakers. Credits stay in one wallet with music and video.