Generate audio
Speech, private voice cloning, and music.
Read as Markdown ↗POST /api/generate/audio
Requires generate scope and Idempotency-Key. Uses the shared media envelope and additionally requires operation: speech, voice_clone, or music. Returns a media task.
Speech and voice cloning
Use model letsgen-voice and parameters.voiceProfileId. Speech text is supplied in prompt (up to 2,000 characters). Select a usable voice through Voices. voice_clone uses your private reference voice; it does not train or publish a voice.
{
"model": "letsgen-voice",
"operation": "speech",
"prompt": "Welcome. Let's make something together.",
"parameters": {"voiceProfileId": "VOICE_ID"},
"maxGems": 20
}Optional speech controls are speed (0.5–2), volume (-6–6 dB), and temperature (0–1, expressiveness). They default to 1, 0, and 0.7 respectively. The same voice profile input is required for voice_clone.
Music
Use model suno-v6 with operation: "music". Simple mode describes the desired song in prompt (up to 5,000 characters).
{
"model": "suno-v6",
"operation": "music",
"prompt": "Warm acoustic instrumental for an early morning walk",
"parameters": {"musicMode": "simple", "instrumental": "on"},
"maxGems": 100
}Advanced mode uses musicMode: "advanced", style (up to 1,000 characters), title (up to 80 characters), optional negativeTags (up to 200), vocalGender (auto, m, f), and duration (off, 30, 60, 120, 180, 240). instrumental accepts on or off.
For an audio cover, upload an owned recording with Upload assets and supply one ID in referenceAssetIds or parameters.coverAudio. Speech/cloning takes a voice profile, not top-level media references. A monthly key cap and request maxGems still apply; model availability comes from discovery.