# Generate audio Speech, private voice cloning, and music. ## POST /api/generate/audio Requires **generate** scope and `Idempotency-Key`. Uses the shared media envelope and additionally requires `operation`: `speech`, `voice_clone`, or `music`. Returns a [media task](/docs/api/tasks). ## Speech and voice cloning Use model `letsgen-voice` and `parameters.voiceProfileId`. Speech text is supplied in `prompt` (up to 2,000 characters). Select a usable voice through [Voices](/docs/api/voices). `voice_clone` uses your private reference voice; it does not train or publish a voice. ```json { "model": "letsgen-voice", "operation": "speech", "prompt": "Welcome. Let's make something together.", "parameters": {"voiceProfileId": "VOICE_ID"}, "maxGems": 20 } ``` Optional speech controls are `speed` (0.5–2), `volume` (-6–6 dB), and `temperature` (0–1, expressiveness). They default to 1, 0, and 0.7 respectively. The same voice profile input is required for `voice_clone`. ## Music Use model `suno-v6` with `operation: "music"`. Simple mode describes the desired song in `prompt` (up to 5,000 characters). ```json { "model": "suno-v6", "operation": "music", "prompt": "Warm acoustic instrumental for an early morning walk", "parameters": {"musicMode": "simple", "instrumental": "on"}, "maxGems": 100 } ``` Advanced mode uses `musicMode: "advanced"`, `style` (up to 1,000 characters), `title` (up to 80 characters), optional `negativeTags` (up to 200), `vocalGender` (`auto`, `m`, `f`), and `duration` (`off`, `30`, `60`, `120`, `180`, `240`). `instrumental` accepts `on` or `off`. For an audio cover, upload an owned recording with [Upload assets](/docs/api/assets) and supply one ID in `referenceAssetIds` or `parameters.coverAudio`. Speech/cloning takes a voice profile, not top-level media references. A monthly key cap and request `maxGems` still apply; model availability comes from discovery.