Text-to-Audio
Stable Audio transforms existing audio samples into new high-quality compositions up to six minutes long at 44.1kHz stereo using text instructions.
Try it out
Grab your API key and head over to
or try Stable Audio for free at stableaudio.com.
This endpoint is asynchronous — it returns a generation id immediately (HTTP 202).
Poll GET /v2beta/audio/results/{id} to retrieve the result.
Credits
Stable Audio 3.0
Flat rate of 26 credits per successful generation.
As always, you will not be charged for failed generations.
Authorizations
Use your Stability API key to authentication requests to this App.
Headers
Your Stability API key, used to authenticate your requests. Although you may have multiple keys in your account, you should use the same key for all requests to this API.
1The content type of the request body. Do not manually specify this header; your HTTP client library will automatically include the appropriate boundary parameter.
1"multipart/form-data"
Specify audio/* to receive the bytes of the audio directly. Otherwise specify application/json to receive the audio as base64 encoded JSON.
audio/*, application/json The name of your application, used to help us communicate app-specific debugging or moderation issues to you.
256"my-awesome-app"
A unique identifier for your end user. Used to help us communicate user-specific debugging or moderation issues to you. Feel free to obfuscate this value to protect user privacy.
256"DiscordUser#9999"
The version of your application, used to help us communicate version-specific debugging or moderation issues to you.
256"1.2.1"
Body
What you wish the output audio to be. A strong, descriptive prompt that clearly defines instruments, moods, styles, and genre will lead to better results.
You can make a prompt as simple or complex as you like. Simple prompts are good for clean output audio. Complex prompts are good for adding texture and depth to the output audio.
10000The model to use for generation.
stable-audio-3requires 26 credits per generation
stable-audio-3 Controls the duration in seconds of the generated audio.
1 <= x <= 380A specific value that is used to guide the 'randomness' of the generation. (Omit this parameter or pass 0 to use a random seed.)
0 <= x <= 4294967294Controls the number of sampling steps.
4 <= x <= 8How strictly the diffusion process adheres to the prompt text (higher values make your audio closer to your prompt). Defaults to 1 if not specified.
1 <= x <= 25Dictates the content-type of the generated audio.
mp3, wav Response
Generation started. Use the returned id to poll for the result.
The id of a generation, typically used for async generations, that can be used to check the status of the generation or retrieve the result.
64"a6dc6c6e20acda010fe14d71f180658f2896ed9b4ec25aa99a6ff06c796987c4"