Audio-to-Audio
Stable Audio transforms existing audio samples into new high-quality compositions up to three minutes long at 44.1kHz stereo using text instructions. Discover techniques for sample transformation in our Audio to Audio Guide to maximize creative control. Read more about the model capabilities here.
Try it out
Grab your API key and head over to
or try Stable Audio 2.0 for free at stableaudio.com.
How to use
Please invoke this endpoint with a POST request.
The headers of the request must include an API key in the authorization field. The body of the request must be
multipart/form-data. The accept header should be set to one of the following:
audio/*to receive the audio in the format specified by theoutput_formatparameter.application/jsonto receive the audio encoded as base64 in a JSON response.
The body of the request should include:
prompt- text to generate the audio from. Check our prompt guide for tipsaudio- the audio to use as the starting point for the generation
Notes:
- We do not allow copyrighted content to be uploaded to our platform.
- Maximum request size is 50Mb.
Optional Parameters:
The body may optionally include:
output_format- the format of the output audioseed- the randomness seed to use for the generationsteps- the number of sampling stepsduration- the number of seconds of the generated audiocfg_scale- controls how strictly the diffusion process adheres to the prompt text (only forstable-audio-2)model- the model to use [stable-audio-2,stable-audio-2.5]strength- controls how much influence theaudioparameter has on the output audio
Note: for more details about these parameters please see the request schema below.
Credits
Stable Audio 2.0
By default, 20 credits per successful generation. The number of credits is determined
by the following formula: credits = 17 + 0.06 * steps.
Examples:
- 50 steps = 20 credits [default]
- 100 steps = 23 credits
Stable Audio 2.5
Requests made using the Stable Audio 2.5 model have a flat rate of 20 credits per successful result.
As always, you will not be charged for failed generations.
Authorizations
Use your Stability API key to authentication requests to this App.
Headers
Your Stability API key, used to authenticate your requests. Although you may have multiple keys in your account, you should use the same key for all requests to this API.
1The content type of the request body. Do not manually specify this header; your HTTP client library will automatically include the appropriate boundary parameter.
1"multipart/form-data"
Specify audio/* to receive the bytes of the audio directly. Otherwise specify application/json to receive the audio as base64 encoded JSON.
audio/*, application/json The name of your application, used to help us communicate app-specific debugging or moderation issues to you.
256"my-awesome-app"
A unique identifier for your end user. Used to help us communicate user-specific debugging or moderation issues to you. Feel free to obfuscate this value to protect user privacy.
256"DiscordUser#9999"
The version of your application, used to help us communicate version-specific debugging or moderation issues to you.
256"1.2.1"
Body
What you wish the output audio to be. A strong, descriptive prompt that clearly defines instruments, moods, styles, and genre will lead to better results.
You can make a prompt as simple or complex as you like. Simple prompts are good for clean output audio. Complex prompts are good for adding texture and depth to the output audio.
Check our prompt guide for tips.
10000The audio to be use as the starting point for the generation.
Supported Formats:
- mp3
- wav
Validation Rule:
- Audio must be between 6 and 190 seconds long
Controls the duration in seconds of the generated audio.
1 <= x <= 190A specific value that is used to guide the 'randomness' of the generation. (Omit this parameter or pass 0 to use a random seed.)
0 <= x <= 4294967294Controls the number of sampling steps.
- For
stable-audio-2: accepts steps between30and100(defaults to50). - For
stable-audio-2.5: accepts steps between4and8(defaults to8).
How strictly the diffusion process adheres to the prompt text (higher values make your audio closer to your prompt).
Defaults to 7 for stable-audio-2 and 1 for stable-audio-2.5 if not specified.
1 <= x <= 25The model to use for generation.
stable-audio-2.5requires 20 credits per generationstable-audio-2requires 20 credits per generation
stable-audio-2.5, stable-audio-2 Dictates the content-type of the generated audio.
mp3, wav Sometimes referred to as denoising, this parameter controls how much influence the
audio parameter has on the generated audio.
A value of 0 would yield audio that is identical to the input.
A value of 1 would be as if you passed in no audio at all.
Minimum value for stable-audio-2.5 is 0.01.
0 <= x <= 1Response
Generation was successful.
The bytes of the generated audio.
The finish-reason and seed will be present as headers.