Skip to main content
POST

Authorizations

authorization
string
header
required

Use your Stability API key to authentication requests to this App.

Headers

authorization
string
required

Your Stability API key, used to authenticate your requests. Although you may have multiple keys in your account, you should use the same key for all requests to this API.

Minimum string length: 1
content-type
string
required

The content type of the request body. Do not manually specify this header; your HTTP client library will automatically include the appropriate boundary parameter.

Minimum string length: 1
Example:

"multipart/form-data"

accept
enum<string>
default:audio/*

Specify audio/* to receive the bytes of the audio directly. Otherwise specify application/json to receive the audio as base64 encoded JSON.

Available options:
audio/*,
application/json
stability-client-id
string

The name of your application, used to help us communicate app-specific debugging or moderation issues to you.

Maximum string length: 256
Example:

"my-awesome-app"

stability-client-user-id
string

A unique identifier for your end user. Used to help us communicate user-specific debugging or moderation issues to you. Feel free to obfuscate this value to protect user privacy.

Maximum string length: 256
Example:

"DiscordUser#9999"

stability-client-version
string

The version of your application, used to help us communicate version-specific debugging or moderation issues to you.

Maximum string length: 256
Example:

"1.2.1"

Body

multipart/form-data
prompt
string
required

What you wish the output audio to be. A strong, descriptive prompt that clearly defines instruments, moods, styles, and genre will lead to better results.

You can make a prompt as simple or complex as you like. Simple prompts are good for clean output audio. Complex prompts are good for adding texture and depth to the output audio.

Maximum string length: 10000
audio
file
required

The audio to be used as the starting point for the generation.

Supported Formats:

  • mp3
  • wav

Validation Rule:

  • Audio must be between 6 and 380 seconds long
model
enum<string>
default:stable-audio-3

The model to use for generation.

  • stable-audio-3 requires 26 credits per generation
Available options:
stable-audio-3
duration
number
default:190

Controls the duration in seconds of the generated audio.

Required range: 1 <= x <= 380
seed
number
default:0

A specific value that is used to guide the 'randomness' of the generation. (Omit this parameter or pass 0 to use a random seed.)

Required range: 0 <= x <= 4294967294
steps
integer
default:8

Controls the number of sampling steps.

Required range: 4 <= x <= 8
cfg_scale
number
default:1

How strictly the diffusion process adheres to the prompt text (higher values make your audio closer to your prompt). Defaults to 1 if not specified.

Required range: 1 <= x <= 25
output_format
enum<string>
default:mp3

Dictates the content-type of the generated audio.

Available options:
mp3,
wav
strength
number
default:1

Sometimes referred to as denoising, this parameter controls how much influence the audio parameter has on the generated audio. A value of 0 would yield audio that is identical to the input. A value of 1 would be as if you passed in no audio at all.

Required range: 0 <= x <= 1

Response

Generation started. Use the returned id to poll for the result.

id
string
required

The id of a generation, typically used for async generations, that can be used to check the status of the generation or retrieve the result.

Required string length: 64
Example:

"a6dc6c6e20acda010fe14d71f180658f2896ed9b4ec25aa99a6ff06c796987c4"