Skip to main content
POST

Authorizations

authorization
string
header
required

Use your Stability API key to authentication requests to this App.

Headers

authorization
string
required

Your Stability API key, used to authenticate your requests. Although you may have multiple keys in your account, you should use the same key for all requests to this API.

Minimum string length: 1
content-type
string
required

The content type of the request body. Do not manually specify this header; your HTTP client library will automatically include the appropriate boundary parameter.

Minimum string length: 1
Example:

"multipart/form-data"

accept
enum<string>
default:audio/*

Specify audio/* to receive the bytes of the audio directly. Otherwise specify application/json to receive the audio as base64 encoded JSON.

Available options:
audio/*,
application/json
stability-client-id
string

The name of your application, used to help us communicate app-specific debugging or moderation issues to you.

Maximum string length: 256
Example:

"my-awesome-app"

stability-client-user-id
string

A unique identifier for your end user. Used to help us communicate user-specific debugging or moderation issues to you. Feel free to obfuscate this value to protect user privacy.

Maximum string length: 256
Example:

"DiscordUser#9999"

stability-client-version
string

The version of your application, used to help us communicate version-specific debugging or moderation issues to you.

Maximum string length: 256
Example:

"1.2.1"

Body

multipart/form-data
prompt
string
required

What you wish the output audio to be. A strong, descriptive prompt that clearly defines instruments, moods, styles, and genre will lead to better results.

You can make a prompt as simple or complex as you like. Simple prompts are good for clean output audio. Complex prompts are good for adding texture and depth to the output audio.

Check our prompt guide for tips.

Maximum string length: 10000
duration
number
default:190

Controls the duration in seconds of the generated audio.

Required range: 1 <= x <= 190
seed
number
default:0

A specific value that is used to guide the 'randomness' of the generation. (Omit this parameter or pass 0 to use a random seed.)

Required range: 0 <= x <= 4294967294
steps
integer

Controls the number of sampling steps.

  • For stable-audio-2: accepts steps between 30 and 100 (defaults to 50).
  • For stable-audio-2.5: accepts steps between 4 and 8 (defaults to 8).
cfg_scale
number

How strictly the diffusion process adheres to the prompt text (higher values make your audio closer to your prompt).

Defaults to 7 for stable-audio-2 and 1 for stable-audio-2.5 if not specified.

Required range: 1 <= x <= 25
model
enum<string>
default:stable-audio-2

The model to use for generation.

  • stable-audio-2.5 requires 20 credits per generation
  • stable-audio-2 requires 20 credits per generation
Available options:
stable-audio-2.5,
stable-audio-2
output_format
enum<string>
default:mp3

Dictates the content-type of the generated audio.

Available options:
mp3,
wav

Response

Generation was successful.

The bytes of the generated audio.

The finish-reason and seed will be present as headers.