Skip to main content
POST

Authorizations

authorization
string
header
required

Use your Stability API key to authentication requests to this App.

Headers

authorization
string
required

Your Stability API key, used to authenticate your requests. Although you may have multiple keys in your account, you should use the same key for all requests to this API.

Minimum string length: 1
content-type
string
required

The content type of the request body. Do not manually specify this header; your HTTP client library will automatically include the appropriate boundary parameter.

Minimum string length: 1
Example:

"multipart/form-data"

accept
enum<string>
default:image/*

Specify image/* to receive the bytes of the image directly. Otherwise specify application/json to receive the image as base64 encoded JSON.

Available options:
image/*,
application/json
stability-client-id
string

The name of your application, used to help us communicate app-specific debugging or moderation issues to you.

Maximum string length: 256
Example:

"my-awesome-app"

stability-client-user-id
string

A unique identifier for your end user. Used to help us communicate user-specific debugging or moderation issues to you. Feel free to obfuscate this value to protect user privacy.

Maximum string length: 256
Example:

"DiscordUser#9999"

stability-client-version
string

The version of your application, used to help us communicate version-specific debugging or moderation issues to you.

Maximum string length: 256
Example:

"1.2.1"

Body

multipart/form-data
prompt
string
required

What you wish to see in the output image. A strong, descriptive prompt that clearly defines elements, colors, and subjects will lead to better results.

Required string length: 1 - 10000
mode
enum<string>
default:text-to-image

Controls whether this is a text-to-image or image-to-image generation, which affects which parameters are required:

  • text-to-image requires only the prompt parameter
  • image-to-image requires the prompt, image, and strength parameters
Available options:
text-to-image,
image-to-image
image
file

The image to use as the starting point for the generation.

Supported formats:

  • jpeg
  • png
  • webp

Supported dimensions:

  • Every side must be at least 64 pixels

Important: This parameter is only valid for image-to-image requests.

strength
number

Sometimes referred to as denoising, this parameter controls how much influence the image parameter has on the generated image. A value of 0 would yield an image that is identical to the input. A value of 1 would be as if you passed in no image at all.

Important: This parameter is only valid for image-to-image requests. For SD 3.5 Flash, the best results for image-to-image generation are achieved with a strength between .94 - .97.

Required range: 0 <= x <= 1
aspect_ratio
enum<string>
default:1:1

Controls the aspect ratio of the generated image. Defaults to 1:1.

Important: This parameter is only valid for text-to-image requests.

Available options:
21:9,
16:9,
3:2,
5:4,
1:1,
4:5,
2:3,
9:16,
9:21
model
enum<string>
default:sd3.5-large

The model to use for generation.

  • sd3.5-large requires 6.5 credits per generation
  • sd3.5-large-turbo requires 4 credits per generation
  • sd3.5-medium requires 3.5 credits per generation
  • sd3.5-flash requires 2.5 credits per generation
  • As of the April 17, 2025, sd3-large, sd3-large-turbo and sd3-medium are re-routed to their sd3.5-[model version] equivalent, at the same price.
Available options:
sd3.5-large,
sd3.5-large-turbo,
sd3.5-medium
seed
number
default:0

A specific value that is used to guide the 'randomness' of the generation. (Omit this parameter or pass 0 to use a random seed.)

Required range: 0 <= x <= 4294967294
output_format
enum<string>
default:png

Dictates the content-type of the generated image.

Available options:
png,
jpeg,
webp
style_preset
enum<string>

Guides the image model towards a particular style.

Available options:
enhance,
anime,
photographic,
digital-art,
comic-book,
fantasy-art,
line-art,
analog-film,
neon-punk,
isometric,
low-poly,
origami,
modeling-compound,
cinematic,
3d-model,
pixel-art,
tile-texture
negative_prompt
string

Keywords of what you do not wish to see in the output image. This is an advanced feature.

Maximum string length: 10000
cfg_scale
number

How strictly the diffusion process adheres to the prompt text (higher values keep your image closer to your prompt). The Large and Medium models use a default of 4. The Turbo and Flash model uses a default of 1.

Required range: 1 <= x <= 10

Response

Generation was successful.

The bytes of the generated image.

The finish-reason and seed will be present as headers.