> ## Documentation Index
> Fetch the complete documentation index at: https://docs.stability.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Audio-to-Audio

> Stable Audio transforms existing audio samples into new high-quality compositions up to six minutes
long at 44.1kHz stereo using text instructions. Discover techniques for sample transformation in our
[Audio to Audio Guide](https://www.stableaudio.com/user-guide/audio-to-audio) to maximize creative control.
Read more about the model capabilities [here](https://stability.ai/news-updates).

### Try it out
Grab your [API key](https://platform.stability.ai/account/keys) and head over to
[![Open Google Colab](https://platform.stability.ai/svg/google-colab.svg)](https://colab.research.google.com/github/stability-ai/stability-sdk/blob/main/nbs/Stable_Audio_API.ipynb)
or try Stable Audio for free at [stableaudio.com](https://stableaudio.com).

This endpoint is asynchronous — it returns a generation `id` immediately (HTTP 202).
Poll `GET /v2beta/audio/results/{id}` to retrieve the result.

> **Note:**
> - We do not allow copyrighted content to be uploaded to our platform.
> - Maximum request size is 100Mb.

### Credits

**Stable Audio 3.0**

Flat rate of 26 credits per successful generation.

As always, you will not be charged for failed generations.



## OpenAPI

````yaml /openapi.json post /v2beta/audio/stable-audio/audio-to-audio
openapi: 3.0.1
info:
  version: v2beta
  title: StabilityAI REST API
  description: >-
    Welcome to the Stability Platform API. As of March 2024, we are building the
    REST v2beta API service to be the primary API service for the Stability
    Platform.

    All AI services on other APIs (gRPC, REST v1, RESTv2alpha) will continue to
    be maintained, however they will not receive

    new features or parameters.


    If you are a REST v2alpha user, we strongly recommend that you adjust the
    URL calls for the specific services that you are using over to the
    equivalent REST v2beta URL. Normally, this means simply replacing "v2alpha"
    with "v2beta". We are not deprecating v2alpha URLs at this time for users
    that are currently using them.


    #### Authentication


    You will need your [Stability API
    key](https://platform.stability.ai/account/keys) in order to make requests
    to this API.

    Make sure you never share your API key with anyone, and you never commit it
    to a public repository. Include this key in

    the `Authorization` header of your requests.


    #### Rate limiting


    This API is rate-limited to 150 requests every 10 seconds. If you exceed
    this limit, you will receive a `429` response

    and be timed out for 60 seconds. If you find this limit too restrictive,
    please reach out to us via [this
    form](https://kb.stability.ai/knowledge-base/kb-tickets/new).


    #### Support


    Please see our [FAQ](https://platform.stability.ai/faq) for answers to
    common questions. If you have any other questions or concerns,

    please reach out to us via [this
    form](https://kb.stability.ai/knowledge-base/kb-tickets/new).


    To see the health of our APIs, please check our [Status
    Page](https://stabilityai.instatus.com/).
servers:
  - url: https://api.stability.ai
security:
  - STABILITY_API_KEY: []
tags:
  - name: Edit
    description: >-
      Tools for editing your own and generated images.


      **[Erase](/docs/api-reference#tag/Edit/paths/~1v2beta~1stable-image~1edit~1erase/post)**


      The Erase service removes unwanted objects, such as blemishes on portraits
      or items on desks, using image masks.


      **[Outpaint](/docs/api-reference#tag/Edit/paths/~1v2beta~1stable-image~1edit~1outpaint/post)**


      The outpaint service inserts additional content in an image to fill in the
      space in any direction, allowing you to "zoom-out" of an image.


      **[Inpaint](/docs/api-reference#tag/Edit/paths/~1v2beta~1stable-image~1edit~1inpaint/post)**


      The Inpaint service modifies images by filling in or replacing specified
      areas with new content based on the content of a "mask" image.


      **[Search and
      Replace](/docs/api-reference#tag/Edit/paths/~1v2beta~1stable-image~1edit~1search-and-replace/post)**


      The Search and Replace service, similar to inpaint, allows to replace
      specified areas with new content, but this time with the help of a prompt
      instead of a mask. The service will automatically segment the object and
      replace it with the object requested in the prompt.


      **[Search and
      Recolor](/docs/api-reference#tag/Edit/paths/~1v2beta~1stable-image~1edit~1search-and-recolor/post)**


      The Search and Recolor service is another derivative of the inpaint
      service and provides the ability to change the color of a specific object
      in an image using a prompt. The Search and Recolor service will
      automatically segment the object and recolor it using the colors requested
      in the prompt.


      **[Remove
      Background](/docs/api-reference#tag/Edit/paths/~1v2beta~1stable-image~1edit~1remove-background/post)**


      The Remove Background service accurately segments the foreground from an
      image to removes the background.
  - name: Upscale
    description: >-
      Tools for increasing the size and resolution of your existing images.


      **[Fast
      Upscaler](/docs/api-reference#tag/Upscale/paths/~1v2beta~1stable-image~1upscale~1fast/post)**


      This service enhances image resolution by 4x using predictive and
      generative AI. This lightweight and fast service (processing in ~1 second)
      is ideal for enhancing the quality of compressed images, making it
      suitable for social media posts and other applications.


      **[Conservative
      Upscaler](/docs/api-reference#tag/Upscale/paths/~1v2beta~1stable-image~1upscale~1conservative/post)**


      This service can upscale images by 20 to 40 times up to a 4 megapixel
      output image with minimal alteration to the original image. The
      Conservative Upscaler can upscale images as small as 64x64 pixels directly
      to a 4 megapixel output. Use this option if you directly need a 4
      megapixel output.


      **[Creative
      Upscaler](/docs/api-reference#tag/Upscale/paths/~1v2beta~1stable-image~1upscale~1creative/post)**


      The service can upscale highly degraded images (lower than 1 megapixel)
      with a creative twist to provide high resolution results.
  - name: Generate
    description: >-
      Tools to generate new images from text, or create variations of existing
      images. Our different services include:


      **[Stable Image
      Ultra](/docs/api-reference#tag/Generate/paths/~1v2beta~1stable-image~1generate~1ultra/post)**:
      Photorealistic, Large-Scale Output


      Our state of the art text to image model based on Stable Diffusion 3.5.
      Stable Image Ultra Produces the highest quality, photorealistic outputs
      perfect for professional print media and large format applications. Stable
      Image Ultra excels at rendering exceptional detail and realism.


      **[Stable Image
      Core](/docs/api-reference#tag/Generate/paths/~1v2beta~1stable-image~1generate~1core/post)**:
      Fast and Affordable


      Optimized for fast and aﬀordable image generation, great for rapidly
      iterating on concepts during ideation. Stable Image Core is the next
      generation model following Stable Diffusion XL.


      **[Stable Diffusion 3.5 Model
      Suite](/docs/api-reference#tag/Generate/paths/~1v2beta~1stable-image~1generate~1sd3/post)**:
      Stability AI's latest base models


      The different versions of our open models are available via API, letting
      you test and adjust speed and quality based on your use case. All model
      versions strike a balance between generation speed and output quality and
      are ideal for creating high-volume, high-quality digital assets like
      websites, newsletters, and marketing materials.
  - name: Control
    description: >-
      Tools for generating precise, controlled variations of existing images or
      sketches.


      **[Sketch](/docs/api-reference#tag/Control/paths/~1v2beta~1stable-image~1control~1sketch/post)**


      This service upgrades sketches to refined outputs with precise control.
      For non-sketch images, it allows detailed manipulation of the final
      appearance by leveraging the contour lines and edges within the image.


      **[Structure](/docs/api-reference#tag/Control/paths/~1v2beta~1stable-image~1control~1structure/post)**


      This service excels in generating images by maintaining the structure of
      an input image, making it especially valuable for advanced content
      creation scenarios such as recreating scenes or rendering characters from
      models.


      **[Style](/docs/api-reference#tag/Control/paths/~1v2beta~1stable-image~1control~1style/post)**


      This service extracts stylistic elements from an input image (control
      image) and uses it to guide the creation of an output image based on the
      prompt. The result is a new image in the same style as the control image.
  - name: Results
    description: Tools for fetching the results of your async generations.
  - name: Stable Audio 2
    description: >-
      Tools to generate music and sound from text or audio, or transform
      existing audio clips into new compositions. Our different services
      include:


      **Stable Audio 2.5**: Fast, Best-Quality, Long-Form Music & Audio
      Generation


      Our most advanced audio generation model, capable of generating up to
      3-minute, 44.1 kHz stereo compositions. Stable Audio 2.5 supports
      text-to-audio, audio-to-audio, and audio-inpaint workflows - allowing
      creators to upload a sound and transform it into new instruments, styles,
      or genres using natural language prompts. Ideal for music production,
      cinematic sound design, and remixing.


      **Stable Audio 2.0**: High-Quality Audio Generation


      Built for text-to-audio and audio-to-audio generation, also capable of
      generating up to 3-minute, 44.1 kHz stereo. Stable Audio 2.0 is great for
      ideation, music demos, and ambient soundscapes. It's optimized for
      creative professionals seeking detailed and extended outputs from simple
      prompts.


      Stable Audio models were exclusively trained on licensed data from the
      [AudioSparx](https://www.audiosparx.com/) music library, honoring opt-out
      requests and ensuring fair compensation for creators. Additionally, Stable
      Audio 2.5 was pre-trained on licensed data from
      [Freesound](https://freesound.org/). Read more about the model
      capabilities [here](https://stability.ai/news/stable-audio-2-0).
    x-displayName: Stable Audio 2.5
  - name: Stable Audio
    description: >-
      Tools to generate music and sound from text or audio, or transform
      existing audio clips into new compositions. Our different services
      include:


      **Stable Audio 3.0**: Fast, Best-Quality, Long-Form Music & Audio
      Generation


      Our most advanced audio generation model, capable of generating up to
      6-minute, 44.1 kHz stereo compositions. Stable Audio 3.0 supports
      text-to-audio, audio-to-audio, and audio-inpaint workflows - allowing
      creators to upload a sound and transform it into new instruments, styles,
      or genres using natural language prompts. Ideal for music production,
      cinematic sound design, and remixing.


      Stable Audio models were exclusively trained on licensed data from the
      [AudioSparx](https://www.audiosparx.com/) music library, honoring opt-out
      requests and ensuring fair compensation for creators. Additionally, Stable
      Audio 3.0 was pre-trained on licensed data from
      [Freesound](https://freesound.org/). Read more about the model
      capabilities [here](https://stability.ai/news/stable-audio-3-0).
    x-displayName: Stable Audio 3.0
paths:
  /v2beta/audio/stable-audio/audio-to-audio:
    post:
      tags:
        - Stable Audio
      summary: Audio-to-Audio
      description: >-
        Stable Audio transforms existing audio samples into new high-quality
        compositions up to six minutes

        long at 44.1kHz stereo using text instructions. Discover techniques for
        sample transformation in our

        [Audio to Audio
        Guide](https://www.stableaudio.com/user-guide/audio-to-audio) to
        maximize creative control.

        Read more about the model capabilities
        [here](https://stability.ai/news-updates).


        ### Try it out

        Grab your [API key](https://platform.stability.ai/account/keys) and head
        over to

        [![Open Google
        Colab](https://platform.stability.ai/svg/google-colab.svg)](https://colab.research.google.com/github/stability-ai/stability-sdk/blob/main/nbs/Stable_Audio_API.ipynb)

        or try Stable Audio for free at
        [stableaudio.com](https://stableaudio.com).


        This endpoint is asynchronous — it returns a generation `id` immediately
        (HTTP 202).

        Poll `GET /v2beta/audio/results/{id}` to retrieve the result.


        > **Note:**

        > - We do not allow copyrighted content to be uploaded to our platform.

        > - Maximum request size is 100Mb.


        ### Credits


        **Stable Audio 3.0**


        Flat rate of 26 credits per successful generation.


        As always, you will not be charged for failed generations.
      parameters:
        - $ref: '#/components/parameters/Authorization'
        - $ref: '#/components/parameters/FormDataContentType'
        - $ref: '#/components/parameters/AcceptAudio'
        - $ref: '#/components/parameters/StabilityClientID'
        - $ref: '#/components/parameters/StabilityClientUserID'
        - $ref: '#/components/parameters/StabilityClientVersion'
      requestBody:
        content:
          multipart/form-data:
            schema:
              type: object
              properties:
                prompt:
                  $ref: '#/components/schemas/AudioPrompt'
                model:
                  type: string
                  enum:
                    - stable-audio-3
                  default: stable-audio-3
                  description: |-
                    The model to use for generation.

                    - `stable-audio-3` requires 26 credits per generation
                duration:
                  type: number
                  minimum: 1
                  maximum: 380
                  default: 190
                  description: Controls the duration in seconds of the generated audio.
                seed:
                  $ref: '#/components/schemas/Seed'
                steps:
                  type: integer
                  minimum: 4
                  maximum: 8
                  default: 8
                  description: Controls the number of sampling steps.
                cfg_scale:
                  type: number
                  minimum: 1
                  maximum: 25
                  default: 1
                  description: >-
                    How strictly the diffusion process adheres to the prompt
                    text (higher values make your audio closer to your prompt).
                    Defaults to 1 if not specified.
                output_format:
                  type: string
                  enum:
                    - mp3
                    - wav
                  default: mp3
                  description: Dictates the `content-type` of the generated audio.
                strength:
                  type: number
                  minimum: 0
                  maximum: 1
                  default: 1
                  description: >-
                    Sometimes referred to as _denoising_, this parameter
                    controls how much influence the

                    `audio` parameter has on the generated audio.

                    A value of 0 would yield audio that is identical to the
                    input.

                    A value of 1 would be as if you passed in no audio at all.
                audio:
                  type: string
                  description: >-
                    The audio to be used as the starting point for the
                    generation.


                    Supported Formats:

                    - mp3

                    - wav


                    Validation Rule:

                    - Audio must be between 6 and 380 seconds long
                  format: binary
                  example: ./some/audio.mp3
              required:
                - prompt
                - audio
      responses:
        '202':
          description: Generation started. Use the returned id to poll for the result.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/AsyncGenerationResponse'
        '400':
          $ref: '#/components/responses/BadRequest'
        '403':
          $ref: '#/components/responses/ContentModerated'
        '422':
          $ref: '#/components/responses/AudioClassificationFailed'
        '429':
          $ref: '#/components/responses/RateLimited'
        '500':
          $ref: '#/components/responses/InternalError'
      x-codeSamples:
        - lang: python
          label: Python
          source: |-
            import requests, time

            response = requests.post(
                f"https://api.stability.ai/v2beta/audio/stable-audio/audio-to-audio",
                headers={"authorization": f"Bearer sk-MYAPIKEY", "accept": "audio/*"},
                files={"audio": open("./input.mp3", "rb")},
                data={
                    "prompt": "Add a lush reverb and layer in warm ambient pads.",
                    "output_format": "mp3",
                    "duration": 30,
                    "strength": 0.5,
                },
            )

            if response.status_code != 202:
                raise Exception(str(response.json()))

            generation_id = response.json()["id"]

            while True:
                result = requests.get(
                    f"https://api.stability.ai/v2beta/audio/results/{generation_id}",
                    headers={"authorization": f"Bearer sk-MYAPIKEY", "accept": "audio/*"},
                )
                if result.status_code == 202:
                    print("Generation in-progress, retrying in 10 seconds...")
                    time.sleep(10)
                elif result.status_code == 200:
                    with open("./output.mp3", "wb") as f:
                        f.write(result.content)
                    break
                else:
                    raise Exception(str(result.json()))
        - lang: javascript
          label: JavaScript
          source: |-
            import fs from "node:fs";
            import axios from "axios";
            import FormData from "form-data";

            const payload = {
              prompt: "Add a lush reverb and layer in warm ambient pads.",
              output_format: "mp3",
              duration: 30,
              strength: 0.5,
              audio: fs.createReadStream("./input.mp3"),
            };

            const response = await axios.postForm(
              `https://api.stability.ai/v2beta/audio/stable-audio/audio-to-audio`,
              axios.toFormData(payload, new FormData()),
              {
                validateStatus: undefined,
                headers: { Authorization: `Bearer sk-MYAPIKEY`, Accept: "audio/*" },
              }
            );

            if (response.status !== 202) {
              throw new Error(`${response.status}: ${JSON.stringify(response.data)}`);
            }

            const { id: generationID } = response.data;

            while (true) {
              const result = await axios.request({
                url: `https://api.stability.ai/v2beta/audio/results/${generationID}`,
                method: "GET",
                validateStatus: undefined,
                responseType: "arraybuffer",
                headers: { Authorization: `Bearer sk-MYAPIKEY`, Accept: "audio/*" },
              });

              if (result.status === 202) {
                console.log("Generation in-progress, retrying in 10 seconds...");
                await new Promise((resolve) => setTimeout(resolve, 10_000));
              } else if (result.status === 200) {
                fs.writeFileSync("./output.mp3", Buffer.from(result.data));
                break;
              } else {
                throw new Error(`${result.status}: ${result.data.toString()}`);
              }
            }
        - lang: terminal
          label: cURL
          source: >-
            generation_id=$(curl -sS -f
            "https://api.stability.ai/v2beta/audio/stable-audio/audio-to-audio"
            \
              -H "authorization: Bearer sk-MYAPIKEY" \
              -H "accept: audio/*" \
              -F prompt="Add a lush reverb and layer in warm ambient pads." \
              -F output_format="mp3" \
              -F duration="30" \
              -F strength="0.5" \
              -F audio=@"./input.mp3" \
              | jq -r '.id')

            while true; do
              http_status=$(curl -sS -f \
                -o "./output.mp3" \
                -w '%{http_code}' \
                -H "authorization: Bearer sk-MYAPIKEY" \
                -H "accept: audio/*" \
                "https://api.stability.ai/v2beta/audio/results/${generation_id}")

              case $http_status in
                202) echo "Still processing. Retrying in 10 seconds..."; sleep 10 ;;
                200) echo "Download complete!"; break ;;
                *) echo "Error: $http_status"; exit 1 ;;
              esac
            done
components:
  parameters:
    Authorization:
      schema:
        type: string
        description: >-
          Your [Stability API key](https://platform.stability.ai/account/keys),
          used to authenticate your requests. Although you may have multiple
          keys in your account, you should use the same key for all requests to
          this API.
        minLength: 1
      required: true
      name: authorization
      in: header
    FormDataContentType:
      schema:
        type: string
        minLength: 1
        description: >-
          The content type of the request body. Do not manually specify this
          header; your HTTP client library will automatically include the
          appropriate boundary parameter.
        example: multipart/form-data
      required: true
      name: content-type
      in: header
    AcceptAudio:
      schema:
        type: string
        default: audio/*
        description: >-
          Specify `audio/*` to receive the bytes of the audio directly.
          Otherwise specify `application/json` to receive the audio as base64
          encoded JSON.
        enum:
          - audio/*
          - application/json
      required: false
      name: accept
      in: header
    StabilityClientID:
      schema:
        type: string
        maxLength: 256
        description: >-
          The name of your application, used to help us communicate app-specific
          debugging or moderation issues to you.
        example: my-awesome-app
      required: false
      name: stability-client-id
      in: header
    StabilityClientUserID:
      schema:
        type: string
        maxLength: 256
        description: >-
          A unique identifier for your end user. Used to help us communicate
          user-specific debugging or moderation issues to you. Feel free to
          obfuscate this value to protect user privacy.
        example: DiscordUser#9999
      required: false
      name: stability-client-user-id
      in: header
    StabilityClientVersion:
      schema:
        type: string
        maxLength: 256
        description: >-
          The version of your application, used to help us communicate
          version-specific debugging or moderation issues to you.
        example: 1.2.1
      required: false
      name: stability-client-version
      in: header
  schemas:
    AudioPrompt:
      type: string
      maxLength: 10000
      description: >-
        What you wish the output audio to be. A strong, descriptive prompt that
        clearly defines

        instruments, moods, styles, and genre will lead to better results.


        You can make a prompt as simple or complex as you like. Simple prompts
        are good for clean

        output audio. Complex prompts are good for adding texture and depth to
        the output audio.
    Seed:
      type: number
      minimum: 0
      maximum: 4294967294
      default: 0
      description: >-
        A specific value that is used to guide the 'randomness' of the
        generation. (Omit this parameter or pass `0` to use a random seed.)
    AsyncGenerationResponse:
      type: object
      properties:
        id:
          $ref: '#/components/schemas/GenerationID'
      required:
        - id
    GenerationID:
      type: string
      minLength: 64
      maxLength: 64
      description: >-
        The `id` of a generation, typically used for async generations, that can
        be used to check the status of the generation or retrieve the result.
      example: a6dc6c6e20acda010fe14d71f180658f2896ed9b4ec25aa99a6ff06c796987c4
    Error:
      type: object
      properties:
        id:
          type: string
          minLength: 1
          description: >-
            A unique identifier associated with this error. Please include this
            in any [support
            tickets](https://kb.stability.ai/knowledge-base/kb-tickets/new)

            you file, as it will greatly assist us in diagnosing the root cause
            of the problem.
          example: a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4
        name:
          type: string
          minLength: 1
          description: >-
            Short-hand name for an error, useful for discriminating between
            errors with the same status code.
          example: bad_request
        errors:
          type: array
          items:
            type: string
          minItems: 1
          description: One or more error messages indicating what went wrong.
          example:
            - 'some-field: is required'
      required:
        - id
        - name
        - errors
    ContentModerationResponse:
      type: object
      properties:
        id:
          type: string
          minLength: 1
          description: >-
            A unique identifier associated with this error. Please include this
            in any [support
            tickets](https://kb.stability.ai/knowledge-base/kb-tickets/new)

            you file, as it will greatly assist us in diagnosing the root cause
            of the problem.
          example: a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4
        name:
          type: string
          minLength: 1
          description: >-
            Our content moderation system has flagged some part of your request
            and subsequently denied it.  You were not charged for this request. 
            While this may at times be frustrating, it is necessary to maintain
            the integrity of our platform and ensure a safe experience for all
            users.


            If you would like to provide feedback, please use the [Support
            Form](https://kb.stability.ai/knowledge-base/kb-tickets/new).
          enum:
            - content_moderation
        errors:
          type: array
          items:
            type: string
          minItems: 1
          description: One or more error messages indicating what went wrong.
          example:
            - 'some-field: is required'
      required:
        - id
        - name
        - errors
      description: Your request was flagged by our content moderation system.
      example:
        id: ed14db44362126aab3cbd25cca51ffe3
        name: content_moderation
        errors:
          - >-
            Your request was flagged by our content moderation system, as a
            result your request was denied and you were not charged.
  responses:
    BadRequest:
      description: Invalid parameter(s), see the `errors` field for details.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
    ContentModerated:
      description: Your request was flagged by our content moderation system.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/ContentModerationResponse'
    AudioClassificationFailed:
      description: >-
        Your request was well-formed, but rejected. See the `errors` field for
        details.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          examples:
            Invalid Language:
              value:
                id: ff54b236a3acdde1522cb1ba641c43ed
                name: invalid_language
                errors:
                  - English is the only supported language for this service.
            Copyrighted Content Detected:
              value:
                id: ff54b236a3acdde1522cb1ba641c43ed
                name: copyrighted_content
                errors:
                  - >-
                    Our system has detected the presence of copyrighted content
                    in your audio. To comply with our guidelines, we cannot
                    process this request. Please upload a different audio file.
    RateLimited:
      description: You have made more than 150 requests in 10 seconds.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            id: rate_limit_exceeded
            name: rate_limit_exceeded
            errors:
              - >-
                You have exceeded the rate limit of 150 requests within a 10
                second period, and have been timed out for 60 seconds.
    InternalError:
      description: >-
        An internal error occurred. If the problem persists [contact
        support](https://kb.stability.ai/knowledge-base/kb-tickets/new).
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            id: 2a1b2d4eafe2bc6ab4cd4d5c6133f513
            name: internal_error
            errors:
              - An unexpected server error has occurred, please try again later.
  securitySchemes:
    STABILITY_API_KEY:
      type: apiKey
      name: authorization
      in: header
      description: >-
        Use your [Stability API key](https://platform.stability.ai/account/keys)
        to authentication requests to this App.

````