Fish TTS API Integration Guide

This interface provides text-to-speech, saved voice invocation, and one-time instant voice cloning. The endpoint is POST https://api.qiyaov.com/fish/tts.

Application Process

To use the Fish TTS API, first obtain your API Token from the qiyaov Console and keep it for later use.

If you are not yet logged in or registered, you will be automatically redirected to the login page and invited to register and log in. After completion, you will automatically return to the current page.

One API Token can invoke all platform services; there is no need to apply separately for each service. Your first application includes free credits for a free trial; when credits are insufficient, you can top up your general balance in the Console.

📘 Complete documentation: Fish TTS API →

Request Headers

Header Required Description
authorization Yes Bearer {token}, where {token} is the key applied for on this platform.
content-type Yes application/json.
accept No application/json.
model No TTS model, optionally s1, s2-pro, or s2.1-pro, defaulting to s2-pro. s2.1-pro is the latest generation, while s2-pro is highly expressive; s1 is more stable and less likely to drift with long text. All three have the same price.

Request Body Fields

Field Type Required Description
text string Yes The text to synthesize, a non-empty string.
format string No Output audio format, optionally mp3 (default), wav, or pcm. Both wav and pcm return WAV containers. opus is not supported and passing it will directly return 400.
reference_id string | string[] No The model ID of a saved or public voice, which can be created via the Fish Model API, or retrieved in Fish Model Query. Cannot be used together with references.
references object[] No One-time instant voice cloning, supporting only one {audio, text} sample: audio is a public HTTPS MP3/WAV URL, and text is the accurate verbatim transcript of the audio. Cannot be used together with reference_id.
sample_rate integer No Sample rate, commonly 16000, 22050, or 44100. format=mp3 defaults to 44100.
mp3_bitrate integer No MP3 bitrate, optionally 64, 128, or 192. Takes effect only when format=mp3.
prosody object No Prosody override, supporting speed (speech rate, with 1.0 being the original rate) and volume (volume gain in dB). For example, {"speed":1.2,"volume":0}.
chunk_length integer No Upstream chunk length, determined by the upstream by default.
temperature number No Sampling temperature, approximately ranging from 0.0 to 1.0.
top_p number No Top-p sampling parameter.
latency string No normal or balanced; if omitted, this interface automatically fills in normal (passing an empty string directly will be rejected by the upstream).
normalize boolean No Whether to normalize the text.
callback_url string No Asynchronous callback address; see “Asynchronous Callback” below for details. This is an extension relative to the official interface.

One-time cloning only accepts HTTPS audio URLs, and does not accept MessagePack, Base64, data URIs, or URLs with credentials. Reference audio is recommended to be 10–270 seconds; Studio uses the more conservative range of 10–60 seconds.

Example 1: Minimal Request (text + format=mp3)

curl -X POST 'https://api.qiyaov.com/fish/tts' \
  -H 'authorization: Bearer {token}' \
  -H 'content-type: application/json' \
  -d '{
    "text": "Hello world.",
    "format": "mp3"
  }'

Response (tested):

{
  "audio_url": "https://cdn.acedata.cloud/assets/examples/fish/e2ffcc06-18da-4a8c-b9aa-9337d0f9ec1d-230825dfc559.mp3"
}

audio_url points to this platform's CDN, and can be downloaded directly with GET or played in <audio>. The final successful response also returns a top-level cost, where amount is the Credits actually deducted for this request; if there is an account discount, list_amount indicates the credits before the discount. The link remains available long-term, but it is still recommended to keep a copy in your own storage.

Example 2: Using a Cloned Voice reference_id

Below uses a public Spanish voice on the Fish platform (_id can be retrieved through Fish Model Query):

curl -X POST 'https://api.qiyaov.com/fish/tts' \
  -H 'authorization: Bearer {token}' \
  -H 'content-type: application/json' \
  -d '{
    "text": "Hermanos míos, hoy es un buen día.",
    "reference_id": "8d2c17a9b26d4d83888ea67a1ee565b2",
    "format": "mp3"
  }'

Response (tested):

{
  "audio_url": "https://cdn.acedata.cloud/assets/examples/fish/b6f161f2-a100-4818-add2-47694f234659-6532864739de.mp3"
}

Example 3: One-Time Instant Voice Cloning (references)

Place the reference audio at a public HTTPS address and provide the accurate original text actually spoken in it. This voice is used only for this synthesis and will not create a long-term model:

curl -X POST 'https://api.qiyaov.com/fish/tts' \
  -H 'authorization: Bearer {token}' \
  -H 'content-type: application/json' \
  -H 'model: s2-pro' \
  -d '{
    "text": "新的旅程从这一刻开始,让我们一起向前。",
    "format": "mp3",
    "references": [{
      "audio": "https://cdn.acedata.cloud/assets/examples/fish/6220d605-39d0-4d43-9e58-0f12949dc9b9-570cfdf96ad9.mp3",
      "text": "春天的清晨,阳光穿过树叶,落在安静的小路上。"
    }]
  }'

Successfully tested response:

{
  "audio_url": "https://cdn.acedata.cloud/assets/examples/fish/995dfe37-b187-474d-8323-b08d6678ed8f-6359be9f8873.mp3",
  "cost": {"amount": 0.007609319999999999, "currency": "credit"}
}
Method Lifecycle Applicable Scenarios
references Current TTS request only Temporary use or different voice each time
reference_id Reusable Public voices or long-term voices created through /fish/model

Example 3: Adjusting Speech Speed / Volume (prosody)

curl -X POST 'https://api.qiyaov.com/fish/tts' \
  -H 'authorization: Bearer {token}' \
  -H 'content-type: application/json' \
  -d '{
    "text": "Faster speech with prosody overrides.",
    "prosody": { "speed": 1.2, "volume": 0 },
    "format": "mp3"
  }'

Response (tested):

{
  "audio_url": "https://cdn.acedata.cloud/assets/examples/fish/5ade0339-5f11-487e-aacc-06a908271706-8e3fcb0e5547.mp3"
}

A speed greater than 1 speeds up, while less than 1 slows down; the unit of volume is dB, where 0 means unchanged, positive values increase gain, and negative values attenuate.

Example 5: Switching Models + Controlling Bitrate

Switch to the stable model through the HTTP header model: s1, and add mp3_bitrate: 128 in the request body to control the MP3 bitrate:

curl -X POST 'https://api.qiyaov.com/fish/tts' \
  -H 'authorization: Bearer {token}' \
  -H 'content-type: application/json' \
  -H 'model: s1' \
  -d '{
    "text": "high bitrate mp3",
    "format": "mp3",
    "mp3_bitrate": 128
  }'

Response (tested):

{
  "audio_url": "https://cdn.acedata.cloud/assets/examples/fish/7e7abf3d-3d72-4c9f-8eb6-8af932d7c96e-6f11bf2f30e9.mp3"
}

Example 6: PCM Raw Waveform

For scenarios where real-time concatenation is needed in the browser, or subsequent processing (mixing, speed adjustment) is needed on the client side, it is recommended to use pcm:

curl -X POST 'https://api.qiyaov.com/fish/tts' \
  -H 'authorization: Bearer {token}' \
  -H 'content-type: application/json' \
  -d '{
    "text": "hi",
    "format": "pcm",
    "sample_rate": 16000
  }'

Response (tested):

{
  "audio_url": "https://cdn.acedata.cloud/assets/examples/fish/64adc04b-c196-4a0f-9070-222ba101ce6c-fc50de38c165.wav"
}

The extension of the link follows the format in the request: mp3 produces .mp3, while wav and pcm produce .wav (WAV container, 16-bit PCM).

Asynchronous Callback (callback_url)

Synthesizing long text at once may take from over ten seconds to dozens of seconds, and retries are required if the connection is interrupted. After passing callback_url in the request body, the API immediately returns {task_id, started_at}. When the upstream processing is actually complete, it sends the complete result back to that URL as POST JSON, with the same task_id and the top-level cost for the final billing of this request in the body. The initial task confirmation has not completed synthesis yet, so it does not include cost.

curl -X POST 'https://api.qiyaov.com/fish/tts' \
  -H 'authorization: Bearer {token}' \
  -H 'content-type: application/json' \
  -d '{
    "text": "今天天气真好,我们一起出去散散步吧。",
    "format": "mp3",
    "callback_url": "https://webhook.site/4815f79f-a40f-4078-ac85-1cc126b6bb34"
  }'

Immediately returns (tested):

{
  "task_id": "79d82713-2897-4eeb-9934-e7544d471aa7",
  "started_at": 1778462584.742
}

Later, callback_url will receive something like:

{
  "task_id": "79d82713-2897-4eeb-9934-e7544d471aa7",
  "audio_url": "https://cdn.acedata.cloud/assets/examples/fish/bd66b8c5-7543-4557-b684-baa72407e336-52f2e79732f2.mp3"
}

You can also actively retrieve results by task_id using the Fish Tasks API; the response.cost in the final-state record is consistent with the cost in the callback. See that document for details.

Error Handling

  • 400 token_mismatched: Request parameters are missing or invalid (most commonly, text is empty, or format is assigned a value other than mp3/wav/pcm).
  • 401 invalid_token: The authentication token does not exist or is invalid.
  • 429 too_many_requests: The account rate limit has been triggered.
  • 500 api_error: Internal server error.

Error response example:

{
  "success": false,
  "error": {
    "code": "api_error",
    "message": "fetch failed"
  },
  "trace_id": "2cf86e86-22a4-46e1-ac2f-032c0f2a4e89"
}

Parameter validation errors will specify the invalid field in the message field, for example:

{
  "status": 400,
  "message": "[{\"type\":\"literal_error\",\"loc\":[\"format\"],\"msg\":\"Input should be 'pcm' or 'mp3'\",\"input\":\"wav\"}]"
}

Conclusion

The minimum cost of integrating Fish TTS is: replace the authentication in existing code calling api.fish.audio/v1/tts with this platform's token, and explicitly include format: "mp3" in the request body. For long-text scenarios, it is recommended to use the callback_url asynchronous callback; for discovering cloned voice reference_ids, please use Fish Model Query together with Fish Model Get.