A preset-voice text-to-speech model for real-time applications
tts-1 is OpenAI's text-to-speech model optimized for real-time use cases, converting prepared text into playable speech. It is suitable for application prompts, content narration, and announcements in voice interactions, offering six preset voices as well as audio format and speech speed options. Compared with tts-1-hd, which prioritizes audio quality, tts-1 is better suited for tasks that prioritize responsiveness.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and interface features
Clarify capacity, input and output, and invocation methods before selecting a model.
Creation method
Text-to-speech: input text, output audio
Native positioning
Optimized for real-time use cases
Preset voices
alloy、echo、fable、onyx、nova、shimmer
Audio format options
mp3、opus、aac、flac、wav、pcm; mp3 by default
Speech speed control
speed numeric parameter, default 1.0
Invocation endpoint
POST /v1/audio/speech; explicitly specify model="tts-1"
Real-time application positioning and preset voices are publicly available native capabilities; use formats, speech speed, and request methods through this platform's speech generation endpoint.
Core capabilities
Learn what tts-1 can bring to your work.
Chosen for timely announcements
The core trade-off of tts-1 is its focus on real-time use rather than making refined audio quality the sole objective. For action feedback, prompts, and narrated conversational responses, you can first prepare concise text, then generate speech and connect it to a player. This positioning is suitable for applications that value responsiveness, but it does not imply a fixed generation time.
Use preset voices to unify the auditory style
Six preset voices allow applications to audition the same text and select a voice that suits the content style, then maintain consistency in subsequent announcements. voice uses fixed names and does not require uploaded recordings. When selecting a model, compare how real copy sounds rather than inferring gender, accent, or emotion solely from voice names.
Configurable audio delivery and speech speed
When generating speech, you can choose the audio format and adjust the playback pace through speed. The default mp3 is convenient for integration with common players, while other formats can be selected according to the delivery pipeline. The returned content is audio data, and applications should save or play it in the selected format rather than parse the result as chat text or a download URL.
Applicable Scenarios
Start with specific tasks to find where the model can be effective.
Application Notifications and Action Guidance
Organize order statuses, operation steps, or device prompts into clear short sentences, select a fixed voice to generate spoken audio, and deliver it to the application player. For prompts used repeatedly, generated audio can be saved on the business side; for dynamic information, first complete text validation, then generate speech, to avoid reading incorrect content directly to users.
Reading Articles and Knowledge Content Aloud
Enter edited article paragraphs, course explanations, or knowledge cards to output playable narration audio. It is recommended to organize content by semantic paragraphs, first listen to a sample containing terminology, numbers, and abbreviations, and then determine the voice and speaking rate. The deliverable is speech corresponding to the text; content rewriting and fact checking should be completed before generation.
Response Playback for Voice Assistants
In a voice assistant, tts-1 can handle the final step of converting response text to speech: first obtain a reply through business logic or a dialogue model, then submit text suitable for spoken expression for generation. It is responsible for speaking predetermined content and does not handle recording recognition or response reasoning, making it suitable for building voice interaction flows with clear responsibilities.
How to Choose This Model
Choose based on task complexity, input materials, and expected results.
Prioritize Real-Time Experience: Choose tts-1
If the task involves frequent short prompts, immediate feedback, or dialogue playback, and response experience is the priority, tts-1 is a better fit. During evaluation, use real text, listen on target devices, and measure the complete playback chain; native real-time optimization does not mean that network, queuing, and playback startup times in an application are guaranteed.
Prioritize Audio Quality: Compare with tts-1-hd
tts-1-hd is a separate model focused on audio quality, not a voice option or output format for tts-1. For repeatedly played narration and polished content, compare both models using the same script, then decide based on listening quality and response requirements. Choosing a more refined format does not mean switching to the HD model; voice, format, and model should be considered separately.
Get Started: Generate Timely Playback for Application Prompts
Prepare the input first, then connect it to the corresponding application flow.
Prepare Input
Prepare short prompt copy, the expected playback duration, and the selected voice, and standardize pronunciations for numbers and proper nouns.
Organize the Call and Follow-Up Flow
Submit model=tts-1, input, and voice using /v1/audio/speech. Save the binary result according to the returned audio format, first listen in the actual player, then incorporate the chosen voice and format into content production.
Practical task example: Generate timely announcements for app prompts
Design tasks directly from the inputs and acceptance priorities below.
Suggested task
Submit model=tts-1, input, and voice via /v1/audio/speech; first preview with the default speed and MP3, then adjust the speaking rate.
Key checks
Check response time, prompt intelligibility, and completeness at the beginning and end of sentences; in the real app, confirm that the audio format is playable and that text prompts appear on failure.
Usage boundaries
Before production use, understand the output quality and capability scope.
tts-1 is a text-to-speech model, not a speech recognition model. Submit the text that needs to be read aloud; recording transcription, automated responses, and copy generation need to be completed in other stages. Do not treat ordinary text input as specialized instructions for precisely controlling emotion, performance style, or character identity.
Selecting a preset voice is not the same as voice cloning. The workflow here uses voice to select a fixed voice; do not design the process around uploading reference recordings, replicating someone's vocal timbre, or precisely imitating an accent. Before production use, preview names, technical terms, and mixed-language text.
Real-time positioning does not mean fixed latency, nor does it guarantee that text of any length can be completed in one request. Long content should be split by meaning, with the application managing concatenation, playback, and retries; after adjusting the speaking rate, preview again to ensure that numbers, abbreviations, and important prompts remain clear.
Frequently Asked Questions
Answers to common questions about using tts-1.
What is the main difference between tts-1 and tts-1-hd?
The main difference is their optimization focus: tts-1 is designed for real-time use cases, while tts-1-hd prioritizes audio quality. For instant prompts and conversational announcements, consider tts-1 first; for polished narration, compare audio samples using the same script, as you cannot determine which model was used based on the file format alone.
Which voices can tts-1 use?
You can choose alloy, echo, fable, onyx, nova, or shimmer, with alloy as the default voice. It is recommended to test each one with actual business scripts before settling on a voice. The names themselves do not promise an accent, emotion, or character, nor do they indicate the ability to replicate any person's voice.
How do I explicitly call tts-1?
Submit the input to be spoken to POST /v1/audio/speech, and explicitly set model to tts-1; add voice, response_format, and speed as needed. Explicitly selecting the model makes the request intent clearer, and the returned result should be handled as binary audio data.
What formats can it output, and can the speed be adjusted?
Audio format options include mp3, opus, aac, flac, wav, and pcm, with mp3 as the default; speed defaults to 1.0. Choose a format based on your player and audio processing workflow, and listen after adjusting the speed rather than treating unspecified numeric ranges as usable limits.
Can tts-1 directly listen to recordings and respond?
You cannot treat tts-1 as a complete voice conversation model. It receives text and generates speech; if user input is a recording, it must first be transcribed, then response text generated, and finally passed to tts-1 for narration. This also makes it easier to review the content before playback and improve the spoken expression of the text.