All models

text-embedding-3-small

OpenAIEmbedding
Get your API key
text-embedding-3-small

Efficient text embedding model for semantic retrieval and knowledge bases

text-embedding-3-small is OpenAI's third-generation text embedding model, designed to convert text into numerical vectors that represent semantic meaning rather than generate chat responses. It is suitable for knowledge base retrieval, similar content matching, and text clustering, and supports adjusting vector dimensions to balance retrieval quality, storage space, and computational overhead, making it a practical starting point for building everyday semantic retrieval systems.

OpenAIModel brand
VectorModel type
VectorTask capability

Specifications and interface features

Clarify capacity, input and output, and invocation methods before choosing a model.

Model type
Text embedding; outputs semantic vectors and does not generate responses
Invocation endpoint
POST /openai/embeddings;model=text-embedding-3-small
Input formats
Non-empty text, text arrays, non-negative integer token arrays, or batches thereof
Batch size
Text batches or token array batches support up to 2048 items
Dimension control
dimensions can shorten vectors; full dimensions are returned when unspecified
Output encoding
float or base64, with float as the default; results include index and token usage

Dimension reduction is a native model capability; the input formats, batch size, and output encodings above correspond to this platform's invocation endpoint.

Core Capabilities

Learn what text-embedding-3-small can bring to your work.

Connect Different Expressions with Semantics

The model maps text content to numerical vectors, providing a foundation for similarity search. Even if user questions and source materials use different wording, related content can be found through vector comparison. It handles semantic representation, while retrieval ranking, filter conditions, and final answers are still completed by the application, making it suitable for use alongside keyword search.

Support Multilingual Retrieval Tasks

In official benchmark results, text-embedding-3-small achieved an average MIRACL score of 44.0%, higher than ada-002's 31.4%; its average MTEB score was 62.3%. These results demonstrate retrieval improvements over the previous version, but specific languages, industry terminology, and text chunking methods can still affect business performance.

Adjust Dimensions for Retrieval Needs

With dimensions, you can shorten output vectors and reduce the number of values in each record, making it easier to balance vector database storage and computational overhead. If not specified, the full dimensions are output. After shortening, recall performance should be reevaluated, and document vectors and query vectors should use the same model and the same dimension configuration.

Use Cases

Start with specific tasks to find where the model can be effective.

Knowledge Base Retrieval

First extract product manuals, help documents, or internal policies into text and split them into chunks, then generate vectors in batches and write them to a retrieval database. When a question is asked, generate a vector for it, find relevant passages, and provide them to an answer model as reference. This model handles the document retrieval stage and does not directly produce knowledge base answers with citations.

Similar Content Search

Convert customer service tickets, question titles, or product descriptions into vectors, compare the similarity of new content with existing records, and return a candidate list. This can be used to find similar issues, link historical resolution cases, or identify similar descriptions; whether items are duplicates or can be merged should be determined using business rules and human review.

Text Clustering and Topic Organization

Vectorize feedback, comments, or short texts in batches, then use clustering algorithms to organize content groups and deliver topic clusters and representative texts. Applications can use the returned index to align with original records, making it easier to review the content in each group. The model provides vector representations; topic naming and summarization must be completed separately.

How to choose this model

Choose based on task complexity, input materials, and expected results.

Evaluate small first for new systems

If the goal is everyday semantic search, similar-text matching, or knowledge-base retrieval, you can start with small to establish a baseline. Compared with ada-002, it performed better in the multilingual retrieval and English task evaluations at official release, and it supports dimension reduction. When migrating existing legacy indexes, regenerate document vectors rather than directly mixing results from different models.

Compare large when accuracy is the priority

If your business is more sensitive to retrieval accuracy, compare small and text-embedding-3-large on the same test set. The large model scored higher in the official release evaluations, but that does not mean every business will see the same benefit. It is recommended to combine real questions, relevant passage annotations, and index configuration to determine whether the improvement justifies using a larger model.

Get started

From a small-scale task to production integration.

01

Prepare tasks and materials

Define the goal, required inputs, and output requirements, using real business examples as a starting point.

02

Try it in the API testing area

Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.

03

Integrate according to the API documentation

Keep the full model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.

Usage limits

Understand output quality and capability scope before production use.

  • It outputs numerical vectors; it does not directly generate natural-language answers, automatically create a vector database, run clustering, or complete retrieval. A complete application still requires steps such as text processing, storage, and similarity calculation, and RAG scenarios also require a separate answer-generation model.
  • Input must be non-empty text or a compliant token array; PDF files cannot be submitted directly as text. A maximum batch size of 2048 refers to the number of input items and does not mean documents of any length can be processed at once; for long documents, first extract the main text and split it semantically.
  • Smaller dimensions are not always better, nor does dimensions mean vectors can be expanded arbitrarily. After adjusting dimensions or changing models, update index configuration and document vectors accordingly; retrieval performance must be validated with real queries, and average evaluation scores cannot replace business acceptance testing.

Frequently Asked Questions

Answers to common questions about using text-embedding-3-small.

Can text-embedding-3-small answer questions directly?

No. It converts questions or materials into vectors for applications to compare semantic similarity. Knowledge-base question answering typically uses it first to retrieve relevant passages, then passes them to a response model to formulate an answer; calling only the embeddings API returns vectors, not explanations, summaries, or conversational replies.

How does it differ from text-embedding-ada-002?

It is a third-generation embedding model, not an alias for ada-002. Official release evaluations show that it has higher average scores on multilingual retrieval and English tasks, and it supports shortening vectors through dimensions. When migrating, rebuild material vectors to avoid comparing results from the two models in the same space.

When should you choose text-embedding-3-large?

When retrieval precision is more important than being lightweight, it is worth comparing large. It scored higher in the official release evaluations, but the choice should still be based on your own question set and document collection. You can first establish a small baseline, then evaluate whether large improves retrieval of relevant passages for key questions.

How do you submit batches and match returned results?

Submit model and input to POST /openai/embeddings. input can use batches of text arrays or token arrays, with up to 2048 items. Each item in the returned data includes index and embedding; use index to align with the original input, and check usage for token consumption.

What do dimensions and encoding_format each control?

dimensions controls vector shortening; if not specified, the full dimensions are returned. encoding_format controls the returned representation and can be float or base64, with float as the default. The former requires evaluating retrieval performance, while the latter adapts to data processing methods; do not interpret the encoding choice as a model accuracy level.

Model information · Updated: 2026-10-01. See the API and pricing sections for calling parameters and billing rules.

Use text-embedding-3-small for your next task

Start with a clear goal and determine from real results whether it is right for your work.