Embeddings overview
Turn text into numerical vectors for semantic search, document retrieval, clustering and RAG. Both models use the OpenAI-compatible POST /v1/embeddings endpoint. The response contains vectors rather than a generated answer.
Turn text into numerical vectors for semantic search, document retrieval, clustering and RAG. Both models use the OpenAI-compatible POST /v1/embeddings endpoint. The response contains vectors rather than a generated answer.
Choose a model
| Model | Default dimensions | Usage guide |
|---|---|---|
text-embedding-3-small |
1536 | text-embedding-3-small |
bge-m3 |
1024 | BGE-M3 |
An OpenAI embedding model with 1536 dimensions by default. Suitable for semantic search and knowledge-base indexing. APIMaster also supports a tested 512-dimensional output through the dimensions parameter.
A multilingual BAAI embedding model with 1024 dimensions. Suitable for multilingual document retrieval. This API returns dense vectors; it does not expose sparse weights or ColBERT vectors. No special query instruction prefix is needed.
Parameters and limits
POST https://apimaster.ai/v1/embeddings
Send a non-empty string or an array of non-empty strings in input. Use encoding_format: float for a numerical array. Do not send messages, max_tokens or stream. The model reference limit is 8192 tokens per text; split long documents and keep batches small. Limits and tokenization can differ by route.
Semantic search and RAG
Split documents into passages, embed and store each passage, then embed the query with the same model and dimension. Rank passages by cosine similarity and pass the retrieved text to a separate chat model. Do not mix vectors from different models or dimensions in one index; rebuild the index when switching models.
