text-embedding-3-small · Send a request
An OpenAI embedding model with 1536 dimensions by default. Suitable for semantic search and knowledge-base indexing. APIMaster also supports a tested 512-dimensional output through the dimensions parameter.
An OpenAI embedding model with 1536 dimensions by default. Suitable for semantic search and knowledge-base indexing. APIMaster also supports a tested 512-dimensional output through the dimensions parameter.
Verified on 2026-10-02: the active route starts at $0.02 per million input tokens. This is APIMaster route pricing, not a universal vendor price; the live marketplace card takes precedence. Live model card
Send a request
POST /v1/embeddings · model: text-embedding-3-small
cURL
export APIMASTER_API_KEY="YOUR_APIMASTER_API_KEY"
curl --fail-with-body "https://apimaster.ai/v1/embeddings" \
-H "Authorization: Bearer $APIMASTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "text-embedding-3-small", "input": "Semantic search retrieves related documents.", "encoding_format": "float"}'
Python · OpenAI SDK
pip install openai
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["APIMASTER_API_KEY"],
base_url="https://apimaster.ai/v1",
)
result = client.embeddings.create(
model="text-embedding-3-small",
input=["Semantic search retrieves related documents.", "猫在窗边休息。"],
encoding_format="float",
)
for item in result.data:
print(item.index, len(item.embedding)) # 1536
print(result.usage.prompt_tokens)
JavaScript · Node.js
const response = await fetch("https://apimaster.ai/v1/embeddings", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.APIMASTER_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "text-embedding-3-small",
input: ["Semantic search retrieves related documents.", "猫在窗边休息。"],
encoding_format: "float",
}),
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const result = await response.json();
console.log(result.data.map((item) => [item.index, item.embedding.length]));
console.log(result.usage.prompt_tokens);
Parameters and limits
Send a non-empty string or an array of non-empty strings in input. Use encoding_format: float for a numerical array. Do not send messages, max_tokens or stream. The model reference limit is 8192 tokens per text; split long documents and keep batches small. Limits and tokenization can differ by route.
Only text-embedding-3-small supports the documented dimensions option here: 1536 by default, or 512 in the tested example. Keep the chosen dimension consistent for documents and queries.
Reduced dimensions
compact = client.embeddings.create(
model="text-embedding-3-small",
input="Semantic search retrieves related documents.",
encoding_format="float",
dimensions=512,
)
print(len(compact.data[0].embedding)) # 512
Read the response
Read data[i].embedding and use data[i].index to match each vector to its input. usage.prompt_tokens and usage.total_tokens describe input usage. A vector is not generated text and is not charged as completion tokens. base64 encoding was also tested, but float is easier to inspect.
Live tests passed for single text, Chinese/English batches, vector lengths and repeated-input consistency. These checks verify API behavior, not retrieval quality for every dataset.
Semantic search and RAG
Split documents into passages, embed and store each passage, then embed the query with the same model and dimension. Rank passages by cosine similarity and pass the retrieved text to a separate chat model. Do not mix vectors from different models or dimensions in one index; rebuild the index when switching models.
Troubleshooting
For 401, check the API key. For 400, check the model ID, non-empty input and optional parameters. For 429 or upstream timeouts, reduce concurrency and retry with backoff. Check the marketplace for route availability and current prices.
