Embeddings API
Overview
Generate vector embeddings for text using state-of-the-art embedding models. Embeddings are useful for semantic search, clustering, recommendations, and similarity comparisons.
Create Embeddings
Endpoint
POST https://api.tokenlio.ai/v1/embeddingsRequest
bash
curl https://api.tokenlio.ai/v1/embeddings \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $TOKENLIO_API_KEY" \
-d '{
"model": "text-embedding-3-small",
"input": "The quick brown fox jumps over the lazy dog"
}'Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Embedding model ID |
input | string or array | Yes | Text to embed (max 8191 tokens per input) |
encoding_format | string | No | Format: float (default) or base64 |
dimensions | integer | No | Output dimensions (model-specific) |
user | string | No | Unique user identifier for abuse monitoring |
Response
json
{
"object": "list",
"data": [
{
"object": "embedding",
"embedding": [
-0.006929283,
-0.005336422,
... // 1536 dimensions
-0.01086957
],
"index": 0
}
],
"model": "text-embedding-3-small",
"usage": {
"prompt_tokens": 9,
"total_tokens": 9
}
}Available Models
| Model | Dimensions | Price per 1M tokens | Max Tokens |
|---|---|---|---|
text-embedding-3-small | 1536 | $0.02 | 8191 |
text-embedding-3-large | 3072 | $0.13 | 8191 |
text-embedding-ada-002 | 1536 | $0.10 | 8191 |
Examples
Python
python
from openai import OpenAI
client = OpenAI(
api_key="your-tokenlio-key",
base_url="https://api.tokenlio.ai/v1"
)
response = client.embeddings.create(
model="text-embedding-3-small",
input="Your text here"
)
embedding = response.data[0].embedding
print(f"Embedding dimension: {len(embedding)}")Batch Embeddings
Embed multiple texts in one request:
python
texts = [
"First document",
"Second document",
"Third document"
]
response = client.embeddings.create(
model="text-embedding-3-small",
input=texts
)
embeddings = [item.embedding for item in response.data]Node.js
javascript
import OpenAI from 'openai';
const client = new OpenAI({
apiKey: 'your-tokenlio-key',
baseURL: 'https://api.tokenlio.ai/v1'
});
const response = await client.embeddings.create({
model: 'text-embedding-3-small',
input: 'Your text here'
});
const embedding = response.data[0].embedding;
console.log(`Embedding dimension: ${embedding.length}`);Use Cases
Semantic Search
- Embed documents:
python
documents = ["Paris is the capital of France", "Berlin is the capital of Germany"]
doc_embeddings = [
client.embeddings.create(model="text-embedding-3-small", input=doc).data[0].embedding
for doc in documents
]- Embed query:
python
query = "What is France's capital?"
query_embedding = client.embeddings.create(
model="text-embedding-3-small",
input=query
).data[0].embedding- Calculate similarity:
python
import numpy as np
def cosine_similarity(a, b):
return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
similarities = [
cosine_similarity(query_embedding, doc_emb)
for doc_emb in doc_embeddings
]
# Get most similar document
best_match_idx = np.argmax(similarities)
print(f"Best match: {documents[best_match_idx]}")Clustering
python
from sklearn.cluster import KMeans
# Embed documents
docs = ["doc1", "doc2", "doc3", ...]
embeddings = [
client.embeddings.create(model="text-embedding-3-small", input=doc).data[0].embedding
for doc in docs
]
# Cluster
kmeans = KMeans(n_clusters=3)
clusters = kmeans.fit_predict(embeddings)Recommendations
python
# User preferences embedding
user_prefs = "I like science fiction movies"
user_emb = client.embeddings.create(
model="text-embedding-3-small",
input=user_prefs
).data[0].embedding
# Item embeddings
items = ["Star Wars", "The Godfather", "Interstellar"]
item_embs = [
client.embeddings.create(model="text-embedding-3-small", input=item).data[0].embedding
for item in items
]
# Rank by similarity
scores = [cosine_similarity(user_emb, item_emb) for item_emb in item_embs]
recommendations = sorted(zip(items, scores), key=lambda x: x[1], reverse=True)Best Practices
Text Preprocessing
Clean text:
- Remove HTML tags
- Normalize whitespace
- Handle special characters
Optimal length:
- Shorter texts (< 512 tokens) work best
- Split long documents into chunks
Meaningful content:
- Avoid embedding UI elements, navigation, etc.
- Focus on semantic content
Caching
Cache embeddings to save costs:
python
import json
import hashlib
def get_embedding_cached(text, cache_file="embeddings_cache.json"):
# Load cache
try:
with open(cache_file) as f:
cache = json.load(f)
except FileNotFoundError:
cache = {}
# Generate cache key
key = hashlib.md5(text.encode()).hexdigest()
if key in cache:
return cache[key]
# Generate embedding
embedding = client.embeddings.create(
model="text-embedding-3-small",
input=text
).data[0].embedding
# Save to cache
cache[key] = embedding
with open(cache_file, 'w') as f:
json.dump(cache, f)
return embeddingBatch Processing
Process in batches to improve efficiency:
python
def embed_batch(texts, batch_size=100):
embeddings = []
for i in range(0, len(texts), batch_size):
batch = texts[i:i+batch_size]
response = client.embeddings.create(
model="text-embedding-3-small",
input=batch
)
embeddings.extend([item.embedding for item in response.data])
return embeddingsDimensions Parameter
Some models support custom dimensions:
python
# Smaller embedding (faster, cheaper, less accurate)
response = client.embeddings.create(
model="text-embedding-3-large",
input="Your text",
dimensions=256 # Default: 3072
)Tradeoff: Lower dimensions = faster search, less storage, lower accuracy.
Storage
Store embeddings efficiently:
Vector Databases
- Pinecone: Managed vector DB
- Weaviate: Open-source vector DB
- Qdrant: Fast similarity search
- Chroma: Embeddings database
Example with NumPy
python
import numpy as np
# Save
embeddings_array = np.array(embeddings)
np.save('embeddings.npy', embeddings_array)
# Load
loaded = np.load('embeddings.npy')Rate Limits
- Default: 3000 requests/minute
- Batch limit: 2048 inputs per request
- Max tokens per input: 8191
Pricing
Charged per token (input text):
| Model | Price per 1M tokens |
|---|---|
| text-embedding-3-small | $0.02 |
| text-embedding-3-large | $0.13 |
| text-embedding-ada-002 | $0.10 |
Example: 1000 documents × 200 tokens each = 200K tokens = $0.004 (small model)
Error Handling
python
from openai import APIError
try:
response = client.embeddings.create(
model="text-embedding-3-small",
input=text
)
except APIError as e:
if e.code == "context_length_exceeded":
# Text too long
print("Text exceeds max tokens")
else:
raiseMigration Guide
From OpenAI
Change base_url only:
python
# Before
client = OpenAI(api_key="sk-...")
# After
client = OpenAI(
api_key="tk-...",
base_url="https://api.tokenlio.ai/v1"
)Everything else stays the same!
Support
Questions about embeddings?
- Email: support@tokenlio.ai
- Docs: /docs/