Available Models & Pricing
Tokenlio aggregates frontier models from leading AI labs with standardized USD ($/1M Tokens) billing. Models are categorized by provider below.
INFO
Looking for real-time latency tests, vendor filters, and ⌘K search? Check out our Live Interactive Models Matrix.
OpenAI
Industry-standard frontier models including GPT-6, GPT-5.6 Sol, o3-mini, and full multimodal series.
| Model ID / Name | Modalities | Context | Input / 1M | Output / 1M | Cache Read | Tier |
|---|---|---|---|---|---|---|
openai/gpt-6-astra | text · image · video · audio | 1M | $15.00 | $45.00 | $3.00 | 20% off ⚡ |
openai/gpt-5.6-sol | text · image · video · audio | 512K | $5.00 | $30.00 | $2.50 | 20% off ⚡ |
openai/gpt-6.1-sol | text · image · video · audio | 512K | $12.00 | $24.00 | $2.40 | 20% off ⚡ |
openai/gpt-5.5-pro | text · image · audio | 256K | $4.00 | $20.00 | $1.00 | 20% off ⚡ |
openai/gpt-5.4-mini | text · image | 128K | $0.20 | $0.80 | $0.05 | Free Tier 🔥 |
openai/o3-mini | text | 200K | $1.10 | $4.40 | $0.55 | 20% off ⚡ |
openai/gpt-4o | text · image · audio | 128K | $2.50 | $10.00 | $1.25 | 20% off ⚡ |
openai/o1 | text · image | 200K | $15.00 | $60.00 | $7.50 | 20% off ⚡ |
openai/gpt-4o-mini | text · image · audio | 128K | $0.15 | $0.60 | $0.07 | Free Tier 🔥 |
openai/gpt-4o-audio | text · audio | 128K | $2.50 | $10.00 | $1.25 | 20% off ⚡ |
Anthropic
Claude series renowned for long-context reasoning, autonomous coding agents, and complex document analysis.
| Model ID / Name | Modalities | Context | Input / 1M | Output / 1M | Cache Read | Tier |
|---|---|---|---|---|---|---|
anthropic/claude-opus-5.5 | text · image · video · audio | 1M | $15.00 | $45.00 | $1.50 | 20% off ⚡ |
anthropic/claude-fable-5.1 | text · image · video · audio | 500K | $8.00 | $36.00 | $0.80 | 50% off 🏷️ |
anthropic/claude-sonnet-4.6 | text · image | 200K | $3.00 | $15.00 | $0.30 | 50% off 🏷️ |
anthropic/claude-haiku-4.5 | text · image | 200K | $0.80 | $4.00 | $0.08 | 20% off ⚡ |
anthropic/claude-3-opus | text · image | 200K | $15.00 | $75.00 | $1.50 | 20% off ⚡ |
DeepSeek
Extreme cost-efficiency MoE models with state-of-the-art reinforcement learning reasoning and math/code optimization.
| Model ID / Name | Modalities | Context | Input / 1M | Output / 1M | Cache Read | Tier |
|---|---|---|---|---|---|---|
deepseek/deepseek-v4-pro | text · image | 128K | $0.55 | $2.19 | $0.11 | Free Tier 🔥 |
deepseek/deepseek-r1 | text | 128K | $0.55 | $2.19 | $0.14 | Free Tier 🔥 |
deepseek/deepseek-chat | text | 128K | $0.14 | $0.28 | $0.0280 | Free Tier 🔥 |
deepseek/deepseek-coder-v2 | text | 128K | $0.14 | $0.28 | $0.0280 | 50% off 🏷️ |
deepseek/deepseek-math-7b | text | 32K | $0.10 | $0.20 | $0.0200 | Free Tier 🔥 |
Google
Ultra-long context windows up to 2M tokens with native video, image, audio, and text multimodality.
| Model ID / Name | Modalities | Context | Input / 1M | Output / 1M | Cache Read | Tier |
|---|---|---|---|---|---|---|
google/gemini-3.8-flash | text · image · video · audio | 2M | $0.95 | $1.90 | $0.19 | 50% off 🏷️ |
google/gemini-3.1-pro | text · image · video · audio | 2M | $1.25 | $5.00 | $0.31 | 20% off ⚡ |
google/gemini-2.5-pro | text · image · video · audio | 2M | $1.25 | $5.00 | $0.31 | 50% off 🏷️ |
google/gemini-2.0-flash | text · image · video · audio | 1M | $0.10 | $0.40 | $0.0250 | 50% off 🏷️ |
google/gemini-1.5-pro | text · image · video · audio | 2M | $1.25 | $5.00 | $0.31 | 20% off ⚡ |
google/gemini-2.0-flash-realtime | text · audio · video | 1M | $0.15 | $0.60 | $0.0300 | 50% off 🏷️ |
Alibaba Qwen
Comprehensive open-weight flagship models excelling in bilingual understanding, coding, and visual chain-of-thought.
| Model ID / Name | Modalities | Context | Input / 1M | Output / 1M | Cache Read | Tier |
|---|---|---|---|---|---|---|
qwen/qwen3.8-max | text · image | 128K | $1.60 | $4.80 | $0.32 | 20% off ⚡ |
qwen/qwen3.7-plus | text · image | 128K | $0.80 | $2.40 | $0.16 | 50% off 🏷️ |
qwen/qwen-2.5-coder-32b | text | 128K | $0.18 | $0.36 | $0.0360 | Free Tier 🔥 |
qwen/qvq-72b-preview | text · image | 32K | $0.80 | $2.40 | $0.16 | 50% off 🏷️ |
qwen/qwen-2.5-72b-instruct | text | 128K | $0.35 | $0.40 | $0.07 | Free Tier 🔥 |
qwen/qwen-2-audio-7b | text · audio | 32K | $0.15 | $0.30 | $0.0300 | Free Tier 🔥 |
Moonshot Kimi
Long-range context understanding and deep web search reinforcement from Moonshot AI.
| Model ID / Name | Modalities | Context | Input / 1M | Output / 1M | Cache Read | Tier |
|---|---|---|---|---|---|---|
moonshotai/kimi-k3 | text · image | 200K+ | $1.80 | $3.60 | $0.36 | 50% off 🏷️ |
moonshotai/kimi-k3-free | text | 128K | $0.000 | $0.000 | $0.0000 | Free Tier 🔥 |
xAI
High-performance real-world knowledge extraction and reasoning powered by xAI compute clusters.
| Model ID / Name | Modalities | Context | Input / 1M | Output / 1M | Cache Read | Tier |
|---|---|---|---|---|---|---|
x-ai/grok-4.5 | text · image | 256K | $2.00 | $6.00 | $0.40 | 50% off 🏷️ |
x-ai/grok-4.3 | text · image | 128K | $2.00 | $8.00 | $0.40 | 20% off ⚡ |
x-ai/grok-vision-beta | text · image | 128K | $5.00 | $15.00 | $1.00 | 20% off ⚡ |
Zhipu GLM
Bilingual Chinese-English flagship LLMs with strong multi-agent collaboration and vision capabilities.
| Model ID / Name | Modalities | Context | Input / 1M | Output / 1M | Cache Read | Tier |
|---|---|---|---|---|---|---|
z-ai/glm-5.3 | text · image · video | 128K | $1.20 | $2.40 | $0.24 | 20% off ⚡ |
z-ai/glm-5.2 | text · image | 128K | $1.00 | $2.00 | $0.20 | 20% off ⚡ |
z-ai/glm-4-flash | text | 128K | $0.060 | $0.060 | $0.0120 | Free Tier 🔥 |
z-ai/glm-4v-plus | text · image · video | 128K | $1.40 | $2.80 | $0.28 | 50% off 🏷️ |
z-ai/glm-4-long | text | 1M | $0.50 | $1.00 | $0.10 | 50% off 🏷️ |
Meta
Open-source frontier weights from Llama 3.3 70B to Llama 3.1 405B and edge multimodal models.
| Model ID / Name | Modalities | Context | Input / 1M | Output / 1M | Cache Read | Tier |
|---|---|---|---|---|---|---|
meta/llama-3.3-70b-instruct | text | 128K | $0.35 | $0.40 | $0.07 | Free Tier 🔥 |
meta/llama-3.1-405b-instruct | text | 128K | $1.80 | $3.60 | $0.36 | 20% off ⚡ |
meta/llama-3.2-90b-vision | text · image | 128K | $0.90 | $0.90 | $0.18 | 50% off 🏷️ |
meta/llama-3.2-11b-vision | text · image | 128K | $0.15 | $0.15 | $0.0300 | 50% off 🏷️ |
Mistral
European premier frontier models including Mistral Large, Codestral for fill-in-the-middle, and Pixtral vision.
| Model ID / Name | Modalities | Context | Input / 1M | Output / 1M | Cache Read | Tier |
|---|---|---|---|---|---|---|
mistral/mistral-large-2411 | text | 128K | $2.00 | $6.00 | $0.40 | 50% off 🏷️ |
mistral/codestral-2501 | text | 256K | $0.30 | $0.90 | $0.06 | 20% off ⚡ |
mistral/pixtral-12b | text · image | 128K | $0.15 | $0.15 | $0.0300 | Free Tier 🔥 |
mistral/pixtral-large-124b | text · image | 128K | $2.00 | $6.00 | $0.40 | 20% off ⚡ |
MiniMax
Up to 4M ultra-long context windows and native chain-of-thought processing at disruptive cost.
| Model ID / Name | Modalities | Context | Input / 1M | Output / 1M | Cache Read | Tier |
|---|---|---|---|---|---|---|
minimax/minimax-m3 | text · audio | 1M | $0.20 | $1.10 | $0.0400 | 50% off 🏷️ |
minimax/minimax-01 | text · image | 4M | $0.20 | $1.10 | $0.0400 | 50% off 🏷️ |
minimax/minimax-text-01 | text | 1M | $0.15 | $0.90 | $0.0300 | Free Tier 🔥 |
minimax/minimax-video-01 | text · video | 32K | $1.50 | $4.50 | $0.30 | 20% off ⚡ |
Cohere
Enterprise-grade RAG, search, and grounded citation generation models.
| Model ID / Name | Modalities | Context | Input / 1M | Output / 1M | Cache Read | Tier |
|---|---|---|---|---|---|---|
cohere/command-r-plus | text | 128K | $2.50 | $10.00 | $0.50 | 20% off ⚡ |
cohere/command-r | text | 128K | $0.50 | $1.50 | $0.10 | 50% off 🏷️ |
NVIDIA
Hardware-accelerated Nemotron models optimized for high-throughput reasoning and agent workloads.
| Model ID / Name | Modalities | Context | Input / 1M | Output / 1M | Cache Read | Tier |
|---|---|---|---|---|---|---|
nvidia/nemotron-3-nano-omni | text · image · audio | 128K | $0.080 | $0.16 | $0.0160 | Free Tier 🔥 |
Amazon
AWS cloud-native Nova series delivering balanced cost, speed, and enterprise reliability.
| Model ID / Name | Modalities | Context | Input / 1M | Output / 1M | Cache Read | Tier |
|---|---|---|---|---|---|---|
amazon/nova-pro-v1 | text · image · video | 300K | $0.80 | $3.20 | $0.20 | 20% off ⚡ |
amazon/nova-lite-v1 | text · image · video | 300K | $0.060 | $0.24 | $0.0150 | 50% off 🏷️ |
amazon/nova-micro-v1 | text | 128K | $0.035 | $0.14 | $0.0088 | Free Tier 🔥 |
Integration Example
curl https://api.tokenlio.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "openai/gpt-6-astra",
"messages": [{"role": "user", "content": "Hello Tokenlio!"}]
}'