Billing & Pricing
Understanding how Sub2API billing works, including pricing structure, payment methods, and cost optimization.
Pricing Model
Sub2API uses pay-as-you-go pricing with no subscriptions or commitments.
How It Works
- Top up your balance with a minimum of $10 USD
- Make API calls using your balance
- Pay per token consumed (input + output)
- Top up again when balance runs low
No monthly fees, no hidden costs, no surprises.
Token-Based Pricing
All models are priced per million tokens:
What is a Token?
Tokens are pieces of words used by AI models. Roughly:
- 1 token ≈ 4 characters in English
- 1 token ≈ ¾ of a word on average
- 100 tokens ≈ 75 words approximately
Example
The sentence "Sub2API makes AI accessible" contains approximately:
- 7 words
- 28 characters
- ~7 tokens
Input vs Output Tokens
- Input tokens: Your prompt and conversation history
- Output tokens: The model's response
Output tokens typically cost more than input tokens because they require generation.
Pricing by Model
See the Models & Pricing page for current rates.
Example Costs
For a typical chat interaction:
Prompt: "Explain quantum computing" (4 tokens)
Response: 150 tokens of explanation
Using GPT-4 Turbo at $0.01/1M input, $0.03/1M output:
- Input cost: 4 × $0.01 / 1,000,000 = $0.00004
- Output cost: 150 × $0.03 / 1,000,000 = $0.0045
- Total: $0.00454 (~half a cent)
Payment Methods
Credit/Debit Cards
Processed securely via Stripe:
- Visa, Mastercard, American Express
- 3D Secure supported
- Instant balance credit
- Receipt emailed automatically
Minimum Top-Up
$10 USD minimum per transaction to cover processing fees.
Supported Currencies
Currently USD only. Your bank may apply conversion fees for non-USD cards.
Top-Up Process
- Go to Billing in the console
- Click Top Up Balance
- Enter amount (minimum $10)
- Complete payment via Stripe
- Balance updates immediately
- Receipt sent to your email
Balance Management
Checking Your Balance
Your current balance is always visible in:
- Console header (top bar)
- Billing page (detailed view)
Low Balance Notifications
Email alerts when balance drops below:
- $5.00 (warning)
- $2.00 (urgent)
- $0.50 (critical)
Insufficient Balance
When balance reaches zero:
- API requests return
402 Payment Required - Service resumes immediately after top-up
- No penalties or reactivation fees
Usage Tracking
Real-Time Monitoring
Track consumption in the console:
Usage Page:
- Total requests
- Tokens consumed (input/output breakdown)
- Cost per model
- Time period filters (day, week, month, custom)
Request Logs:
- Individual request details
- Exact token counts
- Cost per request
- Timestamps and status
Billing History
View all transactions:
- Top-up orders
- Payment confirmations
- Balance changes
- Refunds (if applicable)
Cost Optimization
1. Choose the Right Model
Use the most cost-effective model for your task:
Simple tasks: gpt-3.5-turbo ($0.0015/1M)
Complex tasks: gpt-4-turbo ($0.01/1M)
Long context: claude-3-5-sonnet ($0.003/1M)2. Limit Output Length
Set max_tokens to prevent excessive generation:
response = client.chat.completions.create(
model="gpt-4-turbo",
messages=[...],
max_tokens=200 # Limit response length
)3. Use System Messages
Give clear instructions to get focused responses:
messages = [
{"role": "system", "content": "Be concise. Maximum 2 sentences."},
{"role": "user", "content": "What is AI?"}
]4. Cache Responses
Store and reuse responses for common queries:
from functools import lru_cache
@lru_cache(maxsize=100)
def get_cached_response(prompt):
# Only calls API if prompt is new
return client.chat.completions.create(...)5. Optimize Context
Only send necessary conversation history:
# Keep last 10 messages instead of entire conversation
recent_messages = conversation_history[-10:]6. Use Streaming Wisely
Stream only when needed (user-facing responses). For batch processing, use regular requests.
7. Monitor and Alert
Set up alerts for unusual spending:
def check_daily_spending():
today_cost = get_usage_cost(today)
if today_cost > 10.00: # $10 daily threshold
send_alert(f"Daily spending: ${today_cost}")Workspace Billing
Personal Workspace
- You own and fund the workspace
- Only you can top up
- Only you can view billing
Organization Workspace
- Organization Owner manages billing
- Owner tops up organization balance
- Members cannot see balance or billing
- Usage tracked per member and per key
Billing Isolation
- Personal and organization balances are separate
- API keys are tied to one workspace
- Cannot transfer balance between workspaces
Spending Limits
Per-Key Limits
Set monthly spend limits when creating a key:
- Create API key
- Enable "Monthly Spend Limit"
- Set amount (e.g., $50)
- Key stops working when limit is reached
- Resets at the start of each month (UTC)
Workspace Limits
Organization Owners can set workspace-level monthly limits:
- Go to Organization Settings
- Set "Monthly Spend Limit"
- Applies to all keys in the workspace
- Prevents budget overruns
TIP
Set conservative limits initially, then adjust based on actual usage patterns.
Refunds
Refund Policy
Unused balance may be refunded:
- Within 30 days of top-up
- Minus any amount already consumed
- Minus payment processing fees
- Processed within 7-10 business days
How to Request
Email support@tokenlio.ai with:
- Account email
- Reason for refund
- Refund amount requested
Non-Refundable
- Consumed credits (used for API calls)
- Top-ups older than 30 days
- Amounts below $5 (processing fees exceed refund)
Billing FAQ
How are tokens calculated?
Tokens are counted by the model's tokenizer. The exact count is returned in each API response:
{
"usage": {
"prompt_tokens": 10,
"completion_tokens": 50,
"total_tokens": 60
}
}Are there bulk discounts?
Currently no volume discounts. Contact sales@tokenlio.ai for enterprise pricing.
What about failed requests?
Failed requests are not charged:
- Authentication errors (401)
- Rate limit errors (429)
- Server errors (500, 503)
Only successful completions consume balance.
Can I set a hard budget?
Yes, use key-level or workspace-level monthly spend limits to enforce budgets.
What happens with streaming?
Streaming requests are charged the same as non-streaming. Final token count is the same.
Do retries cost extra?
If you retry a failed request, only successful attempts are charged.
Is there a free tier?
No free tier currently. Minimum top-up is $10, which covers thousands of requests with cost-effective models.
How do I download invoices?
Go to Billing > Payment History and click Download Receipt for each transaction.
Can I pay annually?
No prepaid plans currently. Sub2API uses pay-as-you-go only.
Cost Estimation Tools
Calculate Before You Call
Estimate costs before making requests:
def estimate_cost(prompt, expected_output_tokens, model="gpt-4-turbo"):
# Rough token estimate: 1 token ≈ 4 chars
input_tokens = len(prompt) / 4
output_tokens = expected_output_tokens
# Model pricing (per 1M tokens)
prices = {
"gpt-4-turbo": {"input": 0.01, "output": 0.03},
"gpt-3.5-turbo": {"input": 0.0005, "output": 0.0015},
}
rate = prices.get(model, prices["gpt-4-turbo"])
cost = (input_tokens * rate["input"] + output_tokens * rate["output"]) / 1_000_000
return round(cost, 6)
# Example
prompt = "Write a blog post about AI"
estimated = estimate_cost(prompt, 500)
print(f"Estimated cost: ${estimated}")Monitor Actual Costs
Log costs after each request:
def log_request_cost(response):
usage = response.usage
# Get model pricing
cost = calculate_cost(usage.prompt_tokens, usage.completion_tokens, model)
print(f"Request cost: ${cost:.6f}")
print(f"Tokens: {usage.total_tokens}")Enterprise Billing
For high-volume users:
- Custom pricing available
- Dedicated support
- Monthly invoicing options
- Contract terms
Contact sales@tokenlio.ai for details.