Per token, not per request
Claude-class APIs bill on tokens — roughly three-quarters of a word each. Rates are quoted per million tokens, so the per-token price is just that number divided by a million. On KeyoAPI the four Claude lanes cost:
| Model ID | Input / output per 1M | Per 1K tokens in / out | Best for |
|---|---|---|---|
claude-sonnet-5 | ~$0.88 / $4.41 | $0.00088 / $0.00441 | Everyday production chat |
claude-opus-5 | ~$1.10 / $5.52 | $0.0011 / $0.00552 | Hard agent and coding escalations |
claude-fable-5-1 | ~$7.35 / $36.77 | $0.00735 / $0.03677 | Planner-tier flagship work |
claude-fable-5 | ~$8.82 / $44.12 | $0.00882 / $0.04412 | Flagship chat / agent twin |
Rates are indicative sell prices from the live catalog — confirm the interactive rows on /pricing/claude-sonnet-5, /pricing/claude-opus-5, /pricing/claude-fable-5 and /pricing/claude-fable-5-1 before contracting volume. The full comparison table lives on /claude-api-pricing.
A worked example
Say a support bot sends a 2,000-token prompt and gets a 400-token answer on claude-sonnet-5:
input: 2,000 / 1,000,000 × $0.88 = $0.00176
output: 400 / 1,000,000 × $4.41 = $0.00176
---------
total per conversation ≈ $0.0035
Ten thousand such conversations ≈ $35. The same conversation on claude-fable-5 would run ≈ $0.036 — ten times the cost for flagship quality you may not need on every turn. That ratio, not the headline price, is what per-token budgeting is about: route the routine 90% of traffic to the cheap lane and escalate only the hard cases.
Check the math on real traffic
Every response carries a usage object — read prompt_tokens and completion_tokens and multiply by the table above. Point any OpenAI SDK at the relay and pass a Claude model string:
curl https://www.keyoapi.xyz/v1/chat/completions \
-H "Authorization: Bearer YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"claude-sonnet-5","messages":[{"role":"user","content":"hi"}]}'
Cutting the bill
- Escalate, don't default. Run drafts and classification on
claude-sonnet-5; promote only failing slices toclaude-opus-5or the Fable lanes. - Shift bulk work off Claude entirely. High-volume drafts and classifiers often do fine on
gpt-6-lunaordeepseek-v4.1-flashon the same key — see /openai-api-pricing and /deepseek-api-pricing. - Prototype for free. Fixed $0 catalog IDs ending in
:freebill against your signup gift credit at fair-use limits — draft and smoke-test there, then switchmodel=to a Claude lane: /free-models. - Retry around overload, don't re-send. Claude occasionally returns HTTP 529 overloaded; the retry and fallback playbook at /brand/blog/article/ncx234u2/ keeps those from double-billing you.
One prepaid balance covers every lane above plus GPT-class, DeepSeek, OCR and TTS — no separate Anthropic invoice. Base URL: https://www.keyoapi.xyz/v1.