Anthropic's newest small model for high-volume, latency-sensitive work such as classification, extraction, routing, and subagent tasks. Prompts over 100K tokens are billed at the higher long-prompt rate.
Every rate is 50% of the provider's official list price. USD per million tokens.
| Token type | List price | Tokenless |
|---|---|---|
| Input | $0.10 | $0.05 |
| Output | $0.50 | $0.25 |
| Cache write | $0.13 | $0.06 |
| Cache read | $0.01 | $0.01 |
| Prompts over 100K tokens | ||
| Input | $0.50 | $0.25 |
| Output | $2.50 | $1.25 |
| Cache write | $0.63 | $0.31 |
| Cache read | $0.05 | $0.03 |
A request whose prompt, including cached tokens, is over the threshold is billed entirely at the long-prompt rates, as the provider bills it.
Tokenless speaks the OpenAI Chat Completions format. Point your client at the base URL, set model to haiku-5.5, and you are done.
curl https://api.tokenless.store/api/v1/chat/completions \
-H "Authorization: Bearer $TOKENLESS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "haiku-5.5",
"messages": [{ "role": "user", "content": "Hello" }]
}'