OpenAI's latest mid-tier model, built for agentic workflows and complex coding across a 1M-token context with reasoning effort from low to max. Tool calling runs on the Responses API; Chat Completions serves it without tools.
Every rate is 50% of the provider's official list price. USD per million tokens.
| Token type | List price | Tokenless |
|---|---|---|
| Input | $2.00 | $1.00 |
| Output | $10.00 | $5.00 |
| Cache write | $2.50 | $1.25 |
| Cache read | $0.10 | $0.05 |
| Prompts over 272K tokens | ||
| Input | $4.00 | $2.00 |
| Output | $15.00 | $7.50 |
| Cache write | $5.00 | $2.50 |
| Cache read | $0.20 | $0.10 |
A request whose prompt, including cached tokens, is over the threshold is billed entirely at the long-prompt rates, as the provider bills it.
Tokenless speaks the OpenAI Chat Completions format. Point your client at the base URL, set model to gpt-6.1-sol, and you are done.
curl https://api.tokenless.store/api/v1/chat/completions \
-H "Authorization: Bearer $TOKENLESS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6.1-sol",
"messages": [{ "role": "user", "content": "Hello" }]
}'