Docs menu

Docs

API Credits

Pay-as-you-go, per token, with a discount when input hits the cache.

The DipoleML API is pay-as-you-go. Prepaid credits are debited per request from actual token usage. Credits are independent of every subscription: you do not need a plan to call the API, and a plan never consumes credits.

#Prices

raven Flashraven Max
Input$0.50 / 1M tokens$1.25 / 1M tokens
Cached input$0.05 / 1M tokens$0.13 / 1M tokens
Output$2.50 / 1M tokens$6.25 / 1M tokens

All prices in USD. Output includes reasoning tokens. Cached input is applied automatically when the model reports a cache hit on your prompt prefix; no flag is needed, and repeated prefixes (system prompts, long documents) hit it often.

#The debit formula

Per request: (input − cached) × input rate + cached × cached rate + output × output rate, all per 1M tokens. Worked example on raven Flash with 2,000 fresh input tokens, 8,000 cached input tokens, and 1,000 output tokens:

worked example
fresh input   2,000  × $0.50 / 1M  = $0.0010
cached input  8,000  × $0.05 / 1M  = $0.0004
output        1,000  × $2.50 / 1M  = $0.0025
                                ----------
total                            $0.0039

The exact math, including the per-request prices applied, is written into your credits history in the dashboard, so every debit is auditable after the fact.

#Topping up

  • Credit packs: $5, $10, $20, $50, available in your dashboard.
  • Credits do not expire and are shared by every API key on your account.
  • Your balance and per-request cost history are shown in the dashboard credits panel.

#What happens at zero

The balance is checked before each request. If it cannot cover the request, the call is rejected with 402 and the top-up link, so you find out before spending effort on a dead call:

402 response
{
  "error": {
    "message": "Insufficient API credits: your balance is $0.00. Top up at https://chat.dipoleml.com/dashboard to keep making requests.",
    "type": "insufficient_credits",
    "code": "insufficient_credits"
  }
}

Debits happen after the response completes, from the usage the model actually reported. A request that fails is not debited. The per-minute rate limit still applies to API keys as an abuse guard.