Docs menu
Getting started
API
Account
Reasoning
A thinking channel you can read, and one parameter to tune it.
Both raven models stream a reasoning channel: the model's working, delivered alongside the answer. You can watch the plan form before the answer lands, and you can tune how much effort the model spends with a single parameter.
#reasoning_effort
The reasoning_effort parameter controls the thinking budget:
- omitted (default): the model picks a balanced effort for the task.
- "low": minimal thinking. Faster and cheaper; right for lookup-style questions, transformations, and high-volume pipelines.
- "high": maximum thinking. Right for hard analysis, multi-step planning, and decisions you have to defend.
curl https://api.dipoleml.com/v1/chat/completions \
-H "Authorization: Bearer $RAVEN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "Raven Flash",
"reasoning_effort": "low",
"messages": [{"role": "user", "content": "Reply with exactly: DOCS-OK"}],
"max_tokens": 100
}'#Reading the channel
Reasoning arrives as model output fields, not as chat content:
- Non-streaming: the message carries reasoning (plus reasoning_details) on the text path, or reasoning_content on the vision path.
- Streaming: deltas arrive in the same fields, interleaved with (usually before) the content deltas.
{
"choices": [{
"message": {
"role": "assistant",
"content": "DOCS-OK",
"reasoning": null
},
"finish_reason": "stop"
}],
"usage": {
"prompt_tokens": 609,
"completion_tokens": 6,
"completion_tokens_details": {"reasoning_tokens": 0}
}
}Two rules for parsers: treat the reasoning channel as escaped plain text (it is not markdown and should be rendered verbatim, like a code block), and treat both field names as the same channel. The raven app renders it as a collapsed thinking block you can expand with one click.
#Budgeting
Reasoning tokens count against max_tokens and are billed as output. If the budget is too small, the model spends it all thinking and returns with finish_reason: "length" and a null or short answer. When you constrain output, either raise max_tokens or set reasoning_effort to "low".