Docs menu

Docs

Reasoning

A thinking channel you can read, and one parameter to tune it.

Both raven models stream a reasoning channel: the model's working, delivered alongside the answer. You can watch the plan form before the answer lands, and you can tune how much effort the model spends with a single parameter.

#reasoning_effort

The reasoning_effort parameter controls the thinking budget:

  • omitted (default): the model picks a balanced effort for the task.
  • "low": minimal thinking. Faster and cheaper; right for lookup-style questions, transformations, and high-volume pipelines.
  • "high": maximum thinking. Right for hard analysis, multi-step planning, and decisions you have to defend.
curl
curl https://api.dipoleml.com/v1/chat/completions \
  -H "Authorization: Bearer $RAVEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Raven Flash",
    "reasoning_effort": "low",
    "messages": [{"role": "user", "content": "Reply with exactly: DOCS-OK"}],
    "max_tokens": 100
  }'

#Reading the channel

Reasoning arrives as model output fields, not as chat content:

  • Non-streaming: the message carries reasoning (plus reasoning_details) on the text path, or reasoning_content on the vision path.
  • Streaming: deltas arrive in the same fields, interleaved with (usually before) the content deltas.
response (trimmed)
{
  "choices": [{
    "message": {
      "role": "assistant",
      "content": "DOCS-OK",
      "reasoning": null
    },
    "finish_reason": "stop"
  }],
  "usage": {
    "prompt_tokens": 609,
    "completion_tokens": 6,
    "completion_tokens_details": {"reasoning_tokens": 0}
  }
}

Two rules for parsers: treat the reasoning channel as escaped plain text (it is not markdown and should be rendered verbatim, like a code block), and treat both field names as the same channel. The raven app renders it as a collapsed thinking block you can expand with one click.

#Budgeting

Reasoning tokens count against max_tokens and are billed as output. If the budget is too small, the model spends it all thinking and returns with finish_reason: "length" and a null or short answer. When you constrain output, either raise max_tokens or set reasoning_effort to "low".