Docs menu

Docs

Chat Completions

The core endpoint: POST /v1/chat/completions, standard shapes, both models.

The core endpoint is POST https://api.dipoleml.com/v1/chat/completions. It accepts the standard chat completions shape and returns a standard chat completion, so most existing integrations work by changing the base URL and the model name.

#Request

curl
curl https://api.dipoleml.com/v1/chat/completions \
  -H "Authorization: Bearer $RAVEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Raven Max",
    "messages": [
      {"role": "system", "content": "You are a concise technical editor."},
      {"role": "user", "content": "Tighten this paragraph: ..."}
    ],
    "max_tokens": 1000
  }'
python
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.dipoleml.com/v1",
    api_key=os.environ["RAVEN_API_KEY"],
)

resp = client.chat.completions.create(
    model="Raven Max",
    messages=[
        {"role": "system", "content": "You are a concise technical editor."},
        {"role": "user", "content": "Tighten this paragraph: ..."},
    ],
    max_tokens=1000,
)
print(resp.choices[0].message.content)

Body parameters

ParameterTypeNotes
modelstring, requiredExactly Raven Flash or Raven Max. Unknown names return 404 with the list of valid ids.
messagesarray, requiredStandard roles: system, user, assistant, tool. A user message may carry text and image parts for vision (see Vision).
streambooleanSet true for server-sent events. See Streaming.
max_tokensintegerUpper bound on generated tokens, including reasoning. Both models generate up to 128K output tokens.
reasoning_effortstringlow or high. Omit it for the default effort. See Reasoning.
temperature, top_p, stop, seedvariousStandard sampling parameters pass through to the model.

#Response

The response echoes your requested model name and reports usage, including cached input tokens:

response (trimmed)
{
  "id": "gen-1791206855-fCJ056VkzS7UvRUDZhss",
  "object": "chat.completion",
  "created": 1791206855,
  "model": "Raven Flash",
  "choices": [{
    "index": 0,
    "message": {"role": "assistant", "content": "DOCS-OK"},
    "finish_reason": "stop"
  }],
  "usage": {
    "prompt_tokens": 609,
    "completion_tokens": 6,
    "total_tokens": 615,
    "prompt_tokens_details": {"cached_tokens": 0}
  }
}

finish_reason is stop when the model finished naturally, length when max_tokens cut it off, and tool_calls when the model wants to call a tool. Tool calling follows the standard format: pass tools in the request and answer with role tool messages.

#Listing models

shell
curl https://api.dipoleml.com/v1/models \
  -H "Authorization: Bearer $RAVEN_API_KEY"
response (trimmed)
{
  "object": "list",
  "data": [
    {"id": "Raven Flash", "object": "model", "owned_by": "dipoleml"},
    {"id": "Raven Max", "object": "model", "owned_by": "dipoleml"}
  ]
}

Both models accept text and images, support tool use, and run with up to 1M tokens of context. The service is stateless: there is no conversation endpoint, so you send the full message history on every call, exactly as the standard API shape expects.