Docs menu
Getting started
API
Account
Chat Completions
The core endpoint: POST /v1/chat/completions, standard shapes, both models.
The core endpoint is POST https://api.dipoleml.com/v1/chat/completions. It accepts the standard chat completions shape and returns a standard chat completion, so most existing integrations work by changing the base URL and the model name.
#Request
curl https://api.dipoleml.com/v1/chat/completions \
-H "Authorization: Bearer $RAVEN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "Raven Max",
"messages": [
{"role": "system", "content": "You are a concise technical editor."},
{"role": "user", "content": "Tighten this paragraph: ..."}
],
"max_tokens": 1000
}'import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.dipoleml.com/v1",
api_key=os.environ["RAVEN_API_KEY"],
)
resp = client.chat.completions.create(
model="Raven Max",
messages=[
{"role": "system", "content": "You are a concise technical editor."},
{"role": "user", "content": "Tighten this paragraph: ..."},
],
max_tokens=1000,
)
print(resp.choices[0].message.content)Body parameters
| Parameter | Type | Notes |
|---|---|---|
| model | string, required | Exactly Raven Flash or Raven Max. Unknown names return 404 with the list of valid ids. |
| messages | array, required | Standard roles: system, user, assistant, tool. A user message may carry text and image parts for vision (see Vision). |
| stream | boolean | Set true for server-sent events. See Streaming. |
| max_tokens | integer | Upper bound on generated tokens, including reasoning. Both models generate up to 128K output tokens. |
| reasoning_effort | string | low or high. Omit it for the default effort. See Reasoning. |
| temperature, top_p, stop, seed | various | Standard sampling parameters pass through to the model. |
#Response
The response echoes your requested model name and reports usage, including cached input tokens:
{
"id": "gen-1791206855-fCJ056VkzS7UvRUDZhss",
"object": "chat.completion",
"created": 1791206855,
"model": "Raven Flash",
"choices": [{
"index": 0,
"message": {"role": "assistant", "content": "DOCS-OK"},
"finish_reason": "stop"
}],
"usage": {
"prompt_tokens": 609,
"completion_tokens": 6,
"total_tokens": 615,
"prompt_tokens_details": {"cached_tokens": 0}
}
}finish_reason is stop when the model finished naturally, length when max_tokens cut it off, and tool_calls when the model wants to call a tool. Tool calling follows the standard format: pass tools in the request and answer with role tool messages.
#Listing models
curl https://api.dipoleml.com/v1/models \
-H "Authorization: Bearer $RAVEN_API_KEY"{
"object": "list",
"data": [
{"id": "Raven Flash", "object": "model", "owned_by": "dipoleml"},
{"id": "Raven Max", "object": "model", "owned_by": "dipoleml"}
]
}Both models accept text and images, support tool use, and run with up to 1M tokens of context. The service is stateless: there is no conversation endpoint, so you send the full message history on every call, exactly as the standard API shape expects.