Docs menu

Docs

Vision

Send images alongside text. Routing is automatic.

Both raven models read images. There is no separate vision endpoint: include an image part in a message and the system routes the request for visual understanding automatically. The rest of the request is unchanged.

#Sending an image

An image travels as a content part of type image_url next to your text. The URL can be a base64 data URL (no upload step needed) or a publicly reachable https URL.

shell
B64=$(base64 -w0 image.png)

curl https://api.dipoleml.com/v1/chat/completions \
  -H "Authorization: Bearer $RAVEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d "{
    \"model\": \"Raven Flash\",
    \"messages\": [{
      \"role\": \"user\",
      \"content\": [
        {\"type\": \"text\", \"text\": \"In one sentence, what is in this image?\"},
        {\"type\": \"image_url\", \"image_url\": {\"url\": \"data:image/png;base64,$B64\"}}
      ]
    }]
  }"
python
import base64, os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.dipoleml.com/v1",
    api_key=os.environ["RAVEN_API_KEY"],
)

with open("image.png", "rb") as f:
    b64 = base64.b64encode(f.read()).decode()

resp = client.chat.completions.create(
    model="Raven Flash",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "In one sentence, what is in this image?"},
            {"type": "image_url",
             "image_url": {"url": f"data:image/png;base64,{b64}"}},
        ],
    }],
)
print(resp.choices[0].message.content)

#How routing works

Image handling is automatic and per request. If any message in the request carries an image part, the whole request is served by the vision path; otherwise it takes the text path. You never pick a vision model or change the model name. Mixed requests, where one message has text and another has an image, work the same way.

#Practical notes

  • Supported formats include PNG and JPEG; send the correct MIME type in the data URL.
  • Images count as input tokens and bill at the input rate. Large images use more tokens; downscale before sending when detail is not critical.
  • The vision path has its own quality tuning, so screenshots, diagrams, and photos of documents all read correctly. Ask specific questions and you get specific answers.