Multi-Modality
CompactifAI API’s chat completion endpoint supports multi-modality, empowering you to seamlessly process and generate across text, images, and (for select models) audio and video. This enables richer interactions and more versatile applications—all while maintaining exceptional performance at reduced costs.
Image understanding
Section titled “Image understanding”import requests
API_URL = "https://api.compactif.ai/v1/chat/completions"API_KEY = "your_api_key_here"
headers = { "Authorization": f"Bearer {API_KEY}", "Content-Type": "application/json"}
data = { "model": "qwen-3-8-27b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "What is in this image?" }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/86/170586-120-7E23E561/Taj-Mahal-Agra-India.jpg" } } ] } ], "temperature": 0.7}
response = requests.post(API_URL, headers=headers, json=data)print(response.json()["choices"][0]["message"]["content"])Audio understanding (chat)
Section titled “Audio understanding (chat)”import base64import requests
API_URL = "https://api.compactif.ai/v1/chat/completions"API_KEY = "your_api_key_here"
headers = { "Authorization": f"Bearer {API_KEY}", "Content-Type": "application/json",}
with open("clip.wav", "rb") as f: audio_b64 = base64.b64encode(f.read()).decode("ascii")
data = { "model": "<model-id>", "messages": [ { "role": "user", "content": [ {"type": "text", "text": "What is this audio about?"}, {"type": "input_audio", "input_audio": {"data": audio_b64, "format": "wav"}}, ], } ], "max_tokens": 256,}
response = requests.post(API_URL, headers=headers, json=data)print(response.json()["choices"][0]["message"]["content"])<model-id> is a placeholder for a model whose capabilities.supports_audio is true in GET /v1/models. No current model advertises chat audio, so the request returns 400 until one does.
Audio errors
Section titled “Audio errors”Errors use the standard OpenAI error format.
| Status | Cause | Example error.message |
|---|---|---|
400 |
The model does not support audio input (code: audio_not_supported) |
|
400 |
input_audio.data is not base64, e.g. a data: URL (code: invalid_value, param: messages) |
Invalid content part at index 1: input_audio.data: expected base64-encoded audio bytes. |
400 |
The clip is larger than 50 MB decoded or longer than 30 minutes (code: audio_too_large, param: messages) |
Invalid content part at index 1: audio file is too large: the maximum is 50 MB. |
400 |
format does not match the audio bytes, or the container is not wav/mp3 (code: invalid_value, param: messages) |
Invalid content part at index 1: input_audio.data is wav audio but is labelled format 'mp3'. Correct the 'format' field to match the audio. |
400 |
Audio output requested (code: audio_output_not_supported) |
|
400 |
"audio" in modalities without the audio parameter (code: missing_required_parameter) |
Missing required parameter: 'audio'. It is required when 'modalities' includes 'audio'. |
400 |
The audio output parameter without "audio" in modalities (code: audio_output_not_supported, param: audio) |
Invalid parameter: 'audio'. It is only used when 'modalities' includes 'audio', which no model supports. Remove the 'audio' parameter. |
400 |
An assistant message references an earlier audio response (code: audio_output_not_supported) |
Audio responses are not supported, so there is no previous audio response to reference. |
Video understanding
Section titled “Video understanding”Video follows the OpenAI chat completions standard. The OpenAI schema has no dedicated video content part, so a video is sent as a standard file content part carrying the base64-encoded video. Any OpenAI client or SDK can send it unchanged.
import base64import requests
API_URL = "https://api.compactif.ai/v1/chat/completions"API_KEY = "your_api_key_here"
headers = { "Authorization": f"Bearer {API_KEY}", "Content-Type": "application/json",}
with open("clip.mp4", "rb") as f: video_b64 = base64.b64encode(f.read()).decode("ascii")
data = { "model": "qwen-3-8-27b", "messages": [ { "role": "user", "content": [ {"type": "text", "text": "Describe what happens in this video."}, { "type": "file", "file": { "file_data": f"data:video/mp4;base64,{video_b64}", "filename": "clip.mp4", }, }, ], } ],}
response = requests.post(API_URL, headers=headers, json=data)print(response.json()["choices"][0]["message"]["content"])With the official OpenAI Python SDK, point base_url at the CompactifAI API and send the same file part:
import base64from openai import OpenAI
client = OpenAI(base_url="https://api.compactif.ai/v1", api_key="your_api_key_here")
with open("clip.mp4", "rb") as f: video_b64 = base64.b64encode(f.read()).decode("ascii")
response = client.chat.completions.create( model="qwen-3-8-27b", messages=[ { "role": "user", "content": [ {"type": "text", "text": "Describe what happens in this video."}, { "type": "file", "file": { "file_data": f"data:video/mp4;base64,{video_b64}", "filename": "clip.mp4", }, }, ], } ],)print(response.choices[0].message.content)Video can be combined with the other chat features, such as streaming, tool calling and structured output.
Video tokens and usage
Section titled “Video tokens and usage”The model reads a video as a series of sampled frames, and those frames are counted as prompt tokens. They are included in usage.prompt_tokens and billed like any other prompt tokens. The count grows with duration and resolution: a 60-second 720p clip is roughly 30,000 prompt tokens. Shorten or downscale long videos to reduce cost and latency.
Video errors
Section titled “Video errors”Errors use the standard OpenAI error format.
| Status | Cause | Example error.message |
|---|---|---|
400 |
The model does not support video | Model 'carina-60b' does not support video input. This feature is only supported by certain models. |
400 |
The video is larger than 50 MB (code: video_too_large) |
Video file is too large: the maximum is 50 MB. |
400 |
A non-standard content part such as video_url |
Invalid content part at index 1: type 'video_url' is not supported. Supported types: 'text', 'image_url', 'input_audio', 'file'. |
400 |
More videos than the model accepts | At most 1 video(s) may be provided in one prompt. |
Fields
Section titled “Fields”| Field | Type | Description |
|---|---|---|
model |
string | ID of the compressed model to use |
messages |
array | Array of message objects |
temperature |
number | Controls randomness (0-2) |
top_p |
number | Controls diversity via nucleus sampling |
n |
integer | Number of completions to generate |
max_tokens |
integer | Maximum number of tokens to generate |
stream |
boolean | Whether to stream back partial progress |
stop |
string or array | Sequences where the API will stop generating |
Message Format
Section titled “Message Format”Messages must be an array of objects with the following structure:
{ "role": "user"|"assistant", "content": [ { "type": "text", "text": "What is in this image?" }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/86/170586-120-7E23E561/Taj-Mahal-Agra-India.jpg" } } ]}PS: The multimodal accepts also the by default input(the one mentioned in the basic usage).
Response Example
Section titled “Response Example”{ "id": "chatcmpl-ca2af32f-6ba9-4621-803f-1175312f68ba", "choices": [ { "finish_reason": "stop", "index": 0, "logprobs": null, "message": { "content": "The image depicts a serene, natural landscape featuring a wooden boardwalk that extends into the distance. The boardwalk is surrounded by tall, lush green grasses and various types of vegetation. The sky above is a clear blue with scattered, wispy clouds. In the background, there are clusters of trees with green and autumn-colored leaves, suggesting a transition into the fall season. The overall atmosphere of the image is calm and inviting, ideal for a peaceful walk in nature.", "refusal": null, "role": "assistant", "annotations": null, "audio": null, "function_call": null, "tool_calls": [
], "reasoning_content": null }, "stop_reason": null } ], "created": 1758286672, "model": "qwen-3-8-27b", "object": "chat.completion", "service_tier": null, "system_fingerprint": null, "usage": { "completion_tokens": 97, "prompt_tokens": 2199, "total_tokens": 2296, "completion_tokens_details": null, "prompt_tokens_details": null }, "prompt_logprobs": null, "kv_transfer_params": null}Compatibility
Section titled “Compatibility”The table below covers image input. Audio and video support is per model: check capabilities.supports_audio (chat input_audio) and capabilities.supports_video in GET /v1/models.
| Model Name | Model ID | MultiModal Compatible? |
|---|---|---|
| Qwen 3.8 27B | qwen-3-8-27b |
Yes (image and video) |
| GLM 5.2 | glm-5-2 |
No |
| GLM 5.3 | glm-5-3 |
No |
| Quasar 2 358B | quasar-2-358b |
No |
| Carina 60B | carina-60b |
No |