Skip to content

Multi-Modality

CompactifAI API’s chat completion endpoint supports multi-modality, empowering you to seamlessly process and generate across text, images, and (for select models) audio and video. This enables richer interactions and more versatile applications—all while maintaining exceptional performance at reduced costs.

import requests
API_URL = "https://api.compactif.ai/v1/chat/completions"
API_KEY = "your_api_key_here"
headers = {
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json"
}
data = {
"model": "qwen-3-8-27b",
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "What is in this image?"
},
{
"type": "image_url",
"image_url": {
"url": "https://cdn.britannica.com/86/170586-120-7E23E561/Taj-Mahal-Agra-India.jpg"
}
}
]
}
],
"temperature": 0.7
}
response = requests.post(API_URL, headers=headers, json=data)
print(response.json()["choices"][0]["message"]["content"])
import base64
import requests
API_URL = "https://api.compactif.ai/v1/chat/completions"
API_KEY = "your_api_key_here"
headers = {
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json",
}
with open("clip.wav", "rb") as f:
audio_b64 = base64.b64encode(f.read()).decode("ascii")
data = {
"model": "<model-id>",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "What is this audio about?"},
{"type": "input_audio", "input_audio": {"data": audio_b64, "format": "wav"}},
],
}
],
"max_tokens": 256,
}
response = requests.post(API_URL, headers=headers, json=data)
print(response.json()["choices"][0]["message"]["content"])

<model-id> is a placeholder for a model whose capabilities.supports_audio is true in GET /v1/models. No current model advertises chat audio, so the request returns 400 until one does.

Errors use the standard OpenAI error format.

Status Cause Example error.message
400 The model does not support audio input (code: audio_not_supported)
400 input_audio.data is not base64, e.g. a data: URL (code: invalid_value, param: messages) Invalid content part at index 1: input_audio.data: expected base64-encoded audio bytes.
400 The clip is larger than 50 MB decoded or longer than 30 minutes (code: audio_too_large, param: messages) Invalid content part at index 1: audio file is too large: the maximum is 50 MB.
400 format does not match the audio bytes, or the container is not wav/mp3 (code: invalid_value, param: messages) Invalid content part at index 1: input_audio.data is wav audio but is labelled format 'mp3'. Correct the 'format' field to match the audio.
400 Audio output requested (code: audio_output_not_supported)
400 "audio" in modalities without the audio parameter (code: missing_required_parameter) Missing required parameter: 'audio'. It is required when 'modalities' includes 'audio'.
400 The audio output parameter without "audio" in modalities (code: audio_output_not_supported, param: audio) Invalid parameter: 'audio'. It is only used when 'modalities' includes 'audio', which no model supports. Remove the 'audio' parameter.
400 An assistant message references an earlier audio response (code: audio_output_not_supported) Audio responses are not supported, so there is no previous audio response to reference.

Video follows the OpenAI chat completions standard. The OpenAI schema has no dedicated video content part, so a video is sent as a standard file content part carrying the base64-encoded video. Any OpenAI client or SDK can send it unchanged.

import base64
import requests
API_URL = "https://api.compactif.ai/v1/chat/completions"
API_KEY = "your_api_key_here"
headers = {
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json",
}
with open("clip.mp4", "rb") as f:
video_b64 = base64.b64encode(f.read()).decode("ascii")
data = {
"model": "qwen-3-8-27b",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "Describe what happens in this video."},
{
"type": "file",
"file": {
"file_data": f"data:video/mp4;base64,{video_b64}",
"filename": "clip.mp4",
},
},
],
}
],
}
response = requests.post(API_URL, headers=headers, json=data)
print(response.json()["choices"][0]["message"]["content"])

With the official OpenAI Python SDK, point base_url at the CompactifAI API and send the same file part:

import base64
from openai import OpenAI
client = OpenAI(base_url="https://api.compactif.ai/v1", api_key="your_api_key_here")
with open("clip.mp4", "rb") as f:
video_b64 = base64.b64encode(f.read()).decode("ascii")
response = client.chat.completions.create(
model="qwen-3-8-27b",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "Describe what happens in this video."},
{
"type": "file",
"file": {
"file_data": f"data:video/mp4;base64,{video_b64}",
"filename": "clip.mp4",
},
},
],
}
],
)
print(response.choices[0].message.content)

Video can be combined with the other chat features, such as streaming, tool calling and structured output.

The model reads a video as a series of sampled frames, and those frames are counted as prompt tokens. They are included in usage.prompt_tokens and billed like any other prompt tokens. The count grows with duration and resolution: a 60-second 720p clip is roughly 30,000 prompt tokens. Shorten or downscale long videos to reduce cost and latency.

Errors use the standard OpenAI error format.

Status Cause Example error.message
400 The model does not support video Model 'carina-60b' does not support video input. This feature is only supported by certain models.
400 The video is larger than 50 MB (code: video_too_large) Video file is too large: the maximum is 50 MB.
400 A non-standard content part such as video_url Invalid content part at index 1: type 'video_url' is not supported. Supported types: 'text', 'image_url', 'input_audio', 'file'.
400 More videos than the model accepts At most 1 video(s) may be provided in one prompt.
Field Type Description
model string ID of the compressed model to use
messages array Array of message objects
temperature number Controls randomness (0-2)
top_p number Controls diversity via nucleus sampling
n integer Number of completions to generate
max_tokens integer Maximum number of tokens to generate
stream boolean Whether to stream back partial progress
stop string or array Sequences where the API will stop generating

Messages must be an array of objects with the following structure:

{
"role": "user"|"assistant",
"content": [
{
"type": "text",
"text": "What is in this image?"
},
{
"type": "image_url",
"image_url": {
"url": "https://cdn.britannica.com/86/170586-120-7E23E561/Taj-Mahal-Agra-India.jpg"
}
}
]
}

PS: The multimodal accepts also the by default input(the one mentioned in the basic usage).

{
"id": "chatcmpl-ca2af32f-6ba9-4621-803f-1175312f68ba",
"choices": [
{
"finish_reason": "stop",
"index": 0,
"logprobs": null,
"message": {
"content": "The image depicts a serene, natural landscape featuring a wooden boardwalk that extends into the distance. The boardwalk is surrounded by tall, lush green grasses and various types of vegetation. The sky above is a clear blue with scattered, wispy clouds. In the background, there are clusters of trees with green and autumn-colored leaves, suggesting a transition into the fall season. The overall atmosphere of the image is calm and inviting, ideal for a peaceful walk in nature.",
"refusal": null,
"role": "assistant",
"annotations": null,
"audio": null,
"function_call": null,
"tool_calls": [
],
"reasoning_content": null
},
"stop_reason": null
}
],
"created": 1758286672,
"model": "qwen-3-8-27b",
"object": "chat.completion",
"service_tier": null,
"system_fingerprint": null,
"usage": {
"completion_tokens": 97,
"prompt_tokens": 2199,
"total_tokens": 2296,
"completion_tokens_details": null,
"prompt_tokens_details": null
},
"prompt_logprobs": null,
"kv_transfer_params": null
}

The table below covers image input. Audio and video support is per model: check capabilities.supports_audio (chat input_audio) and capabilities.supports_video in GET /v1/models.

Model Name Model ID MultiModal Compatible?
Qwen 3.8 27B qwen-3-8-27b Yes (image and video)
GLM 5.2 glm-5-2 No
GLM 5.3 glm-5-3 No
Quasar 2 358B quasar-2-358b No
Carina 60B carina-60b No