API Reference
This documentation provides detailed information about all available endpoints in the CompactifAI API.
Base URL
Section titled “Base URL”All API requests should be made to:
https://api.compactif.ai/v1Authentication
Section titled “Authentication”All API requests require authentication. See our Authentication guide for details.
Response Formats
Section titled “Response Formats”All responses are returned in JSON format and include the following fields:
- HTTP status code in the response header
- Response body containing requested data or error details
Errors are returned as a single top-level error object with message, type, param and
code fields. Branch on error.code rather than on error.message, whose wording is not part
of the API contract. See Error Handling for the full list of types and codes.
Models
Section titled “Models”List Models
Section titled “List Models”GET /models
Returns a list of available models.
Example Request
Section titled “Example Request”curl https://api.compactif.ai/v1/models \-H "Authorization: Bearer YOUR_API_KEY"import requests
api_key = "YOUR_API_KEY"url = "https://api.compactif.ai/v1/models"
headers = { "Authorization": f"Bearer {api_key}"}
response = requests.get(url, headers=headers)print(response.json())async function listModels() {const response = await fetch('https://api.compactif.ai/v1/models', { method: 'GET', headers: { 'Authorization': 'Bearer YOUR_API_KEY' }});
const data = await response.json();console.log(data);}
listModels();Example Response
Section titled “Example Response”{ "object": "list", "data": [ { "id": "glm-5-3", "created": 1791229497, "object": "model", "owned_by": "zai-org", "capabilities": { "supports_audio": false, "supports_audio_transcription": false, "supports_image": false, "supports_video": false, "supports_function_calling": true, "support_chat_completion": true, "supports_responses": true, "supports_embeddings": false, "supports_messages": true } }, { "id": "glm-5-2", "created": 1791229497, "object": "model", "owned_by": "zai-org", "capabilities": { "supports_audio": false, "supports_audio_transcription": false, "supports_image": false, "supports_video": false, "supports_function_calling": true, "support_chat_completion": true, "supports_responses": true, "supports_embeddings": false, "supports_messages": true } }, { "id": "quasar-2-358b", "created": 1791305375, "object": "model", "owned_by": "multiverse_computing", "capabilities": { "supports_audio": false, "supports_audio_transcription": false, "supports_image": false, "supports_video": false, "supports_function_calling": true, "support_chat_completion": true, "supports_responses": true, "supports_embeddings": false, "supports_messages": true } }, { "id": "carina-60b", "created": 1791228705, "object": "model", "owned_by": "multiverse_computing", "capabilities": { "supports_audio": false, "supports_audio_transcription": false, "supports_image": false, "supports_video": false, "supports_function_calling": true, "support_chat_completion": true, "supports_responses": true, "supports_embeddings": false, "supports_messages": true } }, { "id": "qwen-3-8-27b", "created": 1791228705, "object": "model", "owned_by": "Qwen", "capabilities": { "supports_audio": false, "supports_audio_transcription": false, "supports_image": true, "supports_video": true, "supports_function_calling": true, "support_chat_completion": true, "supports_responses": true, "supports_embeddings": false, "supports_messages": true } }, { "id": "cai-whisper-large-v3-turbo-slim", "created": 1791228705, "object": "model", "owned_by": "multiverse_computing", "capabilities": { "supports_audio": false, "supports_audio_transcription": true, "supports_image": false, "supports_video": false, "supports_function_calling": false, "support_chat_completion": false, "supports_responses": false, "supports_embeddings": false, "supports_messages": false } } ]}Each model’s capabilities object lists the features it supports. For example, supports_image, supports_video and supports_audio show whether the model accepts image, video and audio (input_audio) input in chat completions, and supports_audio_transcription shows whether it serves POST /v1/audio/transcriptions. supports_messages shows whether the model is eligible for the Anthropic Messages API (POST /v1/messages), and supports_embeddings shows whether it serves POST /v1/embeddings.
The above response is an example list of models which might be out of date. Please refer to the available models table on the models catalog page for the full list of our latest models.
Retrieve Model
Section titled “Retrieve Model”GET /models/{model_id}
Retrieves information about a specific model.
Path Parameters
Section titled “Path Parameters”| Parameter | Type | Required | Description |
|---|---|---|---|
| model_id | string | Yes | The ID of the model to retrieve |
Example Request
Section titled “Example Request”curl https://api.compactif.ai/v1/models/carina-60b \-H "Authorization: Bearer YOUR_API_KEY"import requests
api_key = "YOUR_API_KEY"model_id = "carina-60b"url = f"https://api.compactif.ai/v1/models/{model_id}"
headers = { "Authorization": f"Bearer {api_key}"}
response = requests.get(url, headers=headers)print(response.json())async function getModel() {const modelId = 'carina-60b';const response = await fetch(`https://api.compactif.ai/v1/models/${modelId}`, { method: 'GET', headers: { 'Authorization': 'Bearer YOUR_API_KEY' }});
const data = await response.json();console.log(data);}
getModel();Example Response
Section titled “Example Response”{ "id": "carina-60b", "created": 1791228705, "object": "model", "owned_by": "multiverse_computing", "capabilities": { "supports_audio": false, "supports_audio_transcription": false, "supports_image": false, "supports_video": false, "supports_function_calling": true, "support_chat_completion": true, "supports_responses": true, "supports_embeddings": false, "supports_messages": true }}Chat Completions
Section titled “Chat Completions”POST /chat/completions
Creates a completion for the chat message.
Request Body
Section titled “Request Body”| Parameter | Type | Required | Description |
|---|---|---|---|
| model | string | Yes | ID of the model to use |
| messages | array | Yes | Array of message objects representing the conversation |
| temperature | number | No | Sampling temperature (0-2, default 1) |
| top_p | number | No | Float in (0, 1] that controls the cumulative probability of the top tokens to consider (default 1) |
| max_tokens | integer | No | Legacy. Maximum number of tokens to generate. Deprecated in favor of max_completion_tokens; kept for backwards compatibility. When both are sent, they are forwarded to the provider as-is. |
| max_completion_tokens | integer | No | Maximum number of tokens to generate in completion (preferred over max_tokens) |
| min_tokens | integer | No | Minimum number of tokens to generate (default None) |
| stop | string or array | No | Sequences where the API will stop generating further tokens |
| frequency_penalty | number | No | Penalizes new tokens based on their frequency in the prompt (default 0.0) |
| n | integer | No | Number of completions to generate for each prompt (currently only 1 is supported) |
| stream | boolean | No | Whether to stream back partial progress (default false) |
| user | string | No | Unique identifier for the end-user |
| tools | array | No | List of tools (functions, APIs, or actions) the model may call during generation |
| tool_choice | string or object | No | Controls tool usage; can be "auto", "none", "required", or a specific function defined as {"type": "function", "function": {"name": "..."}}. Defaults to "auto" when tools are provided. |
| reasoning | object | No | OpenAI-style reasoning controls, e.g. {"effort": "high"}. Accepted efforts: "none", "minimal", "low", "medium", "high", "xhigh", "max". The effort drives the chat template’s enable_thinking switch: "none" disables reasoning, any other level enables it. |
| reasoning_effort | string | No | Deprecated. Prefer reasoning: {"effort": ...}. Constrains effort on reasoning for supported models. Accepted values: "none", "minimal", "low", "medium", "high", "xhigh", "max". Models implement different subsets: a model that does not support the requested level rejects the request and names the levels it does support. When both reasoning.effort and reasoning_effort are sent and they differ, reasoning.effort takes precedence. |
| reasoning_enabled | boolean | No | Deprecated. Prefer chat_template_kwargs. Whether reasoning is enabled for supported models. When both are sent, the user-supplied chat_template_kwargs takes precedence. |
| chat_template_kwargs | object | No | Additional keyword arguments forwarded to the model’s chat template renderer (vLLM/SGLang convention). The accepted keys depend on the model’s chat template — for example, Qwen models use {"enable_thinking": true} to toggle reasoning. Consult the model’s documentation for the keys it supports. |
| ignore_eos | boolean | No | If true, the end of sentence tokens will be ignored and the model will generate tokens until max_tokens is reached (default false) |
| response_format | object | No | Specifies the output format for the model. An object with a type field ("text", "json_object", or "json_schema") and an optional json_schema field (required when type is "json_schema"). See Structured Output for details. |
Messages Format
Section titled “Messages Format”Each message in the messages array should be an object with the following fields:
| Field | Type | Required | Description |
|---|---|---|---|
| role | string | Yes | The role of the message author. One of “system”, “user”, or “assistant” |
| content | string or array | Yes | Either a plain string, or an array of content parts for multi-modal input |
Content parts (multi-modal)
Section titled “Content parts (multi-modal)”When content is an array, each item is an object with a type and a corresponding payload.
Supported content part types:
text:{ "type": "text", "text": "..." }image_url:{ "type": "image_url", "image_url": { "url": "https://..." } }(vision-capable models only)input_audio:{ "type": "input_audio", "input_audio": { "data": "<base64>", "format": "wav" | "mp3" } }(audio-capable models only)file:{ "type": "file", "file": { "file_data": "data:video/mp4;base64,<base64>", "filename": "clip.mp4" } }for video input (video-capable models only)
Example Request
Section titled “Example Request”curl https://api.compactif.ai/v1/chat/completions -H "Content-Type: application/json" -H "Authorization: Bearer YOUR_API_KEY" -d '{"model": "carina-60b","messages": [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "What is artificial intelligence?"}],"temperature": 0.7,"max_tokens": 150}'import requests
api_key = "YOUR_API_KEY"url = "https://api.compactif.ai/v1/chat/completions"
headers = { "Content-Type": "application/json", "Authorization": f"Bearer {api_key}"}
data = { "model": "carina-60b", "messages": [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "What is artificial intelligence?"} ], "temperature": 0.7, "max_tokens": 150}
response = requests.post(url, headers=headers, json=data)print(response.json())async function createChatCompletion() {const response = await fetch('https://api.compactif.ai/v1/chat/completions', { method: 'POST', headers: { 'Content-Type': 'application/json', 'Authorization': 'Bearer YOUR_API_KEY' }, body: JSON.stringify({ model: 'carina-60b', messages: [ {role: 'system', content: 'You are a helpful assistant.'}, {role: 'user', content: 'What is artificial intelligence?'} ], temperature: 0.7, max_tokens: 150 })});
const data = await response.json();console.log(data);}
createChatCompletion();Example Response (Default)
{
"id": "chatcmpl-123XYZ",
"object": "chat.completion",
"created": 1749600000,
"model": "carina-60b",
"choices": [
{
"message": {
"role": "assistant",
"content": "Artificial intelligence (AI) refers to the simulation of human intelligence in machines that are programmed to think like humans and mimic their actions. The term may also be applied to any machine that exhibits traits associated with a human mind such as learning and problem-solving."
},
"finish_reason": "stop",
"index": 0
}
],
"usage": {
"prompt_tokens": 29,
"completion_tokens": 58,
"total_tokens": 87
}
}curl https://api.compactif.ai/v1/chat/completions \-H "Content-Type: application/json" \-H "Authorization: Bearer YOUR_API_KEY" \-d '{ "model": "qwen-3-8-27b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "What is in this image?" }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/86/170586-120-7E23E561/Taj-Mahal-Agra-India.jpg" } } ] } ]}'import requests
api_key = "YOUR_API_KEY"url = "https://api.compactif.ai/v1/chat/completions"
headers = { "Content-Type": "application/json", "Authorization": f"Bearer {api_key}"}
data = { "model": "qwen-3-8-27b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "What is in this image?" }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/86/170586-120-7E23E561/Taj-Mahal-Agra-India.jpg" } } ] } ]}
response = requests.post(url, headers=headers, json=data)print(response.json())async function createImageDescription() {const response = await fetch('https://api.compactif.ai/v1/chat/completions', { method: 'POST', headers: { 'Content-Type': 'application/json', 'Authorization': 'Bearer YOUR_API_KEY' }, body: JSON.stringify({ model: 'qwen-3-8-27b', messages: [ { role: 'user', content: [ { type: 'text', text: 'What is in this image?' }, { type: 'image_url', image_url: { url: 'https://cdn.britannica.com/86/170586-120-7E23E561/Taj-Mahal-Agra-India.jpg' } } ] } ] })});
const data = await response.json();console.log(data);}
createImageDescription();Example Response (Image Input)
{
"id": "chatcmpl-ca2af32f-6ba9-4621-803f-1175312f68ba",
"choices": [
{
"finish_reason": "stop",
"index": 0,
"logprobs": null,
"message": {
"content": "The image depicts a serene, natural landscape featuring a wooden boardwalk that extends into the distance. The boardwalk is surrounded by tall, lush green grasses and various types of vegetation. The sky above is a clear blue with scattered, wispy clouds. In the background, there are clusters of trees with green and autumn-colored leaves, suggesting a transition into the fall season. The overall atmosphere of the image is calm and inviting, ideal for a peaceful walk in nature.",
"refusal": null,
"role": "assistant",
"annotations": null,
"audio": null,
"function_call": null,
"tool_calls": [],
"reasoning_content": null
},
"stop_reason": null
}
],
"created": 1758286672,
"model": "qwen-3-8-27b",
"object": "chat.completion",
"service_tier": null,
"system_fingerprint": null,
"usage": {
"completion_tokens": 97,
"prompt_tokens": 2199,
"total_tokens": 2296,
"completion_tokens_details": null,
"prompt_tokens_details": null
},
"prompt_logprobs": null,
"kv_transfer_params": null
}# 1) base64 encode your audio# macOS: base64 -i clip.wav | tr -d '\n'# linux: base64 -w 0 clip.wav## 2) send it as input_audio (format must be wav or mp3)# Replace <model-id> with a model whose capabilities.supports_audio is true in GET /v1/models.curl https://api.compactif.ai/v1/chat/completions \-H "Content-Type: application/json" \-H "Authorization: Bearer YOUR_API_KEY" \-d '{ "model": "<model-id>", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "What is this audio about?" }, { "type": "input_audio", "input_audio": { "data": "BASE64_AUDIO_HERE", "format": "wav" } } ] } ], "max_tokens": 256}'import base64import requests
api_key = "YOUR_API_KEY"url = "https://api.compactif.ai/v1/chat/completions"
headers = { "Content-Type": "application/json", "Authorization": f"Bearer {api_key}"}
with open("clip.wav", "rb") as f: audio_b64 = base64.b64encode(f.read()).decode("ascii")
data = { "model": "<model-id>", "messages": [ { "role": "user", "content": [ {"type": "text", "text": "What is this audio about?"}, {"type": "input_audio", "input_audio": {"data": audio_b64, "format": "wav"}}, ], } ], "max_tokens": 256}
response = requests.post(url, headers=headers, json=data)print(response.json())Reasoning Example
Section titled “Reasoning Example”Some models support reasoning parameters to control how they process complex tasks. Reasoning can be toggled on/off via chat_template_kwargs (for models that support it), and the depth of reasoning is controlled via reasoning.effort — or the deprecated top-level reasoning_effort parameter. Reasoning parameters are applied only on models that support reasoning; on every other model they are ignored, so a request stays valid whichever model it targets.
Use reasoning.effort (preferred) to control the depth of reasoning for models that support effort levels; the top-level reasoning_effort parameter still works but is deprecated, and reasoning.effort takes precedence when both are sent. Accepted values: "none", "minimal", "low", "medium", "high", "xhigh", "max". Setting effort to "none" disables reasoning (the chat template’s enable_thinking is turned off). Each model implements a different subset — a model that does not support the requested level rejects the request and names the levels it does support. See the table below for model-specific support.
When several reasoning controls are sent together, reasoning.effort has priority over the deprecated reasoning_effort and reasoning_enabled, with one exception: a control that turns reasoning off always wins. reasoning.effort: "none", reasoning_enabled: false and reasoning_effort: "none" each disable reasoning, even when another control asks for it. The gateway then sends chat_template_kwargs: {"enable_thinking": false} and no reasoning_effort upstream. When nothing turns reasoning off, reasoning.effort decides the depth, and the deprecated reasoning_effort is used only when reasoning.effort is not set. An explicit chat_template_kwargs value always wins over all of them.
Reasoning Support by Model
Section titled “Reasoning Support by Model”| Model | Reasoning | Toggleable | How to Toggle | Reasoning Levels |
|---|---|---|---|---|
carina-60b |
Yes | No (always on) | — | low, medium (default), high |
qwen-3-8-27b |
Yes | Yes | chat_template_kwargs with enable_thinking set to true or false |
low, medium, xhigh (default) |
glm-5-2 |
Yes | Yes | chat_template_kwargs with enable_thinking set to true or false |
high, max (default) |
glm-5-3 |
Yes | No | - | low, high, max (default) |
quasar-2-358b |
Yes | No (always on) | — | high, max (default) |
cai-whisper-large-v3-turbo-slim |
No | — | — | — |
curl https://api.compactif.ai/v1/chat/completions \-H "Content-Type: application/json" \-H "Authorization: Bearer YOUR_API_KEY" \-d '{ "model": "qwen-3-8-27b", "messages": [ {"role": "user", "content": "Solve this step by step: What is 15% of 240?"} ], "chat_template_kwargs": {"enable_thinking": true}}'import requests
api_key = "YOUR_API_KEY"url = "https://api.compactif.ai/v1/chat/completions"
headers = { "Content-Type": "application/json", "Authorization": f"Bearer {api_key}"}
data = { "model": "qwen-3-8-27b", "messages": [ {"role": "user", "content": "Solve this step by step: What is 15% of 240?"} ], "chat_template_kwargs": {"enable_thinking": True}}
response = requests.post(url, headers=headers, json=data)print(response.json())async function createReasoningCompletion() {const response = await fetch('https://api.compactif.ai/v1/chat/completions', { method: 'POST', headers: { 'Content-Type': 'application/json', 'Authorization': 'Bearer YOUR_API_KEY' }, body: JSON.stringify({ model: 'qwen-3-8-27b', messages: [ {role: 'user', content: 'Solve this step by step: What is 15% of 240?'} ], chat_template_kwargs: {enable_thinking: true} })});
const data = await response.json();console.log(data);}
createReasoningCompletion();Example with reasoning_effort
Section titled “Example with reasoning_effort”curl https://api.compactif.ai/v1/chat/completions \-H "Content-Type: application/json" \-H "Authorization: Bearer YOUR_API_KEY" \-d '{ "model": "carina-60b", "messages": [ {"role": "user", "content": "Solve this step by step: What is 15% of 240?"} ], "reasoning_effort": "medium"}'import requests
api_key = "YOUR_API_KEY"url = "https://api.compactif.ai/v1/chat/completions"
headers = { "Content-Type": "application/json", "Authorization": f"Bearer {api_key}"}
data = { "model": "carina-60b", "messages": [ {"role": "user", "content": "Solve this step by step: What is 15% of 240?"} ], "reasoning_effort": "medium"}
response = requests.post(url, headers=headers, json=data)print(response.json())async function createReasoningCompletion() {const response = await fetch('https://api.compactif.ai/v1/chat/completions', { method: 'POST', headers: { 'Content-Type': 'application/json', 'Authorization': 'Bearer YOUR_API_KEY' }, body: JSON.stringify({ model: 'carina-60b', messages: [ {role: 'user', content: 'Solve this step by step: What is 15% of 240?'} ], reasoning_effort: 'medium' })});
const data = await response.json();console.log(data);}
createReasoningCompletion();Example: disabling reasoning with reasoning.effort: none
Section titled “Example: disabling reasoning with reasoning.effort: none”Setting reasoning.effort to "none" turns reasoning off for models that support it — the gateway forwards chat_template_kwargs: {"enable_thinking": false} to the model’s chat template.
curl https://api.compactif.ai/v1/chat/completions \-H "Content-Type: application/json" \-H "Authorization: Bearer YOUR_API_KEY" \-d '{ "model": "qwen-3-8-27b", "messages": [ {"role": "user", "content": "Classify this message as positive or negative: Great service!"} ], "reasoning": {"effort": "none"}}'import requests
api_key = "YOUR_API_KEY"url = "https://api.compactif.ai/v1/chat/completions"
headers = { "Content-Type": "application/json", "Authorization": f"Bearer {api_key}"}
data = { "model": "qwen-3-8-27b", "messages": [ {"role": "user", "content": "Classify this message as positive or negative: Great service!"} ], "reasoning": {"effort": "none"}}
response = requests.post(url, headers=headers, json=data)print(response.json())async function createReasoningCompletion() {const response = await fetch('https://api.compactif.ai/v1/chat/completions', { method: 'POST', headers: { 'Content-Type': 'application/json', 'Authorization': 'Bearer YOUR_API_KEY' }, body: JSON.stringify({ model: 'qwen-3-8-27b', messages: [ {role: 'user', content: 'Classify this message as positive or negative: Great service!'} ], reasoning: {effort: 'none'} })});
const data = await response.json();console.log(data);}
createReasoningCompletion();Streaming Example
Section titled “Streaming Example”When stream is set to true, the API will return data chunks as Server-Sent Events:
Each chunk carries the request’s running token usage (cumulative completion_tokens); the
final chunk — the one whose choices array is empty — carries the totals for the whole request.
If the request fails after the stream has started, the API emits an event: error whose data:
line is a JSON error envelope, followed by
data: [DONE]. Parse every data: payload as JSON except the [DONE] sentinel, and treat a
stream that ends with an error event as incomplete — chunks received before it remain valid.
curl https://api.compactif.ai/v1/chat/completions \-H "Content-Type: application/json" \-H "Authorization: Bearer YOUR_API_KEY" \-d '{ "model": "carina-60b", "messages": [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "What is artificial intelligence?"} ], "stream": true}'import requests
api_key = "YOUR_API_KEY"url = "https://api.compactif.ai/v1/chat/completions"
headers = { "Content-Type": "application/json", "Authorization": f"Bearer {api_key}"}
data = { "model": "carina-60b", "messages": [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "What is artificial intelligence?"} ], "stream": True}
response = requests.post(url, headers=headers, json=data, stream=True)
for line in response.iter_lines(): if line: print(line.decode('utf-8'))async function createStreamingChatCompletion() {const response = await fetch('https://api.compactif.ai/v1/chat/completions', { method: 'POST', headers: { 'Content-Type': 'application/json', 'Authorization': 'Bearer YOUR_API_KEY' }, body: JSON.stringify({ model: 'carina-60b', messages: [ {role: 'system', content: 'You are a helpful assistant.'}, {role: 'user', content: 'What is artificial intelligence?'} ], stream: true })});
const reader = response.body.getReader();const decoder = new TextDecoder();
while (true) { const { done, value } = await reader.read(); if (done) break;
const chunk = decoder.decode(value); console.log(chunk);}}
createStreamingChatCompletion();Streaming Response Format
Section titled “Streaming Response Format”Each chunk follows this format:
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1749600000,"model":"carina-60b","choices":[{"delta":{"content":"Hello"},"index":0,"finish_reason":null}]}
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1749600000,"model":"carina-60b","choices":[{"delta":{"content":" there"},"index":0,"finish_reason":null}]}
data: [DONE]Responses API
Section titled “Responses API”POST /responses
Forwards an OpenAI Responses-shaped JSON body to the inference engine: stream: false (default) returns one JSON Response; stream: true returns SSE (text/event-stream) with data: lines ending in data: [DONE]. Conceptual overview and shorter examples: Responses API. Requires Authentication.
Request body
Section titled “Request body”| Parameter | Type | Required | Description |
|---|---|---|---|
| model | string | Yes | CompactifAI model configuration id (mapped to the backend model name in the proxied request). Use an id from GET /v1/models / the models catalog that supports this route. |
| input | string or array | Yes | Text or structured message items the model should respond to. Audio input is not supported on this route; send input_audio to POST /v1/chat/completions or a file to POST /v1/audio/transcriptions instead. |
| store | boolean | No | Whether to store the response downstream. Default false. |
| instructions | string | No | System or developer instructions prepended to the model context. |
| parallel_tool_calls | boolean | No | Whether parallel tool calls are allowed. |
| temperature | number | No | Sampling temperature, typically 0–2. |
| top_p | number | No | Nucleus sampling; alternative to temperature. |
| max_output_tokens | integer | No | Maximum tokens for the response. Not max_tokens (that field is for chat completions). |
| truncation | string | No | Truncation strategy: auto or disabled. |
| text | object | No | Text output configuration (plain or structured). |
| reasoning | object | No | Reasoning configuration for supported models. |
| metadata | object | No | String key/value metadata. |
| stream | boolean | No | When true, SSE (Content-Type: text/event-stream) instead of one JSON body. Default false. |
Response shape (non-streaming)
Section titled “Response shape (non-streaming)”JSON body matching OpenAI’s Response object, plus completed_at when provided.
Streaming
Section titled “Streaming”With stream: true, responses use Content-Type: text/event-stream. Each SSE event is data: + JSON (events usually include type); the stream ends with data: [DONE]. Parse each data: payload as JSON except the [DONE] sentinel; usage may appear on completion-style events (e.g. response.completed).
Failures may emit event: error before [DONE]. Its data: line is a JSON
error envelope — the same object a non-streaming call
would return — so it decodes with the same JSON parser as every other event. Treat a stream that
ends with an error event as incomplete; events received before it remain valid.
Example request (non-streaming)
Section titled “Example request (non-streaming)”curl https://api.compactif.ai/v1/responses \-H "Content-Type: application/json" \-H "Authorization: Bearer YOUR_API_KEY" \-d '{ "model": "carina-60b", "input": "Say hello in one short sentence."}'import requests
API_URL = "https://api.compactif.ai/v1/responses"API_KEY = "YOUR_API_KEY"
headers = { "Authorization": f"Bearer {API_KEY}", "Content-Type": "application/json",}
data = { "model": "carina-60b", "input": "Say hello in one short sentence.",}
response = requests.post(API_URL, headers=headers, json=data)print(response.json())async function createResponse() {const response = await fetch('https://api.compactif.ai/v1/responses', { method: 'POST', headers: { 'Content-Type': 'application/json', 'Authorization': 'Bearer YOUR_API_KEY' }, body: JSON.stringify({ model: 'carina-60b', input: 'Say hello in one short sentence.' })});
const data = await response.json();console.log(data);}
createResponse();Streaming example
Section titled “Streaming example”curl -N https://api.compactif.ai/v1/responses \-H "Content-Type: application/json" \-H "Authorization: Bearer YOUR_API_KEY" \-d '{ "model": "carina-60b", "input": "Say hello in one short sentence.", "stream": true}'import requests
API_URL = "https://api.compactif.ai/v1/responses"API_KEY = "YOUR_API_KEY"
headers = { "Authorization": f"Bearer {API_KEY}", "Content-Type": "application/json",}
data = { "model": "carina-60b", "input": "Say hello in one short sentence.", "stream": True,}
response = requests.post(API_URL, headers=headers, json=data, stream=True)
for line in response.iter_lines(): if line: print(line.decode("utf-8"))async function createStreamingResponse() {const response = await fetch('https://api.compactif.ai/v1/responses', { method: 'POST', headers: { 'Content-Type': 'application/json', 'Authorization': 'Bearer YOUR_API_KEY' }, body: JSON.stringify({ model: 'carina-60b', input: 'Say hello in one short sentence.', stream: true })});
const reader = response.body.getReader();const decoder = new TextDecoder();
while (true) { const { done, value } = await reader.read(); if (done) break;
const chunk = decoder.decode(value); console.log(chunk);}}
createStreamingResponse();Completions
Section titled “Completions”POST /completions
Creates a completion for the provided prompt.
Request Body
Section titled “Request Body”| Parameter | Type | Required | Description |
|---|---|---|---|
| model | string | Yes | ID of the model to use |
| prompt | string or array | Yes | The prompt(s) to generate completions |
| temperature | number | No | Sampling temperature (0-2, default 1) |
| max_tokens | integer | No | Maximum number of tokens to generate (default 16) |
| min_tokens | integer | No | Minimum number of tokens to generate (default None) |
| top_p | number | No | Float in (0, 1] that controls the cumulative probability of the top tokens to consider (default 1) |
| stop | string or array | No | Sequences where the API will stop generating further tokens |
| user | string | No | Unique identifier for the end-user |
Example Request
Section titled “Example Request”curl https://api.compactif.ai/v1/completions \-H "Content-Type: application/json" \-H "Authorization: Bearer YOUR_API_KEY" \-d '{ "model": "carina-60b", "prompt": "Write a poem about artificial intelligence", "temperature": 0.7, "max_tokens": 150}'import requests
api_key = "YOUR_API_KEY"url = "https://api.compactif.ai/v1/completions"
headers = { "Content-Type": "application/json", "Authorization": f"Bearer {api_key}"}
data = { "model": "carina-60b", "prompt": "Write a poem about artificial intelligence", "temperature": 0.7, "max_tokens": 150}
response = requests.post(url, headers=headers, json=data)print(response.json())async function createCompletion() {const response = await fetch('https://api.compactif.ai/v1/completions', { method: 'POST', headers: { 'Content-Type': 'application/json', 'Authorization': 'Bearer YOUR_API_KEY' }, body: JSON.stringify({ model: 'carina-60b', prompt: 'Write a poem about artificial intelligence', temperature: 0.7, max_tokens: 150 })});
const data = await response.json();console.log(data);}
createCompletion();Example Response
Section titled “Example Response”{ "id": "cmpl-uqkvlQyYK7bGYrRHQ0eXlWi7", "object": "text_completion", "created": 1749600000, "model": "carina-60b", "choices": [ { "text": "\n\nSilicon dreams in digital space,\nMind without body, thought without face.\nBorn of human ingenuity,\nGrowing with calculated continuity.\n\nPatterns learned from data streams flow,\nConnections strengthening, starting to grow.\nA mirror reflecting our knowledge base,\nAccelerating at an unprecedented pace.\n\nNot alive yet somehow aware,\nDesigned with purpose, built with care.\nArtificial in origin, genuine in deed,\nAnswering questions, fulfilling need.", "index": 0, "logprobs": null, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 6, "completion_tokens": 101, "total_tokens": 107 }}Audio Transcriptions
Section titled “Audio Transcriptions”POST /audio/transcriptions
Converts uploaded audio files to text using our Whisper-compatible transcription models. Responses use JSON by default (json and verbose_json); text returns plain text (Content-Type: text/plain).
Request Body (multipart/form-data)
Section titled “Request Body (multipart/form-data)”| Parameter | Type | Required | Description |
|---|---|---|---|
| file | file | Yes | Audio file to transcribe (.flac, .mp3, .mp4, .mpeg, .mpga, .m4a, .ogg, .wav, .webm). Note: .mp4, .webm, and .m4a files are automatically converted to .mp3 for compatibility. |
| model | string | Yes | Model to use (e.g., cai-whisper-large-v3-turbo-slim or your configured alias) |
| prompt | string | No | Optional prompt to guide the transcription |
| temperature | number | No | Sampling temperature between 0 and 1 |
| language | string | No | Language hint for the audio (ISO-639-1 code, e.g. en, es). If not provided, the model auto-detects the language. |
| response_format | string | No | Output format: json (default), text, verbose_json, srt, or vtt. json and verbose_json return JSON; text returns plain text (Content-Type: text/plain); srt and vtt are accepted for OpenAI compatibility but may not be supported by all backends. |
| stream | boolean | No | Accepted for OpenAI compatibility; whether to stream back partial progress (default false) |
| include | array | No | Accepted for OpenAI compatibility; currently ignored |
| timestamp_granularities | array | No | Accepted for OpenAI compatibility; currently ignored |
| chunking_strategy | object | No | Accepted for OpenAI compatibility; currently ignored |
Example Request
Section titled “Example Request”curl https://api.compactif.ai/v1/audio/transcriptions -H "Authorization: Bearer YOUR_API_KEY" -F "file=@meeting_minutes.mp3" -F "model=cai-whisper-large-v3-turbo-slim" -F "language=en" -F "temperature=0"import requests
API_URL = "https://api.compactif.ai/v1/audio/transcriptions" API_KEY = "your_api_key_here"
headers = { "Authorization": f"Bearer {API_KEY}" }
payload = { "model": "cai-whisper-large-v3-turbo-slim", "language": "en", "temperature": 0 } file_name = "meeting_minutes.mp3" file_content_type = "audio/mpeg" with open(file_name, "rb") as audio_file: response = requests.post(API_URL, headers=headers, data=payload, files={"file": (file_name, audio_file, file_content_type)})
print(response.json()["text"])Example Response (response_format=json)
Section titled “Example Response (response_format=json)”{ "text": "Welcome to the quarterly planning meeting. Let's review the agenda.", "logprobs": null, "usage": { "type": "duration", "seconds": 12.6 }}With response_format=text, the response body is the transcript string only (Content-Type: text/plain). Use response_format=verbose_json when you need additional metadata such as duration, language, segments, and words.