Skip to content

API Reference

This documentation provides detailed information about all available endpoints in the CompactifAI API.

All API requests should be made to:

https://api.compactif.ai/v1

All API requests require authentication. See our Authentication guide for details.

All responses are returned in JSON format and include the following fields:

  • HTTP status code in the response header
  • Response body containing requested data or error details

Errors are returned as a single top-level error object with message, type, param and code fields. Branch on error.code rather than on error.message, whose wording is not part of the API contract. See Error Handling for the full list of types and codes.

GET /models

Returns a list of available models.

cURL
curl https://api.compactif.ai/v1/models \
-H "Authorization: Bearer YOUR_API_KEY"
{
"object": "list",
"data": [
{
"id": "glm-5-3",
"created": 1791229497,
"object": "model",
"owned_by": "zai-org",
"capabilities": {
"supports_audio": false,
"supports_audio_transcription": false,
"supports_image": false,
"supports_video": false,
"supports_function_calling": true,
"support_chat_completion": true,
"supports_responses": true,
"supports_embeddings": false,
"supports_messages": true
}
},
{
"id": "glm-5-2",
"created": 1791229497,
"object": "model",
"owned_by": "zai-org",
"capabilities": {
"supports_audio": false,
"supports_audio_transcription": false,
"supports_image": false,
"supports_video": false,
"supports_function_calling": true,
"support_chat_completion": true,
"supports_responses": true,
"supports_embeddings": false,
"supports_messages": true
}
},
{
"id": "quasar-2-358b",
"created": 1791305375,
"object": "model",
"owned_by": "multiverse_computing",
"capabilities": {
"supports_audio": false,
"supports_audio_transcription": false,
"supports_image": false,
"supports_video": false,
"supports_function_calling": true,
"support_chat_completion": true,
"supports_responses": true,
"supports_embeddings": false,
"supports_messages": true
}
},
{
"id": "carina-60b",
"created": 1791228705,
"object": "model",
"owned_by": "multiverse_computing",
"capabilities": {
"supports_audio": false,
"supports_audio_transcription": false,
"supports_image": false,
"supports_video": false,
"supports_function_calling": true,
"support_chat_completion": true,
"supports_responses": true,
"supports_embeddings": false,
"supports_messages": true
}
},
{
"id": "qwen-3-8-27b",
"created": 1791228705,
"object": "model",
"owned_by": "Qwen",
"capabilities": {
"supports_audio": false,
"supports_audio_transcription": false,
"supports_image": true,
"supports_video": true,
"supports_function_calling": true,
"support_chat_completion": true,
"supports_responses": true,
"supports_embeddings": false,
"supports_messages": true
}
},
{
"id": "cai-whisper-large-v3-turbo-slim",
"created": 1791228705,
"object": "model",
"owned_by": "multiverse_computing",
"capabilities": {
"supports_audio": false,
"supports_audio_transcription": true,
"supports_image": false,
"supports_video": false,
"supports_function_calling": false,
"support_chat_completion": false,
"supports_responses": false,
"supports_embeddings": false,
"supports_messages": false
}
}
]
}

Each model’s capabilities object lists the features it supports. For example, supports_image, supports_video and supports_audio show whether the model accepts image, video and audio (input_audio) input in chat completions, and supports_audio_transcription shows whether it serves POST /v1/audio/transcriptions. supports_messages shows whether the model is eligible for the Anthropic Messages API (POST /v1/messages), and supports_embeddings shows whether it serves POST /v1/embeddings.

The above response is an example list of models which might be out of date. Please refer to the available models table on the models catalog page for the full list of our latest models.

GET /models/{model_id}

Retrieves information about a specific model.

Parameter Type Required Description
model_id string Yes The ID of the model to retrieve
cURL
curl https://api.compactif.ai/v1/models/carina-60b \
-H "Authorization: Bearer YOUR_API_KEY"
{
"id": "carina-60b",
"created": 1791228705,
"object": "model",
"owned_by": "multiverse_computing",
"capabilities": {
"supports_audio": false,
"supports_audio_transcription": false,
"supports_image": false,
"supports_video": false,
"supports_function_calling": true,
"support_chat_completion": true,
"supports_responses": true,
"supports_embeddings": false,
"supports_messages": true
}
}

POST /chat/completions

Creates a completion for the chat message.

Parameter Type Required Description
model string Yes ID of the model to use
messages array Yes Array of message objects representing the conversation
temperature number No Sampling temperature (0-2, default 1)
top_p number No Float in (0, 1] that controls the cumulative probability of the top tokens to consider (default 1)
max_tokens integer No Legacy. Maximum number of tokens to generate. Deprecated in favor of max_completion_tokens; kept for backwards compatibility. When both are sent, they are forwarded to the provider as-is.
max_completion_tokens integer No Maximum number of tokens to generate in completion (preferred over max_tokens)
min_tokens integer No Minimum number of tokens to generate (default None)
stop string or array No Sequences where the API will stop generating further tokens
frequency_penalty number No Penalizes new tokens based on their frequency in the prompt (default 0.0)
n integer No Number of completions to generate for each prompt (currently only 1 is supported)
stream boolean No Whether to stream back partial progress (default false)
user string No Unique identifier for the end-user
tools array No List of tools (functions, APIs, or actions) the model may call during generation
tool_choice string or object No Controls tool usage; can be "auto", "none", "required", or a specific function defined as {"type": "function", "function": {"name": "..."}}. Defaults to "auto" when tools are provided.
reasoning object No OpenAI-style reasoning controls, e.g. {"effort": "high"}. Accepted efforts: "none", "minimal", "low", "medium", "high", "xhigh", "max". The effort drives the chat template’s enable_thinking switch: "none" disables reasoning, any other level enables it.
reasoning_effort string No Deprecated. Prefer reasoning: {"effort": ...}. Constrains effort on reasoning for supported models. Accepted values: "none", "minimal", "low", "medium", "high", "xhigh", "max". Models implement different subsets: a model that does not support the requested level rejects the request and names the levels it does support. When both reasoning.effort and reasoning_effort are sent and they differ, reasoning.effort takes precedence.
reasoning_enabled boolean No Deprecated. Prefer chat_template_kwargs. Whether reasoning is enabled for supported models. When both are sent, the user-supplied chat_template_kwargs takes precedence.
chat_template_kwargs object No Additional keyword arguments forwarded to the model’s chat template renderer (vLLM/SGLang convention). The accepted keys depend on the model’s chat template — for example, Qwen models use {"enable_thinking": true} to toggle reasoning. Consult the model’s documentation for the keys it supports.
ignore_eos boolean No If true, the end of sentence tokens will be ignored and the model will generate tokens until max_tokens is reached (default false)
response_format object No Specifies the output format for the model. An object with a type field ("text", "json_object", or "json_schema") and an optional json_schema field (required when type is "json_schema"). See Structured Output for details.

Each message in the messages array should be an object with the following fields:

Field Type Required Description
role string Yes The role of the message author. One of “system”, “user”, or “assistant”
content string or array Yes Either a plain string, or an array of content parts for multi-modal input

When content is an array, each item is an object with a type and a corresponding payload.

Supported content part types:

  • text: { "type": "text", "text": "..." }
  • image_url: { "type": "image_url", "image_url": { "url": "https://..." } } (vision-capable models only)
  • input_audio: { "type": "input_audio", "input_audio": { "data": "<base64>", "format": "wav" | "mp3" } } (audio-capable models only)
  • file: { "type": "file", "file": { "file_data": "data:video/mp4;base64,<base64>", "filename": "clip.mp4" } } for video input (video-capable models only)
cURL
curl https://api.compactif.ai/v1/chat/completions -H "Content-Type: application/json" -H "Authorization: Bearer YOUR_API_KEY" -d '{
"model": "carina-60b",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is artificial intelligence?"}
],
"temperature": 0.7,
"max_tokens": 150
}'

Example Response (Default)

{
"id": "chatcmpl-123XYZ",
"object": "chat.completion",
"created": 1749600000,
"model": "carina-60b",
"choices": [
  {
    "message": {
      "role": "assistant",
      "content": "Artificial intelligence (AI) refers to the simulation of human intelligence in machines that are programmed to think like humans and mimic their actions. The term may also be applied to any machine that exhibits traits associated with a human mind such as learning and problem-solving."
    },
    "finish_reason": "stop",
    "index": 0
  }
],
"usage": {
  "prompt_tokens": 29,
  "completion_tokens": 58,
  "total_tokens": 87
}
}

Some models support reasoning parameters to control how they process complex tasks. Reasoning can be toggled on/off via chat_template_kwargs (for models that support it), and the depth of reasoning is controlled via reasoning.effort — or the deprecated top-level reasoning_effort parameter. Reasoning parameters are applied only on models that support reasoning; on every other model they are ignored, so a request stays valid whichever model it targets.

Use reasoning.effort (preferred) to control the depth of reasoning for models that support effort levels; the top-level reasoning_effort parameter still works but is deprecated, and reasoning.effort takes precedence when both are sent. Accepted values: "none", "minimal", "low", "medium", "high", "xhigh", "max". Setting effort to "none" disables reasoning (the chat template’s enable_thinking is turned off). Each model implements a different subset — a model that does not support the requested level rejects the request and names the levels it does support. See the table below for model-specific support.

When several reasoning controls are sent together, reasoning.effort has priority over the deprecated reasoning_effort and reasoning_enabled, with one exception: a control that turns reasoning off always wins. reasoning.effort: "none", reasoning_enabled: false and reasoning_effort: "none" each disable reasoning, even when another control asks for it. The gateway then sends chat_template_kwargs: {"enable_thinking": false} and no reasoning_effort upstream. When nothing turns reasoning off, reasoning.effort decides the depth, and the deprecated reasoning_effort is used only when reasoning.effort is not set. An explicit chat_template_kwargs value always wins over all of them.

Model Reasoning Toggleable How to Toggle Reasoning Levels
carina-60b Yes No (always on) — low, medium (default), high
qwen-3-8-27b Yes Yes chat_template_kwargs with enable_thinking set to true or false low, medium, xhigh (default)
glm-5-2 Yes Yes chat_template_kwargs with enable_thinking set to true or false high, max (default)
glm-5-3 Yes No - low, high, max (default)
quasar-2-358b Yes No (always on) — high, max (default)
cai-whisper-large-v3-turbo-slim No — — —
cURL
curl https://api.compactif.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "qwen-3-8-27b",
"messages": [
{"role": "user", "content": "Solve this step by step: What is 15% of 240?"}
],
"chat_template_kwargs": {"enable_thinking": true}
}'
cURL
curl https://api.compactif.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "carina-60b",
"messages": [
{"role": "user", "content": "Solve this step by step: What is 15% of 240?"}
],
"reasoning_effort": "medium"
}'

Example: disabling reasoning with reasoning.effort: none

Section titled “Example: disabling reasoning with reasoning.effort: none”

Setting reasoning.effort to "none" turns reasoning off for models that support it — the gateway forwards chat_template_kwargs: {"enable_thinking": false} to the model’s chat template.

cURL
curl https://api.compactif.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "qwen-3-8-27b",
"messages": [
{"role": "user", "content": "Classify this message as positive or negative: Great service!"}
],
"reasoning": {"effort": "none"}
}'

When stream is set to true, the API will return data chunks as Server-Sent Events:

Each chunk carries the request’s running token usage (cumulative completion_tokens); the final chunk — the one whose choices array is empty — carries the totals for the whole request.

If the request fails after the stream has started, the API emits an event: error whose data: line is a JSON error envelope, followed by data: [DONE]. Parse every data: payload as JSON except the [DONE] sentinel, and treat a stream that ends with an error event as incomplete — chunks received before it remain valid.

cURL
curl https://api.compactif.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "carina-60b",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is artificial intelligence?"}
],
"stream": true
}'

Each chunk follows this format:

data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1749600000,"model":"carina-60b","choices":[{"delta":{"content":"Hello"},"index":0,"finish_reason":null}]}
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1749600000,"model":"carina-60b","choices":[{"delta":{"content":" there"},"index":0,"finish_reason":null}]}
data: [DONE]

POST /responses

Forwards an OpenAI Responses-shaped JSON body to the inference engine: stream: false (default) returns one JSON Response; stream: true returns SSE (text/event-stream) with data: lines ending in data: [DONE]. Conceptual overview and shorter examples: Responses API. Requires Authentication.

Parameter Type Required Description
model string Yes CompactifAI model configuration id (mapped to the backend model name in the proxied request). Use an id from GET /v1/models / the models catalog that supports this route.
input string or array Yes Text or structured message items the model should respond to. Audio input is not supported on this route; send input_audio to POST /v1/chat/completions or a file to POST /v1/audio/transcriptions instead.
store boolean No Whether to store the response downstream. Default false.
instructions string No System or developer instructions prepended to the model context.
parallel_tool_calls boolean No Whether parallel tool calls are allowed.
temperature number No Sampling temperature, typically 0–2.
top_p number No Nucleus sampling; alternative to temperature.
max_output_tokens integer No Maximum tokens for the response. Not max_tokens (that field is for chat completions).
truncation string No Truncation strategy: auto or disabled.
text object No Text output configuration (plain or structured).
reasoning object No Reasoning configuration for supported models.
metadata object No String key/value metadata.
stream boolean No When true, SSE (Content-Type: text/event-stream) instead of one JSON body. Default false.

JSON body matching OpenAI’s Response object, plus completed_at when provided.

With stream: true, responses use Content-Type: text/event-stream. Each SSE event is data: + JSON (events usually include type); the stream ends with data: [DONE]. Parse each data: payload as JSON except the [DONE] sentinel; usage may appear on completion-style events (e.g. response.completed).

Failures may emit event: error before [DONE]. Its data: line is a JSON error envelope — the same object a non-streaming call would return — so it decodes with the same JSON parser as every other event. Treat a stream that ends with an error event as incomplete; events received before it remain valid.

cURL
curl https://api.compactif.ai/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "carina-60b",
"input": "Say hello in one short sentence."
}'
cURL
curl -N https://api.compactif.ai/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "carina-60b",
"input": "Say hello in one short sentence.",
"stream": true
}'

POST /completions

Creates a completion for the provided prompt.

Parameter Type Required Description
model string Yes ID of the model to use
prompt string or array Yes The prompt(s) to generate completions
temperature number No Sampling temperature (0-2, default 1)
max_tokens integer No Maximum number of tokens to generate (default 16)
min_tokens integer No Minimum number of tokens to generate (default None)
top_p number No Float in (0, 1] that controls the cumulative probability of the top tokens to consider (default 1)
stop string or array No Sequences where the API will stop generating further tokens
user string No Unique identifier for the end-user
cURL
curl https://api.compactif.ai/v1/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "carina-60b",
"prompt": "Write a poem about artificial intelligence",
"temperature": 0.7,
"max_tokens": 150
}'
{
"id": "cmpl-uqkvlQyYK7bGYrRHQ0eXlWi7",
"object": "text_completion",
"created": 1749600000,
"model": "carina-60b",
"choices": [
{
"text": "\n\nSilicon dreams in digital space,\nMind without body, thought without face.\nBorn of human ingenuity,\nGrowing with calculated continuity.\n\nPatterns learned from data streams flow,\nConnections strengthening, starting to grow.\nA mirror reflecting our knowledge base,\nAccelerating at an unprecedented pace.\n\nNot alive yet somehow aware,\nDesigned with purpose, built with care.\nArtificial in origin, genuine in deed,\nAnswering questions, fulfilling need.",
"index": 0,
"logprobs": null,
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 6,
"completion_tokens": 101,
"total_tokens": 107
}
}

POST /audio/transcriptions

Converts uploaded audio files to text using our Whisper-compatible transcription models. Responses use JSON by default (json and verbose_json); text returns plain text (Content-Type: text/plain).

Parameter Type Required Description
file file Yes Audio file to transcribe (.flac, .mp3, .mp4, .mpeg, .mpga, .m4a, .ogg, .wav, .webm). Note: .mp4, .webm, and .m4a files are automatically converted to .mp3 for compatibility.
model string Yes Model to use (e.g., cai-whisper-large-v3-turbo-slim or your configured alias)
prompt string No Optional prompt to guide the transcription
temperature number No Sampling temperature between 0 and 1
language string No Language hint for the audio (ISO-639-1 code, e.g. en, es). If not provided, the model auto-detects the language.
response_format string No Output format: json (default), text, verbose_json, srt, or vtt. json and verbose_json return JSON; text returns plain text (Content-Type: text/plain); srt and vtt are accepted for OpenAI compatibility but may not be supported by all backends.
stream boolean No Accepted for OpenAI compatibility; whether to stream back partial progress (default false)
include array No Accepted for OpenAI compatibility; currently ignored
timestamp_granularities array No Accepted for OpenAI compatibility; currently ignored
chunking_strategy object No Accepted for OpenAI compatibility; currently ignored
cURL
curl https://api.compactif.ai/v1/audio/transcriptions -H "Authorization: Bearer YOUR_API_KEY" -F "file=@meeting_minutes.mp3" -F "model=cai-whisper-large-v3-turbo-slim" -F "language=en" -F "temperature=0"
{
"text": "Welcome to the quarterly planning meeting. Let's review the agenda.",
"logprobs": null,
"usage": {
"type": "duration",
"seconds": 12.6
}
}

With response_format=text, the response body is the transcript string only (Content-Type: text/plain). Use response_format=verbose_json when you need additional metadata such as duration, language, segments, and words.