Chat Completion
CompactifAI API’s chat completion endpoint enables you to create dynamic, multi-turn conversations with our advanced compressed language models, offering exceptional performance at significantly reduced costs.
Basic Usage
Section titled “Basic Usage”import requests
API_URL = "https://api.compactif.ai/v1/chat/completions"API_KEY = "your_api_key_here"
headers = { "Authorization": f"Bearer {API_KEY}", "Content-Type": "application/json"}
data = { "model": "carina-60b", "messages": [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "Tell me about quantum computing."} ], "temperature": 0.7}
response = requests.post(API_URL, headers=headers, json=data)print(response.json()["choices"][0]["message"]["content"])Fields
Section titled “Fields”| Field | Type | Description |
|---|---|---|
model |
string | ID of the compressed model to use |
messages |
array | Array of message objects |
temperature |
number | Controls randomness (0-2) |
top_p |
number | Controls diversity via nucleus sampling |
n |
integer | Number of completions to generate |
max_tokens |
integer | Legacy. Maximum number of tokens to generate. Deprecated in favor of max_completion_tokens; kept for backwards compatibility. When both are sent, they are forwarded to the provider as-is. |
max_completion_tokens |
integer | Maximum number of tokens to generate in completion (preferred over max_tokens) |
min_tokens |
integer | Minimum number of tokens to generate |
stream |
boolean | Whether to stream back partial progress |
stream_options |
object | Streaming options. {"include_usage": true} returns the running usage on every chunk; when false (default) only the final chunk carries it. |
stop |
string or array | Sequences where the API will stop generating |
tools |
array | List of tools the model may call during generation |
tool_choice |
string | Controls tool usage ("auto", "none","required", or specific function) |
reasoning_effort |
string | Reasoning effort for models that support it: none, minimal, low, medium, high, xhigh, or max. Deprecated in favour of reasoning.effort, which takes priority; any control that turns reasoning off (none, reasoning_enabled: false) wins over the others |
chat_template_kwargs |
object | Extra keyword arguments forwarded to the model’s chat template (advanced) |
Message Format
Section titled “Message Format”Messages must be an array of objects with the following structure:
{ "role": "system" | "user" | "assistant", "content": "The message content"}Please refer to our API Reference or the OpenAI API reference for more information on the fields.