Skip to content

Changelog

This page lists all notable changes to the CompactifAI API.

  • quasar-2-358b – Quasar 2 358B chat model with tool calling, structured output, the Responses API, and the Anthropic Messages API. See Quasar 2 358B and the models catalog.
  • Video input for qwen-3-8-27b – capabilities.supports_video is now true: send video as a standard file content part in chat completions. See Video understanding.
  • Prompt caching for glm-5-2 – Input tokens served from the cache are billed at 20% of the input price ($0.22 per 1M tokens, versus $1.10 per 1M for non-cached input). See Pricing.
  • Anthropic Messages API support (supports_messages) — carina-60b, glm-5-2, glm-5-3, quasar-2-358b, and qwen-3-8-27b can be called via POST /v1/messages. See Anthropic Messages API.
  • quasar-438b – Removed from the model catalog. Migrate to quasar-2-358b.
  • nemotron-3-nano-omni and mistral-small-3-1 – Deprecated models removed from the catalog.
  • minimal reasoning effort – reasoning.effort and reasoning_effort now accept "minimal" alongside the existing levels. Each model still decides which levels it implements.
  • Prompt caching for glm-5-3 – Prompt caching is now enabled for GLM 5.3. Input tokens served from the cache are billed at 20% of the input price ($0.22 per 1M tokens, versus $1.10 per 1M for non-cached input). See Pricing.
  • No more giving up on a model too early – When a model isn’t immediately available, the API now tries an alternative before returning an error. Previously, if availability hadn’t been checked recently, it could skip that alternative and fail the request.
  • Stalled requests now report 503 instead of 500 – A request where the model never starts responding now returns the retryable 503 first_token_timeout instead of a generic 500 internal_error. Retry on first_token_timeout.
  • No more 404 Model not found for a model that exists – A request for an existing model no longer returns 404 Model not found. When the model cannot start responding, the client now receives a retryable 503 with error.code first_token_timeout or error.code request_deadline_exceeded instead of 500 or 404. Retry on either code.
  • chat_template_kwargs – POST /v1/chat/completions now accepts a top-level chat_template_kwargs object that is forwarded as-is to the model’s chat template renderer (vLLM/SGLang convention). The accepted keys depend on the model’s chat template — for example, Qwen models use {"enable_thinking": true} to toggle reasoning. This is model-agnostic: no catalog entry is required per model.
  • cai-mistral-small-3-1-slim – Removed from the model catalog.
  • gpt-oss-120b – Removed from the model catalog.
  • reasoning_enabled – The top-level reasoning_enabled request field is deprecated in favor of chat_template_kwargs. When both are sent, the user-supplied chat_template_kwargs takes precedence. The reasoning_enabled field will be removed in a future release.

This release contains breaking changes to the error format. If your integration inspects error response bodies, read the migration notes below before upgrading.

  • Standardised error format (breaking) – All errors from POST /v1/chat/completions, POST /v1/responses, POST /v1/audio/transcriptions, GET /v1/models and GET /v1/models/{model} now return a single top-level error object. Responses that previously used {"detail": ...} — including authentication failures, rate limits, request-validation errors and server errors — now use {"error": {"message", "type", "param", "code"}}. See Error Handling.
  • Request-validation errors (breaking) – Validation failures no longer return an array of field errors. They now return one error object whose param identifies the first offending field in dotted/indexed notation (for example messages[0].role), with a count of any remaining errors appended to message.
  • Streaming errors (breaking) – The data: line of an event: error SSE event is now a JSON error envelope rather than a plain-text message, so it can be parsed with the same JSON decoder as every other event. Clients that string-matched the old text will no longer match.
  • Rate limit responses (breaking) – The rate_limit_exceeded identifier has moved from the response body’s detail field to error.code. The Retry-After header is unchanged.
  • param and code are now populated – Both fields were previously always null. param now names the offending request field and code carries a stable machine-readable identifier suitable for programmatic branching. See the error code reference.
  • Removed the router_no_llm_provider error type (breaking) – 503 responses now use the standard server_error type. The router-specific identifier remains available as error.code (no_healthy_llm_provider).
  • Audio transcription status codes (breaking) – A non-boolean stream form value now returns 400 instead of 422. Model provider failures that are not client errors now return 502 instead of forwarding the provider’s own status code.
  • File inputs on /v1/responses – A file resolved from file_id is now attached to the request’s input (as an input_text part for text-decodable files, otherwise an input_file part) instead of being sent as a separate top-level file field. Text-only models can now read uploaded text files. See OpenAI Compatibility.
  • Error messages rewritten – Many messages are clearer, and messages for server-side failures are now deliberately generic so that deployment details are not disclosed. message has always been free-form; do not match on it — branch on error.code or error.type instead.
  1. Read errors from error, not detail. Replace response.json()["detail"] with response.json()["error"]["message"]:

    // Before
    {"detail": "Unexpected server error"}
    // After
    {"error": {"message": "Unexpected server error! Please retry shortly.", "type": "server_error", "param": null, "code": "internal_error"}}
  2. Branch on error.code, not on message text or HTTP status alone. The full list is in the error code reference.

  3. Parse streaming error events as JSON. The data: payload of an event: error is now an object; parse it with the same decoder you use for other events, excluding the data: [DONE] sentinel.

  4. Update rate-limit handling to read error.code == "rate_limit_exceeded" instead of checking detail, and continue honouring the Retry-After header.

  5. If you use POST /v1/completions, no changes are required — it retains the older {"detail": ...} error format. Prefer POST /v1/chat/completions for new integrations.

  • glm-5-2 – GLM 5.2 chat model with tool calling and structured output via POST /v1/chat/completions. Currently available for functional validation only; performance KPIs are not guaranteed during this beta period. See GLM 5.2 and the models catalog.
  • GitHub Copilot – Integration guide for using CompactifAI models directly inside GitHub Copilot via an OpenAI-compatible custom provider. See GitHub Copilot.
  • llama-4-scout, cai-llama-3-3-70b-slim, llama-3-3-70b, llama-3-1-8b, blackstar-10b, gpt-oss-20b, whisper-large-v3, glm-5-2 – These models have been removed. Migrate to an alternative model from the models catalog.
  • nemotron-3-nano-omni – Nemotron 3 Nano Omni multimodal model (text and image inputs via chat completions). See the models catalog and Multi Modality.
  • Cursor and OpenCode – Integration guides for using CompactifAI chat completions, custom model IDs, and the OpenAI-compatible base URL: Cursor and OpenCode.
  • hypernova-60b – Upgraded to version 2605.
  • DeepSeek R1 0528 Slim has been removed.
  • POST /v1/usage/completions (usage statistics API) – This endpoint is deprecated and removed from the API. Usage and billing visibility is now provided through the CompactifAI Dashboard; the route is no longer served.
  • View usage and billing in the CompactifAI Dashboard instead of calling the API. Use the dashboard to monitor consumption, manage API tokens, and review account settings.
  • Whisper Streaming Support – The Whisper audio transcription endpoint now supports streaming, enabling real-time transcription and lower-latency audio processing.
  • GLM-5 (Private Preview) – GLM-5 is now available as a private model in the US region.
  • Agentic Tool Calling Enhancements (Beta) – Improved support for agentic workflows and tool-calling capabilities across the following models:
    • gpt-oss-20b
    • gpt-oss-120b
    • hypernova-60b
  • Deterministic Tool Usage in Agent Loops – When tool_choice=required or tool_choice=<tool_name> is specified, the model consistently returns a tool call, improving reliability in agent-based workflows.
  • Improved Tool Call Reliability – Reduced tool-call hallucinations and improved adherence to defined tool schemas.
  • Improved inference throughput across several models, delivering ~30% higher throughput and better overall serving efficiency for:
    • cai-llama3-1-8b-slim
    • hypernova-60b
    • gpt-oss-120b
    • gpt-oss-20b
    • blackstar-10b
  • Fixed an issue where streaming with tool calling was not supported for some models. The following models now fully support streaming responses with tool calls:
    • gpt-oss-20b
    • gpt-oss-120b
    • hypernova-60b
  • Added tool calling support for hypernova-60b model
  • Fixed a bug where Audio Transcriptions endpoint was not working for all the file mime types specified in our API Reference.
  • Significantly improved the performance of the Audio Transcriptions endpoint using the whisper-large-v3 model, reducing latency and increasing the speed factor from 15x to 100x on a 10 minutes long audio file (The speed of your network connection might affect the speed factor).
  • Fixed a bug where the audio transcription endpoint was not working for audio files with a size greater than 1MB. Now, the endpoint can process audio files up to 25MB in size.
  • Added tool calling support for gpt-oss-20b and gpt-oss-120b
  • Added hypernova-60b
  • Added blackstar-10b model
  • Speech-to-text transcription endpoint /v1/audio/transcriptions with Whisper Large V3 support for multilingual transcription workflows.
  • Feature and API documentation detailing request parameters, Python examples, and guidance for the new speech-to-text capability.
  • Removed the deepseek-r1-0528 model from the API.
  • Multi-modality support for chat completions, enabling image-plus-text inputs across the API.
  • Added mistral-small-3-1 model with full multi-modal understanding and refreshed usage examples.
  • Function tool compatibility has been activated in all models except mistral.
  • Added deepseek-ai/DeepSeek-R1-0528 model, accessible via the deepseek-r1-0528 model ID.
  • Deprecated deepseek-r1.
  • Initial release of the CompactifAI inference API with the following features:
    • Models API endpoint for listing and retrieving available compressed models
    • Chat Completions API endpoint for conversational interactions
    • Completions API endpoint for text generation
  • OpenAI-compatible API design for easy migration and integration
  • Added the following models:
    • cai-llama-3-3-70b-slim
    • cai-mistral-small-3-1-slim
  • HTTPS encryption for all API requests
  • Secure authentication using Bearer token scheme