Changelog
This page lists all notable changes to the CompactifAI API.
2026-10-07
Section titled “2026-10-07”quasar-2-358b– Quasar 2 358B chat model with tool calling, structured output, the Responses API, and the Anthropic Messages API. See Quasar 2 358B and the models catalog.- Video input for
qwen-3-8-27b–capabilities.supports_videois nowtrue: send video as a standardfilecontent part in chat completions. See Video understanding. - Prompt caching for
glm-5-2– Input tokens served from the cache are billed at 20% of the input price ($0.22 per 1M tokens, versus $1.10 per 1M for non-cached input). See Pricing. - Anthropic Messages API support (
supports_messages) —carina-60b,glm-5-2,glm-5-3,quasar-2-358b, andqwen-3-8-27bcan be called viaPOST /v1/messages. See Anthropic Messages API.
Removed
Section titled “Removed”quasar-438b– Removed from the model catalog. Migrate toquasar-2-358b.nemotron-3-nano-omniandmistral-small-3-1– Deprecated models removed from the catalog.
2026-10-05
Section titled “2026-10-05”minimalreasoning effort –reasoning.effortandreasoning_effortnow accept"minimal"alongside the existing levels. Each model still decides which levels it implements.
2026-10-03
Section titled “2026-10-03”- Prompt caching for
glm-5-3– Prompt caching is now enabled for GLM 5.3. Input tokens served from the cache are billed at 20% of the input price ($0.22 per 1M tokens, versus $1.10 per 1M for non-cached input). See Pricing.
2026-09-21
Section titled “2026-09-21”- No more giving up on a model too early – When a model isn’t immediately available, the API now tries an alternative before returning an error. Previously, if availability hadn’t been checked recently, it could skip that alternative and fail the request.
- Stalled requests now report 503 instead of 500 – A request where the model never starts responding now returns the retryable
503first_token_timeoutinstead of a generic500 internal_error. Retry onfirst_token_timeout.
2026-09-16
Section titled “2026-09-16”- No more
404 Model not foundfor a model that exists – A request for an existing model no longer returns404 Model not found. When the model cannot start responding, the client now receives a retryable503witherror.codefirst_token_timeoutorerror.coderequest_deadline_exceededinstead of500or404. Retry on either code.
2026-09-09
Section titled “2026-09-09”Deprecated
Section titled “Deprecated”hypernova-60b– Deprecated. Migrate tocarina-60b, which offers equivalent capabilities. See the models catalog.
2026-08-29
Section titled “2026-08-29”glm-5-3– GLM 5.3 chat model published by zai-org. See GLM 5.3 and the models catalog.
2026-08-25
Section titled “2026-08-25”chat_template_kwargs–POST /v1/chat/completionsnow accepts a top-levelchat_template_kwargsobject that is forwarded as-is to the model’s chat template renderer (vLLM/SGLang convention). The accepted keys depend on the model’s chat template — for example, Qwen models use{"enable_thinking": true}to toggle reasoning. This is model-agnostic: no catalog entry is required per model.
Removed
Section titled “Removed”cai-mistral-small-3-1-slim– Removed from the model catalog.gpt-oss-120b– Removed from the model catalog.
Deprecated
Section titled “Deprecated”reasoning_enabled– The top-levelreasoning_enabledrequest field is deprecated in favor ofchat_template_kwargs. When both are sent, the user-suppliedchat_template_kwargstakes precedence. Thereasoning_enabledfield will be removed in a future release.
2026-08-19
Section titled “2026-08-19”This release contains breaking changes to the error format. If your integration inspects error response bodies, read the migration notes below before upgrading.
Changed
Section titled “Changed”- Standardised error format (breaking) – All errors from
POST /v1/chat/completions,POST /v1/responses,POST /v1/audio/transcriptions,GET /v1/modelsandGET /v1/models/{model}now return a single top-levelerrorobject. Responses that previously used{"detail": ...}— including authentication failures, rate limits, request-validation errors and server errors — now use{"error": {"message", "type", "param", "code"}}. See Error Handling. - Request-validation errors (breaking) – Validation failures no longer return an array of field errors. They now return one
errorobject whoseparamidentifies the first offending field in dotted/indexed notation (for examplemessages[0].role), with a count of any remaining errors appended tomessage. - Streaming errors (breaking) – The
data:line of anevent: errorSSE event is now a JSON error envelope rather than a plain-text message, so it can be parsed with the same JSON decoder as every other event. Clients that string-matched the old text will no longer match. - Rate limit responses (breaking) – The
rate_limit_exceededidentifier has moved from the response body’sdetailfield toerror.code. TheRetry-Afterheader is unchanged. paramandcodeare now populated – Both fields were previously alwaysnull.paramnow names the offending request field andcodecarries a stable machine-readable identifier suitable for programmatic branching. See the error code reference.- Removed the
router_no_llm_providererror type (breaking) – 503 responses now use the standardserver_errortype. The router-specific identifier remains available aserror.code(no_healthy_llm_provider). - Audio transcription status codes (breaking) – A non-boolean
streamform value now returns 400 instead of 422. Model provider failures that are not client errors now return 502 instead of forwarding the provider’s own status code. - File inputs on
/v1/responses– A file resolved fromfile_idis now attached to the request’sinput(as aninput_textpart for text-decodable files, otherwise aninput_filepart) instead of being sent as a separate top-levelfilefield. Text-only models can now read uploaded text files. See OpenAI Compatibility. - Error messages rewritten – Many messages are clearer, and messages for server-side failures are now deliberately generic so that deployment details are not disclosed.
messagehas always been free-form; do not match on it — branch onerror.codeorerror.typeinstead.
Migration
Section titled “Migration”-
Read errors from
error, notdetail. Replaceresponse.json()["detail"]withresponse.json()["error"]["message"]:// Before{"detail": "Unexpected server error"}// After{"error": {"message": "Unexpected server error! Please retry shortly.", "type": "server_error", "param": null, "code": "internal_error"}} -
Branch on
error.code, not on message text or HTTP status alone. The full list is in the error code reference. -
Parse streaming error events as JSON. The
data:payload of anevent: erroris now an object; parse it with the same decoder you use for other events, excluding thedata: [DONE]sentinel. -
Update rate-limit handling to read
error.code == "rate_limit_exceeded"instead of checkingdetail, and continue honouring theRetry-Afterheader. -
If you use
POST /v1/completions, no changes are required — it retains the older{"detail": ...}error format. PreferPOST /v1/chat/completionsfor new integrations.
2026-08-17
Section titled “2026-08-17”qwen-3-8-27b– Qwen 3.8 27B chat model published by Multiverse Computing. See Qwen 3.8 27B and the models catalog.
2026-08-05
Section titled “2026-08-05”quasar-438b– Quasar 438B chat model with tool calling and structured output viaPOST /v1/chat/completions. Currently in beta; capabilities are identical to GLM 5.2. See Quasar 438B and the models catalog.
2026-07-15
Section titled “2026-07-15”glm-5-2– GLM 5.2 chat model with tool calling and structured output viaPOST /v1/chat/completions. Currently available for functional validation only; performance KPIs are not guaranteed during this beta period. See GLM 5.2 and the models catalog.- GitHub Copilot – Integration guide for using CompactifAI models directly inside GitHub Copilot via an OpenAI-compatible custom provider. See GitHub Copilot.
carina-60b– Carina 60B chat model. See Carina 60B and the models catalog.
Removed
Section titled “Removed”llama-4-scout,cai-llama-3-3-70b-slim,llama-3-3-70b,llama-3-1-8b,blackstar-10b,gpt-oss-20b,whisper-large-v3,glm-5-2– These models have been removed. Migrate to an alternative model from the models catalog.
Deprecated
Section titled “Deprecated”2026-05-06
Section titled “2026-05-06”nemotron-3-nano-omni– Nemotron 3 Nano Omni multimodal model (text and image inputs via chat completions). See the models catalog and Multi Modality.- Cursor and OpenCode – Integration guides for using CompactifAI chat completions, custom model IDs, and the OpenAI-compatible base URL: Cursor and OpenCode.
Improvements
Section titled “Improvements”hypernova-60b– Upgraded to version 2605.
2026-04-01
Section titled “2026-04-01”Removed
Section titled “Removed”DeepSeek R1 0528 Slimhas been removed.
2026-03-24
Section titled “2026-03-24”Removed
Section titled “Removed”POST /v1/usage/completions(usage statistics API) – This endpoint is deprecated and removed from the API. Usage and billing visibility is now provided through the CompactifAI Dashboard; the route is no longer served.
Migration
Section titled “Migration”- View usage and billing in the CompactifAI Dashboard instead of calling the API. Use the dashboard to monitor consumption, manage API tokens, and review account settings.
2026-03-10
Section titled “2026-03-10”- Whisper Streaming Support – The Whisper audio transcription endpoint now supports streaming, enabling real-time transcription and lower-latency audio processing.
- GLM-5 (Private Preview) – GLM-5 is now available as a private model in the US region.
- Agentic Tool Calling Enhancements (Beta) – Improved support for agentic workflows and tool-calling capabilities across the following models:
- gpt-oss-20b
- gpt-oss-120b
- hypernova-60b
Improvements
Section titled “Improvements”- Deterministic Tool Usage in Agent Loops – When
tool_choice=requiredortool_choice=<tool_name>is specified, the model consistently returns a tool call, improving reliability in agent-based workflows. - Improved Tool Call Reliability – Reduced tool-call hallucinations and improved adherence to defined tool schemas.
Performance
Section titled “Performance”- Improved inference throughput across several models, delivering ~30% higher throughput and better overall serving efficiency for:
- cai-llama3-1-8b-slim
- hypernova-60b
- gpt-oss-120b
- gpt-oss-20b
- blackstar-10b
Bug Fixes
Section titled “Bug Fixes”- Fixed an issue where streaming with tool calling was not supported for some models. The following models now fully support streaming responses with tool calls:
- gpt-oss-20b
- gpt-oss-120b
- hypernova-60b
2026-02-25
Section titled “2026-02-25”- Added tool calling support for
hypernova-60bmodel
2026-02-16
Section titled “2026-02-16”Bug fixes
Section titled “Bug fixes”- Fixed a bug where Audio Transcriptions endpoint was not working for all the file mime types specified in our API Reference.
- Significantly improved the performance of the Audio Transcriptions endpoint using the
whisper-large-v3model, reducing latency and increasing the speed factor from 15x to 100x on a 10 minutes long audio file (The speed of your network connection might affect the speed factor). - Fixed a bug where the audio transcription endpoint was not working for audio files with a size greater than 1MB. Now, the endpoint can process audio files up to 25MB in size.
2026-01-08
Section titled “2026-01-08”- Added tool calling support for
gpt-oss-20bandgpt-oss-120b
2025-12-23
Section titled “2025-12-23”- Added
hypernova-60b
2025-12-26
Section titled “2025-12-26”- Added
blackstar-10bmodel
2025-10-06
Section titled “2025-10-06”- Speech-to-text transcription endpoint
/v1/audio/transcriptionswith Whisper Large V3 support for multilingual transcription workflows. - Feature and API documentation detailing request parameters, Python examples, and guidance for the new speech-to-text capability.
Models Updates
Section titled “Models Updates”- Removed the
deepseek-r1-0528model from the API.
2025-09-24
Section titled “2025-09-24”- Multi-modality support for chat completions, enabling image-plus-text inputs across the API.
Models Updates
Section titled “Models Updates”- Added
mistral-small-3-1model with full multi-modal understanding and refreshed usage examples.
2025-08-18
Section titled “2025-08-18”- Function tool compatibility has been activated in all models except mistral.
2025-07-01
Section titled “2025-07-01”Models Updates
Section titled “Models Updates”- Added deepseek-ai/DeepSeek-R1-0528 model, accessible via the
deepseek-r1-0528model ID. - Deprecated
deepseek-r1.
2025-06-11
Section titled “2025-06-11”- Initial release of the CompactifAI inference API with the following features:
- Models API endpoint for listing and retrieving available compressed models
- Chat Completions API endpoint for conversational interactions
- Completions API endpoint for text generation
- OpenAI-compatible API design for easy migration and integration
Models Updates
Section titled “Models Updates”- Added the following models:
cai-llama-3-3-70b-slimcai-mistral-small-3-1-slim
Security
Section titled “Security”- HTTPS encryption for all API requests
- Secure authentication using Bearer token scheme