GLM 5.2
Chat completions with tool calling and structured output support.
Welcome to our Models Catalog, where you can explore our collection of cutting-edge compressed language models. CompactifAI API delivers a new class of compressed, ultra-efficient AI models engineered for performance, sustainability, and adaptability. These dramatically reduce compute and energy costs while maintaining enterprise-grade accuracy and speed. It adapts seamlessly to diverse environments, ensuring consistent, high-performance AI. With turnkey software and an intuitive API, it offers cost effective inference enabling enterprises to innovate faster, scale cost-effectively, and unlock competitive edge.
Use the following API model identifiers in your requests:
| Model Name | Original Architecture | Model ID | Available |
|---|---|---|---|
| GLM 5.2 | GLM | glm-5-2 | Yes |
| GLM 5.3 | GLM | glm-5-3 | Yes |
| Quasar 2 358B | GLM | quasar-2-358b | Yes |
| Carina 60B | GPT-OSS | carina-60b | Yes |
| Qwen 3.8 27B | - | qwen-3-8-27b | Yes |
| Whisper Large V3 Turbo Slim by CompactifAI | Whisper Large v3 Turbo | cai-whisper-large-v3-turbo-slim | Yes |
GLM 5.2
Chat completions with tool calling and structured output support.
GLM 5.3
Chat completions with tool calling and structured output support.
Quasar 2 358B
Chat completions with tool calling and structured output support.
Carina 60B
A powerful and lightweight model able to handle complex tasks.
Qwen 3.8 27B
Multimodal chat: image and video inputs, with tool calling and structured output.
Whisper Large V3 Turbo Slim by CompactifAI
Compressed turbo Whisper: ~50% smaller, 90%+ baseline WER retained, lower cost.