Skip to content

Models Catalog

Welcome to our Models Catalog, where you can explore our collection of cutting-edge compressed language models. CompactifAI API delivers a new class of compressed, ultra-efficient AI models engineered for performance, sustainability, and adaptability. These dramatically reduce compute and energy costs while maintaining enterprise-grade accuracy and speed. It adapts seamlessly to diverse environments, ensuring consistent, high-performance AI. With turnkey software and an intuitive API, it offers cost effective inference enabling enterprises to innovate faster, scale cost-effectively, and unlock competitive edge.

  • Dramatic Cost Reduction: Up to 70% lower inference costs compared to uncompressed models through optimized resource utilization
  • Massive Throughput Gains: Process up to 4x more requests per second with compressed models requiring fewer computational resources
  • Low-Latency Inference: Achieve faster response times due to reduced model size and optimized memory usage
  • Minimal Quality Loss: Advanced compression techniques preserve model performance with typically <5% benchmark difference
  • Superior Concurrency: Support significantly more simultaneous users and requests with the same hardware resources
  • Resource Efficiency: Reduced memory footprint and computational requirements enable better hardware utilization

Use the following API model identifiers in your requests:

Model Name Original Architecture Model ID Available
GLM 5.2 GLM glm-5-2 Yes
GLM 5.3 GLM glm-5-3 Yes
Quasar 2 358B GLM quasar-2-358b Yes
Carina 60B GPT-OSS carina-60b Yes
Qwen 3.8 27B - qwen-3-8-27b Yes
Whisper Large V3 Turbo Slim by CompactifAI Whisper Large v3 Turbo cai-whisper-large-v3-turbo-slim Yes