Skip to content

Models Catalog

Welcome to our Models Catalog, where you can explore our collection of cutting-edge compressed language models. CompactifAI API delivers a new class of compressed, ultra-efficient AI models engineered for performance, sustainability, and adaptability. These dramatically reduce compute and energy costs while maintaining enterprise-grade accuracy and speed. It adapts seamlessly to diverse environments, ensuring consistent, high-performance AI. With turnkey software and an intuitive API, it offers cost effective inference enabling enterprises to innovate faster, scale cost-effectively, and unlock competitive edge.

  • Dramatic Cost Reduction: Up to 70% lower inference costs compared to uncompressed models through optimized resource utilization
  • Massive Throughput Gains: Process up to 4x more requests per second with compressed models requiring fewer computational resources
  • Low-Latency Inference: Achieve faster response times due to reduced model size and optimized memory usage
  • Minimal Quality Loss: Advanced compression techniques preserve model performance with typically <5% benchmark difference
  • Superior Concurrency: Support significantly more simultaneous users and requests with the same hardware resources
  • Resource Efficiency: Reduced memory footprint and computational requirements enable better hardware utilization

Use the following API model identifiers in your requests:

Model NameOriginal ArchitectureModel IDAvailable
Mistral Small 3.1 Slim by CompactifAIMistral Small 3.1cai-mistral-small-3-1-slimYes
Mistral Small 3.1-mistral-small-3-1Yes
Nemotron 3 Nano OmniNemotron 3nemotron-3-nano-omniYes
GLM 5.1GLMglm-5-1Yes
GLM 5.1 (Uncensored)GLMcai-glm-5-1Yes
GLM 5.2GLMglm-5-2Yes
Quasar 438B-quasar-438bYes
Openai GPT OSS 120B-gpt-oss-120bYes
Hypernova 60B-hypernova-60bYes
Carina 60B-carina-60bYes
Qwen 3.6 27B (Uncensored)-qwen-3-6-27bYes
Whisper Large V3 Turbo Slim by CompactifAIWhisper Large v3 Turbocai-whisper-large-v3-turbo-slimYes