Skip to content

Rate Limits

Rate limits keep the CompactifAI API fast and predictable for every customer. Your account comes with a generous allowance that refills continuously, so short bursts of traffic are never a problem — you get steady headroom without waiting for a fixed window to roll over.

Your limits are per API key and per model. Each model on your key has its own allowance, so heavy traffic to one model never eats into the budget you have for another.

Your allowance refills continuously — every second. Instead of resetting in one lump at the top of each minute or hour, your allowance grows back at a rate of limit ÷ window. A 100,000 tokens-per-minute budget returns roughly 1,670 tokens every second, so headroom is there the moment you need it.

What that means in practice:

  • Bursts are welcome. You can spend up to your full allowance immediately — at cold start you have your whole budget available.
  • Sustained throughput is the limit itself. Keep running and the refill matches your limit exactly, indefinitely.
  • Unused allowance is never lost. Idle time keeps topping your balance back up to the full limit, so a key that has been quiet for a while comes back fully replenished.

You are charged for what you actually use. Before each request we confirm that you have headroom, then charge the real usage once the response is complete — the prompt tokens and completion tokens the model reported, or the audio seconds transcribed. Streaming requests are charged when the stream finishes and the final token counts are known.

Occasionally a request can be admitted just as your balance runs out, since usage is charged after the response completes. That is normal: the next request waits only long enough for the refill to cover the shortfall.

Every request is measured on two axes: how often you call us, and how much work each call does. All of these limits apply at the same time.

Limit Window What it counts
Request rate per minute, hour and day 1 per API request, whatever its size
Input tokens per minute, hour and day Prompt tokens (everything you send)
Output tokens per minute, hour and day Completion tokens (everything the model returns)
Monthly token budget rolling 30 days Prompt and completion tokens over the trial period
Audio seconds per minute, hour and day Seconds of audio transcribed

Token limits count input and output separately, so a long prompt and a long answer are metered independently.

Every account starts with a 30-day trial. Your trial gives you room to explore at full speed — 30 million tokens over the month in each direction, with a 100k tokens-per-minute budget so you can work in bursts:

Limit Trial value
Requests per minute 60
Input tokens per minute 100,000
Output tokens per minute 100,000
Input tokens per 30 days 30,000,000
Output tokens per 30 days 30,000,000

When your trial ends and you commit to a paid plan, these trial rate limits are lifted and your account runs on standard paid limits. Until then your API access continues with reduced limits.

If you have no headroom left, the API responds with HTTP 429 before contacting any model provider. The response carries a standard error envelope and a Retry-After header telling you exactly how long to wait:

{
"error": {
"message": "Rate limit exceeded. Please retry in 12 seconds.",
"type": "rate_limit_error",
"param": null,
"code": "rate_limit_exceeded"
}
}
Header Example Meaning
Retry-After 12 Seconds until your allowance covers the request again. Always at least 1.

Because your allowance refills every second, the wait is typically short — the value is calculated from your own usage, not from a fixed window boundary.

Branch on error.code === "rate_limit_exceeded" in your integration, and see Error Handling for the full error shape. Streaming requests return the same JSON error before any stream events are emitted, so a 429 is always safe to handle and retry as a whole. A rate-limited request never reaches a model and is never billed.

cURL — inspect Retry-After
curl -i https://api.compactif.ai/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"carina-60b","messages":[{"role":"user","content":"Hello"}]}'
# 429 response includes:
# HTTP/1.1 429 Too Many Requests
# Retry-After: 12
# {"error":{"message":"Rate limit exceeded. Please retry in 12 seconds.","type":"rate_limit_error","param":null,"code":"rate_limit_exceeded"}}
  • Respect Retry-After. It is tailored to your usage — waiting longer than needed only slows you down.
  • Add jitter. Randomise retries slightly so parallel workers do not stampede the API at the same instant.
  • Back off on 5xx too. Use exponential backoff for transient 502/503 errors, and keep retrying only idempotent work.
  • Spread large jobs across models. Because limits are per model, you can run a long job over several models without competing for the same allowance.
  • Watch your usage. The Dashboard shows live usage per key so you can scale before you hit a limit.
  • Treat a truncated stream as incomplete. A stream can end early with an event: error if something goes wrong downstream — retry it like any other failed request.
Endpoint Rate-limited Measured on
POST /v1/chat/completions Yes Request rate, input and output tokens
POST /v1/completions Yes Request rate, input and output tokens (legacy endpoint)
POST /v1/responses Yes Request rate, input and output tokens
POST /v1/audio/transcriptions Yes Request rate, audio seconds
GET /v1/models, GET /v1/models/{model} No Not metered

Do rate-limited requests count against my quota or my bill? No. A 429 is returned before any model is contacted, so no tokens or audio seconds are charged and the request is not billed.

Are limits per API key or per organization? Per API key, and per model within that key. Two keys — or two models on one key — never share an allowance.

Can I burst above my per-minute limit? Yes, up to your full allowance. Your per-minute limit is a budget that refills every second; what you cannot do is sustain traffic above it for long.

What happens when my trial ends? Your account keeps working with reduced limits until you commit to a paid plan. Once you do, the trial rate limits are lifted and your account runs on standard paid limits.

How do I get higher limits? Contact us through the contact page or your account manager with the model and the limit you are hitting — we size limits to real workloads, and enterprise plans are provisioned per model.