Rate Limits
Rate limits keep the CompactifAI API fast and predictable for every customer. Your account comes with a generous allowance that refills continuously, so short bursts of traffic are never a problem — you get steady headroom without waiting for a fixed window to roll over.
How rate limits work
Section titled “How rate limits work”Your limits are per API key and per model. Each model on your key has its own allowance, so heavy traffic to one model never eats into the budget you have for another.
Your allowance refills continuously — every second. Instead of resetting in one lump
at the top of each minute or hour, your allowance grows back at a rate of
limit ÷ window. A 100,000 tokens-per-minute budget returns roughly 1,670 tokens every
second, so headroom is there the moment you need it.
What that means in practice:
- Bursts are welcome. You can spend up to your full allowance immediately — at cold start you have your whole budget available.
- Sustained throughput is the limit itself. Keep running and the refill matches your limit exactly, indefinitely.
- Unused allowance is never lost. Idle time keeps topping your balance back up to the full limit, so a key that has been quiet for a while comes back fully replenished.
You are charged for what you actually use. Before each request we confirm that you have headroom, then charge the real usage once the response is complete — the prompt tokens and completion tokens the model reported, or the audio seconds transcribed. Streaming requests are charged when the stream finishes and the final token counts are known.
Occasionally a request can be admitted just as your balance runs out, since usage is charged after the response completes. That is normal: the next request waits only long enough for the refill to cover the shortfall.
What counts against your limits
Section titled “What counts against your limits”Every request is measured on two axes: how often you call us, and how much work each call does. All of these limits apply at the same time.
| Limit | Window | What it counts |
|---|---|---|
| Request rate | per minute, hour and day | 1 per API request, whatever its size |
| Input tokens | per minute, hour and day | Prompt tokens (everything you send) |
| Output tokens | per minute, hour and day | Completion tokens (everything the model returns) |
| Monthly token budget | rolling 30 days | Prompt and completion tokens over the trial period |
| Audio seconds | per minute, hour and day | Seconds of audio transcribed |
Token limits count input and output separately, so a long prompt and a long answer are metered independently.
Trial limits
Section titled “Trial limits”Every account starts with a 30-day trial. Your trial gives you room to explore at full speed — 30 million tokens over the month in each direction, with a 100k tokens-per-minute budget so you can work in bursts:
| Limit | Trial value |
|---|---|
| Requests per minute | 60 |
| Input tokens per minute | 100,000 |
| Output tokens per minute | 100,000 |
| Input tokens per 30 days | 30,000,000 |
| Output tokens per 30 days | 30,000,000 |
When your trial ends and you commit to a paid plan, these trial rate limits are lifted and your account runs on standard paid limits. Until then your API access continues with reduced limits.
When you reach a limit
Section titled “When you reach a limit”If you have no headroom left, the API responds with HTTP 429 before contacting any
model provider. The response carries a standard error envelope and a Retry-After header
telling you exactly how long to wait:
{ "error": { "message": "Rate limit exceeded. Please retry in 12 seconds.", "type": "rate_limit_error", "param": null, "code": "rate_limit_exceeded" }}| Header | Example | Meaning |
|---|---|---|
Retry-After |
12 |
Seconds until your allowance covers the request again. Always at least 1. |
Because your allowance refills every second, the wait is typically short — the value is calculated from your own usage, not from a fixed window boundary.
Branch on error.code === "rate_limit_exceeded" in your integration, and see
Error Handling for the full error shape. Streaming
requests return the same JSON error before any stream events are emitted, so a 429 is
always safe to handle and retry as a whole. A rate-limited request never reaches a model
and is never billed.
Handling rate limits in your client
Section titled “Handling rate limits in your client”curl -i https://api.compactif.ai/v1/chat/completions \-H "Authorization: Bearer $API_KEY" \-H "Content-Type: application/json" \-d '{"model":"carina-60b","messages":[{"role":"user","content":"Hello"}]}'
# 429 response includes:# HTTP/1.1 429 Too Many Requests# Retry-After: 12# {"error":{"message":"Rate limit exceeded. Please retry in 12 seconds.","type":"rate_limit_error","param":null,"code":"rate_limit_exceeded"}}import randomimport timeimport requests
def chat_with_backoff(payload, max_retries=5): for attempt in range(max_retries): r = requests.post( "https://api.compactif.ai/v1/chat/completions", headers={"Authorization": f"Bearer {API_KEY}", "Content-Type": "application/json"}, json=payload, ) if r.status_code != 429: r.raise_for_status() return r.json() # Wait exactly as long as the API asked, plus a little jitter so # concurrent workers do not retry in lockstep. retry_after = max(int(r.headers.get("Retry-After", "1")), 1) time.sleep(retry_after + random.uniform(0, 0.5)) raise RuntimeError("rate limit retries exhausted")async function chatWithBackoff(payload, maxRetries = 5) {for (let attempt = 0; attempt < maxRetries; attempt++) { const res = await fetch("https://api.compactif.ai/v1/chat/completions", { method: "POST", headers: { Authorization: `Bearer ${API_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify(payload), }); if (res.status !== 429) { if (!res.ok) throw new Error(await res.text()); return res.json(); } const retryAfter = Math.max(parseInt(res.headers.get("Retry-After") ?? "1", 10), 1); const jitter = Math.random() * 500; await new Promise((r) => setTimeout(r, retryAfter * 1000 + jitter));}throw new Error("rate limit retries exhausted");}Best practices
Section titled “Best practices”- Respect
Retry-After. It is tailored to your usage — waiting longer than needed only slows you down. - Add jitter. Randomise retries slightly so parallel workers do not stampede the API at the same instant.
- Back off on 5xx too. Use exponential backoff for transient
502/503errors, and keep retrying only idempotent work. - Spread large jobs across models. Because limits are per model, you can run a long job over several models without competing for the same allowance.
- Watch your usage. The Dashboard shows live usage per key so you can scale before you hit a limit.
- Treat a truncated stream as incomplete. A stream can end early with an
event: errorif something goes wrong downstream — retry it like any other failed request.
Which endpoints are rate-limited?
Section titled “Which endpoints are rate-limited?”| Endpoint | Rate-limited | Measured on |
|---|---|---|
POST /v1/chat/completions |
Yes | Request rate, input and output tokens |
POST /v1/completions |
Yes | Request rate, input and output tokens (legacy endpoint) |
POST /v1/responses |
Yes | Request rate, input and output tokens |
POST /v1/audio/transcriptions |
Yes | Request rate, audio seconds |
GET /v1/models, GET /v1/models/{model} |
No | Not metered |
Do rate-limited requests count against my quota or my bill? No. A 429 is returned before any model is contacted, so no tokens or audio seconds are charged and the request is not billed.
Are limits per API key or per organization? Per API key, and per model within that key. Two keys — or two models on one key — never share an allowance.
Can I burst above my per-minute limit? Yes, up to your full allowance. Your per-minute limit is a budget that refills every second; what you cannot do is sustain traffic above it for long.
What happens when my trial ends? Your account keeps working with reduced limits until you commit to a paid plan. Once you do, the trial rate limits are lifted and your account runs on standard paid limits.
How do I get higher limits? Contact us through the contact page or your account manager with the model and the limit you are hitting — we size limits to real workloads, and enterprise plans are provisioned per model.