Anthropic Messages API
The Anthropic Messages API is Claude’s interface: you send a model, a max_tokens budget, and a messages conversation (plus an optional system prompt, tools, and generation parameters), and receive a Message object—or a stream of Anthropic server-sent events when stream is true. CompactifAI exposes this at POST /v1/messages, served directly by the Anthropic-compatible endpoint.
Use it when your client targets the Anthropic wire format—for example Claude Code pointed at this base URL, or any Anthropic-SDK application. For the OpenAI-style chat surface, continue to use Chat completion (POST /v1/chat/completions).
Endpoint
Section titled “Endpoint”| Path | POST /v1/messages |
| Base URL | https://api.compactif.ai/v1/messages |
Authenticate with a bearer token as described in Authentication. The anthropic-version header is optional here, but sending it—as Anthropic SDKs do—is harmless and recommended for compatibility.
Base URL: point SDKs at the bare host—
https://api.compactif.ai—nothttps://api.compactif.ai/v1. The Anthropic SDK (and Claude Code) append/v1to the base URL, so including it in the base URL results in requests to/v1/v1/messages(404).curlexamples below show the full path since raw HTTP clients don’t append any prefix.
Eligible models
Section titled “Eligible models”Only model configurations with the supports_messages capability in your deployment can call this route. What you see in GET /v1/models and the models catalog is authoritative for your account. A model that declares tool calling or image capabilities can be used with tools and image input, respectively.
If the model is missing or lacks the required capability, the API returns 404 (model_not_found) or 400 with guidance in error.
Request shape
Section titled “Request shape”| Field | Required | Notes |
|---|---|---|
model |
yes | Your CompactifAI model configuration id. |
messages |
yes | Conversation, in order; roles user / assistant. content is a string or a list of blocks (text, image, tool_use, tool_result, …) so multi-turn tool flows round-trip. Every message needs non-empty, non-whitespace content, except a final assistant prefill. |
max_tokens |
yes | Generation budget. Must be ≥ 1 and at most the model’s maximum. |
system |
no | System prompt: a string, or a list of {"type": "text", "text": …} blocks. Takes precedence over any system text inside messages. |
temperature, top_p, top_k |
no | Sampling controls. Deprecated by Anthropic for newer Claude models, but supported here — the backend engine consumes them. top_k: 0 disables top-k; top_p: 0 keeps only the most probable token. |
stop_sequences |
no | Not supported. Do not rely on this field; see Stop sequences. |
stream |
no | true returns SSE (text/event-stream) with Anthropic events. Default false. |
tools |
no | Anthropic-style tools: name, description, input_schema (JSON Schema object). |
tool_choice |
no | Always an object per the Anthropic spec: {"type": "auto"}, {"type": "any"}, {"type": "none"}, or {"type": "tool", "name": …} (all also accept disable_parallel_tool_use). Requires tools unless none. |
metadata |
no | Optional key-value pairs attached to the request. |
Fields you don’t use are ignored; Anthropic parameters not supported(e.g. extended thinking, prompt-cache fields) are silently dropped in the current version rather than rejected.
Stop sequences (not supported)
Section titled “Stop sequences (not supported)”stop_sequences is not supported on this endpoint. The field is accepted but does not behave like Anthropic’s, so do not use it to control where a response ends. Known deviations:
- Matches inside reasoning. The engine checks stop strings against everything the model generates, including its reasoning (
thinking) tokens. If a stop string shows up while the model is thinking, generation ends mid-reasoning and the response can contain athinkingblock and notextblock. Anthropic applies stop sequences to visible output text only. - Wrong
stop_reasonon some models. Depending on the engine version serving the model, a stop-sequence match may be reported asstop_reason: "end_turn"withstop_sequence: null, instead of"stop_sequence"with the matched string.
To end output early, use max_tokens, or trim the response on your side.
Non-streaming example
Section titled “Non-streaming example”from anthropic import Anthropic
# Pass the bare base URL (no trailing /v1): the Anthropic SDK appends /v1 itself.client = Anthropic(base_url="https://api.compactif.ai", api_key="your_api_key_here")
message = client.messages.create( model="quasar-2-358b", max_tokens=1024, system="You are a helpful assistant.", messages=[{"role": "user", "content": "What is the capital of Colombia?"}],)print(message.content[0].text)curl equivalent:
curl -X POST https://api.compactif.ai/v1/messages \ -H "Authorization: Bearer your_api_key_here" \ -H "Content-Type: application/json" \ -H "anthropic-version: 2023-06-01" \ -d '{ "model": "quasar-2-358b", "max_tokens": 1024, "messages": [{"role": "user", "content": "What is the capital of Colombia?"}] }'Response (non-streaming) is the standard Anthropic Message:
{ "id": "msg_01…", "type": "message", "role": "assistant", "model": "quasar-2-358b", "content": [{"type": "text", "text": "The capital of Colombia is Bogotá."}], "stop_reason": "end_turn", "stop_sequence": null, "usage": {"input_tokens": 25, "output_tokens": 14}}Streaming example
Section titled “Streaming example”Set "stream": true. The response is a stream of Anthropic events—message_start, content_block_start, content_block_delta, content_block_stop, message_delta, message_stop—emitted as SSE event: + data: pairs. The final message_delta carries the output-token count; message_start carries the input-token count. The stream ends at message_stop (there is no [DONE] sentinel).
Streaming usage is terminal-only (no continuous usage).
/v1/messagesreports usage just twice:input_tokensonmessage_startand the finaloutput_tokenson the terminalmessage_delta. UnlikePOST /v1/chat/completionswithstream_options.continuous_usage_stats, there are no intermediate, running usage updates during the stream. This is a downstream dependency incompatibility—the model backend currently does not attach running usage to intermediatemessage_deltaevents—not a gateway limitation, and it will be lifted once the backend is upgraded. Clients that need mid-generation token accounting should usePOST /v1/chat/completionsin the meantime.
curl -N -X POST https://api.compactif.ai/v1/messages \ -H "Authorization: Bearer your_api_key_here" \ -H "Content-Type: application/json" \ -d '{ "model": "quasar-2-358b", "max_tokens": 256, "stream": true, "messages": [{"role": "user", "content": "Say hello."}] }'Tool use
Section titled “Tool use”Send tools in the Anthropic tool shape; the model may answer with a tool_use content block and stop_reason: "tool_use". Feed the result back as a tool_result block in the next user message to continue the loop—same pattern as Claude:
{ "model": "quasar-2-358b", "max_tokens": 512, "messages": [ {"role": "user", "content": "What is 2+2?"} ], "tools": [ { "name": "calculator", "description": "Evaluate an arithmetic expression", "input_schema": { "type": "object", "properties": {"expression": {"type": "string"}}, "required": ["expression"] } } ]}Image input
Section titled “Image input”Send an image content block in a user message (model must have image capability):
{ "role": "user", "content": [ {"type": "text", "text": "What do you see in this image?"}, { "type": "image", "source": {"type": "url", "url": "https://example.com/photo.jpg"} } ]}Base64 sources work too: {"type": "base64", "media_type": "image/jpeg", "data": "…"}.
Counting tokens
Section titled “Counting tokens”POST /v1/messages/count_tokens counts the tokens a Message would use — across
messages, system, and tools — without generating anything. The request
body is the subset of the create-message body that affects the count:
model, messages, and optional system, tools, thinking, tool_choice.
Sampling parameters and max_tokens are not needed.
curl https://api.compactif.ai/v1/messages/count_tokens \ -H "Authorization: Bearer $KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "my-anthropic-model", "messages": [{"role": "user", "content": "Hello, world"}], "system": "You are a helpful assistant." }'Response:
{"input_tokens": 29}With the Python SDK, use client.messages.count_tokens(...).
Token counting emits no usage events and debits no tokens from your wallet — only the standard request rate limit applies.
Errors
Section titled “Errors”Errors use the Anthropic error shape, so Anthropic SDKs raise their typed exceptions:
{"type": "error", "error": {"type": "invalid_request_error", "message": "…", "param": "max_tokens"}, "request_id": "…"}error.type follows the status: invalid_request_error (400), authentication_error (401), billing_error (402), permission_error (403), not_found_error (404), rate_limit_error (429), api_error (5xx). code and param are included when known. A failure after a stream has started arrives as an event: error with the same body. Upstream validation failures from the model backend (e.g. bad image data) are relayed with their original status code and message.