Skip to content

Anthropic Messages API

The Anthropic Messages API is Claude’s interface: you send a model, a max_tokens budget, and a messages conversation (plus an optional system prompt, tools, and generation parameters), and receive a Message object—or a stream of Anthropic server-sent events when stream is true. CompactifAI exposes this at POST /v1/messages, served directly by the Anthropic-compatible endpoint.

Use it when your client targets the Anthropic wire format—for example Claude Code pointed at this base URL, or any Anthropic-SDK application. For the OpenAI-style chat surface, continue to use Chat completion (POST /v1/chat/completions).

Path POST /v1/messages
Base URL https://api.compactif.ai/v1/messages

Authenticate with a bearer token as described in Authentication. The anthropic-version header is optional here, but sending it—as Anthropic SDKs do—is harmless and recommended for compatibility.

Base URL: point SDKs at the bare host—https://api.compactif.ai—not https://api.compactif.ai/v1. The Anthropic SDK (and Claude Code) append /v1 to the base URL, so including it in the base URL results in requests to /v1/v1/messages (404). curl examples below show the full path since raw HTTP clients don’t append any prefix.

Only model configurations with the supports_messages capability in your deployment can call this route. What you see in GET /v1/models and the models catalog is authoritative for your account. A model that declares tool calling or image capabilities can be used with tools and image input, respectively.

If the model is missing or lacks the required capability, the API returns 404 (model_not_found) or 400 with guidance in error.

Field Required Notes
model yes Your CompactifAI model configuration id.
messages yes Conversation, in order; roles user / assistant. content is a string or a list of blocks (text, image, tool_use, tool_result, …) so multi-turn tool flows round-trip. Every message needs non-empty, non-whitespace content, except a final assistant prefill.
max_tokens yes Generation budget. Must be ≥ 1 and at most the model’s maximum.
system no System prompt: a string, or a list of {"type": "text", "text": …} blocks. Takes precedence over any system text inside messages.
temperature, top_p, top_k no Sampling controls. Deprecated by Anthropic for newer Claude models, but supported here — the backend engine consumes them. top_k: 0 disables top-k; top_p: 0 keeps only the most probable token.
stop_sequences no Not supported. Do not rely on this field; see Stop sequences.
stream no true returns SSE (text/event-stream) with Anthropic events. Default false.
tools no Anthropic-style tools: name, description, input_schema (JSON Schema object).
tool_choice no Always an object per the Anthropic spec: {"type": "auto"}, {"type": "any"}, {"type": "none"}, or {"type": "tool", "name": …} (all also accept disable_parallel_tool_use). Requires tools unless none.
metadata no Optional key-value pairs attached to the request.

Fields you don’t use are ignored; Anthropic parameters not supported(e.g. extended thinking, prompt-cache fields) are silently dropped in the current version rather than rejected.

stop_sequences is not supported on this endpoint. The field is accepted but does not behave like Anthropic’s, so do not use it to control where a response ends. Known deviations:

  • Matches inside reasoning. The engine checks stop strings against everything the model generates, including its reasoning (thinking) tokens. If a stop string shows up while the model is thinking, generation ends mid-reasoning and the response can contain a thinking block and no text block. Anthropic applies stop sequences to visible output text only.
  • Wrong stop_reason on some models. Depending on the engine version serving the model, a stop-sequence match may be reported as stop_reason: "end_turn" with stop_sequence: null, instead of "stop_sequence" with the matched string.

To end output early, use max_tokens, or trim the response on your side.

from anthropic import Anthropic
# Pass the bare base URL (no trailing /v1): the Anthropic SDK appends /v1 itself.
client = Anthropic(base_url="https://api.compactif.ai", api_key="your_api_key_here")
message = client.messages.create(
model="quasar-2-358b",
max_tokens=1024,
system="You are a helpful assistant.",
messages=[{"role": "user", "content": "What is the capital of Colombia?"}],
)
print(message.content[0].text)

curl equivalent:

Terminal window
curl -X POST https://api.compactif.ai/v1/messages \
-H "Authorization: Bearer your_api_key_here" \
-H "Content-Type: application/json" \
-H "anthropic-version: 2023-06-01" \
-d '{
"model": "quasar-2-358b",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "What is the capital of Colombia?"}]
}'

Response (non-streaming) is the standard Anthropic Message:

{
"id": "msg_01…",
"type": "message",
"role": "assistant",
"model": "quasar-2-358b",
"content": [{"type": "text", "text": "The capital of Colombia is Bogotá."}],
"stop_reason": "end_turn",
"stop_sequence": null,
"usage": {"input_tokens": 25, "output_tokens": 14}
}

Set "stream": true. The response is a stream of Anthropic events—message_start, content_block_start, content_block_delta, content_block_stop, message_delta, message_stop—emitted as SSE event: + data: pairs. The final message_delta carries the output-token count; message_start carries the input-token count. The stream ends at message_stop (there is no [DONE] sentinel).

Streaming usage is terminal-only (no continuous usage). /v1/messages reports usage just twice: input_tokens on message_start and the final output_tokens on the terminal message_delta. Unlike POST /v1/chat/completions with stream_options.continuous_usage_stats, there are no intermediate, running usage updates during the stream. This is a downstream dependency incompatibility—the model backend currently does not attach running usage to intermediate message_delta events—not a gateway limitation, and it will be lifted once the backend is upgraded. Clients that need mid-generation token accounting should use POST /v1/chat/completions in the meantime.

Terminal window
curl -N -X POST https://api.compactif.ai/v1/messages \
-H "Authorization: Bearer your_api_key_here" \
-H "Content-Type: application/json" \
-d '{
"model": "quasar-2-358b",
"max_tokens": 256,
"stream": true,
"messages": [{"role": "user", "content": "Say hello."}]
}'

Send tools in the Anthropic tool shape; the model may answer with a tool_use content block and stop_reason: "tool_use". Feed the result back as a tool_result block in the next user message to continue the loop—same pattern as Claude:

{
"model": "quasar-2-358b",
"max_tokens": 512,
"messages": [
{"role": "user", "content": "What is 2+2?"}
],
"tools": [
{
"name": "calculator",
"description": "Evaluate an arithmetic expression",
"input_schema": {
"type": "object",
"properties": {"expression": {"type": "string"}},
"required": ["expression"]
}
}
]
}

Send an image content block in a user message (model must have image capability):

{
"role": "user",
"content": [
{"type": "text", "text": "What do you see in this image?"},
{
"type": "image",
"source": {"type": "url", "url": "https://example.com/photo.jpg"}
}
]
}

Base64 sources work too: {"type": "base64", "media_type": "image/jpeg", "data": "…"}.

POST /v1/messages/count_tokens counts the tokens a Message would use — across messages, system, and tools — without generating anything. The request body is the subset of the create-message body that affects the count: model, messages, and optional system, tools, thinking, tool_choice. Sampling parameters and max_tokens are not needed.

Terminal window
curl https://api.compactif.ai/v1/messages/count_tokens \
-H "Authorization: Bearer $KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "my-anthropic-model",
"messages": [{"role": "user", "content": "Hello, world"}],
"system": "You are a helpful assistant."
}'

Response:

{"input_tokens": 29}

With the Python SDK, use client.messages.count_tokens(...).

Token counting emits no usage events and debits no tokens from your wallet — only the standard request rate limit applies.

Errors use the Anthropic error shape, so Anthropic SDKs raise their typed exceptions:

{"type": "error", "error": {"type": "invalid_request_error", "message": "…", "param": "max_tokens"}, "request_id": "…"}

error.type follows the status: invalid_request_error (400), authentication_error (401), billing_error (402), permission_error (403), not_found_error (404), rate_limit_error (429), api_error (5xx). code and param are included when known. A failure after a stream has started arrives as an event: error with the same body. Upstream validation failures from the model backend (e.g. bad image data) are relayed with their original status code and message.