Model ID
glm-5-2
Model ID
glm-5-2
Publisher
zai-org
GLM 5.2 is exposed through POST /v1/chat/completions using the standard OpenAI-compatible chat payload. It supports function / tool calling (tools, tool_choice, assistant tool_calls) and works with structured outputs via response_format where applicable.
| Specification | Value |
|---|---|
| API ID | glm-5-2 |
| Capability | Supported |
|---|---|
Chat completions (/v1/chat/completions) |
Yes |
| Tool / function calling | Yes |
Structured output (response_format) |
Yes |
Responses API (/v1/responses) |
Yes |
Anthropic Messages API (/v1/messages) |
Yes |
| Prompt caching | Yes |
| Reasoning | Yes |
Prompt caching is enabled for GLM 5.2. Input tokens served from the cache are billed at 20% of the input price (currently $0.22 per 1M tokens, versus $1.10 per 1M for non-cached input). See Pricing for details.
| Property | Value |
|---|---|
| Reasoning supported | Yes |
| Toggleable | Yes |
| How to toggle | chat_template_kwargs with {"enable_thinking": true} or {"enable_thinking": false} |
| Reasoning effort levels | high, max (default) |
| How to set effort | reasoning_effort parameter |
Reasoning is always enabled for this model — you cannot turn it off via chat_template_kwargs. Control the depth of reasoning by setting reasoning_effort to "high" or "max" (default). A model that does not support the requested level rejects the request and names the levels it does support.
For a full example, see Reasoning Example.
Use glm-5-2 as the model field in chat completion requests. For tool usage patterns and examples, see Tool Calling and Structured Output.