Skip to content

GLM 5.2

Model ID

glm-5-2

Publisher

zai-org

GLM 5.2 is exposed through POST /v1/chat/completions using the standard OpenAI-compatible chat payload. It supports function / tool calling (tools, tool_choice, assistant tool_calls) and works with structured outputs via response_format where applicable.

Specification Value
API ID glm-5-2
Capability Supported
Chat completions (/v1/chat/completions) Yes
Tool / function calling Yes
Structured output (response_format) Yes
Responses API (/v1/responses) Yes
Anthropic Messages API (/v1/messages) Yes
Prompt caching Yes
Reasoning Yes

Prompt caching is enabled for GLM 5.2. Input tokens served from the cache are billed at 20% of the input price (currently $0.22 per 1M tokens, versus $1.10 per 1M for non-cached input). See Pricing for details.

Property Value
Reasoning supported Yes
Toggleable Yes
How to toggle chat_template_kwargs with {"enable_thinking": true} or {"enable_thinking": false}
Reasoning effort levels high, max (default)
How to set effort reasoning_effort parameter

Reasoning is always enabled for this model — you cannot turn it off via chat_template_kwargs. Control the depth of reasoning by setting reasoning_effort to "high" or "max" (default). A model that does not support the requested level rejects the request and names the levels it does support.

For a full example, see Reasoning Example.

Use glm-5-2 as the model field in chat completion requests. For tool usage patterns and examples, see Tool Calling and Structured Output.