Skip to content
Dashboard

GLM 4.5

GLM 4.5 is Z.AI's full-scale model released July 28, 2025, unifying reasoning, coding, and agentic capabilities in a single endpoint. Available through AI Gateway with built-in observability and intelligent provider routing. Your use subject to Z.AI's Terms & Privacy Policies.

ReasoningTool UseImplicit Caching
index.ts
import { streamText } from 'ai'
const result = streamText({
model: 'zai/glm-4.5',
prompt: 'Why is the sky blue?'
})

Playground

Try out GLM 4.5 by Z.AI. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

zai logo
zai logo

GLM 4.5

Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.

Provider
Context
Max Output
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
ZDR
No Training
Release Date
128K96K
0.9s
97tps
$0.60/M
$2.20/M
Read:$0.11/M
Write:
07/28/2025
Throughput

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the docs for more info.

Latency

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the docs for more info.

Uptime

Direct request success rate on AI Gateway and per-provider. Visit the docs for more info.

More models by Z.AI

Model
Context
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
Providers
ZDR
No Training
Release Date
1M
2.3s
168tps
$2.10/M
$6.60/M
Read:$0.21/M
Write:
fireworks logo
wafer logo
06/23/2026
1M
0.6s
338tps
$1.40/M$0.90/M
Fast $2.10/M
$4.40/M$2.84/M
Fast $6.60/M
Read:$0.26/M$0.14/M
Write:
alibaba logo
baseten logo
crusoe logo
+11
06/16/2026
205K
0.8s
61tps
$1.40/M
$4.40/M
Read:$0.26/M
Write:
deepinfra logo
fireworks logo
novita logo
+1
04/07/2026
203K
3.3s
47tps
$1.20/M
$4/M
Read:$0.24/M
Write:
zai logo
03/15/2026
203K
0.2s
105tps
$0.80/M
$2.56/M
Read:$0.16/M
Write:
bedrock logo
deepinfra logo
novita logo
+2
02/12/2026
205K
0.1s
839tps
$2.25/M
$2.75/M
Read:$2.25/M
Write:
baseten logo
bedrock logo
cerebras logo
+3
12/22/2025

About GLM 4.5

GLM 4.5 was released July 28, 2025 as Z.AI's full-scale large language model designed to unify reasoning, coding, and agentic capabilities. It represents the full-scale offering in the GLM-4.5 generation, targeting workloads where broad competence across analytical and generative tasks matters more than narrow specialization.

The model supports configurable thinking, letting you enable or disable chain-of-thought reasoning depending on the task. This flexibility is useful in agentic pipelines where some steps benefit from deliberation and others need fast, direct responses. GLM 4.5 operates within a context window of 128K tokens, handling long documents, extended conversations, and multi-file code analysis in a single pass.

Z.AI positions GLM 4.5 alongside other widely used closed-source models. For teams evaluating alternatives across providers, it offers a distinct cost-performance point. Through AI Gateway, you access GLM 4.5 with a unified API, automatic retries, and provider routing without managing separate accounts.

What To Consider When Choosing a Provider

  • Configuration: GLM 4.5 supports a context window of 128K tokens and up to 96K tokens per request. For reasoning-heavy tasks with thinking enabled, budget extra output tokens for chain-of-thought traces that precede the final answer.
  • Configuration: Test both thinking-enabled and thinking-disabled modes. Thinking mode improves accuracy on complex reasoning but increases latency and token usage. Disable it for straightforward generation tasks.
  • Configuration: When using AI Gateway, configure fallback providers to maintain availability. GLM 4.5 is available through Z.AI.
  • Zero Data Retention: AI Gateway does not currently support Zero Data Retention for this model. See the documentation for models that support ZDR.
  • Authentication: AI Gateway authenticates requests using an API key or OIDC token. You do not need to manage provider credentials directly.

When to Use GLM 4.5

Best for

  • General-purpose reasoning and coding: Unified capability across math, logic, and code generation reduces the need for task-specific models
  • Agentic workflows: Multi-step planning, tool use, and configurable thinking run within a single model
  • Long-document analysis: The context window of 128K tokens fits contracts, research papers, or large codebases
  • Production deployments: Built-in observability and automatic retries through AI Gateway reduce operational overhead
  • Cost-conscious teams: Compare listed rates against alternative providers when evaluating total spend

Consider alternatives when

  • Lightweight high-volume alternative: GLM-4.5-Air offers reduced latency and cost for less demanding workloads
  • Vision or multimodal input: GLM-4.5V adds image understanding on top of the GLM-4.5 foundation
  • Code generation focus: GLM-4.6 and later models include targeted coding improvements
  • Deeper reasoning and planning: GLM-5 introduces multiple thinking modes and improved long-range planning

Conclusion

GLM 4.5 is Z.AI's full-capability model in the GLM-4.5 generation, balancing reasoning depth, coding proficiency, and agentic flexibility. For teams that need a single model covering a broad range of tasks with configurable thinking, it's available through AI Gateway with unified billing and observability.