Robots Atlas>ROBOTS ATLAS
GLM-5.3

GLM-5.3

5.3 · Family: GLM
Z.ai MoE language model (753B parameters) with a 1M context window and controllable reasoning, specialized in coding and cybersecurity.
✓ Active✓ Public access⚖ Open weightsLLMReasoning modelTool-using model📁 GLM
Context window
1M
tokens
Parameters
753B (MoE)
parameters
Max output
128,000
tokens
Release date
17 February 2026
Access:APIDownloadHostedDeployment:☁ Cloud💻 Local

Overview

GLM-5.3 is a Mixture-of-Experts language model built by Z.ai (Zhipu AI), the next iteration of the GLM family. It has 753 billion total parameters and is optimized for complex coding and long-horizon tasks.

Architecture

The model uses an MoE architecture with a DSA sparse-attention mechanism (tagged glm_moe_dsa). The context window reaches 1 million tokens and the maximum response length is 128K tokens. GLM-5.3 supports controllable reasoning via a reasoning_effort parameter with low, high and max settings.

Capabilities and results

Compared with GLM-5.2 the model delivers about a 50% improvement on Z.ai Code Bench and sets state of the art on CyberGym for vulnerability discovery, with markedly stronger exploitation capabilities. In benchmarks it scores 88.2 on Terminal-Bench 2.1, 28.3 on Terminal-Bench 3.0, 66.9 on DeepSWE, 84.5 on CyberGym and 54.4 on ExploitBench. It was released on February 17, 2026.

Classification
LLMReasoning modelTool-using model
Family: GLM
Access & deployment
APIDownloadHosted
CloudLocal
Weights: Open weights
Key parameters
📏 Context: 1M
🧩 Parameters: 753B (MoE)
Tools · ✓ Fine-tuning
📥 Input: text, structured data, documents

Technical specification

Context window
1M
tokens
Parameters
753B (MoE)
parameters
Max output tokens
128,000
tokens per response
License
GLM Model License (open weights) — dostęp produkcyjny przez api.z.ai / GLM Coding Plan
Hardware requirements
Open-weight 753B MoE model — self-hosting requires multi-GPU infrastructure (recommended min. 8-16x H100/H200 for 1M context handling, ~200GB+ VRAM after INT8/FP8 quantization). Z.ai serves the model commercially via api.z.ai as the main production option.
Features:Tool useFine-tuning
Modalities
⬇ Input
textstructured_datadocuments
⬆ Output
textcodestructured_data

Capabilities and applications

Native model capabilities
Advanced reasoning
The ability to perform multi-step, structured reasoning: analysing problems, planning steps, and drawing conclusions from hypotheses. Reasoning-first models (e.g. GPT-5.1 Thinking) dedicate a portion of inference to chains of thought before responding.
Category: reasoning
Extended thinking mode
A reasoning-model variant with a larger inference budget: more thinking cycles, higher answer precision at the cost of response time. Choice between 'standard' and 'extended' thinking is left to the user (e.g. the selector in GPT-5.2 Pro).
Category: reasoning
Adaptive reasoning effort
The model decides how much 'thinking' to allocate to a given query: simple questions are answered quickly, complex problems receive more inference cycles. A GPT-5.1 feature (both Instant and Thinking) that shortens time on easy tasks and extends it for hard ones.
Category: reasoning
Multi-step reasoning
Carrying out multi-step chains of reasoning across long, complex tasks.
Category: reasoning
Mathematical reasoning
The model's ability to solve mathematical tasks requiring multi-step reasoning — equations, proofs, combinatorics, geometry, calculus and competition-level problems.
Category: reasoning
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Agentic coding
Multi-hour, multi-step programming tasks performed autonomously by the model: cloning a repository, running tests, iterating on fixes, integrating with CLI tools. Characteristic of Codex variants (GPT-5.1-Codex-Mini, Codex-Max).
Category: coding
Agentic capability
The model's ability to autonomously plan and execute multi-step tasks by sequentially using tools, maintaining context, and adapting to intermediate results.
Category: planning
Tool use
The model's ability to call external functions, APIs and tools during a conversation: calculator, search engine, code editor, database. The model decides when and how to use a tool and interprets its result.
Category: planning
Long context
Support for large context windows — tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Prompt caching
Cost-performance optimisation: repeated prompt fragments (e.g. system prompt, long documentation) are cached server-side and cheaper in subsequent calls. Significantly reduces cost for applications with long contexts.
Category: other
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language
MCP support
Native support for the Model Context Protocol - the model can integrate with external MCP servers, invoke their tools, and use their data sources without a dedicated wrapper.
Category: other

Benchmark results

5 benchmarks
Terminal-Bench 2.0
accuracy · Terminal-Bench 2.1
88.2%
📄 Oficjalna karta modelu (Hugging Face)
Terminal-Bench 3.0
accuracy · Terminal-Bench 3.0
28.3%
📄 Oficjalna karta modelu (Hugging Face)
DeepSWE
accuracy · DeepSWE
66.9%
📄 Oficjalna karta modelu (Hugging Face)
CyberGym
success rate · CyberGym (vulnerability discovery)
84.5%
📄 Oficjalna karta modelu (Hugging Face)
ExploitBench
success rate · ExploitBench
54.4%
📄 Oficjalna karta modelu (Hugging Face)

Pricing

Technical architecture