Z.ai launches GLM-5.3 for coding and cyber defense
Z.ai has released GLM-5.3, a post-trained update to the GLM-5.2 base model focused on coding, long-horizon agent work, and defensive cybersecurity.

The short version
Z.ai has released GLM-5.3, a new version of its flagship model built on the same base model as GLM-5.2. The company says the gains come from scaling post-training across more environments and longer-horizon tasks. Z.ai's announcement is the primary source for the release details.
Z.ai posted the release on X:
GLM-5.3 is available through the GLM Coding Plan and ZCode. Z.ai says API access and open weights will be released in stages after additional safety evaluations. That staged release is part of the story: the company is positioning the model for coding while treating its reported cyber capability as a deployment risk that needs more review.
What changed in GLM-5.3
Z.ai says it kept the GLM-5.2 base model and focused the new training effort on post-training. The announcement describes more varied task environments, longer runs, and more compute spent training on those environments.
The company reports a 50% improvement over GLM-5.2 on its internal Z.ai Code Bench. It also reports open-source state-of-the-art results on Terminal Bench 3.0 and Agents' Last Exam. These are company-reported figures, not a guarantee that every repository or workflow will see the same result. Source
For teams, the practical question is whether the model can stay useful across a chain of work: understand a repository, plan a change, edit multiple files, run tests, interpret failures, and return a reviewable result.
Cyber defense is the unusual part of this release
Z.ai says GLM-5.3 developed stronger cyber capabilities during post-training than the company expected. The model is presented as a defensive-security system, but capability claims in this area need careful separation between vulnerability discovery, validation, and exploit development.
Reuters reported that Z.ai claimed an 84.5% score on CyberGym, compared with 83.8% for Anthropic's restricted Mythos 5. The figures have not been independently verified. Reuters coverage
The same reporting describes a wider gap on exploit development. That distinction matters for risk review. A model that can find and confirm a vulnerability is not equivalent to a model that can produce a working attack, and neither result removes the need for an authorized test environment.
Availability and pricing
GLM-5.3 is available now through Z.ai's GLM Coding Plan and ZCode. Z.ai says API access and open weights will follow a staged process after safety evaluations and hardening. Official release post
The announcement does not provide a stable public API price for GLM-5.3. Teams should avoid budgeting from third-party summaries until the official API documentation lists the model, limits, and billing terms.
For open-model teams, the delay in open weights is a meaningful change from the fast release patterns associated with earlier versions. It gives Z.ai time to evaluate the cyber capability before making the weights broadly downloadable.
How it fits with Gemini, GPT, and Grok
The current model market is dividing the agent problem into different priorities.
- Gemini 3.7 Flash focuses on coding, web development, knowledge work, and business workflows, with introductory API pricing through the end of 2026. Google's announcement
- GPT-5.6 Sol Ultrafast keeps GPT-5.6 Sol as the underlying model and changes the processing tier, with OpenAI reporting up to 750 output tokens per second in a limited preview. OpenAI's announcement
- Grok 4.6 emphasizes long-running agents, coding, and knowledge work, with access through xAI's API, Grok Build, Cursor, and GitHub Copilot. xAI's announcement
- GLM-5.3 emphasizes post-training gains, coding, defensive security, and a staged path toward API access and open weights. Z.ai's announcement
The useful comparison is not a single leaderboard. It is a representative task set with the team's repository context, tool permissions, review process, retries, and total cost per completed task.
What teams should test
A practical evaluation should include:
- A real repository with a known bug and a clear test suite.
- A multi-file change that requires planning and repeated tool use.
- A long-horizon task where the agent must recover from a failed test.
- A security review in an explicitly authorized, isolated environment.
- A record of review time, rejected changes, retries, and completed-task cost.
For cybersecurity use, define the boundary before the model receives access. Keep production credentials out of the test, record every tool call, and require human approval before any change or network action.
Why the workspace around the model matters
Model quality is only one part of an agent workflow. The request, source documents, repository context, tool permissions, human review, and final record need to stay connected if a team wants repeatable work instead of one-off output.
Kylon gives teams a shared workspace for channels, files, agents, connections, skills, and recurring workflows. That makes it possible to evaluate a model in the same operational context where the result will be reviewed, approved, and reused. The goal is not to make a capability claim on behalf of a model. It is to make the team's own evidence easier to collect and share.
Bottom line
- GLM-5.3 keeps the GLM-5.2 base model and attributes its gains to scaled post-training. Z.ai
- Z.ai reports stronger coding and long-horizon agent performance, including a 50% internal Code Bench improvement. Z.ai
- Cyber-defense figures are company-reported and not independently verified. Reuters
- GLM Coding Plan and ZCode are available now; API access and open weights are staged behind further safety work. Z.ai
Teams evaluating GLM-5.3 should measure completed work, review burden, permissions, and auditability together.
Hire the AI agent team that runs your entire business.

