GLM-5.3 Flash: Cost-Optimized Model Cuts Costs by 17x
GLM-5.3 Flash, a cost-optimized model developed by Z.ai, has been found to significantly reduce costs while maintaining high performance levels in coding workloads.
A recent analysis on DeepSWE, a benchmark for software engineering tasks, compared the performance of GLM-5.3 Flash with its flagship counterpart, GLM-5.3. The results showed that GLM-5.3 Flash achieved a 17x cost reduction while retaining 94% of task coverage.
The cost savings are substantial: at $0.24 per rollout, GLM-5.3 Flash delivered 264 solves per $100 in the DeepSWE tests, far outpacing the 17 solves achieved by GLM-5.3 at $3.99 per rollout.
However, the reduced reliability of GLM-5.3 Flash manifests as higher flakiness on tasks requiring retries, and it struggled to convert extended runs into successful solutions compared to its flagship counterpart (46% effort payoff vs. 61%).