GLM-5.3 Crushes Fable on DeepSWE Benchmark with Lower Costs
The DeepSWE benchmark has revealed that GLM-5.3 outperforms Claude Fable 5 in coding tasks, while also being significantly cheaper.
In a head-to-head comparison on the DeepSWE benchmark, which evaluates AI models on 113 original software engineering tasks, GLM-5.3 achieved comparable first-attempt accuracy to Fable at a cost of $3.99 per rollout compared to Fable's $21.63.
The results show that while both models performed similarly on pass@1, the metric for solving tasks on the first attempt, with Fable narrowly leading at 69.7% versus GLM-5.3's 69.0%, GLM-5.3 pulled ahead in pass@2 and pass@4 metrics.
GLM-5.3 offers a broader task coverage than Fable, achieving 81.1% success rates on pass@2 and 87.6% on pass@4 compared to Fable's 77.1% and 84.1%, respectively.