AI roundup
GLM-5.3-Flash drops: 320B/18B MoE, 1M context
Thu 27 August 2026
Formerly ‘Ox Alpha’: MIT weights, beats Opus 4.8 on agentic coding, 1/10th cost of GLM-5.3
| Parameters | 320B total / 18B active MoE |
| Context Window | 1M tokens |
| DeepSWE v1.1 | 63.4 (Opus 4.8: 58.0) |
| GDPval-AA Elo | 1773 (Opus 4.8: 1582) |
| Pricing | $0.15 in / $0.50 out per 1M |
| Weights | MIT License on Hugging Face |
Key takeaways
- Beats Claude Opus 4.8 on DeepSWE (+5.4 points) and GDPval-AA (+191 Elo). Code Bench 29.0 vs Opus 4.8’s 29.5—near parity at 10× lower list price than flagship GLM-5.3.
- Efficiency gains IndexPool sparse/linear attention cuts attention computation 3.01× and KV cache 4.44× vs GLM-5.3. Layer count down to 45 from 92.
- Day-0 logistics CoreWeave, Baseten, and Cline integration live at launch. Chat template updated post-release (Zixuan Li)—early HF downloaders must re-pull weights.
Source: smol.ai AI News