Glm 5.2

Blog posts tagged “GLM-5.2”


Self hosting a 744B param LLM with only 25 GB RAM

GLM-5.2 Local LLM Mixture of Experts

A single-file C engine called Colibrì streams GLM-5.2's 744B mixture-of-experts weights off an NVMe drive to run the full model on 25 GB of RAM at 0.05-2 tokens/second. Here's how it works, what Hacker News made of it, and how to check on a queued run from your phone with Pinggy.

Best Open Source Self-Hosted LLMs for Coding in 2026

Open Source LLM Self-Hosted AI GLM-5.2

Discover the best open source LLMs for coding and development that you can self-host. Compare Kimi K3, GLM-5.2, MiniMax M3, DeepSeek-V4-Pro-Max, Qwen3.6, Muse Glimmer 30B, Devstral 2, MiMo-V2.5-Pro, and more with benchmarks, hardware requirements, and deployment guides.