Local llm

Blog posts tagged “Local LLM”


Self hosting a 744B param LLM with only 25 GB RAM

GLM-5.2 Colibrì Local LLM

A single-file C engine called Colibrì streams GLM-5.2's 744B mixture-of-experts weights off an NVMe drive to run the full model on 25 GB of RAM at 0.05-2 tokens/second. Here's how it works, what Hacker News made of it, and how to check on a queued run from your phone with Pinggy.

Picking the Right Hardware to Run LLMs Locally in 2026

Local LLM Self-Hosted AI GPU for LLM

A practical hardware guide for self-hosting LLMs in 2026. Compare consumer GPUs, Apple Silicon, enterprise cards, and pre-built AI workstations. Find the right setup for 7B to 405B models at every budget.

Best Open-Source Alternatives to ChatGPT in 2026

Open Source AI ChatGPT Alternatives Self-Hosted AI

Discover the best open-source alternatives to ChatGPT in 2026. Compare Open WebUI, Jan, GPT4All, LibreChat, LobeChat, AnythingLLM, and more self-hosted AI chat platforms you can run locally with full privacy.

Self-Host AI Agents Using n8n and Pinggy

N8n Ai Self-Hosted

Learn how to self-host n8n's AI Starter Kit and access your AI workflows remotely with Pinggy. Step-by-step guide to privacy-focused AI solutions with local LLMs.