Self hosted ai

Blog posts tagged “Self-Hosted AI”


Self-Host OmniRoute: A Free AI Gateway for 500+ Models and 290+ Providers

OmniRoute AI Gateway Self-Hosted AI

OmniRoute is a free MIT-licensed AI gateway you run yourself: one OpenAI-compatible endpoint in front of 290+ providers and 500+ models. We ran v3.8.48 in Docker, got 99 models resolving with zero configuration, tested combos, compression, MCP, and the CLI, then shared the whole thing over a public HTTPS URL with Pinggy.

Best Open Source Self-Hosted Text-to-Speech Models in 2026

Text to Speech Self-Hosted AI Kokoro TTS

A guide to the best open-weight text-to-speech models you can self-host in 2026, ranked by Artificial Analysis Speech Arena Elo. Compare Breeze TTS 2, Fish Audio S2 Pro, Step Audio EditX, Voxtral TTS, Qwen3-TTS, VoxCPM2, GLM-TTS, Higgs Audio V3, Kokoro 82M and Chatterbox on quality, licensing, hardware and deployment.

Self hosting a 744B param LLM with only 25 GB RAM

GLM-5.2 Local LLM Mixture of Experts

A single-file C engine called Colibrì streams GLM-5.2's 744B mixture-of-experts weights off an NVMe drive to run the full model on 25 GB of RAM at 0.05-2 tokens/second. Here's how it works, what Hacker News made of it, and how to check on a queued run from your phone with Pinggy.

Picking the Right Hardware to Run LLMs Locally in 2026

Local LLM Self-Hosted AI Apple Silicon

A practical hardware guide for self-hosting LLMs in 2026. Compare consumer GPUs, Apple Silicon, enterprise cards, and pre-built AI workstations. Find the right setup for 7B to 405B models at every budget.

Best Open Source Self-Hosted LLMs for Coding in 2026

Open Source LLM Self-Hosted AI Local LLM

Discover the best open source LLMs for coding and development that you can self-host. Compare Kimi K3, Qwen3.8-Max, GLM-5.3, GLM-5.3-Flash, DeepSeek-V4-Pro, MiniMax M3, Qwen3.8-27B, Muse Glimmer 30B, Nemotron 3.5 Lightning and more with benchmarks, hardware requirements, and deployment guides.