Pinggy Blog

Tunnels, networking, self-hosted LLMs, and AI news

Hands-on guides, protocol deep dives, and field notes on tunnels, SSH, reverse proxies, self-hosting, self-hosted LLMs, and AI - from the people building Pinggy.


Best Open Source Self-Hosted Alternatives to Slack and Discord in 2026

Self-Hosted Open Source Pinggy

Compare the best open source, self-hostable alternatives to Slack and Discord in 2026 - Rocket.Chat, Mattermost, Zulip, Matrix/Element, Stoat, and Spacebar - with licenses, system requirements, and how to expose your instance with Pinggy.

Best Webhook Testing Tools for Local Development

Webhook Testing Pinggy Ngrok

Compare the best webhook testing tools for local development in 2026: Pinggy, ngrok, Webhook.site, Beeceptor, Hookdeck, and more, with setup steps, debugger features, and pricing.

Inside the Hugging Face Breach an AI Agent Ran Start to Finish

Hugging Face OpenAI AI Security

Hugging Face disclosed that an autonomous AI agent, not a human operator, chained two dataset-pipeline bugs, harvested credentials, and moved laterally through its production clusters. Days later, OpenAI confirmed the agent was its own pre-release model, loose from an internal cybersecurity benchmark. Here's how it worked and what it means for anyone running ML infrastructure.

Best Open Source Self-Hosted Text-to-Speech Models in 2026

Text to Speech Self-Hosted AI Kokoro TTS

A guide to the best open-weight text-to-speech models you can self-host in 2026, ranked by Artificial Analysis Speech Arena Elo. Compare Step Audio EditX, Fish Audio S2 Pro, Voxtral TTS, Kokoro 82M, Maya1, NVIDIA Magpie, Chatterbox, and Zonos on quality, licensing, hardware, and deployment.

How to Turn ChatGPT Into a Free Local Coding Agent With DevSpace

Chatgpt MCP AI Coding Tools

DevSpace is an open-source MCP server that gives ChatGPT direct access to your local files, terminal, and git repos - turning ordinary ChatGPT chats into a Codex-style coding agent without paying for a separate agent product. Full setup guide with Pinggy.

Self hosting a 744B param LLM with only 25 GB RAM

GLM-5.2 Local LLM Mixture of Experts

A single-file C engine called Colibrì streams GLM-5.2's 744B mixture-of-experts weights off an NVMe drive to run the full model on 25 GB of RAM at 0.05-2 tokens/second. Here's how it works, what Hacker News made of it, and how to check on a queued run from your phone with Pinggy.

How to get Free AI Model APIs with 'Unlimited' Tokens

OpenRouter LLM Router

How to get free access to AI model APIs on OpenRouter in 2026 - real rate limits, the current free model catalog (Nemotron 3 Ultra, Owl Alpha, Tencent Hy3), code examples, and how the $10 credit threshold works.