Intelligent LLM gateway and VRAM-aware router for Ollama, llama.cpp, and OpenAI with semantic caching and auto-failover.
Alternative to
Self-hosted mixed-vendor GPU inference cluster manager with speculative decoding proxy for faster LLM serving.
Mixed-vendor GPU inference cluster manager with speculative decoding
Target audience: Developers and hobbyists with spare GPUs who want to build a cost-effective, self-hosted LLM inference cluster.
Other open-source tools that share tags with tightwad.
Intelligent LLM gateway and VRAM-aware router for Ollama, llama.cpp, and OpenAI with semantic caching and auto-failover.
Alternative to
Run Hermes Agent and Claude Code locally on llama.cpp with zero API costs.
Alternative to
Open-source AI engine to run LLMs, vision, voice, image, video on any hardware, no GPU required.
Alternative to
93+ production-ready AI projects with tutorials on LLMs, RAG, and agents.
QwenPaw: Your personal AI assistant, deployable locally or in the cloud, with memory, multi-agent support, and multi-cha
Alternative to
AI-powered news radar generating daily bilingual briefings in English & Chinese.