100100 Ideas AI
Submit Tool

tightwad

Open source

Self-hosted mixed-vendor GPU inference cluster manager with speculative decoding proxy for faster LLM serving.

33MIT

Mixed-vendor GPU inference cluster manager with speculative decoding

Pros

  • +Pools heterogeneous GPUs (CUDA, ROCm, CPU) into one cluster
  • +Speculative decoding accelerates inference up to 2-3x
  • +OpenAI-compatible API endpoint, drop-in replacement
  • +Supports multiple modes: combined, proxy, multi-drafter consensus

Cons

  • −Requires technical setup and hardware management
  • −Network latency can limit speedup in some configurations
  • −Limited to llama.cpp backends and specific model support

Target audience: Developers and hobbyists with spare GPUs who want to build a cost-effective, self-hosted LLM inference cluster.

Related tools

Other open-source tools that share tags with tightwad.

Horizon

MCPMIT

AI-powered news radar generating daily bilingual briefings in English & Chinese.

9k·Other·Open source