Intelligent LLM gateway and VRAM-aware router for Ollama, llama.cpp, and OpenAI with semantic caching and auto-failover.
Alternative to
Run Hermes Agent and Claude Code locally on llama.cpp with zero API costs.
Run Hermes Agent + Claude Code locally on llama.cpp — zero API costs. A 4h / 7M-token session that would have cost $94 on Claude Opus 4.7
Target audience: Developers and AI enthusiasts who want to run autonomous coding agents locally without cloud API costs.
Other open-source tools that share tags with hermes-claude-code-local.
Intelligent LLM gateway and VRAM-aware router for Ollama, llama.cpp, and OpenAI with semantic caching and auto-failover.
Alternative to
Self-hosted mixed-vendor GPU inference cluster manager with speculative decoding proxy for faster LLM serving.
Run IBM Granite 4.0 locally on Raspberry Pi 5 with Ollama.This is a privacy-first AI. Your data never leaves your device
Open-source AI engine to run LLMs, vision, voice, image, video on any hardware, no GPU required.
Alternative to
93+ production-ready AI projects with tutorials on LLMs, RAG, and agents.
QwenPaw: Your personal AI assistant, deployable locally or in the cloud, with memory, multi-agent support, and multi-cha
Alternative to