Hermes Atlas
Models, providers & proxies

Hermes Local Stack

KaiFelixBennett/hermes-claude-code-local

Run Hermes Agent and Claude Code on a local llama.cpp model with no API costs

In short

Hermes Local Stack is a setup that wires Hermes Agent to a local llama.cpp server, and optionally bridges Claude Code through a LiteLLM proxy, so agentic coding runs on your own hardware. It includes setup scripts for Windows, Linux and macOS.

What Hermes Local Stack does

Hermes Local Stack wires Hermes Agent directly to llama.cpp so the agent plans tasks, calls tools and edits files on your own hardware with no API costs. Claude Code can optionally be bridged through a local LiteLLM proxy so it also talks to the local model instead of Anthropic's API. The author reports one 4-hour autonomous session of 7,256,671 tokens that would have cost $94.34 on Claude Opus 4.7.

The core path is Hermes to llama.cpp, and local Claude Code tasks go from Hermes through a claude-code skill to Claude Code, LiteLLM and llama.cpp. A setup script installs Hermes Agent and llama.cpp (Metal on Apple Silicon, Vulkan or CPU on Linux), downloads a model if you have none, and points Hermes at the local server. It includes Telegram control with voice messages, MTP and speculative decoding with a Qwen3.6-27B-MTP GGUF, self-healing scripts that restart LiteLLM, and model tuning notes under docs/models.

Key features

  • Hermes Agent connected directly to llama.cpp for fully local agent loops
  • Optional Claude Code bridge through a local LiteLLM proxy
  • Telegram control of Hermes, including voice messages
  • MTP and speculative decoding with a Qwen3.6-27B-MTP GGUF
  • Self-healing scripts that restart LiteLLM if it crashes
  • Per-model tuning notes under docs/models

When to use it

  • Run long autonomous coding sessions on local hardware instead of paying per token
  • Point Claude Code at a local model through LiteLLM
  • Control a local Hermes agent from your phone over Telegram

Who it is for: Developers with capable local hardware who want Hermes Agent and Claude Code to run without cloud API costs.

How it fits with Hermes Agent

Built around Hermes Agent: Hermes talks directly to llama.cpp, and Claude Code is reached from Hermes through a skill.

How to install Hermes Local Stack

These commands are copied from the project's README. Check the repository for the latest steps before you run them.

curl -fsSL https://raw.githubusercontent.com/KaiFelixBennett/hermes-claude-code-local/main/setup.sh | bash
cd ~/hermes-claude-code-local && make start

Note: The README notes Claude Code can use 60k+ tokens of overhead before the first message on a 64k context window, and recommends Hermes direct until you run 128k+ context reliably.

FAQ

What is Hermes Local Stack?

Hermes Local Stack is a setup that connects Hermes Agent to a local llama.cpp server and optionally bridges Claude Code through LiteLLM. Agent loops, tool calls and file edits then run on your own hardware.

How do I install Hermes Local Stack?

On Linux or macOS, run the setup.sh script from the README with curl and bash, then start it with make start from the cloned folder. On Windows, run ./setup_hermes_local.ps1 in PowerShell.

Is Hermes Local Stack free and open source?

The repository has no license file, so default copyright applies and you should check with the author before reuse. The source code is public on GitHub.

Similar models for Hermes Agent

All models

Related guides: How to install Hermes Agent