Self-Hosted Honcho for Hermes Agent
elkimek/honcho-self-hosted
Run the Honcho memory layer on your own server for Hermes Agent with OpenRouter or Venice, no code changes
Self-Hosted Honcho for Hermes Agent is a set of three config files that run Plastic Labs' Honcho memory layer on your own machine instead of their cloud. Hermes Agent then talks to it at localhost:8000.
What Self-Hosted Honcho for Hermes Agent does
Hermes Agent uses Honcho for its cross-session memory layer, and by default that runs on Plastic Labs' managed cloud with their Neuromancer models. This repository lets you host Honcho yourself with three configuration files on top of upstream Honcho, so no fork is required. It runs the full stack on your machine: the API, the Deriver, PostgreSQL with pgvector, and Redis.
LLM calls for memory work go to any OpenAI-compatible provider, with an optional backup provider, and the README names OpenRouter, Venice, Routstr, Together and Ollama as options. Your conversation data and user profile stay on your own hardware, and only inference requests leave it. The README compares three setups: managed cloud, self-hosted with an API, and self-hosted with a local model through Ollama or vLLM. It notes that general-purpose models may give slightly less precise deductive reasoning than Neuromancer, and that the main benefit here is data ownership. Setup is described as taking about 3 minutes.
Key features
- Runs Honcho's API, Deriver, PostgreSQL and Redis on your own machine
- Three config files on top of upstream Honcho, with no fork
- Primary and optional backup LLM provider through any OpenAI-compatible API
- Hermes Agent connects to the self-hosted API at localhost:8000
- Optional fully local setup with Ollama, vLLM or llama.cpp via LLM_VLLM_BASE_URL
When to use it
- Keeping Hermes Agent's cross-session user model off a third-party cloud
- Running Hermes memory on a VPS or home server with OpenRouter or Venice for inference
- Building a fully local memory stack with a model on your own network
Who it is for: Hermes Agent users who want its Honcho memory layer to live on their own server.
How it fits with Hermes Agent
It is built for Hermes Agent's Honcho-based memory layer and replaces the default Plastic Labs cloud endpoint with a self-hosted one.
Requirements: Ubuntu 22.04+ Linux server (tested on 22.04 with 6GB RAM) and an OpenAI-compatible LLM provider such as OpenRouter or Venice.
FAQ
What is Self-Hosted Honcho for Hermes Agent?
It is a configuration repository for running Plastic Labs' Honcho memory layer on your own server, so Hermes Agent's cross-session memory is stored locally instead of in the managed cloud.
Does Self-Hosted Honcho for Hermes Agent work with Hermes Agent?
Yes, that is its purpose. Hermes Agent talks to the self-hosted Honcho API at localhost:8000 and needs no code changes.
Is Self-Hosted Honcho for Hermes Agent free and open source?
The repository is licensed under GPL-3.0. Running it costs only the LLM API usage you choose, or hardware if you use a local model.
Similar memory for Hermes Agent
All memoryShared encrypted memory and messaging for Claude Code, Codex, OpenClaw and Hermes Agent on your machine
chandra447 Pi Hermes MemoryHermes-style persistent memory, session search and learning loop for the Pi coding agent
itechmeat Open Second BrainLocal-first Hermes Agent memory in an Obsidian vault, with nightly dream passes that learn preferences
text2future FlowixLocal Markdown notebook that serves as shared memory for Codex, Claude Code, OpenCode and Hermes
ReflexioAI ReflexioSelf-improvement harness that turns user corrections into persisted agent playbooks and user profiles
Signet-AI SignetShared memory layer that syncs memories, identity files and transcripts across agents and models
Related guides: SOUL.md for Hermes Agent: what it is and how to write one · Run multiple Hermes agents with profiles