Hermes Atlas
Memory & knowledge

Self-Hosted Honcho for Hermes Agent

elkimek/honcho-self-hosted

Run the Honcho memory layer on your own server for Hermes Agent with OpenRouter or Venice, no code changes

In short

Self-Hosted Honcho for Hermes Agent is a set of three config files that run Plastic Labs' Honcho memory layer on your own machine instead of their cloud. Hermes Agent then talks to it at localhost:8000.

What Self-Hosted Honcho for Hermes Agent does

Hermes Agent uses Honcho for its cross-session memory layer, and by default that runs on Plastic Labs' managed cloud with their Neuromancer models. This repository lets you host Honcho yourself with three configuration files on top of upstream Honcho, so no fork is required. It runs the full stack on your machine: the API, the Deriver, PostgreSQL with pgvector, and Redis.

LLM calls for memory work go to any OpenAI-compatible provider, with an optional backup provider, and the README names OpenRouter, Venice, Routstr, Together and Ollama as options. Your conversation data and user profile stay on your own hardware, and only inference requests leave it. The README compares three setups: managed cloud, self-hosted with an API, and self-hosted with a local model through Ollama or vLLM. It notes that general-purpose models may give slightly less precise deductive reasoning than Neuromancer, and that the main benefit here is data ownership. Setup is described as taking about 3 minutes.

Key features

  • Runs Honcho's API, Deriver, PostgreSQL and Redis on your own machine
  • Three config files on top of upstream Honcho, with no fork
  • Primary and optional backup LLM provider through any OpenAI-compatible API
  • Hermes Agent connects to the self-hosted API at localhost:8000
  • Optional fully local setup with Ollama, vLLM or llama.cpp via LLM_VLLM_BASE_URL

When to use it

  • Keeping Hermes Agent's cross-session user model off a third-party cloud
  • Running Hermes memory on a VPS or home server with OpenRouter or Venice for inference
  • Building a fully local memory stack with a model on your own network

Who it is for: Hermes Agent users who want its Honcho memory layer to live on their own server.

How it fits with Hermes Agent

It is built for Hermes Agent's Honcho-based memory layer and replaces the default Plastic Labs cloud endpoint with a self-hosted one.

Requirements: Ubuntu 22.04+ Linux server (tested on 22.04 with 6GB RAM) and an OpenAI-compatible LLM provider such as OpenRouter or Venice.

FAQ

What is Self-Hosted Honcho for Hermes Agent?

It is a configuration repository for running Plastic Labs' Honcho memory layer on your own server, so Hermes Agent's cross-session memory is stored locally instead of in the managed cloud.

Does Self-Hosted Honcho for Hermes Agent work with Hermes Agent?

Yes, that is its purpose. Hermes Agent talks to the self-hosted Honcho API at localhost:8000 and needs no code changes.

Is Self-Hosted Honcho for Hermes Agent free and open source?

The repository is licensed under GPL-3.0. Running it costs only the LLM API usage you choose, or hardware if you use a local model.

Similar memory for Hermes Agent

All memory

Related guides: SOUL.md for Hermes Agent: what it is and how to write one · Run multiple Hermes agents with profiles