Hermes Atlas
Plugins & extensions

Hermes Warm Compaction

Elevatormusic/hermes-warm-compaction

Hermes Agent context engine where the main model writes the compaction handoff on its cached prompt prefix

In short

Hermes Warm Compaction is a context engine plugin for Hermes Agent that speeds up compaction by having the main model write the handoff summary on the cached prefix of its last request. It uses documented plugin APIs only.

What Hermes Warm Compaction does

At each compaction, the plugin re-sends the session's last main-model request, adds the rows that came after it and one handoff instruction at the end. Because the server can reuse the cached prefix, it reads only the new rows. The reply is a Markdown handoff with five headings: Goal, User instructions, Current state, Key facts and Next step. The older rows are replaced with the handoff, and a verbatim tail of recent rows is kept. It works for manual /compress and automatic compaction.

If the warm request cannot run or its reply fails the plugin's gate, the same attempt falls back to a model-based summary through the Hermes auxiliary route. If neither works, it preserves the original history and reports an error. The speedup needs the same model and a server with a prefix cache. The README tables hosted APIs with prompt caching, vLLM, SGLang, TensorRT-LLM, llama.cpp, MLX, Ollama and LM Studio. It runs on unpatched Hermes, with no host patch or runtime wrapping.

Key features

  • Warm request that reuses the prefix cache of the last main-model request
  • Five-heading Markdown handoff: Goal, User instructions, Current state, Key facts, Next step
  • Verbatim tail of recent rows kept after compaction
  • Model-based fallback summary through the auxiliary route
  • Works for /compress and automatic compaction
  • One warm_compaction log line per compaction in logs/agent.log

When to use it

  • Shortening compaction in long sessions on a local vLLM, SGLang or llama.cpp server
  • Compacting on a hosted API that offers prompt caching
  • Checking cache behavior through the cached_tokens value in the log

Who it is for: Hermes Agent users with long sessions who run a model server or API that supports prompt prefix caching.

How it fits with Hermes Agent

Built for Hermes Agent as a context engine plugin selected with context.engine, using only documented plugin APIs.

How to install Hermes Warm Compaction

These commands are copied from the project's README. Check the repository for the latest steps before you run them.

hermes plugins install Elevatormusic/hermes-warm-compaction#warm_compaction --enable
hermes config set context.engine warm_compaction

Requirements: Hermes Agent 45871e10 (2026-10-02) or later; for the warm path a chat_completions, codex_responses or anthropic_messages route; a server with prefix caching and the same main model for the speedup

Note: The speedup depends on a server that reuses its prefix cache, and the README says it was tested only on DGX and LM Studio.

FAQ

What is Hermes Warm Compaction?

Hermes Warm Compaction is a context engine plugin for Hermes Agent. At each compaction the main model writes the handoff summary on the cached prefix of its last request, so the server reads only the new rows.

How do I install Hermes Warm Compaction?

Run hermes plugins install Elevatormusic/hermes-warm-compaction#warm_compaction --enable, then hermes config set context.engine warm_compaction. Start a new Hermes session afterward.

Does Hermes Warm Compaction need a prefix cache?

Yes, for the speedup. The warm request is faster only when the server reuses its prefix cache, and the fallback summary can use a different model.

Similar plugins for Hermes Agent

All plugins

Related guides: How to install Hermes Agent · How to run Hermes Agent securely