hermes-compression-eval
NousResearch/hermes-compression-eval Published by Nous Research
Offline probe-based evaluation harness for the Hermes Agent ContextCompressor
hermes-compression-eval is an offline evaluation harness from Nous Research for the ContextCompressor in hermes-agent. It runs conversation fixtures through the compressor and has a judge model score the results on six dimensions.
What hermes-compression-eval does
hermes-compression-eval is an offline evaluation harness for agent/context_compressor.py in hermes-agent, which decides what survives when a session exceeds the context-window threshold. It runs a real conversation fixture through ContextCompressor.compress(), asks the compressor model to answer probe questions from the compressed state, and has a judge model score each answer from 0 to 5 on six dimensions: accuracy, context_awareness, artifact_trail, completeness, continuity and instruction_following.
The methodology is adapted from Factory's December 2025 write-up Evaluating Compression, without its scoreboard framing. The aim is a signal between a green test suite and a bad summary in production: edit the compressor prompt, rerun the eval, and compare per-dimension scores with a saved baseline using --compare-to. It ships three scrubbed session fixtures (feature-impl, debug, config-build), three probe banks of 10 to 11 probes covering recall, artifact, continuation and decision, a scrub_fixtures.py pipeline for turning real session files into public-safe fixtures, and 33 hermetic unit tests for the non-LLM paths. Each run writes a report.md that is ready to paste into a PR body.
Key features
- Six-dimension judge rubric scored 0 to 5 per probe answer
- Three scrubbed session fixtures with matching probe banks
- Baseline comparison with per-dimension deltas through --compare-to
- scrub_fixtures.py to turn real ~/.hermes/sessions files into public-safe fixtures
- run_eval.py command line built on Fire
- 33 hermetic unit tests for the non-LLM code paths
When to use it
- Checking whether a prompt change in context_compressor.py lowers summary quality
- Building new fixtures from your own Hermes sessions
- Attaching a generated report.md to a pull request
Who it is for: Hermes Agent contributors who modify the context compressor and want measured before-and-after scores.
How it fits with Hermes Agent
An official Nous Research repository that imports ContextCompressor and agent.redact from a hermes-agent checkout.
How to install hermes-compression-eval
These commands are copied from the project's README. Check the repository for the latest steps before you run them.
git clone https://github.com/NousResearch/hermes-compression-eval.git
cd hermes-compression-eval
pip install -r requirements.txtRequirements: A hermes-agent checkout located through HERMES_AGENT_ROOT, ~/.hermes/hermes-agent/ or a sibling directory; the openai and fire packages; and a configured LLM provider for the compressor and judge models
Note: Runs are LLM-graded and non-deterministic, with about 30 probe pairs across the three fixtures at default settings, so they use provider credits and are not suited to CI.
FAQ
What is hermes-compression-eval?
It is an offline evaluation harness for the ContextCompressor in hermes-agent. A judge model scores answers to probe questions asked from the compressed conversation state.
How do I install hermes-compression-eval?
Clone the repository and run pip install -r requirements.txt. The harness also needs a hermes-agent checkout, found through HERMES_AGENT_ROOT, the default ~/.hermes/hermes-agent/ location or a sibling directory.
What do I need to run hermes-compression-eval?
You need a hermes-agent checkout, the openai and fire Python packages, and an LLM provider for both the compressor and judge models. Each probe costs one continuation call and one grading call.
Similar official for Hermes Agent
All officialHermes Kanban demo where four agent profiles make a video explaining Hermes Kanban from one command
NousResearch Hermes Telegram BusinessTelegram Business Mode secretary-bot plugin where every drafted reply needs owner approval
NousResearch Hermes Example PluginsReference plugins from Nous Research showing each Hermes Agent plugin surface in minimal code
NousResearch Hermes Memory WikiHermes dashboard plugin with a browsable session-history wiki and a read-only memory audit panel
NousResearch Forge API DemoNous Research demo CLI for Forge API reasoning completions, with Hermes-3-70b as orchestrator
NousResearch hermes-plugin-snykHermes Agent plugin that runs Snyk code, dependency, container and IaC scans through Snyk's MCP server
Related guides: What is Hermes Agent? · How to install Hermes Agent