Hermes MemConflict Benchmark
EngTurtle/hermes-memconflict
Benchmark comparing self-hostable memory providers for Hermes Agent on the MemConflict dataset
Hermes MemConflict Benchmark is a benchmark project that compares self-hostable long-term memory providers for Hermes Agent when a user's stored facts conflict. It helps Hermes users choose a memory backend using one shared harness and scorer.
What Hermes MemConflict Benchmark does
Hermes MemConflict Benchmark compares self-hostable long-term memory providers for Hermes Agent on the MemConflict benchmark. The benchmark tests whether a provider retrieves and uses the stored memory that is temporally valid, factually correct and contextually applicable when a user's facts conflict across multi-session dialogues. The harness uses MemConflict Step4_4.jsonl, which has 30 personas and 3,750 questions.
Every provider runs the same harness contract and emits a model answer and the retrieved memories for each question. One shared, provider-agnostic scorer judges the output, and the headline metric is macro answer accuracy averaged evenly across conflict categories. Adapters cover Mnemosyne, Hindsight, mem0, Supermemory, Honcho, OpenViking and RetainDB, each in the best-effort configuration a real deployment would use. Results are published on an interactive report site with a browser for the MemConflict dialogues, and the docs record design decisions and troubleshooting notes.
Key features
- Shared harness with a fixed dataset, answer model, judge model, top-K, prompts and scorer
- Adapters for Mnemosyne, Hindsight, mem0, Supermemory, Honcho, OpenViking and RetainDB
- Macro answer accuracy as the headline metric across conflict categories
- Interactive benchmark report and conversation browser published as a site
- Container stack and docs covering decisions, troubleshooting and the benchmark matrix
When to use it
- Pick a self-hosted memory provider for a Hermes Agent deployment based on measured results
- Reproduce the MemConflict runs with the provided container stack
- Study how memory systems handle contradictory facts across sessions
Who it is for: Hermes Agent users and memory-provider developers who want evidence-based comparisons of long-term memory backends.
How it fits with Hermes Agent
Built around Hermes Agent: the project selects a memory provider for Hermes to deploy and pins hermes-agent as a submodule.
FAQ
What is Hermes MemConflict Benchmark?
It is a project that compares self-hostable long-term memory providers for Hermes Agent on the MemConflict benchmark. It scores how well each provider handles stored facts that conflict across multi-session dialogues.
Does Hermes MemConflict Benchmark work with Hermes Agent?
Yes. The project exists to choose a memory provider for Hermes to deploy, and its external folder pins hermes-agent as a submodule next to the MemConflict dataset.
Is Hermes MemConflict Benchmark free and open source?
Yes. The repository is released under the MIT license.
Similar research for Hermes Agent
All researchBenchmark report ranking local LLMs for Hermes Agent-style tool use on DGX Spark class hardware
fox-in-the-box-ai Hermes Best ModelsMonthly benchmark that tests LLMs on Hermes-style agent tasks such as tool calling and delegation
digitalspaceport Multi Agent Benchmark ToolAsync benchmark that simulates many agents hitting one OpenAI-compatible endpoint such as vLLM
am423 HermesBenchBenchmark for local models running inside the Hermes Agent harness, with traces and hardware telemetry
beardthelion Hermes Skill DistillationHackathon environment that turns Hermes Agent task runs into scored trajectories for Hermes 4 training
vcruz305 hermes-agentic-benchAgentic test batteries that check whether local models can chain Hermes tools and stop before the cap
Related guides: What is Hermes Agent?