Hermes Agent Meta-Harness
howdymary/hermes-agent-metaharness
Outer-loop optimizer that searches over Hermes' benchmark harness, not model weights
Hermes Agent Meta-Harness is a standalone outer-loop optimizer that treats hermes-agent as the execution backend for benchmark candidates, searching over harness code rather than model weights to improve coding-benchmark performance.
What Hermes Agent Meta-Harness does
The project is a direct application of the Meta-Harness paper's core argument, that system quality depends on the harness code deciding what context is collected, stored and shown to the model, not only on model weights. hermes-agent owns the inner runtime (candidate protocol, benchmark integration and archive writing) for TBLite and TB2 coding benchmarks, while hermes-agent-metaharness owns the outer loop: candidate evaluation, archive analysis, baseline reuse, frontier tracking and search.
Search is intentionally conservative in the current release: it generates deterministic wrapper candidates around a seed candidate rather than rewriting Hermes' core, and a simple JSON-backed frontier with cross-platform locking tracks the best candidates found so far. The current scope targets verifiable coding benchmarks specifically, not general production chat behavior, and the production runtime never self-modifies.
Key features
- Outer-loop search over benchmark harness code, not model weights
- Paired baseline-vs-candidate evaluation with task-set comparability checks
- Archive parsing for manifest, summary and per-task JSON records
- JSON-backed frontier tracking with cross-platform locking
- Deterministic wrapper-mutation search with persisted dry-run summaries
When to use it
- Benchmarking whether a change to Hermes' harness code improves TBLite or TB2 scores
- Comparing a new harness candidate against a reused baseline or the current frontier-best
- Running a conservative, wrapper-only harness search without touching Hermes' core
Who it is for: Researchers and Hermes contributors evaluating harness-level changes against coding benchmarks like TBLite and TB2.
How it fits with Hermes Agent
Hermes Agent Meta-Harness treats hermes-agent as its required inner execution runtime, checking that a given Hermes checkout exposes the Meta-Harness benchmark surface before running any evaluation.
How to install Hermes Agent Meta-Harness
These commands are copied from the project's README. Check the repository for the latest steps before you run them.
git clone https://github.com/howdymary/hermes-agent-metaharness.git
cd hermes-agent-metaharness
pip install -e ".[dev]"Requirements: A hermes-agent checkout, pointed to by HERMES_AGENT_REPO, a sibling ../hermes-agent directory, or ~/.hermes/hermes-agent
Note: Candidate search is currently limited to deterministic wrapper mutations around a seed candidate, not full harness rewriting.
FAQ
What is Hermes Agent Meta-Harness?
It is an outer-loop optimizer, inspired by the Meta-Harness paper, that searches over Hermes' benchmark harness code to improve coding-benchmark results, rather than changing model weights.
Does it work with Hermes Agent?
Yes, it requires a hermes-agent checkout as its execution backend and checks that the checkout exposes the Meta-Harness benchmark surface before running.
How do I install Hermes Agent Meta-Harness?
Clone the repository and run pip install -e ".[dev]", then point it at a Hermes checkout with the HERMES_AGENT_REPO environment variable or a sibling directory.
Similar research for Hermes Agent
All researchEvaluate whether agent skills help by running tasks with and without them, on Hermes and other agents
KhanCold MerchantBench365-day simulated e-commerce benchmark for LLM agents, with a Hermes adapter
Raidriar7170 Hermes SkillEvalSkill-routing evaluation and release-gate toolkit for SKILL.md agent skills
Q00 RLM-ForgeRecursive Language Model runtime for Hermes Agent with Ouroboros recursion and TraceGuard evidence gating
MiaAI-Lab Best Local Model for Agentic Workflows 2026Benchmark report ranking local LLMs for Hermes Agent-style tool use on DGX Spark class hardware
EngTurtle Hermes MemConflict BenchmarkBenchmark comparing self-hostable memory providers for Hermes Agent on the MemConflict dataset
Related guides: What is Hermes Agent?