Hermes SkillEval
Raidriar7170/hermes-skilleval
Skill-routing evaluation and release-gate toolkit for SKILL.md agent skills
Hermes SkillEval is a Python toolkit that tests whether an agent picks the right SKILL.md skill, including near-miss skills with the wrong responsibility, and blocks risky router changes before release. It targets Hermes-style agent skills, and the README also names Claude Code, Codex and Cursor.
What Hermes SkillEval does
Hermes SkillEval addresses a problem that appears when an agent has many similar SKILL.md skills: choosing the right one while rejecting close neighbours that have a different job. It parses SKILL.md files into an index, loads a benchmark with gold and negative labels, and compares five router types: keyword, hybrid, embedding, verification-gated and cross-encoder. Confusion mining finds hard negatives, which feed a SentenceTransformer router trained on a frozen MiniLM model.
A reusable GitHub Action and a release selector act as a gate, so a candidate router that regresses does not become the default. In the README's pilot on a self-built set of 16 held-out tasks, Recall@1 for the trained router rose from 12/16 to 16/16. The default decision still stays KEEP_BASELINE because the blind validation did not run. The author labels the work a model-only pilot, not a public benchmark.
Key features
- SKILL.md parser and portable skill index
- Benchmark format with gold and negative skill labels
- Five router types behind one evaluation interface
- Hard-negative mining and SentenceTransformer router training
- Paired multi-seed evaluation with failure slices
- GitHub Action and release selector that block regressing routers
When to use it
- Checking whether an agent picks the right skill when many skills look alike
- Gating a new skill router in CI before it becomes the default
- Finding which pairs of skills get confused with each other
Who it is for: Developers who maintain large collections of SKILL.md skills, including Hermes-style ones, and want measurable routing quality.
How it fits with Hermes Agent
The project is named for Hermes and describes its benchmark as Hermes-style. The README names Claude Code, Codex and Cursor as target agents and does not describe a runtime integration with Hermes Agent.
Note: The README calls the work a model-only pilot with no independent human review, and says the default router decision remains KEEP_BASELINE and the result is not release eligible.
FAQ
What is Hermes SkillEval?
Hermes SkillEval is a Python toolkit for evaluating skill routing. It measures whether an agent selects the correct SKILL.md skill, mines confusing neighbours as hard negatives and gates router releases in CI.
Does Hermes SkillEval work with Hermes Agent?
It works with Hermes-style SKILL.md skills, but the README does not describe a direct runtime integration with Hermes Agent. It names Claude Code, Codex and Cursor as the agents it targets.
Is Hermes SkillEval free and open source?
Yes. The repository is released under the MIT license.
Similar research for Hermes Agent
All researchEvaluate whether agent skills help by running tasks with and without them, on Hermes and other agents
KhanCold MerchantBench365-day simulated e-commerce benchmark for LLM agents, with a Hermes adapter
Q00 RLM-ForgeRecursive Language Model runtime for Hermes Agent with Ouroboros recursion and TraceGuard evidence gating
howdymary Hermes Agent Meta-HarnessOuter-loop optimizer that searches over Hermes' benchmark harness, not model weights
MiaAI-Lab Best Local Model for Agentic Workflows 2026Benchmark report ranking local LLMs for Hermes Agent-style tool use on DGX Spark class hardware
EngTurtle Hermes MemConflict BenchmarkBenchmark comparing self-hostable memory providers for Hermes Agent on the MemConflict dataset
Related guides: What is Hermes Agent?