Hermes Atlas
Research, training & evaluation · works with Hermes Agent

Hermes SkillEval

Raidriar7170/hermes-skilleval

Skill-routing evaluation and release-gate toolkit for SKILL.md agent skills

In short

Hermes SkillEval is a Python toolkit that tests whether an agent picks the right SKILL.md skill, including near-miss skills with the wrong responsibility, and blocks risky router changes before release. It targets Hermes-style agent skills, and the README also names Claude Code, Codex and Cursor.

What Hermes SkillEval does

Hermes SkillEval addresses a problem that appears when an agent has many similar SKILL.md skills: choosing the right one while rejecting close neighbours that have a different job. It parses SKILL.md files into an index, loads a benchmark with gold and negative labels, and compares five router types: keyword, hybrid, embedding, verification-gated and cross-encoder. Confusion mining finds hard negatives, which feed a SentenceTransformer router trained on a frozen MiniLM model.

A reusable GitHub Action and a release selector act as a gate, so a candidate router that regresses does not become the default. In the README's pilot on a self-built set of 16 held-out tasks, Recall@1 for the trained router rose from 12/16 to 16/16. The default decision still stays KEEP_BASELINE because the blind validation did not run. The author labels the work a model-only pilot, not a public benchmark.

Key features

  • SKILL.md parser and portable skill index
  • Benchmark format with gold and negative skill labels
  • Five router types behind one evaluation interface
  • Hard-negative mining and SentenceTransformer router training
  • Paired multi-seed evaluation with failure slices
  • GitHub Action and release selector that block regressing routers

When to use it

  • Checking whether an agent picks the right skill when many skills look alike
  • Gating a new skill router in CI before it becomes the default
  • Finding which pairs of skills get confused with each other

Who it is for: Developers who maintain large collections of SKILL.md skills, including Hermes-style ones, and want measurable routing quality.

How it fits with Hermes Agent

The project is named for Hermes and describes its benchmark as Hermes-style. The README names Claude Code, Codex and Cursor as target agents and does not describe a runtime integration with Hermes Agent.

Note: The README calls the work a model-only pilot with no independent human review, and says the default router decision remains KEEP_BASELINE and the result is not release eligible.

FAQ

What is Hermes SkillEval?

Hermes SkillEval is a Python toolkit for evaluating skill routing. It measures whether an agent selects the correct SKILL.md skill, mines confusing neighbours as hard negatives and gates router releases in CI.

Does Hermes SkillEval work with Hermes Agent?

It works with Hermes-style SKILL.md skills, but the README does not describe a direct runtime integration with Hermes Agent. It names Claude Code, Codex and Cursor as the agents it targets.

Is Hermes SkillEval free and open source?

Yes. The repository is released under the MIT license.

Similar research for Hermes Agent

All research

Related guides: What is Hermes Agent?