Hermes Atlas
Research, training & evaluation · works with Hermes Agent

Caliper

edonadei/caliper

Evaluate whether agent skills help by running tasks with and without them, on Hermes and other agents

In short

Caliper is a Python evaluation tool that runs the same tasks with and without a skill, MCP server or rule, then reports whether it helps and what it costs in tokens. It supports Claude Code, Codex, Pi and Hermes as engines.

What Caliper does

A Caliper spec is an .eval.yaml file next to your skill. It lists the skills to install, including neighbours that might compete for the same prompts, and the tasks to run. Each task has a prompt, optional setup and at least one check: expect is graded by an LLM judge, assert runs locally as Python with a 30-second limit, and activates asserts which skills the agent chose to load without needing a judge.

Skills are installed where the agent looks rather than pasted into the prompt. The --ablate flag re-runs tasks without the skill so you can see whether it earns its context, and activation is scored separately from whether the work succeeded. Reports lead with single-run success rate instead of pass@k, and full results are saved as JSON under .caliper/results. The engine is chosen at run time among Claude Code, Codex, Pi and Hermes.

Key features

  • Spec file (.eval.yaml) with tasks checked by an LLM judge, local Python asserts or skill-activation checks
  • Installs skills where the agent looks instead of pasting them into the prompt
  • --ablate flag re-runs tasks without the skill to test whether it earns its context
  • Activation scoreboard scored separately from task success
  • Neighbourhoods of competing skills, with assertions about which one should win each prompt
  • Engine chosen at run time among Claude Code, Codex, Pi and Hermes

When to use it

  • Checking that a SKILL.md still works after a model update
  • Measuring a skill's token cost against the bare agent
  • Finding out whether two skills compete for the same prompts

Who it is for: Authors of agent skills and teams who want evidence that a skill, MCP server or rule actually helps.

How it fits with Hermes Agent

Hermes is one of four selectable engines (Claude Code, Codex, Pi and Hermes), and the repository carries the hermes-skill topic.

How to install Caliper

These commands are copied from the project's README. Check the repository for the latest steps before you run them.

npx skills@latest add edonadei/caliper
pipx install caliper-eval

FAQ

What is Caliper?

Caliper is a Python tool for testing agent skills. It runs the same tasks with and without a skill and reports whether the skill helps and what it costs in tokens.

Does Caliper work with Hermes Agent?

Yes, Hermes is one of four supported engines alongside Claude Code, Codex and Pi. A spec never names an engine, and you pick one at run time.

How do I install Caliper?

Run pipx install caliper-eval to install the command-line tool. To add it as a skill for your agent, run npx skills@latest add edonadei/caliper.

Similar research for Hermes Agent

All research

Related guides: What is Hermes Agent?