Caliper
edonadei/caliper
Evaluate whether agent skills help by running tasks with and without them, on Hermes and other agents
Caliper is a Python evaluation tool that runs the same tasks with and without a skill, MCP server or rule, then reports whether it helps and what it costs in tokens. It supports Claude Code, Codex, Pi and Hermes as engines.
What Caliper does
A Caliper spec is an .eval.yaml file next to your skill. It lists the skills to install, including neighbours that might compete for the same prompts, and the tasks to run. Each task has a prompt, optional setup and at least one check: expect is graded by an LLM judge, assert runs locally as Python with a 30-second limit, and activates asserts which skills the agent chose to load without needing a judge.
Skills are installed where the agent looks rather than pasted into the prompt. The --ablate flag re-runs tasks without the skill so you can see whether it earns its context, and activation is scored separately from whether the work succeeded. Reports lead with single-run success rate instead of pass@k, and full results are saved as JSON under .caliper/results. The engine is chosen at run time among Claude Code, Codex, Pi and Hermes.
Key features
- Spec file (.eval.yaml) with tasks checked by an LLM judge, local Python asserts or skill-activation checks
- Installs skills where the agent looks instead of pasting them into the prompt
- --ablate flag re-runs tasks without the skill to test whether it earns its context
- Activation scoreboard scored separately from task success
- Neighbourhoods of competing skills, with assertions about which one should win each prompt
- Engine chosen at run time among Claude Code, Codex, Pi and Hermes
When to use it
- Checking that a SKILL.md still works after a model update
- Measuring a skill's token cost against the bare agent
- Finding out whether two skills compete for the same prompts
Who it is for: Authors of agent skills and teams who want evidence that a skill, MCP server or rule actually helps.
How it fits with Hermes Agent
Hermes is one of four selectable engines (Claude Code, Codex, Pi and Hermes), and the repository carries the hermes-skill topic.
How to install Caliper
These commands are copied from the project's README. Check the repository for the latest steps before you run them.
npx skills@latest add edonadei/caliper
pipx install caliper-eval FAQ
What is Caliper?
Caliper is a Python tool for testing agent skills. It runs the same tasks with and without a skill and reports whether the skill helps and what it costs in tokens.
Does Caliper work with Hermes Agent?
Yes, Hermes is one of four supported engines alongside Claude Code, Codex and Pi. A spec never names an engine, and you pick one at run time.
How do I install Caliper?
Run pipx install caliper-eval to install the command-line tool. To add it as a skill for your agent, run npx skills@latest add edonadei/caliper.
Similar research for Hermes Agent
All researchVersioned archive of system prompts from agent CLIs including Claude Code, Codex and Hermes
InternLM WildClawBenchBenchmark of 60 hand-built real-world tasks run across OpenClaw, Claude Code, Codex CLI and Hermes Agent
KhanCold MerchantBench365-day simulated e-commerce benchmark for LLM agents, with a Hermes adapter
Raidriar7170 Hermes SkillEvalSkill-routing evaluation and release-gate toolkit for SKILL.md agent skills
Q00 RLM-ForgeRecursive Language Model runtime for Hermes Agent with Ouroboros recursion and TraceGuard evidence gating
howdymary Hermes Agent Meta-HarnessOuter-loop optimizer that searches over Hermes' benchmark harness, not model weights
Related guides: What is Hermes Agent?