Hermes Atlas
Research, training & evaluation

HermesBench

am423/hermes-bench-tool-call

Benchmark for local models running inside the Hermes Agent harness, with traces and hardware telemetry

In short

HermesBench is a benchmark that scores local models on the tool-calling patterns Hermes Agent users actually hit, capturing full traces and hardware telemetry for each task run.

What HermesBench does

HermesBench measures how well a local model works inside the real Hermes Agent harness, not just how well it generates text. The suite has 61 tasks: 48 core tasks across 11 categories (terminal smoke, file read, patch, search, write, process, todo, execute_code, web_lookup, memory and error_recovery), 3 real-world integration tasks and 10 small coding tasks. Each run launches the actual AIAgent from ~/.hermes/hermes-agent/ in a subprocess with a custom tmux_isolated environment backend.

Every task run produces three artifacts: trace.jsonl with each message, tool call, result, reasoning and token IDs; trace.cast, an asciinema recording of the terminal session; and stats.jsonl, 5 Hz hardware telemetry covering CPU, GPU, RAM, NVMe and host power. Verifiers are deterministic and use only the standard library, with no LLM-as-judge. Per-model summaries report pass rate, tool efficiency, token efficiency, wall clock, recovery rate and format compliance, plus GPU power and temperature. An export command turns traces into SFT JSONL with loss masks.

Key features

  • 61 tasks across terminal, file, search, memory, error recovery and coding categories
  • Runs the real Hermes AIAgent with a tmux_isolated environment backend
  • Captures trace.jsonl, an asciinema trace.cast and 5 Hz hardware telemetry per task
  • Deterministic standard-library verifiers with no LLM-as-judge
  • Report command that builds REPORT.md and a video timeline
  • export-sft command that converts traces into SFT JSONL with loss masks

When to use it

  • Comparing local models on how well they chain Hermes tools
  • Collecting full reasoning and tool traces as fine-tuning data
  • Checking thermals and power draw while a model works through agent tasks

Who it is for: People running local models with Hermes Agent who want reproducible tool-calling scores and trace data.

How it fits with Hermes Agent

Built around Hermes Agent: it drives the real Hermes AIAgent through a purpose-built benchmark CLI named hermesbench.

How to install HermesBench

These commands are copied from the project's README. Check the repository for the latest steps before you run them.

git clone https://github.com/am423/hermes-bench-tool-call.git
cd hermes-bench-tool-call
./scripts/bootstrap.sh
source .venv/bin/activate

Requirements: A Hermes Agent install at ~/.hermes/hermes-agent/ and a local model to benchmark

Note: The project describes itself as benchmark v0.1, and the repository has no license file.

FAQ

What is HermesBench?

HermesBench is a benchmark for local models running inside the Hermes Agent harness. It runs 61 tasks and records traces, terminal recordings and hardware telemetry for each one.

How do I install HermesBench?

Clone the repository, run ./scripts/bootstrap.sh to create the virtual environment, then activate it with source .venv/bin/activate. Running hermesbench doctor --install fixes missing Python dependencies.

What does HermesBench record for each task run?

It records a trace.jsonl conversation trace, a trace.cast asciinema recording and a stats.jsonl file of 5 Hz hardware telemetry. The trace includes every tool call, result, reasoning and token IDs.

Similar research for Hermes Agent

All research

Related guides: What is Hermes Agent?