HermesBench
am423/hermes-bench-tool-call
Benchmark for local models running inside the Hermes Agent harness, with traces and hardware telemetry
HermesBench is a benchmark that scores local models on the tool-calling patterns Hermes Agent users actually hit, capturing full traces and hardware telemetry for each task run.
What HermesBench does
HermesBench measures how well a local model works inside the real Hermes Agent harness, not just how well it generates text. The suite has 61 tasks: 48 core tasks across 11 categories (terminal smoke, file read, patch, search, write, process, todo, execute_code, web_lookup, memory and error_recovery), 3 real-world integration tasks and 10 small coding tasks. Each run launches the actual AIAgent from ~/.hermes/hermes-agent/ in a subprocess with a custom tmux_isolated environment backend.
Every task run produces three artifacts: trace.jsonl with each message, tool call, result, reasoning and token IDs; trace.cast, an asciinema recording of the terminal session; and stats.jsonl, 5 Hz hardware telemetry covering CPU, GPU, RAM, NVMe and host power. Verifiers are deterministic and use only the standard library, with no LLM-as-judge. Per-model summaries report pass rate, tool efficiency, token efficiency, wall clock, recovery rate and format compliance, plus GPU power and temperature. An export command turns traces into SFT JSONL with loss masks.
Key features
- 61 tasks across terminal, file, search, memory, error recovery and coding categories
- Runs the real Hermes AIAgent with a tmux_isolated environment backend
- Captures trace.jsonl, an asciinema trace.cast and 5 Hz hardware telemetry per task
- Deterministic standard-library verifiers with no LLM-as-judge
- Report command that builds REPORT.md and a video timeline
- export-sft command that converts traces into SFT JSONL with loss masks
When to use it
- Comparing local models on how well they chain Hermes tools
- Collecting full reasoning and tool traces as fine-tuning data
- Checking thermals and power draw while a model works through agent tasks
Who it is for: People running local models with Hermes Agent who want reproducible tool-calling scores and trace data.
How it fits with Hermes Agent
Built around Hermes Agent: it drives the real Hermes AIAgent through a purpose-built benchmark CLI named hermesbench.
How to install HermesBench
These commands are copied from the project's README. Check the repository for the latest steps before you run them.
git clone https://github.com/am423/hermes-bench-tool-call.git
cd hermes-bench-tool-call
./scripts/bootstrap.sh
source .venv/bin/activateRequirements: A Hermes Agent install at ~/.hermes/hermes-agent/ and a local model to benchmark
Note: The project describes itself as benchmark v0.1, and the repository has no license file.
FAQ
What is HermesBench?
HermesBench is a benchmark for local models running inside the Hermes Agent harness. It runs 61 tasks and records traces, terminal recordings and hardware telemetry for each one.
How do I install HermesBench?
Clone the repository, run ./scripts/bootstrap.sh to create the virtual environment, then activate it with source .venv/bin/activate. Running hermesbench doctor --install fixes missing Python dependencies.
What does HermesBench record for each task run?
It records a trace.jsonl conversation trace, a trace.cast asciinema recording and a stats.jsonl file of 5 Hz hardware telemetry. The trace includes every tool call, result, reasoning and token IDs.
Similar research for Hermes Agent
All researchAsync benchmark that simulates many agents hitting one OpenAI-compatible endpoint such as vLLM
beardthelion Hermes Skill DistillationHackathon environment that turns Hermes Agent task runs into scored trajectories for Hermes 4 training
vcruz305 hermes-agentic-benchAgentic test batteries that check whether local models can chain Hermes tools and stop before the cap
shuklabhay llm-synesthesiaPipeline that reads emotions from Hermes-4.3-36B hidden states and renders them as a live WebGL visual
ctala AI Benchmarks AlternativosOpen Spanish-language benchmark of LLMs for business and agent use, scored by an independent Phi-4 judge
ohikava BitGN ECOM AgentHermes-based agent for the BitGN E-commerce benchmark, locked to one MCP tool channel
Related guides: What is Hermes Agent?