Hermes Best Models
fox-in-the-box-ai/hermes-best-models
Monthly benchmark that tests LLMs on Hermes-style agent tasks such as tool calling and delegation
Hermes Best Models is a Python evaluation script and monthly benchmark that runs LLMs through 25 agent tasks to find which models suit Hermes agents. It covers tool calling, multi-step reasoning, delegation and failure handling, with results published on the Fox in the Box blog.
What Hermes Best Models does
The benchmark runs 25 tasks in five categories: simple QA (5), multi-step tool use (7, using real function calls through the OpenRouter tools API), delegation and reasoning (5), reasoning (5) and failure modes (3), which checks refusal and clarification. Models run in parallel threads, and the README expects a run of 15 to 25 minutes costing $2 to $5. The May 2026 run covered 19 models at a total cost of $2.70.
Scores combine pass rate at 50 percent, latency at 20 percent, token efficiency at 15 percent and throughput at 15 percent, normalized across the field. The README says scores within 5 points are roughly tied because 25 tasks is a sample. You add or remove models in config.yaml using OpenRouter slugs, and raw results land in eval_results as timestamped JSON. The benchmark comes from Fox in the Box, a self-hosted, privacy-first AI agent platform.
Key features
- 25 tasks across simple QA, tool use, delegation, reasoning and failure modes
- A composite score of pass rate, latency, token efficiency and throughput
- Models run in parallel through OpenRouter slugs set in config.yaml
- Timestamped raw and aggregated JSON results
- A monthly refresh with the ranking published on the Fox in the Box blog
When to use it
- Picking a model for a Hermes agent based on tool-calling and delegation results
- Adding a new model to the field and re-running the benchmark yourself
- Estimating the cost and latency trade-off between models on agent workloads
Who it is for: People choosing or comparing LLMs for Hermes agents who want task-based results instead of general knowledge benchmarks.
How it fits with Hermes Agent
It is built to evaluate models for Hermes agents, and its README and title frame the 25 tasks as real Hermes agent workloads.
How to install Hermes Best Models
These commands are copied from the project's README. Check the repository for the latest steps before you run them.
git clone https://github.com/fox-in-the-box-ai/hermes-best-models.git
cd hermes-best-models
python3 -m venv venv && source venv/bin/activate
pip install -r requirements.txtRequirements: Python 3, git and an OpenRouter API key added to config.yaml; a full run costs about $2 to $5 depending on the models included.
Note: The README says scores within 5 points are roughly tied since 25 tasks is a sample, and running the benchmark needs an OpenRouter key and costs about $2 to $5.
FAQ
What is Hermes Best Models?
Hermes Best Models is a monthly LLM benchmark for Hermes agents. It tests tool calling, multi-step reasoning, delegation and failure-mode handling instead of general knowledge questions.
How is a model scored in Hermes Best Models?
The composite score runs from 0 to 100: 50 percent pass rate, 20 percent latency, 15 percent token efficiency and 15 percent throughput. The README treats scores within 5 points as roughly tied.
How do I reproduce the Hermes Best Models results?
Clone the repository, create a virtual environment, install requirements.txt, copy config.example.yaml to config.yaml, add an OpenRouter API key, and run python3 eval_harness.py.
Similar research for Hermes Agent
All researchBenchmark report ranking local LLMs for Hermes Agent-style tool use on DGX Spark class hardware
EngTurtle Hermes MemConflict BenchmarkBenchmark comparing self-hostable memory providers for Hermes Agent on the MemConflict dataset
digitalspaceport Multi Agent Benchmark ToolAsync benchmark that simulates many agents hitting one OpenAI-compatible endpoint such as vLLM
am423 HermesBenchBenchmark for local models running inside the Hermes Agent harness, with traces and hardware telemetry
beardthelion Hermes Skill DistillationHackathon environment that turns Hermes Agent task runs into scored trajectories for Hermes 4 training
vcruz305 hermes-agentic-benchAgentic test batteries that check whether local models can chain Hermes tools and stop before the cap
Related guides: What is Hermes Agent?