Hermes Atlas
Research, training & evaluation

Hermes Best Models

fox-in-the-box-ai/hermes-best-models

Monthly benchmark that tests LLMs on Hermes-style agent tasks such as tool calling and delegation

In short

Hermes Best Models is a Python evaluation script and monthly benchmark that runs LLMs through 25 agent tasks to find which models suit Hermes agents. It covers tool calling, multi-step reasoning, delegation and failure handling, with results published on the Fox in the Box blog.

What Hermes Best Models does

The benchmark runs 25 tasks in five categories: simple QA (5), multi-step tool use (7, using real function calls through the OpenRouter tools API), delegation and reasoning (5), reasoning (5) and failure modes (3), which checks refusal and clarification. Models run in parallel threads, and the README expects a run of 15 to 25 minutes costing $2 to $5. The May 2026 run covered 19 models at a total cost of $2.70.

Scores combine pass rate at 50 percent, latency at 20 percent, token efficiency at 15 percent and throughput at 15 percent, normalized across the field. The README says scores within 5 points are roughly tied because 25 tasks is a sample. You add or remove models in config.yaml using OpenRouter slugs, and raw results land in eval_results as timestamped JSON. The benchmark comes from Fox in the Box, a self-hosted, privacy-first AI agent platform.

Key features

  • 25 tasks across simple QA, tool use, delegation, reasoning and failure modes
  • A composite score of pass rate, latency, token efficiency and throughput
  • Models run in parallel through OpenRouter slugs set in config.yaml
  • Timestamped raw and aggregated JSON results
  • A monthly refresh with the ranking published on the Fox in the Box blog

When to use it

  • Picking a model for a Hermes agent based on tool-calling and delegation results
  • Adding a new model to the field and re-running the benchmark yourself
  • Estimating the cost and latency trade-off between models on agent workloads

Who it is for: People choosing or comparing LLMs for Hermes agents who want task-based results instead of general knowledge benchmarks.

How it fits with Hermes Agent

It is built to evaluate models for Hermes agents, and its README and title frame the 25 tasks as real Hermes agent workloads.

How to install Hermes Best Models

These commands are copied from the project's README. Check the repository for the latest steps before you run them.

git clone https://github.com/fox-in-the-box-ai/hermes-best-models.git
cd hermes-best-models
python3 -m venv venv && source venv/bin/activate
pip install -r requirements.txt

Requirements: Python 3, git and an OpenRouter API key added to config.yaml; a full run costs about $2 to $5 depending on the models included.

Note: The README says scores within 5 points are roughly tied since 25 tasks is a sample, and running the benchmark needs an OpenRouter key and costs about $2 to $5.

FAQ

What is Hermes Best Models?

Hermes Best Models is a monthly LLM benchmark for Hermes agents. It tests tool calling, multi-step reasoning, delegation and failure-mode handling instead of general knowledge questions.

How is a model scored in Hermes Best Models?

The composite score runs from 0 to 100: 50 percent pass rate, 20 percent latency, 15 percent token efficiency and 15 percent throughput. The README treats scores within 5 points as roughly tied.

How do I reproduce the Hermes Best Models results?

Clone the repository, create a virtual environment, install requirements.txt, copy config.example.yaml to config.yaml, add an OpenRouter API key, and run python3 eval_harness.py.

Similar research for Hermes Agent

All research

Related guides: What is Hermes Agent?