Hermes Atlas
Research, training & evaluation

hermes-agentic-bench

vcruz305/hermes-agentic-bench

Agentic test batteries that check whether local models can chain Hermes tools and stop before the cap

In short

hermes-agentic-bench is a set of small test batteries that score local models on chaining Hermes Agent tools, recovering from failed calls and stopping before the consecutive-tool cap.

What hermes-agentic-bench does

hermes-agentic-bench targets the questions static leaderboards miss: can a model chain Hermes tools, recover from a failed call, and stop before the consecutive-tool cap. The main script, hermes_native_battery.py, runs hermes chat -q with real toolsets and scores n_tools, HIT_CAP and pass from the CLI footer, tool previews or the Hermes session database, which is needed when -Q hides the footer. Tasks that need a tool set min_tools, and residual tasks catch dummy no-tool calls and refuse-delete cases.

Two contract tests run without a Hermes process, against any OpenAI-compatible endpoint such as llama.cpp's llama-server or vLLM. hermes_loop_gate.py has 20 scripted-tool tasks for loops, duplicates and unparsed ATEM or XML output, and simulated_battery.py is an older six-scenario smoke test. generate_report.py compares result JSON files. The two layers exist because a Muse Glimmer versus Bonsai comparison showed simulated and native results disagreeing. The native battery's file tools are not sandboxed, and the destructive test is opt-in.

Key features

  • Native battery that runs real hermes chat sessions with toolsets
  • Scores n_tools, HIT_CAP and pass from the CLI footer or session database
  • 20-task scripted loop gate that needs no Hermes process
  • Older six-scenario simulated smoke test
  • Report generator that compares result JSON files
  • Opt-in destructive test

When to use it

  • Checking whether a new local model stops before the tool cap in Hermes
  • Separating model weaknesses from harness problems with simulated versus native runs
  • Comparing several local models in one generated report

Who it is for: People who run local models with Hermes Agent and want honest, repeatable tool-use checks.

How it fits with Hermes Agent

Built around Hermes Agent: the native battery drives the real Hermes CLI, with your model registered as a provider in config.yaml.

How to install hermes-agentic-bench

These commands are copied from the project's README. Check the repository for the latest steps before you run them.

git clone https://github.com/vcruz305/hermes-agentic-bench && cd hermes-agentic-bench
pip install -r requirements.txt

Requirements: Python 3.9+; an OpenAI-compatible endpoint for the simulated tests; Hermes Agent with your model registered as a provider for the native battery

Note: The native battery's file tools are not sandboxed, and the destructive test is opt-in.

FAQ

What is hermes-agentic-bench?

hermes-agentic-bench is a set of test batteries for local models used with Hermes Agent. It checks tool chaining, error recovery and whether the model stops before the consecutive-tool cap.

How do I install hermes-agentic-bench?

Clone the repository, then run pip install -r requirements.txt. The native battery additionally needs Hermes Agent installed with your model registered as a provider.

What do I need to run hermes-agentic-bench?

You need Python 3.9 or newer. The simulated tests need an OpenAI-compatible chat endpoint such as llama-server or vLLM, and the native battery needs Hermes Agent.

Similar research for Hermes Agent

All research

Related guides: What is Hermes Agent?