hermes-agentic-bench
vcruz305/hermes-agentic-bench
Agentic test batteries that check whether local models can chain Hermes tools and stop before the cap
hermes-agentic-bench is a set of small test batteries that score local models on chaining Hermes Agent tools, recovering from failed calls and stopping before the consecutive-tool cap.
What hermes-agentic-bench does
hermes-agentic-bench targets the questions static leaderboards miss: can a model chain Hermes tools, recover from a failed call, and stop before the consecutive-tool cap. The main script, hermes_native_battery.py, runs hermes chat -q with real toolsets and scores n_tools, HIT_CAP and pass from the CLI footer, tool previews or the Hermes session database, which is needed when -Q hides the footer. Tasks that need a tool set min_tools, and residual tasks catch dummy no-tool calls and refuse-delete cases.
Two contract tests run without a Hermes process, against any OpenAI-compatible endpoint such as llama.cpp's llama-server or vLLM. hermes_loop_gate.py has 20 scripted-tool tasks for loops, duplicates and unparsed ATEM or XML output, and simulated_battery.py is an older six-scenario smoke test. generate_report.py compares result JSON files. The two layers exist because a Muse Glimmer versus Bonsai comparison showed simulated and native results disagreeing. The native battery's file tools are not sandboxed, and the destructive test is opt-in.
Key features
- Native battery that runs real hermes chat sessions with toolsets
- Scores n_tools, HIT_CAP and pass from the CLI footer or session database
- 20-task scripted loop gate that needs no Hermes process
- Older six-scenario simulated smoke test
- Report generator that compares result JSON files
- Opt-in destructive test
When to use it
- Checking whether a new local model stops before the tool cap in Hermes
- Separating model weaknesses from harness problems with simulated versus native runs
- Comparing several local models in one generated report
Who it is for: People who run local models with Hermes Agent and want honest, repeatable tool-use checks.
How it fits with Hermes Agent
Built around Hermes Agent: the native battery drives the real Hermes CLI, with your model registered as a provider in config.yaml.
How to install hermes-agentic-bench
These commands are copied from the project's README. Check the repository for the latest steps before you run them.
git clone https://github.com/vcruz305/hermes-agentic-bench && cd hermes-agentic-bench
pip install -r requirements.txtRequirements: Python 3.9+; an OpenAI-compatible endpoint for the simulated tests; Hermes Agent with your model registered as a provider for the native battery
Note: The native battery's file tools are not sandboxed, and the destructive test is opt-in.
FAQ
What is hermes-agentic-bench?
hermes-agentic-bench is a set of test batteries for local models used with Hermes Agent. It checks tool chaining, error recovery and whether the model stops before the consecutive-tool cap.
How do I install hermes-agentic-bench?
Clone the repository, then run pip install -r requirements.txt. The native battery additionally needs Hermes Agent installed with your model registered as a provider.
What do I need to run hermes-agentic-bench?
You need Python 3.9 or newer. The simulated tests need an OpenAI-compatible chat endpoint such as llama-server or vLLM, and the native battery needs Hermes Agent.
Similar research for Hermes Agent
All researchAsync benchmark that simulates many agents hitting one OpenAI-compatible endpoint such as vLLM
am423 HermesBenchBenchmark for local models running inside the Hermes Agent harness, with traces and hardware telemetry
beardthelion Hermes Skill DistillationHackathon environment that turns Hermes Agent task runs into scored trajectories for Hermes 4 training
shuklabhay llm-synesthesiaPipeline that reads emotions from Hermes-4.3-36B hidden states and renders them as a live WebGL visual
ctala AI Benchmarks AlternativosOpen Spanish-language benchmark of LLMs for business and agent use, scored by an independent Phi-4 judge
ohikava BitGN ECOM AgentHermes-based agent for the BitGN E-commerce benchmark, locked to one MCP tool channel
Related guides: What is Hermes Agent?