HermesAgent-20
stevibe/HermesAgent-20
BenchLocal benchmark of 20 scenarios scoring models as controllers of a real Hermes Agent runtime
HermesAgent-20 is a BenchLocal Bench Pack that measures how well a model performs as the controller inside the real Hermes Agent runtime. It runs 20 public scenarios and scores them from runtime artifacts and traces, not mocked tool calls.
What HermesAgent-20 does
HermesAgent-20 is a benchmark package for the BenchLocal desktop app, which handles provider setup, model selection, run history and side-by-side comparison. Its 20 public scenarios, with canonical IDs HA-01 through HA-20, run against a real Hermes runtime pinned to a specific repository revision. Scoring checks concrete side effects such as files, memory state, cron state, delivery logs, browser exports and approval traces, using deterministic artifacts, runtime state and Hermes trace invariants rather than string matching.
The design splits the work in two. A thin BenchLocal host runtime needs only Node.js and forwards each scenario to a required Docker verifier. BenchLocal keeps provider secrets and exposes an OpenAI-compatible inference endpoint for the selected model, reachable from the container. Inside the verifier, which holds the pinned Hermes checkout and its Python dependencies, a temporary Hermes config points Hermes at that endpoint, a temporary workspace and HERMES_HOME are created, and the real Hermes CLI runs the scenario. The repository includes the scenarios, orchestration code, METHODOLOGY.md, a BenchLocal adapter and the verifier runtime.
Key features
- 20 public scenarios with IDs HA-01 through HA-20
- Runs the real Hermes CLI pinned to a specific repository revision
- Verifies files, memory state, cron state, delivery logs, browser exports and approval traces
- Deterministic scoring from artifacts, runtime state and trace invariants
- Docker verifier with BenchLocal-managed model routing and credentials
When to use it
- Comparing several models as Hermes controllers side by side in BenchLocal
- Checking whether a local or hosted model handles Hermes memory, cron and approval flows
- Reading METHODOLOGY.md to see how trace-based scoring works
Who it is for: People choosing or evaluating models to drive Hermes Agent who want results from a real runtime instead of mocked tool calls.
How it fits with Hermes Agent
Built around Hermes Agent: each scenario launches the real Hermes CLI inside a Docker verifier pinned to a specific Hermes revision.
Requirements: The BenchLocal desktop app and Docker for the verifier; the host runtime needs Node.js
FAQ
What is HermesAgent-20?
HermesAgent-20 is a BenchLocal Bench Pack with 20 scenarios that test how well a model works as the controller inside the real Hermes Agent runtime. It scores runtime artifacts and traces rather than mocked tool calls.
How do I run HermesAgent-20?
Download BenchLocal from its latest release, install HermesAgent-20 from the Bench Pack registry inside BenchLocal, add one or more models, select HermesAgent-20 and start a run.
What do I need to run HermesAgent-20?
You need the BenchLocal desktop app and Docker, because the verifier that runs Hermes is a required Docker service. BenchLocal manages provider selection and credentials.
Similar research for Hermes Agent
All researchPipeline that reads emotions from Hermes-4.3-36B hidden states and renders them as a live WebGL visual
ctala AI Benchmarks AlternativosOpen Spanish-language benchmark of LLMs for business and agent use, scored by an independent Phi-4 judge
ohikava BitGN ECOM AgentHermes-based agent for the BitGN E-commerce benchmark, locked to one MCP tool channel
CardSorting Raiden (Sky Circuit)Benchmark repository for long-horizon agentic development with Hermes, JoyZoning and JSDP
Ondemand-OSS Agentic Harness ArenaBlind arena that compares agent harnesses, including Hermes, on the same task and model, ranked by Elo
teknium1 Hermes and Jev Play MinecraftHermes Agent plans, Jev picks bounded actions and Mineflayer executes in Minecraft
Related guides: What is Hermes Agent?