Hermes Atlas
Research, training & evaluation

HermesAgent-20

stevibe/HermesAgent-20

BenchLocal benchmark of 20 scenarios scoring models as controllers of a real Hermes Agent runtime

In short

HermesAgent-20 is a BenchLocal Bench Pack that measures how well a model performs as the controller inside the real Hermes Agent runtime. It runs 20 public scenarios and scores them from runtime artifacts and traces, not mocked tool calls.

What HermesAgent-20 does

HermesAgent-20 is a benchmark package for the BenchLocal desktop app, which handles provider setup, model selection, run history and side-by-side comparison. Its 20 public scenarios, with canonical IDs HA-01 through HA-20, run against a real Hermes runtime pinned to a specific repository revision. Scoring checks concrete side effects such as files, memory state, cron state, delivery logs, browser exports and approval traces, using deterministic artifacts, runtime state and Hermes trace invariants rather than string matching.

The design splits the work in two. A thin BenchLocal host runtime needs only Node.js and forwards each scenario to a required Docker verifier. BenchLocal keeps provider secrets and exposes an OpenAI-compatible inference endpoint for the selected model, reachable from the container. Inside the verifier, which holds the pinned Hermes checkout and its Python dependencies, a temporary Hermes config points Hermes at that endpoint, a temporary workspace and HERMES_HOME are created, and the real Hermes CLI runs the scenario. The repository includes the scenarios, orchestration code, METHODOLOGY.md, a BenchLocal adapter and the verifier runtime.

Key features

  • 20 public scenarios with IDs HA-01 through HA-20
  • Runs the real Hermes CLI pinned to a specific repository revision
  • Verifies files, memory state, cron state, delivery logs, browser exports and approval traces
  • Deterministic scoring from artifacts, runtime state and trace invariants
  • Docker verifier with BenchLocal-managed model routing and credentials

When to use it

  • Comparing several models as Hermes controllers side by side in BenchLocal
  • Checking whether a local or hosted model handles Hermes memory, cron and approval flows
  • Reading METHODOLOGY.md to see how trace-based scoring works

Who it is for: People choosing or evaluating models to drive Hermes Agent who want results from a real runtime instead of mocked tool calls.

How it fits with Hermes Agent

Built around Hermes Agent: each scenario launches the real Hermes CLI inside a Docker verifier pinned to a specific Hermes revision.

Requirements: The BenchLocal desktop app and Docker for the verifier; the host runtime needs Node.js

FAQ

What is HermesAgent-20?

HermesAgent-20 is a BenchLocal Bench Pack with 20 scenarios that test how well a model works as the controller inside the real Hermes Agent runtime. It scores runtime artifacts and traces rather than mocked tool calls.

How do I run HermesAgent-20?

Download BenchLocal from its latest release, install HermesAgent-20 from the Bench Pack registry inside BenchLocal, add one or more models, select HermesAgent-20 and start a run.

What do I need to run HermesAgent-20?

You need the BenchLocal desktop app and Docker, because the verifier that runs Hermes is a required Docker service. BenchLocal manages provider selection and credentials.

Similar research for Hermes Agent

All research

Related guides: What is Hermes Agent?