Hermes Atlas
Research, training & evaluation · works with Hermes Agent

Agentic Harness Arena

Ondemand-OSS/harness-arena

Blind arena that compares agent harnesses, including Hermes, on the same task and model, ranked by Elo

In short

Agentic Harness Arena is a blind, LMSYS Chatbot Arena-style benchmark that compares agent harnesses instead of models. Hermes is one of its built-in harnesses, run on the same task and model as Claude Code, Codex CLI, OpenClaw, opencode and OnDemand.

What Agentic Harness Arena does

Agentic Harness Arena applies the Chatbot Arena idea one layer up, to the CLI, tools, prompts, permissions and execution environment wrapped around a model. A task with a prompt, rubric and expected deliverables is run through each selected harness with the same model and provider config, each in its own isolated workdir. Judges see the deliverables as anonymous Response A, B and C, score every one, and only then are identities revealed.

Each signed-in user judges a task once, and every verdict feeds a public Elo leaderboard that is recomputed from the full score history on each read. Built-in harnesses are Claude Code, Codex CLI, OpenClaw, Hermes, opencode and the hosted OnDemand API, and any other agent can be added as a webhook-backed harness from the Setup UI without a code change. The backend is FastAPI with MongoDB and covers dataset import and run orchestration, and a live arena is available at harness-arena.ai.

Key features

  • Same task, same model and isolated workspaces for every harness
  • Blind judging with anonymous Response A, B and C labels
  • One vote per person per task
  • Elo leaderboard recomputed from the full score history
  • Built-in harnesses including Hermes, plus webhook-backed custom harnesses
  • FastAPI and MongoDB backend with dataset import

When to use it

  • Comparing Hermes against Claude Code or Codex CLI on the same task and model
  • Adding an in-house agent as a webhook harness and seeing how it ranks
  • Judging coding-agent outputs blind and tracking Elo ratings

Who it is for: People evaluating coding-agent harnesses who want blind human judging instead of self-reported results.

How it fits with Hermes Agent

Lists Hermes as a built-in local CLI harness next to Claude Code, Codex CLI, OpenClaw, opencode and OnDemand.

Requirements: An API key for the model provider; self-hosting uses the FastAPI and MongoDB backend

FAQ

What is Agentic Harness Arena?

Agentic Harness Arena is a blind benchmark that compares agent harnesses rather than models. It runs the same task and model through each harness and lets people score the results before identities are revealed.

Does Agentic Harness Arena work with Hermes Agent?

Yes. Hermes is listed as a built-in local CLI harness, so it can be compared with Claude Code, Codex CLI, OpenClaw, opencode and OnDemand without extra setup beyond an API key.

Is Agentic Harness Arena free and open source?

Yes. The repository is released under the MIT license.

Similar research for Hermes Agent

All research

Related guides: What is Hermes Agent?