Agentic Harness Arena
Ondemand-OSS/harness-arena
Blind arena that compares agent harnesses, including Hermes, on the same task and model, ranked by Elo
Agentic Harness Arena is a blind, LMSYS Chatbot Arena-style benchmark that compares agent harnesses instead of models. Hermes is one of its built-in harnesses, run on the same task and model as Claude Code, Codex CLI, OpenClaw, opencode and OnDemand.
What Agentic Harness Arena does
Agentic Harness Arena applies the Chatbot Arena idea one layer up, to the CLI, tools, prompts, permissions and execution environment wrapped around a model. A task with a prompt, rubric and expected deliverables is run through each selected harness with the same model and provider config, each in its own isolated workdir. Judges see the deliverables as anonymous Response A, B and C, score every one, and only then are identities revealed.
Each signed-in user judges a task once, and every verdict feeds a public Elo leaderboard that is recomputed from the full score history on each read. Built-in harnesses are Claude Code, Codex CLI, OpenClaw, Hermes, opencode and the hosted OnDemand API, and any other agent can be added as a webhook-backed harness from the Setup UI without a code change. The backend is FastAPI with MongoDB and covers dataset import and run orchestration, and a live arena is available at harness-arena.ai.
Key features
- Same task, same model and isolated workspaces for every harness
- Blind judging with anonymous Response A, B and C labels
- One vote per person per task
- Elo leaderboard recomputed from the full score history
- Built-in harnesses including Hermes, plus webhook-backed custom harnesses
- FastAPI and MongoDB backend with dataset import
When to use it
- Comparing Hermes against Claude Code or Codex CLI on the same task and model
- Adding an in-house agent as a webhook harness and seeing how it ranks
- Judging coding-agent outputs blind and tracking Elo ratings
Who it is for: People evaluating coding-agent harnesses who want blind human judging instead of self-reported results.
How it fits with Hermes Agent
Lists Hermes as a built-in local CLI harness next to Claude Code, Codex CLI, OpenClaw, opencode and OnDemand.
Requirements: An API key for the model provider; self-hosting uses the FastAPI and MongoDB backend
FAQ
What is Agentic Harness Arena?
Agentic Harness Arena is a blind benchmark that compares agent harnesses rather than models. It runs the same task and model through each harness and lets people score the results before identities are revealed.
Does Agentic Harness Arena work with Hermes Agent?
Yes. Hermes is listed as a built-in local CLI harness, so it can be compared with Claude Code, Codex CLI, OpenClaw, opencode and OnDemand without extra setup beyond an API key.
Is Agentic Harness Arena free and open source?
Yes. The repository is released under the MIT license.
Similar research for Hermes Agent
All researchPipeline that reads emotions from Hermes-4.3-36B hidden states and renders them as a live WebGL visual
ctala AI Benchmarks AlternativosOpen Spanish-language benchmark of LLMs for business and agent use, scored by an independent Phi-4 judge
ohikava BitGN ECOM AgentHermes-based agent for the BitGN E-commerce benchmark, locked to one MCP tool channel
stevibe HermesAgent-20BenchLocal benchmark of 20 scenarios scoring models as controllers of a real Hermes Agent runtime
CardSorting Raiden (Sky Circuit)Benchmark repository for long-horizon agentic development with Hermes, JoyZoning and JSDP
teknium1 Hermes and Jev Play MinecraftHermes Agent plans, Jev picks bounded actions and Mineflayer executes in Minecraft
Related guides: What is Hermes Agent?