Hermes Atlas
Skills & skill packs · works with Hermes Agent

Agent Arena

zhjai/agent-arena

Evidence-first multi-agent debate skill for second opinions on code and architecture decisions

In short

Agent Arena is a skill that has agents such as Codex and Claude Code analyze a problem independently, critique each other and preserve dissent, and it is designed to work with Hermes Agent among others.

What Agent Arena does

Agent Arena follows a fixed protocol: independent generation, claim extraction, evidence checking, critique, revision, blind judging and synthesis. Its stated principle is independent, heterogeneous agents first, debate later, evidence before consensus and dissent preserved. It targets architecture decisions, implementation plan review, pull request review, bug root-cause analysis and RAG claim verification, and says it is not for simple lookups or routine small edits.

Two skills are included: agent-arena, the main review protocol, and deliberative-analysis, a lighter companion against overconfidence. It can pit different model families against each other, such as GLM-backed Claude Code against Codex. The README is explicit that this is a protocol and instruction skill with an optional local result parser, not an executable orchestrator. It does not install, authenticate or call any agent, so cross-agent runs depend on your host agent and local CLIs. The license is MIT.

Key features

  • Seven-step protocol from independent generation to synthesis
  • Evidence ledgers, claim extraction and blind judge scoring
  • Dissent preserved in the final synthesis
  • Companion deliberative-analysis skill for lighter checks
  • Support for mixed model backends such as GLM, DeepSeek, Qwen and Kimi through compatible proxies

When to use it

  • Getting a second opinion on an architecture decision before committing
  • Red-teaming an implementation plan or a pull request
  • Comparing competing root-cause hypotheses for a bug

Who it is for: Engineers who run more than one coding agent and want structured cross-examination of high-stakes decisions.

How it fits with Hermes Agent

The README lists Hermes Agent among the agents the skill is designed for, those that support custom skills, instructions or delegation, and states the project is not affiliated with Hermes Agent.

Requirements: An agent that supports custom skills or instructions, plus the local CLIs, authentication and permissions for each agent taking part in a debate

Note: It is an instruction protocol with an optional result parser, not an executable orchestrator, so cross-agent runs depend on your host agent and local CLIs.

FAQ

What is Agent Arena?

Agent Arena is an evidence-first multi-agent debate skill. Agents analyze a problem independently, critique each other, verify claims and a judge scores the results, with dissent kept in the output.

Does Agent Arena work with Hermes Agent?

It is designed for agents that support custom skills, custom instructions or tool-driven delegation, and the README names Hermes Agent in that list. The project states it is not affiliated with Hermes Agent.

Does Agent Arena run the other agents for me?

No. The README says it is a protocol and instruction skill with an optional local result parser, not an executable orchestrator. Running Codex, Claude Code or others depends on your host agent, local CLIs and approvals.

Similar skills for Hermes Agent

All skills

Related guides: What is Hermes Agent? · SOUL.md for Hermes Agent: what it is and how to write one