Agent Arena
zhjai/agent-arena
Evidence-first multi-agent debate skill for second opinions on code and architecture decisions
Agent Arena is a skill that has agents such as Codex and Claude Code analyze a problem independently, critique each other and preserve dissent, and it is designed to work with Hermes Agent among others.
What Agent Arena does
Agent Arena follows a fixed protocol: independent generation, claim extraction, evidence checking, critique, revision, blind judging and synthesis. Its stated principle is independent, heterogeneous agents first, debate later, evidence before consensus and dissent preserved. It targets architecture decisions, implementation plan review, pull request review, bug root-cause analysis and RAG claim verification, and says it is not for simple lookups or routine small edits.
Two skills are included: agent-arena, the main review protocol, and deliberative-analysis, a lighter companion against overconfidence. It can pit different model families against each other, such as GLM-backed Claude Code against Codex. The README is explicit that this is a protocol and instruction skill with an optional local result parser, not an executable orchestrator. It does not install, authenticate or call any agent, so cross-agent runs depend on your host agent and local CLIs. The license is MIT.
Key features
- Seven-step protocol from independent generation to synthesis
- Evidence ledgers, claim extraction and blind judge scoring
- Dissent preserved in the final synthesis
- Companion deliberative-analysis skill for lighter checks
- Support for mixed model backends such as GLM, DeepSeek, Qwen and Kimi through compatible proxies
When to use it
- Getting a second opinion on an architecture decision before committing
- Red-teaming an implementation plan or a pull request
- Comparing competing root-cause hypotheses for a bug
Who it is for: Engineers who run more than one coding agent and want structured cross-examination of high-stakes decisions.
How it fits with Hermes Agent
The README lists Hermes Agent among the agents the skill is designed for, those that support custom skills, instructions or delegation, and states the project is not affiliated with Hermes Agent.
Requirements: An agent that supports custom skills or instructions, plus the local CLIs, authentication and permissions for each agent taking part in a debate
Note: It is an instruction protocol with an optional result parser, not an executable orchestrator, so cross-agent runs depend on your host agent and local CLIs.
FAQ
What is Agent Arena?
Agent Arena is an evidence-first multi-agent debate skill. Agents analyze a problem independently, critique each other, verify claims and a judge scores the results, with dissent kept in the output.
Does Agent Arena work with Hermes Agent?
It is designed for agents that support custom skills, custom instructions or tool-driven delegation, and the README names Hermes Agent in that list. The project states it is not affiliated with Hermes Agent.
Does Agent Arena run the other agents for me?
No. The README says it is a protocol and instruction skill with an optional local result parser, not an executable orchestrator. Running Codex, Claude Code or others depends on your host agent, local CLIs and approvals.
Similar skills for Hermes Agent
All skillsHermes Agent skill that generates images through a logged-in Codex CLI using ChatGPT OAuth
ilang-ai iLang OpenClawText-only skills and OpenClaw plugins from iLang, aimed at OpenClaw, Hermes and other AI agents
xiaohei-info Oh My Agent SkillsInstallable Hermes-compatible skills and workflow patterns for common agent failures
Igloo302 Interest RadarPlatform-agnostic skill that judges how relevant a link or piece of content is to you
alblez hermes-skillsA collection of reusable Hermes Agent skills, including local Qwen3-TTS and upstream contribution
genshin-hermes Hormozi $100M Frameworks SkillHermes Agent skill packaging Alex Hormozi's business frameworks from 13 books
Related guides: What is Hermes Agent? · SOUL.md for Hermes Agent: what it is and how to write one