WildClawBench
InternLM/WildClawBench
Benchmark of 60 hand-built real-world tasks run across OpenClaw, Claude Code, Codex CLI and Hermes Agent
WildClawBench is an end-to-end benchmark that tests AI agents on 60 original, real-world tasks inside isolated Docker containers. Hermes Agent is one of the four agent harnesses the README lists as running the full suite.
What WildClawBench does
WildClawBench is an end-to-end benchmark for AI agents built from 60 original, hand-made tasks. Examples in the README include clipping goal highlights from a football match video, negotiating meeting times over multi-round emails, hunting for contradictions in search results, writing inference scripts for undocumented codebases and catching privacy leaks. Tasks are grouped by agency, multimodal work, long-horizon workflows, coding and safety.
Each task runs in its own Docker container with the same image, data and grading code, and ground truth and graders are injected only after the agent finishes. The README says OpenClaw, Claude Code, Codex CLI and Hermes Agent all execute the same 60 tasks under the same grading, which lets you separate harness scaffolding from model ability. The evaluation code lives in the repository, the data ships as three Hugging Face datasets, and a paper is on arXiv.
Key features
- 60 original tasks across agency, multimodal, long-horizon, coding and safety categories
- One Docker container per task with the same image and grading code
- Ground truth and graders injected after the agent finishes
- The same task suite for OpenClaw, Claude Code, Codex CLI and Hermes Agent
- Three Hugging Face datasets: WildClawBench, WildClawBench-Harbor and WildClawBench-Trajectories
- Project page, technical report and arXiv paper
When to use it
- Compare Hermes Agent with other harnesses on the same real-world tasks
- Measure how much of an agent's score comes from its harness rather than the model
- Check safety behavior such as prompt injection defense and credential leak detection
Who it is for: Researchers and agent developers who need a reproducible benchmark for comparing agent harnesses and models.
How it fits with Hermes Agent
Hermes Agent is one of four agent harnesses the README lists as running the full task suite, so the benchmark can be used to compare Hermes Agent with OpenClaw, Claude Code and Codex CLI.
FAQ
What is WildClawBench?
WildClawBench is an agent benchmark of 60 original tasks that tests whether an AI agent can do real work end to end. Tasks run in isolated Docker containers with hidden graders.
Does WildClawBench work with Hermes Agent?
Yes. The README lists Hermes Agent as one of four harnesses, with OpenClaw, Claude Code and Codex CLI, that execute the same 60 tasks under the same grading.
Is WildClawBench free and open source?
Yes. The repository is released under the MIT license, and the benchmark data is published as Hugging Face datasets.
Similar research for Hermes Agent
All researchMentor-reviewed workflow loops for local agents such as Hermes Agent, plus a shared experience dataset
WEIFENG2333 PhistoryVersioned archive of system prompts from agent CLIs including Claude Code, Codex and Hermes
edonadei CaliperEvaluate whether agent skills help by running tasks with and without them, on Hermes and other agents
KhanCold MerchantBench365-day simulated e-commerce benchmark for LLM agents, with a Hermes adapter
Raidriar7170 Hermes SkillEvalSkill-routing evaluation and release-gate toolkit for SKILL.md agent skills
Q00 RLM-ForgeRecursive Language Model runtime for Hermes Agent with Ouroboros recursion and TraceGuard evidence gating
Related guides: What is Hermes Agent?