Hermes Atlas
Research, training & evaluation · works with Hermes Agent

WildClawBench

InternLM/WildClawBench

Benchmark of 60 hand-built real-world tasks run across OpenClaw, Claude Code, Codex CLI and Hermes Agent

In short

WildClawBench is an end-to-end benchmark that tests AI agents on 60 original, real-world tasks inside isolated Docker containers. Hermes Agent is one of the four agent harnesses the README lists as running the full suite.

What WildClawBench does

WildClawBench is an end-to-end benchmark for AI agents built from 60 original, hand-made tasks. Examples in the README include clipping goal highlights from a football match video, negotiating meeting times over multi-round emails, hunting for contradictions in search results, writing inference scripts for undocumented codebases and catching privacy leaks. Tasks are grouped by agency, multimodal work, long-horizon workflows, coding and safety.

Each task runs in its own Docker container with the same image, data and grading code, and ground truth and graders are injected only after the agent finishes. The README says OpenClaw, Claude Code, Codex CLI and Hermes Agent all execute the same 60 tasks under the same grading, which lets you separate harness scaffolding from model ability. The evaluation code lives in the repository, the data ships as three Hugging Face datasets, and a paper is on arXiv.

Key features

  • 60 original tasks across agency, multimodal, long-horizon, coding and safety categories
  • One Docker container per task with the same image and grading code
  • Ground truth and graders injected after the agent finishes
  • The same task suite for OpenClaw, Claude Code, Codex CLI and Hermes Agent
  • Three Hugging Face datasets: WildClawBench, WildClawBench-Harbor and WildClawBench-Trajectories
  • Project page, technical report and arXiv paper

When to use it

  • Compare Hermes Agent with other harnesses on the same real-world tasks
  • Measure how much of an agent's score comes from its harness rather than the model
  • Check safety behavior such as prompt injection defense and credential leak detection

Who it is for: Researchers and agent developers who need a reproducible benchmark for comparing agent harnesses and models.

How it fits with Hermes Agent

Hermes Agent is one of four agent harnesses the README lists as running the full task suite, so the benchmark can be used to compare Hermes Agent with OpenClaw, Claude Code and Codex CLI.

FAQ

What is WildClawBench?

WildClawBench is an agent benchmark of 60 original tasks that tests whether an AI agent can do real work end to end. Tasks run in isolated Docker containers with hidden graders.

Does WildClawBench work with Hermes Agent?

Yes. The README lists Hermes Agent as one of four harnesses, with OpenClaw, Claude Code and Codex CLI, that execute the same 60 tasks under the same grading.

Is WildClawBench free and open source?

Yes. The repository is released under the MIT license, and the benchmark data is published as Hugging Face datasets.

Similar research for Hermes Agent

All research

Related guides: What is Hermes Agent?