Hermes Atlas
22 repos · updated October 10, 2026

Research, training and evaluation projects around Hermes Agent

Hermes Agent is also a research platform: it can generate batches of trajectories for training tool-calling models. This category covers benchmarks, datasets, training code and evaluation harnesses connected to Hermes.

Top 12 research by GitHub stars

ray-r-ren Agent Apprenticeship

Mentor-reviewed workflow loops for local agents such as Hermes Agent, plus a shared experience dataset

★ 1.6kPythonResearch
WEIFENG2333 Phistory

Versioned archive of system prompts from agent CLIs including Claude Code, Codex and Hermes

★ 661PythonResearch
InternLM WildClawBench

Benchmark of 60 hand-built real-world tasks run across OpenClaw, Claude Code, Codex CLI and Hermes Agent

★ 531PythonResearch
edonadei Caliper

Evaluate whether agent skills help by running tasks with and without them, on Hermes and other agents

★ 209PythonResearch
KhanCold MerchantBench

365-day simulated e-commerce benchmark for LLM agents, with a Hermes adapter

★ 171PythonResearch
Raidriar7170 Hermes SkillEval

Skill-routing evaluation and release-gate toolkit for SKILL.md agent skills

★ 125PythonResearch
Q00 RLM-Forge

Recursive Language Model runtime for Hermes Agent with Ouroboros recursion and TraceGuard evidence gating

★ 121PythonResearch
howdymary Hermes Agent Meta-Harness

Outer-loop optimizer that searches over Hermes' benchmark harness, not model weights

★ 118PythonResearch
MiaAI-Lab Best Local Model for Agentic Workflows 2026

Benchmark report ranking local LLMs for Hermes Agent-style tool use on DGX Spark class hardware

★ 36HTMLResearch
EngTurtle Hermes MemConflict Benchmark

Benchmark comparing self-hostable memory providers for Hermes Agent on the MemConflict dataset

★ 33PythonResearch
fox-in-the-box-ai Hermes Best Models

Monthly benchmark that tests LLMs on Hermes-style agent tasks such as tool calling and delegation

★ 26PythonResearch
digitalspaceport Multi Agent Benchmark Tool

Async benchmark that simulates many agents hitting one OpenAI-compatible endpoint such as vLLM

★ 17PythonResearch

All 22 research repos

Click a column to sort
ray-r-ren/agent-apprenticeshipMentor-reviewed workflow loops for local agents such as Hermes Agent, plus a shared experience dataset Python 1,618 Jul 6, 2026
WEIFENG2333/phistoryVersioned archive of system prompts from agent CLIs including Claude Code, Codex and Hermes Python 661 Oct 9, 2026
InternLM/WildClawBenchBenchmark of 60 hand-built real-world tasks run across OpenClaw, Claude Code, Codex CLI and Hermes Agent Python 531 Sep 18, 2026
edonadei/caliperEvaluate whether agent skills help by running tasks with and without them, on Hermes and other agents Python 209 Oct 10, 2026
KhanCold/merchantbench365-day simulated e-commerce benchmark for LLM agents, with a Hermes adapter Python 171 Oct 9, 2026
Raidriar7170/hermes-skillevalSkill-routing evaluation and release-gate toolkit for SKILL.md agent skills Python 125 Sep 26, 2026
Q00/rlm-forgeRecursive Language Model runtime for Hermes Agent with Ouroboros recursion and TraceGuard evidence gating Python 121 May 10, 2026
howdymary/hermes-agent-metaharnessOuter-loop optimizer that searches over Hermes' benchmark harness, not model weights Python 118 Jul 11, 2026
MiaAI-Lab/Best-Local-Model_Agentic-Workflows_2026Benchmark report ranking local LLMs for Hermes Agent-style tool use on DGX Spark class hardware HTML 36 Oct 4, 2026
EngTurtle/hermes-memconflictBenchmark comparing self-hostable memory providers for Hermes Agent on the MemConflict dataset Python 33 Aug 10, 2026
fox-in-the-box-ai/hermes-best-modelsMonthly benchmark that tests LLMs on Hermes-style agent tasks such as tool calling and delegation Python 26 May 22, 2026
digitalspaceport/Multi-Agent-Benchmark-ToolAsync benchmark that simulates many agents hitting one OpenAI-compatible endpoint such as vLLM Python 17 Mar 30, 2026
am423/hermes-bench-tool-callBenchmark for local models running inside the Hermes Agent harness, with traces and hardware telemetry Python 16 Jun 24, 2026
beardthelion/hermes-skill-distillationHackathon environment that turns Hermes Agent task runs into scored trajectories for Hermes 4 training Python 16 Mar 15, 2026
vcruz305/hermes-agentic-benchAgentic test batteries that check whether local models can chain Hermes tools and stop before the cap Python 15 Aug 14, 2026
shuklabhay/llm-synesthesiaPipeline that reads emotions from Hermes-4.3-36B hidden states and renders them as a live WebGL visual Python 14 Apr 5, 2026
ctala/ai-benchmarks-alternativosOpen Spanish-language benchmark of LLMs for business and agent use, scored by an independent Phi-4 judge Python 13 Sep 29, 2026
ohikava/ecom-agentHermes-based agent for the BitGN E-commerce benchmark, locked to one MCP tool channel Python 12 Jun 12, 2026
stevibe/HermesAgent-20BenchLocal benchmark of 20 scenarios scoring models as controllers of a real Hermes Agent runtime JavaScript 11 May 15, 2026
CardSorting/raidenBenchmark repository for long-horizon agentic development with Hermes, JoyZoning and JSDP Makefile 11 May 29, 2026
Ondemand-OSS/harness-arenaBlind arena that compares agent harnesses, including Hermes, on the same task and model, ranked by Elo Python 11 Sep 15, 2026
teknium1/hermes-and-jev-play-minecraftHermes Agent plans, Jev picks bounded actions and Mineflayer executes in Minecraft JavaScript 10 Sep 21, 2026

Research FAQ

Can Hermes Agent generate training data?

Yes. Hermes supports batch trajectory generation and trajectory compression, which researchers use to train the next generation of tool-calling models.

Are the Hermes language models the same as Hermes Agent?

No. Hermes is also the name of Nous Research's model family. Hermes Agent is the agent software, and it works with many models, not only Hermes models.

Related guides: What is Hermes Agent?