Research, training and evaluation projects around Hermes Agent
Hermes Agent is also a research platform: it can generate batches of trajectories for training tool-calling models. This category covers benchmarks, datasets, training code and evaluation harnesses connected to Hermes.
Top 12 research by GitHub stars
Mentor-reviewed workflow loops for local agents such as Hermes Agent, plus a shared experience dataset
WEIFENG2333 PhistoryVersioned archive of system prompts from agent CLIs including Claude Code, Codex and Hermes
InternLM WildClawBenchBenchmark of 60 hand-built real-world tasks run across OpenClaw, Claude Code, Codex CLI and Hermes Agent
edonadei CaliperEvaluate whether agent skills help by running tasks with and without them, on Hermes and other agents
KhanCold MerchantBench365-day simulated e-commerce benchmark for LLM agents, with a Hermes adapter
Raidriar7170 Hermes SkillEvalSkill-routing evaluation and release-gate toolkit for SKILL.md agent skills
Q00 RLM-ForgeRecursive Language Model runtime for Hermes Agent with Ouroboros recursion and TraceGuard evidence gating
howdymary Hermes Agent Meta-HarnessOuter-loop optimizer that searches over Hermes' benchmark harness, not model weights
MiaAI-Lab Best Local Model for Agentic Workflows 2026Benchmark report ranking local LLMs for Hermes Agent-style tool use on DGX Spark class hardware
EngTurtle Hermes MemConflict BenchmarkBenchmark comparing self-hostable memory providers for Hermes Agent on the MemConflict dataset
fox-in-the-box-ai Hermes Best ModelsMonthly benchmark that tests LLMs on Hermes-style agent tasks such as tool calling and delegation
digitalspaceport Multi Agent Benchmark ToolAsync benchmark that simulates many agents hitting one OpenAI-compatible endpoint such as vLLM
All 22 research repos
Click a column to sort| ray-r-ren/agent-apprenticeshipMentor-reviewed workflow loops for local agents such as Hermes Agent, plus a shared experience dataset | Python | 1,618 | Jul 6, 2026 |
| WEIFENG2333/phistoryVersioned archive of system prompts from agent CLIs including Claude Code, Codex and Hermes | Python | 661 | Oct 9, 2026 |
| InternLM/WildClawBenchBenchmark of 60 hand-built real-world tasks run across OpenClaw, Claude Code, Codex CLI and Hermes Agent | Python | 531 | Sep 18, 2026 |
| edonadei/caliperEvaluate whether agent skills help by running tasks with and without them, on Hermes and other agents | Python | 209 | Oct 10, 2026 |
| KhanCold/merchantbench365-day simulated e-commerce benchmark for LLM agents, with a Hermes adapter | Python | 171 | Oct 9, 2026 |
| Raidriar7170/hermes-skillevalSkill-routing evaluation and release-gate toolkit for SKILL.md agent skills | Python | 125 | Sep 26, 2026 |
| Q00/rlm-forgeRecursive Language Model runtime for Hermes Agent with Ouroboros recursion and TraceGuard evidence gating | Python | 121 | May 10, 2026 |
| howdymary/hermes-agent-metaharnessOuter-loop optimizer that searches over Hermes' benchmark harness, not model weights | Python | 118 | Jul 11, 2026 |
| MiaAI-Lab/Best-Local-Model_Agentic-Workflows_2026Benchmark report ranking local LLMs for Hermes Agent-style tool use on DGX Spark class hardware | HTML | 36 | Oct 4, 2026 |
| EngTurtle/hermes-memconflictBenchmark comparing self-hostable memory providers for Hermes Agent on the MemConflict dataset | Python | 33 | Aug 10, 2026 |
| fox-in-the-box-ai/hermes-best-modelsMonthly benchmark that tests LLMs on Hermes-style agent tasks such as tool calling and delegation | Python | 26 | May 22, 2026 |
| digitalspaceport/Multi-Agent-Benchmark-ToolAsync benchmark that simulates many agents hitting one OpenAI-compatible endpoint such as vLLM | Python | 17 | Mar 30, 2026 |
| am423/hermes-bench-tool-callBenchmark for local models running inside the Hermes Agent harness, with traces and hardware telemetry | Python | 16 | Jun 24, 2026 |
| beardthelion/hermes-skill-distillationHackathon environment that turns Hermes Agent task runs into scored trajectories for Hermes 4 training | Python | 16 | Mar 15, 2026 |
| vcruz305/hermes-agentic-benchAgentic test batteries that check whether local models can chain Hermes tools and stop before the cap | Python | 15 | Aug 14, 2026 |
| shuklabhay/llm-synesthesiaPipeline that reads emotions from Hermes-4.3-36B hidden states and renders them as a live WebGL visual | Python | 14 | Apr 5, 2026 |
| ctala/ai-benchmarks-alternativosOpen Spanish-language benchmark of LLMs for business and agent use, scored by an independent Phi-4 judge | Python | 13 | Sep 29, 2026 |
| ohikava/ecom-agentHermes-based agent for the BitGN E-commerce benchmark, locked to one MCP tool channel | Python | 12 | Jun 12, 2026 |
| stevibe/HermesAgent-20BenchLocal benchmark of 20 scenarios scoring models as controllers of a real Hermes Agent runtime | JavaScript | 11 | May 15, 2026 |
| CardSorting/raidenBenchmark repository for long-horizon agentic development with Hermes, JoyZoning and JSDP | Makefile | 11 | May 29, 2026 |
| Ondemand-OSS/harness-arenaBlind arena that compares agent harnesses, including Hermes, on the same task and model, ranked by Elo | Python | 11 | Sep 15, 2026 |
| teknium1/hermes-and-jev-play-minecraftHermes Agent plans, Jev picks bounded actions and Mineflayer executes in Minecraft | JavaScript | 10 | Sep 21, 2026 |
Research FAQ
Can Hermes Agent generate training data?
Yes. Hermes supports batch trajectory generation and trajectory compression, which researchers use to train the next generation of tool-calling models.
Are the Hermes language models the same as Hermes Agent?
No. Hermes is also the name of Nous Research's model family. Hermes Agent is the agent software, and it works with many models, not only Hermes models.
Related guides: What is Hermes Agent?