Hermes Atlas
Research, training & evaluation

Hermes Skill Distillation

beardthelion/hermes-skill-distillation

Hackathon environment that turns Hermes Agent task runs into scored trajectories for Hermes 4 training

In short

Hermes Skill Distillation is a hackathon project whose RealWorldTaskEnv runs real-world tasks through Hermes Agent and exports scored trajectories for fine-tuning Hermes 4.

What Hermes Skill Distillation does

The project's idea is a closed learning loop: Hermes Agent runs tasks, trajectories are captured, a judge scores them, Atropos fine-tunes Hermes 4, and the better model makes a better agent. Its RealWorldTaskEnv is a Hermes environment, a subclass of HermesAgentBaseEnv, with a battery of 30 tasks: 8 coding, 6 web research, 6 file operations, 5 sysadmin and 5 data analysis.

Rewards combine three components. Completion carries a weight of 0.6 and is checked through ToolContext verification, such as whether a file exists or tests pass. Efficiency carries 0.2 and equals one minus turns used divided by the turn limit. Recovery carries 0.2 and is judged by an LLM judge on how gracefully the agent handled errors. The environment runs in three modes: process writes SFT-ready JSONL for Atropos, evaluate runs a benchmark, and serve connects to a live Atropos API server for GRPO training. A demo script compares Hermes 4-14B before and after fine-tuning on held-out tasks. The default config uses a Modal backend.

Key features

  • RealWorldTaskEnv, a Hermes environment subclassing HermesAgentBaseEnv
  • 30 tasks across coding, web research, file ops, sysadmin and data analysis
  • Three-part reward: completion, efficiency and error recovery
  • process mode that exports SFT-ready JSONL for Atropos
  • serve mode for live RL training through an Atropos API server
  • Before and after comparison script for Hermes 4-14B

When to use it

  • Generating agentic training data from real tool-using tasks
  • Evaluating a Hermes 4 model on a fixed battery of everyday tasks
  • Experimenting with reinforcement learning on Hermes Agent trajectories

Who it is for: Researchers and hobbyists who fine-tune Hermes models and want task-grounded training trajectories.

How it fits with Hermes Agent

Built around Hermes Agent and Hermes 4: it installs hermes-agent and Atropos from GitHub and defines a Hermes environment that records the agent's tool use.

How to install Hermes Skill Distillation

These commands are copied from the project's README. Check the repository for the latest steps before you run them.

pip install git+https://github.com/NousResearch/hermes-agent.git
pip install git+https://github.com/NousResearch/atropos.git

Requirements: Hermes Agent and Atropos installed from GitHub

Note: It is a hackathon demo concept, and the repository has no license file.

FAQ

What is Hermes Skill Distillation?

Hermes Skill Distillation is a hackathon project that generates agentic training trajectories from real-world tasks run by Hermes Agent. It scores them and exports data for fine-tuning Hermes 4.

How do I install Hermes Skill Distillation?

Install hermes-agent and Atropos with the two pip install git+ commands from the README. Then run the real_world_task_env.py script in process, evaluate or serve mode.

Does Hermes Skill Distillation work with Hermes Agent?

Yes. It is built on Hermes Agent and defines a Hermes environment, RealWorldTaskEnv, that runs tasks and scores the resulting trajectories.

Similar research for Hermes Agent

All research

Related guides: What is Hermes Agent?