Hermes Atlas
Automations & use cases

Hermes Autoresearch

AtlasOmnia/hermes-autoresearch

Archived harness that loops an agent over a Git repo, keeping changes that raise a score

In short

Hermes Autoresearch is an archived Python harness that repeatedly lets an agent change a Git repository, scores each change, and keeps improvements as local commits or reverts them. The project has moved into the hermes-loops monorepo.

What Hermes Autoresearch does

Hermes Autoresearch is a reusable, Karpathy-style autoresearch harness: a control loop for testing changes to code, configuration, prompts or other repository-based experiments. A configurable proposal command lets an agent edit the repository, then a configurable evaluator command prints JSON to stdout with a numeric score and optional accepted, reason and metrics fields. An improvement of at least min_improvement is kept as a local commit; otherwise the change is reverted.

The loop is agent-command and provider agnostic and stops on max_trials, max_seconds, target_score or after N trials without improvement. An allowlist of repo-relative paths reverts edits outside it, and results are logged to runs/experiments.tsv and runs/experiments.jsonl. Commits go to the current branch with messages such as autoresearch: accept trial 3 score=12.5, and nothing is pushed. The README warns that commands run with your user's permissions and that the allowlist is not an operating-system sandbox.

Key features

  • Proposal and evaluator commands that you configure yourself
  • Evaluator JSON contract with a required numeric score
  • Git keep or revert decisions with local commits only
  • Path allowlist that reverts edits outside permitted files
  • Stop conditions: max trials, time budget, target score and no-improvement streak
  • TSV and JSONL experiment logs

When to use it

  • Iterating on a prompt file while an evaluator scores each version
  • Letting an agent tune configuration or code until a test score stops improving
  • Logging every trial's decision and score for later review

Who it is for: Developers who want to run automated, score-driven experiments with an agent on a Git repository.

How it fits with Hermes Agent

Described as a harness for Hermes Agent, though the proposal command is configurable and agent agnostic, so any agent command can be plugged in.

How to install Hermes Autoresearch

These commands are copied from the project's README. Check the repository for the latest steps before you run them.

python -m venv .venv
source .venv/bin/activate
python -m pip install .
python -m hermes_autoresearch.cli --config examples/example_config.json

Requirements: Python, a clean Git repository on a dedicated experiment branch with runs/ ignored, and your own proposal and evaluator commands

Note: The repository is archived and read-only; the project moved to the hermes-loops monorepo under autoresearch/.

FAQ

What is Hermes Autoresearch?

Hermes Autoresearch is a Python harness that lets an agent repeatedly change a Git repository and measures each change with an evaluator score. Improvements are kept as local commits and the rest are reverted.

Does Hermes Autoresearch work with Hermes Agent?

It is described as a harness for Hermes Agent, but the proposal command is configurable and agent agnostic, so you decide which agent command it runs. The repository is archived, and the project continues inside AtlasOmnia's hermes-loops monorepo.

Is Hermes Autoresearch free and open source?

Yes, it is released under the MIT license, but the repository is archived and read-only.

Similar automations for Hermes Agent

All automations

Related guides: Build agent teams with the Hermes kanban board · How to run Hermes Agent securely