Hermes Atlas
Developer tools & SDKs · works with Hermes Agent

autocontext

greyhaven-ai/autocontext

Self-improving harness that runs agent tasks against evaluation and keeps playbooks, traces and datasets

In short

autocontext is a harness for improving agents through repeated evaluated runs, with a CLI-first skill export for Hermes Agent. It keeps the lessons from each run as playbooks, traces and reports so later runs start from them.

What autocontext does

You give autocontext a goal, and it runs the task against evaluation, keeps the useful lessons and discards dead ends. Each run leaves traces, reports, playbooks, datasets and optional local-model training artifacts on the filesystem, so you can inspect, diff, replay or export them. The Python package is autocontext and the CLI is autoctx, with an npm CLI published as autoctx.

For Hermes Agent users, the README documents exporting a CLI-first skill with the autoctx hermes export-skill command. The same tool also offers a Pi extension and an MCP server for Claude Code, Cursor or other MCP clients. Providers include Anthropic, OpenAI-compatible endpoints, OpenRouter, Claude CLI, Codex and Pi, and the docs cover self-hosted models on vLLM or Ollama. Kernel campaigns add provider-generation receipts and bounded paid-call accounting.

Key features

  • autoctx solve runs a goal against evaluation for a set number of iterations
  • Filesystem-first output per run: trace.jsonl, generation analysis, report.md and artifacts
  • Knowledge folder holding a playbook, hints and tools for the next run
  • Exports a CLI-first skill for Hermes Agent with autoctx hermes export-skill
  • MCP server mode started with autoctx serve mcp
  • Provider choices including Anthropic, OpenAI-compatible, OpenRouter, Claude CLI, Codex and Pi

When to use it

  • Iteratively improving an agent on a defined task, such as customer-support replies for billing disputes
  • Giving Hermes Agent an exported skill that runs autocontext from the command line
  • Keeping replayable traces and datasets from agent runs for later review or training

Who it is for: Developers who want to measure and improve agent performance on a task and keep the resulting playbooks and traces.

How it fits with Hermes Agent

Supports Hermes Agent as one of several agent entry points, through an exported CLI-first skill. It is a general tool, and the repository carries the hermes-agent topic.

How to install autocontext

These commands are copied from the project's README. Check the repository for the latest steps before you run them.

uv tool install autocontext==0.19.1
bun add -g autoctx@0.19.0
uv run autoctx hermes export-skill --with-references --json

Requirements: The npm CLI and TUI need Node.js 22.19.0 or newer; provider variables are listed in .env.example

FAQ

What is autocontext?

autocontext is a harness for agent improvement. You give it a goal, it runs the task against evaluation, keeps useful lessons, discards dead ends and leaves traces, reports and playbooks for the next run.

Does autocontext work with Hermes Agent?

Yes, the README lists Hermes as an agent entry point. You export a CLI-first skill with uv run autoctx hermes export-skill --with-references --json.

How do I install autocontext?

Install the Python CLI with uv tool install autocontext==0.19.1, or the TypeScript CLI with bun add -g autoctx@0.19.0. For Hermes Agent, then run the hermes export-skill command shown in the README.

Similar dev tools for Hermes Agent

All dev tools

Related guides: How to install Hermes Agent · Run multiple Hermes agents with profiles