RunbookHermes
Tommy-yw/RunbookHermes
Hermes-based AIOps agent for evidence-driven incident response and approval-gated remediation
RunbookHermes is an incident-response agent built on Hermes Agent that gathers metrics, logs and traces before proposing a root cause, and gates risky actions behind approval. It gives SRE teams a Hermes-based system that also turns each incident into reusable runbook knowledge.
What RunbookHermes does
RunbookHermes keeps the Hermes Agent runtime, including its tool registry, memory provider, skills, gateway, multi-provider model support and context compression, and adds an AIOps and SRE layer on top. When an alert arrives from the web console, Alertmanager, Feishu, WeCom or an API call, it first collects evidence from Prometheus metrics, Loki logs, traces and deploy history, then produces a root cause analysis that cites that evidence. Memory is treated as a weak prior and cannot replace current evidence.
Risky actions such as rollback, restart, scaling changes, traffic switching, config changes and cache flushes pass through an action policy, approval, a checkpoint and a dry run before controlled execution, followed by recovery verification and an audit timeline. After an incident it writes back a summary, a fault pattern, a generated SKILL.md, a RAG document, an eval case and a training trajectory, so later incidents start with more context.
Key features
- Evidence stack built from metrics, logs, traces, deploy history and service profiles
- Root cause guard that requires current evidence before a conclusion
- Action policy with approval, checkpoint and dry run before execution
- Recovery verification and an audit timeline after each action
- Incident intake from the web console, Alertmanager, Feishu, WeCom and API
- Post-incident memory, RAG documents and generated runbook skills
When to use it
- Triage a production alert by collecting metrics, logs and traces before choosing a fix
- Require human approval and a dry run before a rollback or restart
- Build a team runbook library from past incidents
- Export incident data as eval cases and training trajectories
Who it is for: SRE and operations teams that want an approval-gated incident response agent built on Hermes Agent.
How it fits with Hermes Agent
It is a Hermes-native system that reuses Hermes Agent's runtime, memory, skills and gateway and adds an AIOps domain layer, so it is a derived application rather than a plugin.
Note: The documentation is mainly in Chinese, and the project is a modified Hermes Agent rather than an add-on for a stock install.
FAQ
What is RunbookHermes?
RunbookHermes is an AIOps and SRE incident-response agent built on Hermes Agent. It collects evidence first, proposes a root cause, gates dangerous actions behind approval, and records what it learns after each incident.
Does RunbookHermes work with Hermes Agent?
Yes. It is built on Hermes Agent and keeps the agent loop, tool registry, memory provider, skills, gateway and model providers, then adds incident-response features on top.
Is RunbookHermes free and open source?
Yes. The repository is released under the MIT license.
Similar automations for Hermes Agent
All automationsElectron and CLI tool for batch video publishing to Chinese platforms, callable by agents
indranilbanerjee Digital Marketing ProOpen-source AI marketing plugin with 164 skills and 24 agents that installs on Hermes Agent and others
unifapi-agent UnifAPI AgentsRead-only marketing agents as SKILL.md skills: SEO audits, AI visibility, local SEO and KOL pricing
markfulton AI EmployeesEight open source business roles with 60 scheduled routines that run on the AI agent you already use
TradingAi666 TzFilm Douyin ToolHourly Douyin creator-data export for macOS with a Hermes Agent skill and Telegram push
sharbelxyz Nova YouTube AgentYouTube growth agent for Hermes Agent: channel analysis, competitor scans, ideas, scripts and performance logs
Related guides: Build agent teams with the Hermes kanban board · How to run Hermes Agent securely