Hermes Atlas
Security & sandboxing

Hermes Katana

claudlos/hermes-katana

Security layer for Hermes Agent with taint tracking, a policy engine and outbound secret scrubbing

In short

Hermes Katana is a defense-in-depth security layer for AI agents that tracks where text came from, scans for prompt injection and unsafe commands, and applies YAML policies before tool dispatch. It also scrubs outbound secrets and keeps a tamper-evident audit trail.

What Hermes Katana does

Hermes Katana sits between an agent runtime such as Hermes and its inputs, tool outputs and MCP servers, passing each item through a middleware chain of seven layers. A taint tracker tags every value with its origin, and flow analysis blocks untrusted data from reaching critical sinks. Input scanning checks for 30 or more injection patterns plus encodings, and output scanning looks for ANSI, markdown and homograph tricks. A declarative policy engine returns ALLOW, DENY or ESCALATE per tool call.

Decisions are written to a SHA-256 hash-chained JSONL audit log. A mitmproxy-based HTTPS proxy scrubs secrets from outbound traffic, and an AES-256-GCM vault keeps secrets with a master key in the OS keyring. Character-level provenance is inspired by Google DeepMind's CaMeL paper, and injection classifiers include a distilled MiniLM by default and DeBERTa-v3-large for higher accuracy. A Proving Ground harness tests attack effectiveness across models, and the project documents false-positive and adversarial regression gates. The CLI is katana, with commands such as doctor, policy, vault, scan and setup.

Key features

  • Character-level taint tracking that tags every value with its origin
  • Declarative per-tool policies with ALLOW, DENY and ESCALATE decisions
  • Input and output scanners for injection patterns, ANSI, markdown and homograph tricks
  • SHA-256 hash-chained JSONL audit trail
  • mitmproxy HTTPS proxy that scrubs secrets from outbound traffic
  • AES-256-GCM vault with the master key in the OS keyring

When to use it

  • Blocking prompt injection in tool output before it triggers a dangerous command
  • Keeping API keys out of outbound requests an agent makes
  • Requiring human approval for tool calls that policy marks as risky
  • Auditing which agent decisions were allowed, denied or escalated

Who it is for: Developers and operators who run Hermes Agent or other LLM agents with tool access and need injection defense, policy control and an audit trail.

How it fits with Hermes Agent

The repository is tagged hermes-agent and its architecture shows Hermes as the agent runtime that Katana wraps, though the toolkit is described for LLM agents in general.

How to install Hermes Katana

These commands are copied from the project's README. Check the repository for the latest steps before you run them.

git clone https://github.com/claudlos/hermes-katana.git
cd hermes-katana
python -m pip install -e ".[security]"
katana doctor

Requirements: Python with pip. The base install needs no model downloads, and optional model artifacts are fetched from Hugging Face through katana setup.

FAQ

What is Hermes Katana?

Hermes Katana is a defense-in-depth security toolkit for AI agents. It tracks text provenance, scans for prompt injection, enforces YAML policies before tool calls and scrubs secrets from outbound traffic.

How do I install Hermes Katana?

Clone the repository, run python -m pip install -e ".[security]", and check prerequisites with katana doctor. You can then activate a policy with katana policy use balanced and test the scanner with katana scan.

Is Hermes Katana free and open source?

Yes. The repository is public and released under the MIT license.

Similar security for Hermes Agent

All security

Related guides: How to run Hermes Agent securely