Hermes Atlas
Security & sandboxing

hermes-firewall

jooray/hermes-firewall

Prompt-injection gate plugin that scans web, MCP, email and image content before Hermes Agent sees it

In short

hermes-firewall is a prompt-injection gate plugin for Hermes Agent that scans untrusted content before the model sees it. Each web page, MCP or tool result, fetched email, file or image is either passed unchanged or withheld.

What hermes-firewall does

hermes-firewall works in two stages. First, local extraction written in standard-library Python pulls out what a human would not see as well as the visible text: hidden HTML elements and comments, sentence-like attribute values, JSON-LD, script strings, invisible Unicode tag characters and decoded base64. Images are run through OCR with Tesseract, or Apple Vision on macOS when ocrmac is installed, and their EXIF, XMP, comments and bytes after the JPEG end marker are read too.

Second, the extracted text is scored by Jev, a decision model on Venice's Decisions API, and each result is passed unchanged or withheld. Text that the harness itself wrote into a result, such as approval denials and tool notices, is removed before scoring. For local use, setting PROMPT_FIREWALL_BACKEND=nimble scores with Nimble in Ollama 0.35 or newer (about 11 GB of memory), and the local backend runs RSI-Jev v6.1-VL 4B. The repository also contains a benchmark.

Key features

  • Scans web pages, MCP and tool results, fetched email, files and images
  • Extracts hidden HTML, JSON-LD, invisible Unicode and decoded base64 before scoring
  • OCR for images plus EXIF, XMP and trailing-byte metadata
  • Venice-hosted Jev scoring by default, with local Nimble and RSI-Jev backends
  • Strips Hermes' own notices from results before scoring
  • Includes a benchmark of the gate

When to use it

  • Protecting an agent that browses the web from instructions hidden in page markup
  • Checking fetched email and MCP results for planted instructions before the model reads them
  • Running the gate on your own machine through Ollama to keep content local

Who it is for: Hermes Agent users whose agents read untrusted web pages, email, files or images and who want a filter in front of the model.

How it fits with Hermes Agent

Built for Hermes Agent as a plugin installed with hermes plugins install; it sits between untrusted tool results and the model.

How to install hermes-firewall

These commands are copied from the project's README. Check the repository for the latest steps before you run them.

hermes plugins install jooray/hermes-firewall/hermes-plugin/prompt-firewall --enable

Requirements: A Venice API key in a file, optionally Tesseract for images, and a Hermes restart. The local backends need Ollama 0.35+, and Nimble needs about 11 GB of memory

FAQ

What is hermes-firewall?

hermes-firewall is a prompt-injection gate for Hermes Agent. It scans untrusted content such as web pages, MCP results, email, files and images before the model sees it, and either passes or withholds each result.

Does hermes-firewall work with Hermes Agent?

Yes, it is a plugin built for Hermes Agent. It inspects web, MCP, email, file and image results before they reach the model.

How do I install hermes-firewall?

Run hermes plugins install jooray/hermes-firewall/hermes-plugin/prompt-firewall --enable, put a Venice API key in a file, optionally install Tesseract for image OCR, and restart Hermes. INSTALL.md in the repository describes the full steps.

Similar security for Hermes Agent

All security

Related guides: How to run Hermes Agent securely