hermes-firewall
jooray/hermes-firewall
Prompt-injection gate plugin that scans web, MCP, email and image content before Hermes Agent sees it
hermes-firewall is a prompt-injection gate plugin for Hermes Agent that scans untrusted content before the model sees it. Each web page, MCP or tool result, fetched email, file or image is either passed unchanged or withheld.
What hermes-firewall does
hermes-firewall works in two stages. First, local extraction written in standard-library Python pulls out what a human would not see as well as the visible text: hidden HTML elements and comments, sentence-like attribute values, JSON-LD, script strings, invisible Unicode tag characters and decoded base64. Images are run through OCR with Tesseract, or Apple Vision on macOS when ocrmac is installed, and their EXIF, XMP, comments and bytes after the JPEG end marker are read too.
Second, the extracted text is scored by Jev, a decision model on Venice's Decisions API, and each result is passed unchanged or withheld. Text that the harness itself wrote into a result, such as approval denials and tool notices, is removed before scoring. For local use, setting PROMPT_FIREWALL_BACKEND=nimble scores with Nimble in Ollama 0.35 or newer (about 11 GB of memory), and the local backend runs RSI-Jev v6.1-VL 4B. The repository also contains a benchmark.
Key features
- Scans web pages, MCP and tool results, fetched email, files and images
- Extracts hidden HTML, JSON-LD, invisible Unicode and decoded base64 before scoring
- OCR for images plus EXIF, XMP and trailing-byte metadata
- Venice-hosted Jev scoring by default, with local Nimble and RSI-Jev backends
- Strips Hermes' own notices from results before scoring
- Includes a benchmark of the gate
When to use it
- Protecting an agent that browses the web from instructions hidden in page markup
- Checking fetched email and MCP results for planted instructions before the model reads them
- Running the gate on your own machine through Ollama to keep content local
Who it is for: Hermes Agent users whose agents read untrusted web pages, email, files or images and who want a filter in front of the model.
How it fits with Hermes Agent
Built for Hermes Agent as a plugin installed with hermes plugins install; it sits between untrusted tool results and the model.
How to install hermes-firewall
These commands are copied from the project's README. Check the repository for the latest steps before you run them.
hermes plugins install jooray/hermes-firewall/hermes-plugin/prompt-firewall --enableRequirements: A Venice API key in a file, optionally Tesseract for images, and a Hermes restart. The local backends need Ollama 0.35+, and Nimble needs about 11 GB of memory
FAQ
What is hermes-firewall?
hermes-firewall is a prompt-injection gate for Hermes Agent. It scans untrusted content such as web pages, MCP results, email, files and images before the model sees it, and either passes or withholds each result.
Does hermes-firewall work with Hermes Agent?
Yes, it is a plugin built for Hermes Agent. It inspects web, MCP, email, file and image results before they reach the model.
How do I install hermes-firewall?
Run hermes plugins install jooray/hermes-firewall/hermes-plugin/prompt-firewall --enable, put a Venice API key in a file, optionally install Tesseract for image OCR, and restart Hermes. INSTALL.md in the repository describes the full steps.
Similar security for Hermes Agent
All securityHermes plugin that has a second-lab model review each task before a subagent starts writing code
intentframe IntentFrame for Hermes AgentIntentFrame security plugin that checks Hermes terminal, code, file and cron tool calls against policy
mauricemohr88-debug Hermes Plugin GuardStatic security scanner for Hermes Agent plugins that never imports or runs the plugin code
wnstify Hermes Agent Hardening PatternsHardened Docker Compose and SSH sandbox patterns for self-hosting Hermes Agent with Honcho
aibuild-lab Skills GuardThreat scanner and trust matrix for AI skill files, ported from Hermes Agent's skills_guard
0xtbug RecatLocal workspace for reviewing security findings reported by Hermes and other agents
Related guides: How to run Hermes Agent securely