Hermes Mistral Cache Affinity
kedf/hermes-mistral-cache-affinity
Hermes plugin that adds an x-affinity header for Mistral requests to improve prompt cache hits
Hermes Mistral Cache Affinity is a Hermes plugin, built as llm_request middleware, that adds an x-affinity header carrying Hermes's existing prompt_cache_key to Chat Completions requests sent to Mistral. The aim is steadier cache routing and more prompt cache hits.
What Hermes Mistral Cache Affinity does
The plugin acts only when a request is in chat_completions mode, goes over HTTPS to api.mistral.ai or api.eu.mistral.ai, and carries a non-empty prompt_cache_key. It then adds x-affinity with that key to extra_headers. It does not change messages, tools, the system prompt or the cache key, and it returns a copy rather than mutating the original request. An explicit x-affinity header in any letter case is preserved, and for other providers, hosts, schemes or API modes it does nothing.
The README reports that Mistral sessions plateaued around 25 percent cache hit before the plugin and reached 93 to 95 percent in observed sessions, citing verification/live-report.json and agent.log cache lines. It describes the change as a routing improvement, not a guarantee, since a prefix change, expiry or provider behavior can still cause a miss. The plugin performs no network access, logging or secret handling of its own. Tests include synthetic unit tests, integration tests that need a hermes-agent checkout, and an opt-in live test against the Mistral API.
Key features
- Adds x-affinity from the existing prompt_cache_key
- Limited to HTTPS Chat Completions requests to the two Mistral hosts
- Leaves messages, tools, system prompt and cache key unchanged
- Preserves an explicit x-affinity header
- Unit, integration and opt-in live tests with privacy-preserving reports
When to use it
- Improving prompt cache hit rates when Hermes talks to Mistral
- Cutting repeated prefill cost in long Hermes sessions on Mistral models
- Choosing the EU or global Mistral endpoint through the provider base_url
Who it is for: Hermes Agent users who run Mistral models and want better prompt cache reuse.
How it fits with Hermes Agent
It is a Hermes plugin that hooks the llm_request middleware and is enabled per profile in config.yaml.
Note: The README presents it as a cache routing improvement, not a cache-hit guarantee, and the repository has no license file.
FAQ
What is Hermes Mistral Cache Affinity?
Hermes Mistral Cache Affinity is a Hermes plugin that adds an x-affinity header to Mistral Chat Completions requests. The header carries the prompt_cache_key Hermes already generates, to stabilize cache routing.
Does Hermes Mistral Cache Affinity work with Hermes Agent?
Yes. It is an llm_request middleware plugin for Hermes. Copy it into the plugins directory of the Hermes profile and add mistral-cache-affinity under plugins.enabled in config.yaml.
Does Hermes Mistral Cache Affinity change my prompts?
No. It changes neither the messages, tools, system prompt nor cache key. It only adds an x-affinity header, and an existing x-affinity header is left as it is.
Similar plugins for Hermes Agent
All pluginsHermes plugin and skills that turn approved conversations into evidence-backed daily Markdown journal notes
fquresh Hermes Quote SelectionHermes Desktop plugin that quotes selected agent output into the chat input as an inline chip
sammcf model-disciplineHermes plugin for bounded, runbook-driven workflows that suit small local models
LongshotAI Signal RushTerminal arcade game with an embeddable widget for agent CLI downtime, tagged as a Hermes plugin
kxlion delegate-task-advancedHermes plugin for named subagents with per-call skills, toolsets and model selection
CocaKova Hermes Syntax OracleHermes plugin that appends compile-verified fixes for Python SyntaxErrors to tool results
Related guides: How to install Hermes Agent · How to run Hermes Agent securely