HermesAgent-MultiModel
InfiniteWhispers/HermesAgent-MultiModel
Fully local multi-model setup for Hermes Agent with Ollama, Mixture-of-Agents routing and a tuning guide
HermesAgent-MultiModel is a configuration and documentation repository for running Hermes Agent fully locally on Ollama with a roster of eight models and Mixture-of-Agents orchestration, tuned for 16 GB consumer GPUs.
What HermesAgent-MultiModel does
The repository routes tasks to specialized local models for planning, coding, reasoning and vision through Hermes. Its eight-model roster includes gpt-oss-20b as the default tool-calling model, gemma4-heretic for fast general work, qwen3-14b for planning, qwythos-9b for reasoning, qwen25-coder for scripts, ornith-9b, an embeddings model and a vision model. In Mixture-of-Agents mode, several models are queried in parallel and an aggregator synthesizes the answers. Everything runs on a local Ollama backend with no API keys and 64K context windows.
The repo contains a .hermes/config.yaml with Hermes providers and routing, and docs/optimization-guide.md, an 11-section field guide covering hardware, models and tuning. Setup steps cover installing Ollama, pulling and building models from Modelfiles, editing ~/.hermes/config.yaml, and tuning Ollama through systemd with OLLAMA_FLASH_ATTENTION=1 and OLLAMA_KV_CACHE_TYPE=q8_0. The README notes that Hermes uses a qwen3-14b-think alias rather than the raw qwen3:14b tag. It targets RTX 4080 or 5080 class GPUs.
Key features
- Eight-model local roster for tools, planning, reasoning, coding, embeddings and vision
- Mixture-of-Agents mode that queries models in parallel and aggregates the results
- Ollama backend with no API keys and 64K context windows
- Sample .hermes/config.yaml with providers and routing
- 11-section optimization guide covering hardware, models and tuning
- Ollama systemd tuning for flash attention and KV cache type
When to use it
- Running Hermes Agent entirely offline on a single consumer GPU
- Routing planning, coding and vision tasks to different local models
- Tuning Ollama memory use for 64K context
Who it is for: Hermes Agent users with a 16 GB VRAM GPU who want a private, local-only multi-model setup.
How it fits with Hermes Agent
Built for Hermes Agent: it supplies Hermes provider and routing configuration so that Hermes uses local Ollama models.
How to install HermesAgent-MultiModel
These commands are copied from the project's README. Check the repository for the latest steps before you run them.
curl -fsSL https://ollama.com/install.sh | sh
ollama pull qwen2.5-coder:7b
ollama pull nomic-embed-text:latest
ollama pull qwen3-vl:8bRequirements: A GPU such as an RTX 4080 or 5080 with at least 16 GB VRAM, 16+ CPU cores, 48+ GB RAM (64 GB ideal), 100 GB free NVMe storage, and Ubuntu 22.04+, WSL2 on Windows 11, or macOS 12+
FAQ
What is HermesAgent-MultiModel?
HermesAgent-MultiModel is a fully local multi-model setup for Hermes Agent. It combines Ollama, a roster of eight models and Mixture-of-Agents orchestration, with a configuration file and a tuning guide.
Does HermesAgent-MultiModel work with Hermes Agent?
Yes, it is built around Hermes Agent. You copy its provider block into ~/.hermes/config.yaml so Hermes routes tasks to the local Ollama models.
What do I need to run HermesAgent-MultiModel?
The README calls for a GPU with 16 GB VRAM or more, such as an RTX 4080 or 5080, 16+ CPU cores, 48+ GB RAM and 100 GB of free NVMe storage. It runs on Ubuntu 22.04+, WSL2 on Windows 11 or macOS 12+.
Similar models for Hermes Agent
All modelsAdaptive local-inference provider that fits the best sustainable model to your hardware
tuxevil tuxevil-rotatorOpenAI-compatible proxy that rotates free-tier LLM accounts with per-model quota routing
piyush-tyagi-13 llm-keypoolFree-tier LLM API key pool with rotation, 429 cooldowns and a local OpenAI-compatible proxy
nujovich Hermes TelemetryBudget enforcement and observability plugin for Hermes Agent that stops runs before they overspend
open-world-project model-routerHermes Agent plugin that routes each turn to the cheapest of five model tiers and escalates when needed
KaiFelixBennett Hermes Local StackRun Hermes Agent and Claude Code on a local llama.cpp model with no API costs
Related guides: How to install Hermes Agent