Hermes Atlas
Models, providers & proxies

HermesAgent-MultiModel

InfiniteWhispers/HermesAgent-MultiModel

Fully local multi-model setup for Hermes Agent with Ollama, Mixture-of-Agents routing and a tuning guide

In short

HermesAgent-MultiModel is a configuration and documentation repository for running Hermes Agent fully locally on Ollama with a roster of eight models and Mixture-of-Agents orchestration, tuned for 16 GB consumer GPUs.

What HermesAgent-MultiModel does

The repository routes tasks to specialized local models for planning, coding, reasoning and vision through Hermes. Its eight-model roster includes gpt-oss-20b as the default tool-calling model, gemma4-heretic for fast general work, qwen3-14b for planning, qwythos-9b for reasoning, qwen25-coder for scripts, ornith-9b, an embeddings model and a vision model. In Mixture-of-Agents mode, several models are queried in parallel and an aggregator synthesizes the answers. Everything runs on a local Ollama backend with no API keys and 64K context windows.

The repo contains a .hermes/config.yaml with Hermes providers and routing, and docs/optimization-guide.md, an 11-section field guide covering hardware, models and tuning. Setup steps cover installing Ollama, pulling and building models from Modelfiles, editing ~/.hermes/config.yaml, and tuning Ollama through systemd with OLLAMA_FLASH_ATTENTION=1 and OLLAMA_KV_CACHE_TYPE=q8_0. The README notes that Hermes uses a qwen3-14b-think alias rather than the raw qwen3:14b tag. It targets RTX 4080 or 5080 class GPUs.

Key features

  • Eight-model local roster for tools, planning, reasoning, coding, embeddings and vision
  • Mixture-of-Agents mode that queries models in parallel and aggregates the results
  • Ollama backend with no API keys and 64K context windows
  • Sample .hermes/config.yaml with providers and routing
  • 11-section optimization guide covering hardware, models and tuning
  • Ollama systemd tuning for flash attention and KV cache type

When to use it

  • Running Hermes Agent entirely offline on a single consumer GPU
  • Routing planning, coding and vision tasks to different local models
  • Tuning Ollama memory use for 64K context

Who it is for: Hermes Agent users with a 16 GB VRAM GPU who want a private, local-only multi-model setup.

How it fits with Hermes Agent

Built for Hermes Agent: it supplies Hermes provider and routing configuration so that Hermes uses local Ollama models.

How to install HermesAgent-MultiModel

These commands are copied from the project's README. Check the repository for the latest steps before you run them.

curl -fsSL https://ollama.com/install.sh | sh
ollama pull qwen2.5-coder:7b
ollama pull nomic-embed-text:latest
ollama pull qwen3-vl:8b

Requirements: A GPU such as an RTX 4080 or 5080 with at least 16 GB VRAM, 16+ CPU cores, 48+ GB RAM (64 GB ideal), 100 GB free NVMe storage, and Ubuntu 22.04+, WSL2 on Windows 11, or macOS 12+

FAQ

What is HermesAgent-MultiModel?

HermesAgent-MultiModel is a fully local multi-model setup for Hermes Agent. It combines Ollama, a roster of eight models and Mixture-of-Agents orchestration, with a configuration file and a tuning guide.

Does HermesAgent-MultiModel work with Hermes Agent?

Yes, it is built around Hermes Agent. You copy its provider block into ~/.hermes/config.yaml so Hermes routes tasks to the local Ollama models.

What do I need to run HermesAgent-MultiModel?

The README calls for a GPU with 16 GB VRAM or more, such as an RTX 4080 or 5080, 16+ CPU cores, 48+ GB RAM and 100 GB of free NVMe storage. It runs on Ubuntu 22.04+, WSL2 on Windows 11 or macOS 12+.

Similar models for Hermes Agent

All models

Related guides: How to install Hermes Agent