oMLX
jundot/omlx
LLM inference server for Apple Silicon with continuous batching and SSD caching, run from the menu bar
oMLX is a local LLM inference server for Apple Silicon Macs with an OpenAI-compatible endpoint, managed from the macOS menu bar. Its README lists Hermes Agent among the clients it can connect to.
What oMLX does
oMLX is an LLM inference server for Apple Silicon Macs that is managed from the macOS menu bar. It uses continuous batching and a tiered KV cache that persists across a hot in-memory tier and a cold SSD tier, so past context stays cached and reusable across requests even when context changes mid-conversation. The author built it to pin everyday models in memory, auto-swap heavier ones on demand and set context limits.
Any OpenAI-compatible client can connect to http://localhost:8000/v1, and a built-in chat UI sits at http://localhost:8000/admin/chat. The server discovers LLMs, VLMs, embedding models and rerankers from subdirectories of the model directory. The README lists integrations for OpenClaw, OpenCode, Codex, Hermes Agent, Copilot and DeepSeek Harness. It installs as a macOS app with in-app updates, through Homebrew or from source, with optional MCP support. Model families such as GLM-5.2, MiniMax M3 and Qwen3.5 need native custom kernels, which require full Xcode when built locally.
Key features
- Continuous batching with a hot in-memory and cold SSD KV cache
- Menu bar app plus CLI commands such as omlx start, stop, restart and serve
- OpenAI-compatible endpoint at localhost:8000/v1
- Discovers LLMs, VLMs, embedding models and rerankers from a model directory
- Built-in chat UI at /admin/chat
- Optional MCP support
When to use it
- Serving a local model to Hermes Agent on a Mac
- Running local models for coding agents with a cache that survives context changes
- Pinning everyday models in memory and swapping heavier ones on demand
Who it is for: Apple Silicon Mac users who want to run local models behind an OpenAI-compatible endpoint for agents such as Hermes.
How it fits with Hermes Agent
The README names Hermes Agent among the clients oMLX can connect to, alongside OpenClaw, OpenCode, Codex, Copilot and DeepSeek Harness, and points to its Integrations section for the steps.
How to install oMLX
These commands are copied from the project's README. Check the repository for the latest steps before you run them.
brew tap jundot/omlx https://github.com/jundot/omlx
brew install jundot/omlx/omlx --with-custom-kernelRequirements: macOS 15.0+ (Sequoia), Python 3.11-3.13, and Apple Silicon (M1/M2/M3/M4/M5)
FAQ
What is oMLX?
oMLX is an LLM inference server for Apple Silicon that uses continuous batching and tiered KV caching. It is managed from the macOS menu bar and serves an OpenAI-compatible API.
Does oMLX work with Hermes Agent?
Yes. The README lists Hermes Agent among the clients that can connect to oMLX and refers to its Integrations section. Any OpenAI-compatible client can use the endpoint at http://localhost:8000/v1.
What do I need to run oMLX?
You need macOS 15.0 or later, Python 3.11 to 3.13 and an Apple Silicon Mac (M1 through M5). Building the native custom kernels locally also requires full Xcode with the Metal toolchain.
Similar models for Hermes Agent
All modelsCLI that reports token usage and costs from local coding agent data, including Hermes Agent
mnfst Manifest LLM GatewayOpen-source LLM gateway with one OpenAI-compatible endpoint, model routing, fallbacks and cost tracking
lemony-ai cascadeflowIn-process runtime that routes agent model calls by cost, latency, quality and policy
Soju06 codex-lbLoad balancer and proxy that pools ChatGPT accounts behind OpenAI-compatible endpoints with a dashboard
Javis603 Token MonitorDesktop widget that tracks token usage and costs across 43+ AI coding tools, including Hermes Agent
xiufengsun TokenTrackerLocal-first dashboard for AI token usage and cost across dozens of coding tools, with native apps
Related guides: How to install Hermes Agent