Hermes Atlas
Desktop apps, web UIs & dashboards

Hermes Voice Assistant

l0cut15/hermes-voice-assistant

Push-to-talk voice assistant on M5Stack CoreS3 using whisper.cpp, a Hermes agent and Kokoro TTS

In short

Hermes Voice Assistant is firmware for the M5Stack Core S3 SE that turns it into a standalone WiFi push-to-talk voice assistant. It transcribes speech with a local whisper.cpp server, asks a Hermes agent, and speaks the answer with Kokoro TTS.

What Hermes Voice Assistant does

Hermes Voice Assistant is a standalone WiFi push-to-talk voice assistant for the M5Stack Core S3 SE. You hold the touch screen to speak. The device records 16 kHz PCM into PSRAM, sends it to a local whisper.cpp server for transcription, sends the text to a Hermes agent's API server, then passes the reply to a Kokoro TTS server and plays the audio on its built-in speaker.

All three servers run on a machine on your local network, and the Hermes agent uses whatever LLM provider you configure in ~/.hermes/config.yaml, cloud or local. The README gives Docker setups for whisper.cpp and Kokoro, plus a macOS bare-process option for whisper.cpp that it says transcribes about 70 times faster than real time on Apple Silicon. Hermes needs its API server enabled on port 7237, bound to 0.0.0.0 so the device can reach it, with a key that matches hermes_key in secrets.json. The firmware is built and flashed with PlatformIO, and WiFi credentials and server hosts live in a secrets.json file stored in LittleFS.

Key features

  • Push-to-talk by holding the touch screen
  • Local speech-to-text with whisper.cpp over HTTP
  • Hermes agent reached through its API server, with any LLM that Hermes supports
  • Kokoro TTS replies played on the device speaker
  • Docker setups for whisper.cpp and Kokoro, plus a native macOS whisper option
  • Secrets kept in LittleFS and flashed once with the uploadfs target

When to use it

  • A desk voice assistant that talks to a local Hermes agent
  • Keeping speech recognition and synthesis on your own network while choosing any LLM through Hermes
  • A hardware project for an M5Stack Core S3 SE with a touchscreen and speaker

Who it is for: Makers with an M5Stack Core S3 SE who want a local push-to-talk voice front end for Hermes Agent.

How it fits with Hermes Agent

Built for Hermes Agent: the firmware calls the Hermes API server on port 7237 with a shared key.

How to install Hermes Voice Assistant

These commands are copied from the project's README. Check the repository for the latest steps before you run them.

cp data/secrets.example.json data/secrets.json
python -m platformio run
python -m platformio run --target upload
python -m platformio run --target uploadfs

Requirements: An M5Stack Core S3 SE, a local machine running whisper.cpp (port 7124), Hermes Agent with its API server enabled (port 7237) and Kokoro TTS in Docker (port 7235), plus VS Code with PlatformIO, Python 3.10 to 3.13 and a USB-C cable

FAQ

What is Hermes Voice Assistant?

It is firmware for the M5Stack Core S3 SE that makes the device a standalone WiFi voice assistant. You hold the screen to speak, and the reply from your Hermes agent is played through the built-in speaker.

Does Hermes Voice Assistant work with Hermes Agent?

Yes, it is built for it. Hermes must have its API server enabled on port 7237 and reachable from the device, and any LLM that Hermes supports can be used.

What do I need to run Hermes Voice Assistant?

You need an M5Stack Core S3 SE, a local whisper.cpp server, a Hermes Agent with its API server on, and a Kokoro TTS container. To build the firmware you also need VS Code with PlatformIO and Python 3.10 to 3.13.

Similar interfaces for Hermes Agent

All interfaces

Related guides: Connect Hermes agents on several machines with Hermes Desktop · How to install Hermes Agent