Hermes Atlas
Memory & knowledge

Multimodal Wiki

kigner/multimodal-wiki

Hermes skill that builds an interlinked markdown wiki from text, web pages, PDFs, screenshots and audio

In short

Multimodal Wiki is a Hermes skill that builds a personal knowledge base in the style of Karpathy's LLM Wiki pattern, ingesting text, web pages, PDFs, screenshots and audio into one markdown vault with provenance. It forks Hermes's built-in llm-wiki skill and adds vision and audio ingestion.

What Multimodal Wiki does

Multimodal Wiki is a Hermes skill that builds a compounding personal knowledge base in the style of Andrej Karpathy's LLM Wiki pattern. It is a self-contained fork of Hermes' built-in llm-wiki skill, extended with screenshot ingestion through a vision model, audio ingestion through whisper, and a stronger contract from raw sources to the compiled layer. Text, web pages, PDFs, screenshots and audio all end up in one interlinked markdown vault with provenance.

The skill is instructions only, so you configure it yourself: set WIKI_PATH in the Hermes .env, point auxiliary.vision at a vision-capable model for screenshots, and set stt to the local provider with faster-whisper for audio, where large-v3 downloads about 3 GB. On first run you ask it to create a wiki and it builds SCHEMA.md, index.md, log.md and the raw tree. It supports ingest, query and lint workflows, and scripts/audit.py runs six audit checks in one pass. Obsidian can open the same vault.

Key features

  • Ingests text, web pages, PDFs, screenshots and audio into one markdown vault
  • Screenshot ingestion through a vision-capable model set as auxiliary.vision
  • Audio transcription with faster-whisper, saved as verbatim transcripts
  • Note-level provenance with copy-pasteable raw source paths
  • Ingest, query and lint workflows, plus an audit script with six checks
  • Optional Obsidian graph view by pointing OBSIDIAN_VAULT_PATH at the vault

When to use it

  • Dropping a screenshot or voice memo into chat and having it filed in your wiki
  • Asking questions that are answered from the compiled wiki with source paths
  • Auditing a growing vault for broken provenance or unindexed files

Who it is for: Hermes users who keep a personal knowledge base and want screenshots and audio ingested alongside text and web pages.

How it fits with Hermes Agent

A Hermes skill installed under the Hermes skills tree that supersedes the built-in llm-wiki skill; the README suggests disabling the built-in one so both do not activate.

Requirements: WIKI_PATH set in the Hermes .env; a vision-capable model for auxiliary.vision to ingest screenshots; faster-whisper with a local stt configuration for audio

Note: The repository has no license file, and the skill is instructions only, so you must configure the vault path, vision model and speech-to-text yourself.

FAQ

What is Multimodal Wiki?

Multimodal Wiki is a Hermes skill that maintains a compounding personal wiki from text, web pages, PDFs, screenshots and audio. It is a fork of Hermes' built-in llm-wiki skill with extra vision and audio ingestion.

Does Multimodal Wiki work with Hermes Agent?

Yes, it is built as a Hermes skill. Copy the multimodal-wiki folder into your Hermes skills tree, restart Hermes, and consider disabling the built-in llm-wiki so the two do not both activate.

What do I need to run Multimodal Wiki?

You need WIKI_PATH set in your Hermes .env, a vision-capable model configured as auxiliary.vision for screenshots, and a local speech-to-text setup with faster-whisper for audio.

Similar memory for Hermes Agent

All memory

Related guides: SOUL.md for Hermes Agent: what it is and how to write one · Run multiple Hermes agents with profiles