Hermes Atlas
Skills & skill packs

Multimodal LLM Wiki

nknishio/LLM-Wiki-Multimodal-TWM-Agent-Skill

Hermes skill that turns PDFs, decks, spreadsheets and images into a cross-linked Markdown wiki

In short

Multimodal LLM Wiki is an agent skill that extends the built-in Hermes llm-wiki skill so it can ingest PDFs, slide decks, spreadsheets and images. It captions every embedded image so it becomes searchable text.

What Multimodal LLM Wiki does

Multimodal LLM Wiki extends the Hermes Agent built-in llm-wiki skill with a multimodal ingest path. The base skill reads only URLs and pasted text. This one converts PDF, PPTX, DOCX, XLSX, CSV and HTML files to Markdown, extracts embedded images while dropping tiny noise images, and has the agent write a factual caption for each. Captions are stored as ordinary Markdown text, so search finds them, and images stay referenced inline so they still render in Obsidian.

Captioning uses the agent's own multimodal reading, so no extra API key or vision service is needed. A SHA-256 cache means an identical image is described once, even across different decks. The bundled llm-wiki CLI is pure Python, using PyMuPDF, python-pptx, python-docx, pandas and markdownify, and it never calls an LLM. The wiki is a directory of Markdown files in three layers: raw sources, agent-owned pages linked with wikilinks, and a schema file. The README says it is available to all employees at Taiwan Mobile.

Key features

  • Converts PDF, PPTX, DOCX, XLSX, CSV and HTML to Markdown
  • Extracts embedded images with a built-in noise filter
  • Captions images through the agent's own multimodal read, with no extra API key
  • SHA-256 caption cache so each image is described once
  • Pure-Python llm-wiki CLI with convert, extract-images and caption get or put commands
  • Plain Markdown output with wikilinks that opens in Obsidian or any editor

When to use it

  • Turn a folder of research papers and slide decks into a searchable wiki
  • Make charts and diagrams in documents findable by text search
  • Re-ingest a revised deck without paying again to caption unchanged images

Who it is for: Hermes Agent users who keep research in binary documents and want an agent-maintained Markdown knowledge base.

How it fits with Hermes Agent

It builds on the Hermes llm-wiki skill and is tagged hermes-skill, shipped as a SKILL.md folder that drops into an agent's skills directory.

Note: The README describes it as a proof of concept, developed and smoke-tested in Claude Code.

FAQ

What is Multimodal LLM Wiki?

Multimodal LLM Wiki is an agent skill that converts PDFs, slide decks, spreadsheets and images into a cross-linked Markdown knowledge base, captioning embedded images so they become searchable.

Does Multimodal LLM Wiki work with Hermes Agent?

Yes. It extends the built-in Hermes llm-wiki skill. The README says it was developed and smoke-tested in Claude Code and is designed to integrate into Taiwan Mobile's internal Hermes agent.

Is Multimodal LLM Wiki free and open source?

Yes, it is open source under an MIT license.

Similar skills for Hermes Agent

All skills

Related guides: What is Hermes Agent? · SOUL.md for Hermes Agent: what it is and how to write one