Hermes Atlas
Developer tools & SDKs · works with Hermes Agent

ModLens

liustack/modlens

Vision plugin that lets text-only coding agents read pasted images as structured JSON evidence

In short

ModLens is a vision plugin and skill that gives text-only models the ability to read images, returning structured evidence such as OCR text, layout and semantics. It is tagged hermes-agent, but the opening part of the README documents DeepSeek Harness, Claude Code, Codex, OpenCode and Pi.

What ModLens does

ModLens turns a pasted image into structured JSON evidence that a text-only model can quote: full transcription, layout regions in reading order, and entity and relation lists. On a text-only model, a pasted image is saved as a private temp file and the modlens_read_image tool takes over. In DeepSeek Harness it can also add (modlens vision) entries to the model selector for eligible DeepSeek, GLM and MiMo Pro routes, so the thumbnail stays visible while the image is converted at request time.

It is built as a single skill folder on skill-based agents and a single plugin on DeepSeek Harness, with no hooks, wrappers or local proxy daemon, so uninstalling means deleting a folder. For the vision backend it reuses setup already present in Claude Code, Codex, OpenCode and Pi, can use the Antigravity CLI as a no-key channel or a Gemini key, and accepts API keys from OpenAI-compatible providers. Comma-separated keys rotate on auth, rate-limit or quota failures.

Key features

  • Structured JSON evidence from images: transcription, layout regions, entities and relations
  • Paste images directly into the chat instead of saving a file first
  • Auto-discovered (modlens vision) model entries for eligible text-only routes
  • Reuses vision setups already present in Claude Code, Codex, OpenCode and Pi
  • Key rotation on auth, rate-limit and quota failures
  • No hooks, wrappers or local proxy daemon; uninstall by deleting a folder

When to use it

  • Let DeepSeek or GLM text-only models read screenshots and diagrams
  • Extract text and layout from an image through OCR evidence
  • Add image understanding to a coding agent without changing its harness config

Who it is for: Developers using text-only models in coding agents who need those agents to read screenshots and other images.

How it fits with Hermes Agent

Tagged hermes-agent on GitHub, but the opening part of the README documents DeepSeek Harness, Claude Code, Codex, OpenCode and Pi and does not describe a Hermes Agent setup.

How to install ModLens

These commands are copied from the project's README. Check the repository for the latest steps before you run them.

npx -y @deepseek-ai/dsh plugin --profile web add @liustack/modlens@3.26.6

Requirements: A vision backend: an existing Claude Code, Codex, OpenCode or Pi setup, the Antigravity CLI, a Gemini key, or an API key from an OpenAI-compatible provider

FAQ

What is ModLens?

ModLens is a vision plugin that gives text-only models sight. You paste an image and it returns structured JSON evidence such as OCR text, layout and semantics.

Does ModLens work with Hermes Agent?

The repository is tagged hermes-agent, but the opening part of the README documents DeepSeek Harness and says it was verified in Claude Code, Codex, Pi and OpenCode. It does not describe a Hermes Agent setup.

How do I install ModLens?

On DeepSeek Harness, run npx -y @deepseek-ai/dsh plugin --profile web add @liustack/modlens@3.26.6. For other harnesses the README offers installation through skills.sh and points to a setup guide.

Similar dev tools for Hermes Agent

All dev tools

Related guides: How to install Hermes Agent · Run multiple Hermes agents with profiles