DGX Spark Single-Stack AI Agent
marksunner/dgx-spark-single-stack
Documented single-box setup running Qwen 3.5 122B, Hermes Agent and Honcho on one NVIDIA DGX Spark
DGX Spark Single-Stack AI Agent is a documented, reproducible configuration for running a local LLM, Hermes Agent, Honcho conversation memory and Uptime Kuma monitoring on one 128GB NVIDIA DGX Spark.
What DGX Spark Single-Stack AI Agent does
The repository documents a four-part stack on one machine. Qwen 3.5 122B serves inference through vLLM from a hybrid INT4+FP8 checkpoint with MTP speculative decoding and uses about 67 GB of GPU memory. Hermes Agent runs on CPU as the agent runtime with tool use and Discord or chat integration, Honcho stores conversation memory in PostgreSQL and Redis, and Uptime Kuma provides health monitoring. The README's key point is that the GPU and CPU workloads do not compete, because vLLM reserves GPU memory exclusively.
The docs include a from-scratch walkthrough estimated at about two hours, a vLLM flag reference, steps for building the hybrid checkpoint, Hermes and Honcho setup, thinking-mode controls, troubleshooting and a benchmark methodology. The author reports 41 to 47 tokens per second for a single user and says one Spark outperformed two Sparks running Ray.
Key features
- vLLM launch command for Qwen 3.5 122B with MTP speculative decoding
- Hermes Agent as the CPU-side agent runtime
- Optional Honcho conversation memory on PostgreSQL and Redis
- Uptime Kuma health monitoring
- Documentation for the hybrid checkpoint, thinking modes, troubleshooting and benchmarks
When to use it
- Build a fully local Hermes Agent setup on a single DGX Spark
- Reproduce the author's vLLM configuration and hybrid INT4+FP8 checkpoint
- Compare single-Spark and dual-Spark inference results
Who it is for: Owners of an NVIDIA DGX Spark who want a worked example of running a large local model with Hermes Agent on one box.
How it fits with Hermes Agent
Hermes Agent is the agent runtime in the stack, and the repository includes a docs/hermes-setup.md guide for installing it.
Requirements: A 128GB NVIDIA DGX Spark; the vLLM step uses Docker with GPU access
Note: The repository holds documentation and configuration notes rather than installable software, and the benchmark figures are reported by the author.
FAQ
What is DGX Spark Single-Stack AI Agent?
It is a set of documentation and configuration for running an autonomous AI agent stack on one NVIDIA DGX Spark. The stack combines Qwen 3.5 122B, Hermes Agent, Honcho memory and Uptime Kuma.
Does DGX Spark Single-Stack work with Hermes Agent?
Yes, Hermes Agent is the agent runtime in the stack. It runs on CPU while vLLM holds the GPU memory, and the repository includes a Hermes setup guide.
Is DGX Spark Single-Stack free and open source?
The repository has no license file, so default copyright applies and you should check with the author before reuse. The README says the documentation is shared freely and that vLLM, Hermes, Honcho and Qwen have their own licenses.
Similar deployment for Hermes Agent
All deploymentWindows-native fork of Hermes Agent 0.13.0 with path, process and terminal fixes, no WSL required
tecno-consultores llm-labDocker Compose stack with profiles for n8n, Open WebUI, Ollama, Qdrant, OpenCode and Hermes Agent
JackTheGit Hermes Autonomous Server GuideGuide to running Hermes Agent headless on a Linux server with a systemd gateway and cron jobs
catamsp Hermes Agent for Termux (proot-distro)Installer scripts that run Hermes Agent on Android inside Termux using a proot-distro Debian or Ubuntu
JiaDe-Wu Hermes Agent on AWS with BedrockOne-click CloudFormation deployment of Hermes Agent on AWS using Amazon Bedrock and IAM authentication
oablab ecsctlkubectl-style Rust CLI for deploying and managing AI agent services such as Hermes on Amazon ECS
Related guides: How to run Hermes Agent securely · How to install Hermes Agent · Connect Hermes agents on several machines with Hermes Desktop