Hermes Atlas
Research, training & evaluation · works with Hermes Agent

MerchantBench

KhanCold/merchantbench

365-day simulated e-commerce benchmark for LLM agents, with a Hermes adapter

In short

MerchantBench is a benchmark that puts an LLM agent in charge of a simulated online store for 365 days. A separate Hermes runtime lets Hermes Agent be run and evaluated on it.

What MerchantBench does

MerchantBench tests whether an agent can keep a coherent strategy over a long horizon rather than finish one bounded task. The agent must source products, manage listings and prices, control cash flow and adapt to changing market conditions. The simulation couples supplier changes that are visible right away with order outcomes that arrive later, so earlier decisions have to be revised as evidence accumulates. Demand is simulated as individual orders that move through procurement, fulfillment, delivery, settlement and after-sales. The project comes from Alibaba Group (1688) and Zhejiang University.

The repository holds a Flask simulator with a dashboard, an agent SDK with reference baselines, a Docker-based evaluation framework, a batch runner and a test suite. A rule-based baseline runs without an API key. The Hermes-specific runtime is released separately in the KhanCold/hermes-agent repository, and the two repositories are meant to be cloned side by side.

Key features

  • 365-day simulated store with order-level dynamics
  • Upstream supplier simulation combined with delayed downstream order outcomes
  • Flask simulator with a dashboard and a browser playground for human players
  • Rule-based baseline that runs without an API key, plus an LLM-driven ReAct baseline
  • Docker-based evaluation framework and a batch runner for model sweeps
  • Hermes adapter released in a separate repository

When to use it

  • Comparing how different models handle long-horizon operations as a Hermes Agent backend
  • Running a deterministic baseline to check a new agent submission against the simulator
  • Studying how agents react to delayed feedback in supplier and order data

Who it is for: Researchers and agent developers who want to evaluate long-term coherence of LLM agents, including Hermes Agent, on a business-operations task.

How it fits with Hermes Agent

A general agent benchmark with a Hermes integration. The README says the Hermes-specific runtime lives in the KhanCold/hermes-agent repository and was released on 2026-08-11.

How to install MerchantBench

These commands are copied from the project's README. Check the repository for the latest steps before you run them.

python3.11 -m venv .venv
.venv/bin/python -m pip install --upgrade pip
.venv/bin/python -m pip install -r requirements.txt
PYTHONPATH=env:agent .venv/bin/python -m pytest tests/

Requirements: Python 3.10+ (3.11 recommended); Docker only for containerized evaluation; an OpenAI-compatible API key only for the LLM-driven ReAct or Hermes agents

Note: Real-world business data is not included; the repository provides a synthetic-data generator with 1,000 products and 200 suppliers by default.

FAQ

What is MerchantBench?

MerchantBench is a 365-day, order-level benchmark for LLM agents that run the seller side of an online store. It measures whether an agent's decisions stay coherent as delayed results come in.

Does MerchantBench work with Hermes Agent?

Yes, through a separate repository. The README says the MerchantBench-specific Hermes runtime is in KhanCold/hermes-agent and should be cloned next to merchantbench.

Is MerchantBench free and open source?

Yes. The repository is licensed under Apache-2.0 and includes the simulator, agent SDK, baselines and evaluation framework. Real-world business data is not included.

Similar research for Hermes Agent

All research

Related guides: What is Hermes Agent?