Hermes Atlas
Research, training & evaluation · works with Hermes Agent

Multi Agent Benchmark Tool

digitalspaceport/Multi-Agent-Benchmark-Tool

Async benchmark that simulates many agents hitting one OpenAI-compatible endpoint such as vLLM

In short

Multi Agent Benchmark Tool is an asynchronous Python benchmark that emulates several autonomous agents sending requests to a single OpenAI-compatible endpoint. It helps people planning to run multiple local agents, including Hermes-style ones, tune their inference stack.

What Multi Agent Benchmark Tool does

Multi Agent Benchmark Tool is a lightweight asynchronous Python benchmark that emulates several autonomous agents sending requests to one OpenAI-compatible endpoint, such as vLLM or llama.cpp. It measures time to first token, time per output token, request latency including p95, prefill and decode throughput and requests per second. Realistic behavior includes streamed tool calls, follow-up requests after tool output and multimodal image-URL inputs.

Its author describes the workload as OpenClaw or Hermes-style, and the README states that this refers to the prompting and interaction pattern, not a formal compatibility claim with any Hermes framework. Options include -n for the number of agents, --max-turns, --target-total-rps, --min-per-agent-rps, --enable-tools, --tool-followup and --include-usage for accurate streaming token counts. It helps tune an inference stack before running several local agents against it.

Key features

  • Measures TTFT, TPOT, latency percentiles including p95, throughput and RPS
  • Simulated tool calls with streamed aggregation and follow-up requests
  • Multimodal user inputs using image URLs
  • Pacing controls with jitter to avoid synchronized request spikes
  • Usage-based token counts when the server supports streaming usage
  • Timeouts and SDK-level retries

When to use it

  • Checking how many concurrent local agents a vLLM server can handle
  • Comparing inference settings with tool-calling workloads
  • Sizing hardware before running several Hermes or OpenClaw agents on local models

Who it is for: People who run local LLM servers and want to measure them under multi-agent load.

How it fits with Hermes Agent

It is a general benchmark; the README says its Hermes-style label describes the prompting pattern and is not a formal compatibility claim with any Hermes framework.

How to install Multi Agent Benchmark Tool

These commands are copied from the project's README. Check the repository for the latest steps before you run them.

pip install --upgrade openai

Requirements: Python 3.9+, the openai Python SDK and an OpenAI-compatible endpoint

Note: The README states that Hermes-style describes the workload pattern and is not a formal compatibility claim with Hermes Agent.

FAQ

What is Multi Agent Benchmark Tool?

It is an asynchronous Python benchmark that simulates multiple agents sending requests to one OpenAI-compatible endpoint. It reports time to first token, time per output token, latency and throughput.

Does Multi Agent Benchmark Tool work with Hermes Agent?

It does not run Hermes Agent. The README says Hermes-style refers only to the prompting and interaction pattern of the workload, not a formal compatibility claim.

What do I need to run Multi Agent Benchmark Tool?

You need Python 3.9 or newer, the openai Python SDK, and an OpenAI-compatible server such as vLLM or llama.cpp. Run python mabt.py with --base-url, --model and -n for the agent count.

Similar research for Hermes Agent

All research

Related guides: What is Hermes Agent?