cascadeflow
lemony-ai/cascadeflow
In-process runtime that routes agent model calls by cost, latency, quality and policy
cascadeflow is an in-process runtime layer for AI agents that decides per step which model to call and whether to continue, based on cost, latency, quality, budget and policy. It has a dedicated Hermes Agent integration for cascading models by skill, task complexity and subagent topic.
What cascadeflow does
cascadeflow sits inside the agent execution loop rather than at the HTTP boundary like an external proxy. It can make a model decision at each step, gate budgets per tool call, and apply four runtime actions: allow, switch_model, deny_tool and stop. It also accumulates information from every model call, tool result and quality score, and records decisions so they can be audited. It is available as a Python package and as TypeScript packages, with integrations for LangChain, OpenAI Agents SDK, CrewAI, PydanticAI, Google ADK, n8n, Vercel AI SDK, OpenClaw and Hermes Agent.
The Hermes Agent integration supports per-skill model cascading, task-complexity cascading, topic-aware subagent cascading, an observe-mode rollout and auditable decisions. According to the README it does not take over provider credentials, base URLs, fallback chains or API modes, so those settings are left alone. Setup details are on the integration page of the cascadeflow documentation.
Key features
- Per-step model decisions inside the agent loop
- Per-tool-call budget gating and four runtime actions: allow, switch_model, deny_tool and stop
- Hermes Agent integration with per-skill, task-complexity and subagent-topic cascading
- Observe-mode rollout and auditable decisions
- Python and TypeScript packages
- Integrations for LangChain, CrewAI, OpenAI Agents SDK, Google ADK, n8n and more
When to use it
- Route simple tasks to smaller models and escalate harder ones in a Hermes Agent setup
- Try cascading in observe mode before letting it change model choices
- Cap spend per tool call and stop a runaway agent loop
Who it is for: Developers who run agents in production and want runtime control over model choice, spend and policy.
How it fits with Hermes Agent
Has a dedicated Hermes Agent integration listed in its README, and supports Hermes among many other agent frameworks.
How to install cascadeflow
These commands are copied from the project's README. Check the repository for the latest steps before you run them.
pip install cascadeflow
npm install @cascadeflow/core FAQ
What is cascadeflow?
cascadeflow is an in-process intelligence layer for AI agents. It optimizes cost, latency, quality, budget and policy decisions inside the agent loop instead of at the HTTP boundary.
Does cascadeflow work with Hermes Agent?
Yes, it provides a Hermes Agent integration for per-skill, task-complexity and subagent-topic model cascading, with an observe mode and auditable decisions. It does not take over provider credentials, base URLs, fallback chains or API modes.
How do I install cascadeflow?
Run pip install cascadeflow for Python or npm install @cascadeflow/core for TypeScript. The Hermes Agent integration is described on the cascadeflow documentation site.
Similar models for Hermes Agent
All modelsLLM inference server for Apple Silicon with continuous batching and SSD caching, run from the menu bar
ccusage ccusageCLI that reports token usage and costs from local coding agent data, including Hermes Agent
mnfst Manifest LLM GatewayOpen-source LLM gateway with one OpenAI-compatible endpoint, model routing, fallbacks and cost tracking
Soju06 codex-lbLoad balancer and proxy that pools ChatGPT accounts behind OpenAI-compatible endpoints with a dashboard
Javis603 Token MonitorDesktop widget that tracks token usage and costs across 43+ AI coding tools, including Hermes Agent
xiufengsun TokenTrackerLocal-first dashboard for AI token usage and cost across dozens of coding tools, with native apps
Related guides: How to install Hermes Agent