Turbofit
SouthpawIN/turbofit
Adaptive local-inference provider that fits the best sustainable model to your hardware
Turbofit is a first-class Hermes Agent model provider and adaptive local-inference runtime that inventories a machine's compute and memory, recommends an evidence-backed model, launches a native backend and exposes one OpenAI-compatible endpoint.
What Turbofit does
TurboFit Check scans whether the machine has dedicated VRAM, unified or integrated memory, or RAM only, then applies either an automatic or a manually selected compatible model tier from the TurboFit List, a ranked, benchmark-scored lineup of five models suitable for running Hermes Agent locally, from a 5.5 GiB ternary MoE model up to a 160 GiB DeepSeek V4.1 Flash deployment. The fit rule is explicit: model size plus KV cache must stay within the tier's size budget, with hardware tiers documented separately for GPU VRAM, system RAM and Apple Silicon unified memory.
Configured in Hermes as provider: custom:turbofit with model: auto, Turbofit then picks and launches the matching native backend itself. A fuller model zoo beyond the five-model TurboFit List, along with benchmark scores (AIME26, GPQA-D, DeepSWE, LiveCodeBench v6 and TB4.0) and hardware-tier tables, is documented in the project's wiki.
Key features
- Hardware inventory (dedicated VRAM, unified memory or RAM-only) with an evidence-backed model recommendation
- Five-model TurboFit List ranked by capability, from a 5.5 GiB model up to 160 GiB
- One OpenAI-compatible endpoint exposed regardless of which native backend is launched
- Separate hardware-tier tables for GPU VRAM, system RAM and Apple Silicon unified memory
- Benchmark scores (AIME26, GPQA-D, DeepSWE, LiveCodeBench v6, TB4.0) per model
When to use it
- Letting Hermes Agent pick the best model a given machine can run, without manual tuning
- Running Hermes fully locally on hardware ranging from under 8GB VRAM up to 256GB RAM
- Comparing local model options for Hermes by benchmark score before committing to one
Who it is for: Hermes Agent users who want to run models locally and have the right model for their hardware chosen automatically rather than guessed.
How it fits with Hermes Agent
Turbofit is described as a first-class Hermes Agent provider, configured directly as provider: custom:turbofit inside Hermes' own provider configuration.
FAQ
What is Turbofit?
Turbofit is a first-class Hermes Agent model provider that inventories a machine's hardware and automatically runs the best local model its memory can sustain, behind one OpenAI-compatible endpoint.
Does Turbofit work with Hermes Agent?
Yes, it is described as a first-class Hermes Agent provider, configured directly in Hermes as provider: custom:turbofit with model: auto.
What do I need to run Turbofit?
You need a machine with dedicated VRAM, unified memory or sufficient RAM for one of the TurboFit List's five tiers, since the fit rule requires model size plus KV cache to stay within that tier's size budget.
Similar models for Hermes Agent
All modelsContext reuse layer that cuts redundant tokens before they reach the model
dyedd LensSelf-hosted multi-protocol LLM gateway with one base URL across many providers
Socialpranker agentburnLocal profiler that shows which 5-hour window burned your usage, and where tokens went
tuxevil tuxevil-rotatorOpenAI-compatible proxy that rotates free-tier LLM accounts with per-model quota routing
piyush-tyagi-13 llm-keypoolFree-tier LLM API key pool with rotation, 429 cooldowns and a local OpenAI-compatible proxy
InfiniteWhispers HermesAgent-MultiModelFully local multi-model setup for Hermes Agent with Ollama, Mixture-of-Agents routing and a tuning guide
Related guides: How to install Hermes Agent