Once your agent gets busy, your bill is decided less by which single model you pick and more by how you route work across models. The whole principle fits in one line: reason on a strong model, orchestrate on a cheap one. Get that split right and a heavy agent workload costs surprisingly little. Get it wrong and you pay frontier prices for glue work.
This is an advanced piece. It assumes you have already connected a provider or two earlier in the Agentic AI series.
Why agent loops are so token-hungry
An agent is not one prompt β it is a loop. And here is the thing that drives the cost: every tool result comes back into the context on the next turn. Read a file, search the web, run a command β each result is re-sent as input on the following step, and the conversation only grows. So in an agent loop, input tokens dominate and keep climbing, which means the input price β and especially the cache discount on it β matters far more than the output price.
A concrete example from my own use: one real session ran roughly 235 turns and around 19 million tokens, and it cost well under a dollar β because it ran on DeepSeek V4 Flash, a budget-tier model built for exactly this kind of high-volume agent work. The same volume on a flagship model would have been an order of magnitude more. That is the whole game: push volume to the cheap tier, keep reasoning in the main agent.
The ladder
Think in two rungs rather than one model for everything:
- Strong model (frontier tier) β deep reasoning, research, hard synthesis, code that has to be correct. It is called relatively few times, so its high per-token price barely registers.
- Cheap model (budget tier) β orchestration, summarising, formatting, retrieval, “glue,” and subagents fanning out over sources. This is called constantly, so its low per-token price is what actually controls the bill.

Every major provider gives you both rungs β a budget and a frontier variant of the same family. I am deliberately not quoting numbers here, because model prices change constantly; check the provider’s live pricing page before you commit. The ratio is the durable point: within a family the frontier model is typically several times the price of the budget one, which is exactly why routing beats picking.
How Hermes routes work across models
Hermes gives you four mechanisms to put the right model on the right job:
- Main model + auxiliary slots. Beyond the main model the agent “thinks” with, Hermes has 11 auxiliary slots for side jobs β context compression, image analysis, web-page summarising, approval scoring, MCP tool routing, session-title generation, skill search. Each is independently overridable, so you can drop the cheapest usable model onto all the side work while the main model stays strong.
- Mixture of Agents (MoA). A virtual provider that runs several cheap reference models (they see only plain text, no tools, so they stay cheap) under one strong aggregator that writes the real answer. Many cheap perspectives, one strong voice.
- Delegation. Subagents run on a pinned cheap model while the planner stays on frontier β so the volume runs cheap while the reasoning stays strong.
- Fallback chains, provider-agnostic. Hermes speaks to 35+ providers (local and hosted), lets you switch models mid-session with
/model, and chains fallbacks withhermes fallback add|remove|listso a failed provider automatically rolls to the next. Under the hood those commands edit the top-levelfallback_providerslist in~/.hermes/config.yaml, so if you would rather script it, edit that list directly. (The older singlefallback_modelkey still works, butfallback_providerstakes priority.)
The cost levers that actually move the needle
- Prompt caching is the big one. Cached input costs a fraction of fresh input β typically 50β90% less, and some providers go even higher. The exact discount depends on the provider and on whether caching is automatic or opt-in, so check your provider’s docs. In a long agent loop, the cache hit is the real working price, not the headline input rate.
- Do not switch models mid-session carelessly. This is the trap. A model change partway through a session β an explicit
/model, an automatic fallback, or a credential-pool rotation β invalidates the prompt cache. The next turn re-reads the entire conversation at full input price instead of the cached rate, and on a long session that single re-read can wipe out everything you saved by going cheaper. Switch early in a conversation, not deep into one. - Input is cheap, output is dear. Design prompts so the agent uses tools and answers tersely rather than generating long essays.
- Offload the auxiliaries. Moving the side jobs off the main model onto a budget one is free money β those slots do not need frontier reasoning.

A sane routing table
Mapping roles to tiers, without a single price attached:
| Role | Tier | Call volume | Why |
|---|---|---|---|
| Deep research / hard reasoning | Frontier | Low | Few calls, high value per token |
| Orchestration / coordination | Budget or mid | High | Simple instruction-following |
| Bulk / parallel subagents | Budget | Very high | Volume is the cost, so cheap pays off most here |
| Auxiliary side jobs | Cheapest usable | Constant | No reasoning required |
| MoA reference calls | Budget, no tools | Medium | Multiple perspectives that only read text |
One more route: a single subscription
If juggling several providers and keys sounds like a chore, a subscription route such as the Nous Portal bundles a large catalogue of models behind one plan and a single key, and applies a discount on token-billed providers on top. It is a convenience-versus-control trade-off; I am not quoting plan prices here, so check the current tiers before deciding.
The takeaway is not “find the cheapest model.” It is: stop optimising the sticker price of one model, and start optimising the flow of work across a ladder. Put your main agent on the strongest model you can justify for the turns it personally takes, and push every bit of volume down to the cheapest tier that can still do the job. For the numbers, always trust the providers’ live pricing pages over any figure written down months ago.
This article is part of the Agentic AI series.
Liked this? There's more where it came from.
Get our digest β articles worth your time, no spam, unsubscribe in one click.
Subscribe to Factnetize →
[…] up: once your agent is doing real work like this, the bill starts to matter. Model routing and cost control for your AI agent shows how to route work across a strong and a cheap model so a heavy workload stays […]