🌱 Welcome to Factnetize β€” honest, well-researched articles on tech, health and AI. Read more →
Agentic AI

Model routing and cost control for your AI agent

In a busy agent, cost is decided by how you route work across models, not which one you pick. The principle: reason on a strong model, orchestrate on a cheap one, plus the caching trap that quietly wipes out your savings.

Model routing and cost control for your AI agent

Once your agent gets busy, your bill is decided less by which single model you pick and more by how you route work across models. The whole principle fits in one line: reason on a strong model, orchestrate on a cheap one. Get that split right and a heavy agent workload costs surprisingly little. Get it wrong and you pay frontier prices for glue work.

This is an advanced piece. It assumes you have already connected a provider or two earlier in the Agentic AI series.

Why agent loops are so token-hungry

An agent is not one prompt β€” it is a loop. And here is the thing that drives the cost: every tool result comes back into the context on the next turn. Read a file, search the web, run a command β€” each result is re-sent as input on the following step, and the conversation only grows. So in an agent loop, input tokens dominate and keep climbing, which means the input price β€” and especially the cache discount on it β€” matters far more than the output price.

A concrete example from my own use: one real session ran roughly 235 turns and around 19 million tokens, and it cost well under a dollar β€” because it ran on DeepSeek V4 Flash, a budget-tier model built for exactly this kind of high-volume agent work. The same volume on a flagship model would have been an order of magnitude more. That is the whole game: push volume to the cheap tier, keep reasoning in the main agent.

The ladder

Think in two rungs rather than one model for everything:

  • Strong model (frontier tier) β€” deep reasoning, research, hard synthesis, code that has to be correct. It is called relatively few times, so its high per-token price barely registers.
  • Cheap model (budget tier) β€” orchestration, summarising, formatting, retrieval, “glue,” and subagents fanning out over sources. This is called constantly, so its low per-token price is what actually controls the bill.
Routing ladder diagram with three tiers: frontier tier for deep research, correct code and the main agent's turns; budget or mid tier for orchestration, parallel subagents and MoA reference models; cheapest usable model for auxiliary jobs like compression, session titles and tool routing
Put each job on the cheapest tier that can still do it: few expensive calls at the top, high-volume work at the bottom.

Every major provider gives you both rungs β€” a budget and a frontier variant of the same family. I am deliberately not quoting numbers here, because model prices change constantly; check the provider’s live pricing page before you commit. The ratio is the durable point: within a family the frontier model is typically several times the price of the budget one, which is exactly why routing beats picking.

How Hermes routes work across models

Hermes gives you four mechanisms to put the right model on the right job:

  • Main model + auxiliary slots. Beyond the main model the agent “thinks” with, Hermes has 11 auxiliary slots for side jobs β€” context compression, image analysis, web-page summarising, approval scoring, MCP tool routing, session-title generation, skill search. Each is independently overridable, so you can drop the cheapest usable model onto all the side work while the main model stays strong.
  • Mixture of Agents (MoA). A virtual provider that runs several cheap reference models (they see only plain text, no tools, so they stay cheap) under one strong aggregator that writes the real answer. Many cheap perspectives, one strong voice.
  • Delegation. Subagents run on a pinned cheap model while the planner stays on frontier β€” so the volume runs cheap while the reasoning stays strong.
  • Fallback chains, provider-agnostic. Hermes speaks to 35+ providers (local and hosted), lets you switch models mid-session with /model, and chains fallbacks with hermes fallback add|remove|list so a failed provider automatically rolls to the next. Under the hood those commands edit the top-level fallback_providers list in ~/.hermes/config.yaml, so if you would rather script it, edit that list directly. (The older single fallback_model key still works, but fallback_providers takes priority.)

The cost levers that actually move the needle

  • Prompt caching is the big one. Cached input costs a fraction of fresh input β€” typically 50–90% less, and some providers go even higher. The exact discount depends on the provider and on whether caching is automatic or opt-in, so check your provider’s docs. In a long agent loop, the cache hit is the real working price, not the headline input rate.
  • Do not switch models mid-session carelessly. This is the trap. A model change partway through a session β€” an explicit /model, an automatic fallback, or a credential-pool rotation β€” invalidates the prompt cache. The next turn re-reads the entire conversation at full input price instead of the cached rate, and on a long session that single re-read can wipe out everything you saved by going cheaper. Switch early in a conversation, not deep into one.
  • Input is cheap, output is dear. Design prompts so the agent uses tools and answers tersely rather than generating long essays.
  • Offload the auxiliaries. Moving the side jobs off the main model onto a budget one is free money β€” those slots do not need frontier reasoning.
Line chart of input cost per turn over a 60-turn agent session: with the same model the cost rises slowly as the cache stays warm, while a model switch at turn 40 spikes that turn to about eight times a normal turn because the whole conversation is re-read at full price
Illustrative model, relative units: switching models mid-session forces one full-price re-read of the entire context.

A sane routing table

Mapping roles to tiers, without a single price attached:

Role Tier Call volume Why
Deep research / hard reasoning Frontier Low Few calls, high value per token
Orchestration / coordination Budget or mid High Simple instruction-following
Bulk / parallel subagents Budget Very high Volume is the cost, so cheap pays off most here
Auxiliary side jobs Cheapest usable Constant No reasoning required
MoA reference calls Budget, no tools Medium Multiple perspectives that only read text

One more route: a single subscription

If juggling several providers and keys sounds like a chore, a subscription route such as the Nous Portal bundles a large catalogue of models behind one plan and a single key, and applies a discount on token-billed providers on top. It is a convenience-versus-control trade-off; I am not quoting plan prices here, so check the current tiers before deciding.

The takeaway is not “find the cheapest model.” It is: stop optimising the sticker price of one model, and start optimising the flow of work across a ladder. Put your main agent on the strongest model you can justify for the turns it personally takes, and push every bit of volume down to the cheapest tier that can still do the job. For the numbers, always trust the providers’ live pricing pages over any figure written down months ago.

This article is part of the Agentic AI series.

John Lock
Written by

John Lock

Liked this? There's more where it came from.

Get our digest β€” articles worth your time, no spam, unsubscribe in one click.

Subscribe to Factnetize →

1 comment

  1. October 8, 2026 at 6:24 pm

    […] up: once your agent is doing real work like this, the bill starts to matter. Model routing and cost control for your AI agent shows how to route work across a strong and a cheap model so a heavy workload stays […]

    Reply

Leave a comment

Your email address will not be published. Required fields are marked *

Weekly Β· No spam

Get smarter,
one Sunday at a time.

Join our weekly digest β€” the articles worth your time, plus one thing that made us think differently.