Agents Router

LLM Gateway covers where AI calls go today, now that they’re direct provider calls rather than routed through a shared proxy. This page covers the one piece that’s still centralized: the small router that lets the orchestrator agent reach the weather, time, and currency agents through one address instead of three.

Until 2026-08-25 this page described LiteLLM — a much bigger shared proxy that every AI call in the fleet routed through, with its own model registry, per-app virtual keys, and spending budgets. It was removed to reclaim the memory it cost on this fleet’s single small machine; see AI Provider Access & Cost Control. What’s below is a genuinely smaller, narrower component, not LiteLLM under a new name.

What it is

A single lightweight process hosting the weather, time, and currency agents as three sub-applications behind one path-prefixed address, instead of three separate deployments. It does no reasoning, no model routing, and holds no registry or budget of its own — it’s the same three agents' own code, just packaged to run as one process and reached at one internal address.

How it fits together

Diagram

Nothing outside the cluster can reach this router, and nothing inside the cluster can reach it except the orchestrator — enforced at the network-policy level, not by any key or token on the call itself. That’s a deliberate simplification from the old gateway’s per-app virtual-key model: there’s exactly one legitimate caller, so a credential on top of network policy would be redundant.

Why bundle three agents into one process

Before this router existed, weather, time, and currency each ran as their own standalone deployment — three processes, three sets of infrastructure, for genuinely tiny workloads. Bundling them behind one process cuts that overhead to one deployment while leaving each agent’s own code and behavior untouched; from the orchestrator’s point of view, a delegated call looks exactly the same as it did when each agent ran standalone.

What it deliberately doesn’t do

Unlike the shared gateway it partially replaced, this router doesn’t do reasoning, doesn’t hold a model registry, doesn’t scope or budget per-app credentials, and doesn’t trace calls. The orchestrator’s own reasoning, and every other AI/TTS call in the fleet, goes straight to the outside provider instead — see LLM Gateway.