AI Provider Access & Cost Control
Every LLM and text-to-speech call made anywhere in this fleet — by any agent, by the podcast pipeline, by any other app that needs one — goes straight to the outside provider, with each app holding its own scoped credential, rather than all of them routing through one shared internal proxy.
Who this is for
Whoever’s operating this fleet and wants to understand the tradeoff being made here: this setup gives up centralized spend visibility and shared rate limiting in exchange for real memory savings on a single small machine. Also anyone onboarding a new app that needs AI access — the answer is "get your own provider credential," not "get a gateway key."
Why direct access, not a shared gateway
This fleet used to route every AI/TTS call through one internal gateway, for exactly the reasons a gateway usually earns its keep: one place to see total spend, shared rate limiting, and a credential rotation touching one deployment instead of every app. That gateway was removed on 2026-08-25 — running it cost roughly 1.5-2Gi of steady-state memory on a 2-vCPU node, a real fraction of this machine’s total capacity, for a proxy whose actual traffic was light. On a single-operator, single-machine setup, that trade wasn’t worth it.
What this costs in return: there’s no longer one place to answer "what did we spend, on what, and why" — each app’s own logs are the only record of its own usage. There’s no shared rate limiting or per-app spending ceiling enforced by the platform anymore; the podcast pipeline (the one app in this fleet that regularly spends real money) is trusted to stay within its own provider account’s limits rather than a gateway-enforced budget. Call tracing that used to happen automatically for every call has also gone quiet — see Langfuse.
Concretely, this buys back meaningful memory headroom on a machine that doesn’t have much to spare, at the cost of the operational visibility a shared gateway used to give for free.
A narrower router remains, for one specific job
One small piece of the old gateway’s job — letting the orchestrator agent delegate to the weather/time/currency agents without a caller needing to know they’re three separate processes — is still centralized, through a lightweight internal router. It does no reasoning and no model access of its own; see The Agent Fleet for what it actually does and why it’s a genuinely different thing from an AI gateway.
Architecture
The full current shape — direct provider access, and the narrower delegation router that remains — lives in LLM Gateway.