Langfuse

Observability covers logs and tracing as two separate systems. This page covers the tracing half specifically — Langfuse, the component that used to receive and group every AI call the fleet made.

Deployment shape

Langfuse still runs as its own release in its own namespace, with its own database holding every trace it’s ever recorded. It’s a genuinely separate system from the fleet’s log collection (Observability) — different data, different storage, different UI — not two views onto the same thing.

Unlike most of the fleet’s internal tools, Langfuse doesn’t sit behind the shared login gate — it has a real login of its own, reusing the same underlying identity provider everything else in the fleet uses. Putting the shared gate in front of an app that already logs people in properly would just be a second, redundant login. See Authentication's "app’s own login" pattern.

Tracing has gone quiet

Langfuse is deployed and still holds every trace recorded before 2026-08-25, but nothing feeds it new traces today. It kept its infrastructure when the fleet’s shared LLM gateway (LiteLLM) was removed, but LiteLLM’s success/failure callbacks were the only thing sending calls to Langfuse — with that gateway gone, no other app in this fleet currently forwards anything to it.

Before the gateway was removed, every AI/TTS call anywhere in the fleet carried a shared session identifier, and the gateway sent each call to Langfuse as it happened — so a whole podcast episode’s script-writing call and every narration call it triggered would show up together as one connected trace. That mechanism no longer exists: every app now calls its outside provider directly (see LLM Gateway), and none of them currently re-implement sending their own call data to Langfuse.

This is a known, accepted gap from the gateway removal, not a separate outage — revisit it if a direct-provider tracing integration (for example, if a provider ships its own OpenTelemetry export) becomes worth wiring up.

Why a shared identifier wasn’t automatic, historically

Back when tracing worked, a single request that fanned out into several AI calls — like the orchestrator agent delegating to three smaller agents at once — needed each of those delegated calls to carry the same identifier as the request that triggered them, not a fresh one each. That was a deliberate convention every agent in the fleet followed, not something Langfuse or the gateway enforced on its own. Some app code still sets this identifier out of habit; it just has nowhere to go today.