Observability

Two separate systems cover this fleet: Langfuse, which used to trace AI calls specifically (see Langfuse — it currently receives no new data), and log collection, which covers everything else — logs, across every app, searchable and kept around for two weeks instead of disappearing when a pod restarts.

There’s no metrics stack (no CPU/memory/latency graphs) — only logs. This was added to solve a real, immediate problem (debugging something that needed log history that no longer existed), not as part of a bigger observability plan. A metrics stack is a known, deliberately deferred next step.
Diagram

The two systems genuinely don’t overlap — general logs flow this path, AI call tracing was a completely separate one. See Langfuse for why that second path currently sits idle.

Why this exists

Before this existed, debugging anything meant reading whatever log output happened to still be sitting in a pod’s own short-lived buffer — often just the last day or so. This gives persistent, searchable log history across the whole fleet instead.

Dashboards

Three dashboards give an at-a-glance view: overall fleet health across every app, a combined view of the small agents plus the podcast pipeline, and a view of the front door — the ingress layer and login gate. All three are built from log data, since there’s no metrics stack yet to draw from.

What a real metrics stack would add

Not built yet. If it is, it would sit alongside the existing log collection rather than replace it, giving actual CPU/memory/latency graphs instead of only log-derived signals.

See Langfuse for the tracing system’s own deployment and login, and why it currently sits idle.