Observability
Two separate systems cover this fleet: Langfuse, which used to trace AI calls specifically (see Langfuse — it currently receives no new data), and log collection, which covers everything else — logs, across every app, searchable and kept around for two weeks instead of disappearing when a pod restarts.
| There’s no metrics stack (no CPU/memory/latency graphs) — only logs. This was added to solve a real, immediate problem (debugging something that needed log history that no longer existed), not as part of a bigger observability plan. A metrics stack is a known, deliberately deferred next step. |
The two systems genuinely don’t overlap — general logs flow this path, AI call tracing was a completely separate one. See Langfuse for why that second path currently sits idle.
Why this exists
Before this existed, debugging anything meant reading whatever log output happened to still be sitting in a pod’s own short-lived buffer — often just the last day or so. This gives persistent, searchable log history across the whole fleet instead.
Dashboards
Three dashboards give an at-a-glance view: overall fleet health across every app, a combined view of the small agents plus the podcast pipeline, and a view of the front door — the ingress layer and login gate. All three are built from log data, since there’s no metrics stack yet to draw from.
What a real metrics stack would add
Not built yet. If it is, it would sit alongside the existing log collection rather than replace it, giving actual CPU/memory/latency graphs instead of only log-derived signals.
Read next
See Langfuse for the tracing system’s own deployment and login, and why it currently sits idle.