Reliability
Pod restarts on master over the last 24h and why —
OOM-killed, crash-looped, killed by a liveness probe, or the whole node taken away (spot preemption /
GCP auto-repair / host maintenance). Deployment-first, scoped to
tilda · neo4j-system · weaviate. From Cloud Monitoring +
the Kubernetes events & audit logs.
Incident timeline
Every node disruption in the window, correlated to the pods it took down and any user-facing 5xx that followed — so "why did this happen?" is answered here, not by drilling logs.
correlating node events → pods → failures…
Health & restarts
gathering restart, health & resource signals…