Loading…
Loading…
Practice debugging production systems using the evidence an engineer would actually have during an incident. No memorized commands. No guessing. Just investigation.
Checkout API is returning 503s
22% error rate, 12s latency, and climbing.
Terraform plan wants to destroy resources nobody meant to touch
8% of tracked resources have drifted from state.
Payments pods stuck in CrashLoopBackOff after a config rollout
Payments service is down. Pods keep restarting.
A deploy shipped two feature flags that were never tested together
Checkout is erroring for a subset of users after today's deploy.
The support bot started confidently answering with the wrong docs
Retrieval is returning irrelevant chunks after last night's reindex.