Community guidelines
Be specific and constructive. No vendor spam — promoting your own product belongs in a listing. Anyone can read; posting needs a free account.
during an outage somebody edited a deployment directly to restore service. Argo CD later put the broken value back because Git still had it. i understand why, but telling on call staff never touch the cluster doesn't feel realistic. How do you handle emergency changes without fighting GitOps?
We pause sync for the application, make the emergency change, then immediately open a pull request with the same change. One person owns the clock until Git and the cluster match again. The problem was not kubectl. The problem was an undocumented second source of truth.
add a break glass command that records who paused sync and posts to the incident channel. Manual edits expire after a short window. If the fix can't be represented in Git, that's a signal the deployment model is missing something
edit: we had no easy way to pause only one app, so the responder disabled the controller. That made everyone nervous. i like the expiring break glass idea
Practice it before the next outage. The runbook should include pause, patch, verify, commit, resume and confirm no diff. GitOps is not a ban on emergencies. It is a promise that the emergency state gets reconciled deliberately instead of becoming folklore.