Community guidelines
Be specific and constructive. No vendor spam — promoting your own product belongs in a listing. Anyone can read; posting needs a free account.
A recent auth outage took down our application even though compute was healthy in two regions. We now have pressure to go active active everywhere. The app can run in two regions, but identity, queues and the database still have single region assumptions. Where do you draw the line before resilience becomes a second product?
start with the business outage target. Active active is expensive in engineering attention, not only cloud spend. We use active passive for stateful systems and active active at the edge. That meets our recovery time without turning every write into a distributed systems project
Also list the shared dependencies. We once had two regions and one global AWS Secrets Manager endpoint. Looked redundant on the diagram, failed as one system. Identity and DNS deserve the same failure testing as compute.
Our target is two hours, but people are reacting to a six hour incident and asking for zero. I think we need to put a price on the extra nines
exactly. Run a game day with the current design and see if you can hit two hours. Fix the slow manual steps first. If the business then wants near zero recovery, show the database and consistency work required. Resilience is a product decision with a permanent operating cost