InfoQ Homepage Resilience Content on InfoQ
-
How to Run on Three Clouds at Once, and When Not to
Ross McFarlane and Kevin Holditch explain how Form3 engineered a triple active multi-cloud architecture across AWS, GCP, and Azure to deliver ultra-resilient, regulatory-compliant payment systems.
-
Understanding Progressive Collapse: How To Avoid A Cascading Failure
Sam Newman explains how to apply civil engineering’s concept of progressive collapse to digital systems, sharing actionable techniques to mitigate cascading failures in complex cloud architecture.
-
Compiling Workflows into Databases: the Architecture That Shouldn't Work (But Does)
Jeremy Edberg & Qian Li explain how to replace complex external orchestrators with DBOS Transact, an open-source library that implements durable workflow execution directly inside your database.
-
Chaos Engineering GPU Clusters
Bryan Oliver explains how to apply chaos engineering to massive GPU clusters. Learn how to handle hardware variability, NUMA nodes, and network faults to secure your AI infrastructure.
-
Enhancing Reliability Using Service-Level Prioritized Load Shedding at Netflix
Anirudh Mendiratta and Benjamin Fedorka explain how Netflix handles massive traffic storms using service-level prioritized load shedding and client-side attempt budgets to protect critical path APIs.
-
How Netflix Shapes our Fleet for Efficiency and Reliability
Joseph Lynch and Argha C. discuss how Netflix balances hardware supply and software demand. They explain techniques like risk-adjusted net value, buffer management, and priority-based load shedding.
-
Systems Thinking for Building Resilient Engineering Organizations
Michelle Alexander discusses how engineering leaders can build effective organizations by focusing on intentionality, prioritization, and operational excellence.
-
Timeouts, Retries and Idempotency In Distributed Systems
Sam Newman explains the three foundational principles of distributed systems: timeouts, retries, and idempotency. He shares practical advice on how to implement each to build more resilient software.
-
Built to Outlast: Cultivating a Culture of Resilience
Kathleen Vignos explains key strategies for software leaders to navigate uncertainty and build lasting careers.
-
Slack's Migration to a Cellular Architecture
Cooper Bethea explains the journey of converting Slack's monolithic production services to cellular, highlighting the challenges and key success factors.
-
Designing Cloud Applications for Elasticity and Resilience
The panelists explore elasticity and resilience, discussing how architects can design systems that withstand workload variations, user traffic fluctuations, and infrastructure failures.
-
Resilience and Chaos Engineering in a Kubernetes World
The panelists discuss the tools, knowledge, and resources that can help achieve faster incident response and recovery times.