InfoQ Homepage Distributed Systems Content on InfoQ
-
The Agent Harness: Control Planes, Invariants, and Approval Boundaries for Production AI Agents
OpenAI’s Vinoth Govindarajan explains how to build reliable AI agent harnesses in production, focusing on state ownership, mutation ordering, scoped authority, and user-visible proof of action.
-
From DVDs to Global Streaming: How Netflix’s Commerce Architecture Actually Evolved
Kasia Trapszo explains how Netflix evolved its commerce platform, making pragmatic architectural trade-offs through global expansion, regulatory shifts, domain splits, and live-event scaling.
-
Understanding Progressive Collapse: How To Avoid A Cascading Failure
Sam Newman explains how to apply civil engineering’s concept of progressive collapse to digital systems, sharing actionable techniques to mitigate cascading failures in complex cloud architecture.
-
Parting the Clouds: the Rise of Disaggregated Systems
Murat Demirbas explains how cloud economics drive compute-storage disaggregation in modern database architectures. He shares insights on network bottlenecks, Paxos origins, and future tech.
-
Latency: the Race to Zero...Are We There Yet?
Amir Langer explains the critical role of predictable low latency in fintech, sharing lessons from the past and modern techniques like kernel bypass and Aeron to push systems toward zero latency.
-
The Ideal Micro-Frontends Platform
Luca Mezzalira discusses the evolution of micro-frontends, explaining how to scale engineering teams through technical decentralization, business subdomain boundaries, and organizational autonomy.
-
Fast Eventual Consistency: inside Corrosion, the Distributed System Powering Fly.io
Somtochi Onyekwere discusses Corrosion, Fly.io's eventual consistent service discovery tool. She explains how the networking team leverages SQLite and CRDTs to replicate data across 800+ nodes.
-
Scaling Cloud and Distributed Applications: Lessons and Strategies from chase.com, #1 Banking Portal in the US
Durai Arasan shares how Chase.com achieved a 71% latency reduction. He explains strategies for efficient scaling, multi-region resilience, and automated "repaving" to secure large-scale systems.
-
WASM in the Enterprise: Secure, Portable, and Ready for Business
Andrea Peruffo discusses WebAssembly's role on the server, showcasing how Chicory enables secure, portable, and performant polyglot applications, sharing key enterprise use cases.
-
Test Smarter, Not Harder: Achieving Confidence in Complex Distributed Systems
Elias Nogueira explains how to build robust test suites for microservices by solving three common problems: testing with multiple databases, mocking dependencies, and managing asynchronous events.
-
Timeouts, Retries and Idempotency In Distributed Systems
Sam Newman explains the three foundational principles of distributed systems: timeouts, retries, and idempotency. He shares practical advice on how to implement each to build more resilient software.
-
Beyond Durability: Database Resilience and Entropy Reduction with Write-Ahead Logging at Netflix
Prudhviraj Karumanchi and Vidhya Arvind share how Netflix built a Write-Ahead Log to guarantee data durability and reliability, tackling issues like data loss, corruption, and replication at scale.