BT

Facilitating the Spread of Knowledge and Innovation in Professional Software Development

Write for InfoQ

Topics

Choose your language

InfoQ Homepage News Cursor Uses S3 WAL to Scale Git Storage to More than 300 Pushes per Second

Cursor Uses S3 WAL to Scale Git Storage to More than 300 Pushes per Second

Listen to this article -  0:00

Cursor has developed Continuity, a Git storage architecture that uses an S3-backed write-ahead log (WAL) as the source of truth rather than relying on replica coordination as the consistency mechanism. Cursor reports linear read scaling with up to 100 replicas in synthetic tests and more than 300 pushes per second using S3 Express One Zone. The architecture targets both large repositories with heavy CI workloads and the large number of smaller repositories created by coding agents.

Git hosting at scale depends on local Git repositories and packfiles. GitHub's Spokes architecture maintains multiple NVMe replicas and uses three-phase commit for reference updates. The design keeps replicas synchronized and allows reads to be served from any replica, but coordination overhead increases as replica counts grow.

Continuity changes that model by making an S3-backed WAL the durable source of truth. Cursor stores pushed data in S3 and records the corresponding reference update in the WAL. A push is acknowledged only after the required data has been persisted, providing durability before acknowledgment. Cursor also batches operations to reduce the impact of S3 PUT latency on throughput.

Continuity architecture (Source: Cursor Blog Post)

Cursor engineer Vicent Martí described local NVMe repositories as warm caches rather than authoritative copies. Rendezvous hashing selects preferred nodes, while atomic compare and swap operations on S3 allow any server to accept a push. A repository can be materialized from the WAL when a local copy is unavailable. UDP gossip propagates WAL updates, while conditional S3 reads verify replica state. Cursor reports these reads take less than 10 milliseconds on average and says lost gossip does not affect correctness because S3 remains the source of truth.

The storage model has prompted comparisons with database systems. Maksim Al Dandan, a senior software engineer, described the approach as treating Git storage like a database, pointing to push consistency, force-push transactions, and point-in-time recovery as questions relevant to the model.

Continuity Push Flow (Source: Cursor Blog Post)

Casey Lee, CTO at Liatrio and a former AWS engineer, highlighted the architectural shift, describing the WAL in S3 as the source of truth while local NVMe serves as a cache. Lee also noted that Cursor's published performance figures had not been independently verified.

The Continuity model changes how replication and compaction scale. Cursor says monorepos can support hundreds of replicas for CI, while idle repositories can be materialized on demand. Only the primary performs Git compaction, with replicas downloading resulting packs from S3, trading bandwidth for CPU.

In synthetic tests using Cursor's everysphere monorepo, Cursor reports up to 120 pushes per second with S3 Standard and more than 300 with S3 Express One Zone. At the higher rate, Git compaction became the bottleneck. Cursor says tested pushes were linearizable and persisted to external storage before acknowledgment, while clones remained fully consistent.

The architectural shift moves the consistency boundary from the replica layer to durable object storage. Instead of synchronously coordinating more replicas, Continuity lets local repositories converge independently on state stored in S3, trading replica coordination for storage validation, asynchronous propagation, and extra bandwidth.

About the Author

Rate this Article

Adoption
Style

BT