Transcript
Kasia Trapszo: Let me take you back, so 2016, just 10 years ago, Reed Hastings, who was the CEO of Netflix at the time, was given the keynote speech at CES, which is the Consumer Electronics Show. It was a pretty good keynote. He talked about the history of television, the history of streaming, Netflix, of course, all of that. The real point of the keynote was not to talk about the history of television, it was to announce that Netflix would become available globally. We were launching in 130 additional countries, all at the same time. It was a very big moment for us, and it was very exciting. For those of us working on the commerce platform, it was mildly terrifying. This was because this wasn't just the announcement that we were launching, this was the launch. As Reed was saying, Netflix is now available. A bunch of us engineers were sitting in a cafeteria in Building B in Los Gatos, waiting for our cue to turn our individual feature toggles to actually do the launch.
Because launching Netflix globally was not like one big red button, it was a bunch of buttons, a bunch of feature toggles, owned by a bunch of different engineering teams. We all needed to press it all roughly at the same time for this thing to be feasible. We're all sitting there, Reed says the line, and everyone scrambles, and just like that, signups were enabled, streaming open, payment routing was enabled. Then we just watched the dashboards. Somehow this whole thing worked. We started seeing signups from countries that we'd never supported before. We're processing payments, and people were watching Netflix. It was a very cool moment to be a part of. It was also slightly terrifying, because the commerce platform that powered that launch, it wasn't built for a global streaming company. It was designed for a small mail DVDs business based in the U.S. This talk is the story of how that system evolved. It was through a series of very real and sometimes very messy tradeoffs, as the business kept expanding.
Background
I'm Kasia Trapszo. I am a principal engineer at Netflix. I'm mostly working on the commerce platform. I joined Netflix roughly about 11 years ago. When I joined Netflix, it was fewer than 2000 employees globally. Less than half of that was engineering. Seems really silly to say that these days, we've grown quite a bit. I joined the payments team. When I joined, the team was in the middle of migrating our payment services into our cloud infrastructure. Since then, I've started working more broadly on the commerce platform, which is our billing system, entitlement, payment processing. I have had the front row seat to how that system has evolved as Netflix grew into the streaming platform it is today.
Architecture Evolution is not Straightforward
One of the things you learn pretty quickly when working on something like this is that architecture evolution is never straightforward. You don't get to redesign the system every couple of years. It evolves because the business evolves. We get new markets, new payment methods, new business models. With all of those changes, we get new usage patterns. Most of the time, the engineers don't get to just stop the system and redesign it from scratch. You have to evolve it while it's running. Can you imagine sitting down with product, and like, we'll get you those features, but we might have to shut down subscriptions for the next couple of weeks. Not a conversation anybody wants to have. What I want to share today is not a story about a perfect architecture. It's a story about how a real system evolves over time.
How it Started
Before Netflix started to expand globally, the system assumptions were incredibly simple. As I said before, we started as a U.S.-based DVD by mail business, which meant we had an incredibly simple payments model. Most consumers in the U.S. pay with credit cards, and credit cards, at least in the U.S., are incredibly easy to use online. If all your consumers are paying with credit cards, recurring billing is incredibly straightforward. If you want to get a little bit more adventurous as we did, you can introduce PayPal. PayPal does a pretty good job of abstracting away a whole bunch of different payment methods behind a simple real-time facade. Now we got our credit cards, we got PayPal, and both support real-time authorization. Both give us immediate feedback on whether the payments are valid, which means that we also give immediate confirmation. Architecturally, the commerce system at the time was incredibly simple.
You can think of it as four core components. We had payments and billing, and then we had our membership and entitlement systems. Those systems were owned by just two engineering teams. Payments and billing was essentially one central system that handled everything from charging a customer to storing the payment methods. It also performed the monthly subscription renewals with retries, and then produced a simple successful failure signal that was fed straight into the entitlement system. If a charge succeeded, customers got access. If it failed and then the retries were exhausted, the access would stop. Very clean coupling between payment state and entitlement state. The retry model itself was also highly optimized for recurring subscriptions. We get our return codes from processor, as one does, classify them, and then use those to decide how and when to retry a failed payment. Everything was fairly well-tuned for this predictable monthly billing model.
Because, as I mentioned, we're all using credit cards and PayPal, everything was real-time authorized. Because we had the real-time feedback, our fraud attempts could mostly be filtered at authorization, so entitlement decisions could happen instantly. Architecturally, this system was extremely optimized for predictable recurring subscriptions with real-time payment validation. Customer pays, customer gets access. Payment fails, access stops. Retry succeeds, access resumes. Simple, easy to reason about, fairly easy to operate. Of course, as long as the entire world behaves like the U.S. payment system, it worked really well.
Assumption: Architecture Reflects Reality
Before we truly started to expand internationally, something else happened that I'm pretty sure most people don't remember. In 2011, we announced that Netflix would split into two separate businesses. DVD by mail would become this greatly-named company called Qwikster. Streaming would stay as Netflix. If any of you actually remember this? Anybody remember this? You also remember how well this landed. The consumer backlash was significant. People were absolutely not happy about managing two separate accounts, two separate bills, two separate websites for more money than before. About three weeks later, we revised the decision. No more Qwikster. Everything stays under a single Netflix brand. You might wonder, what does this have to do with the commerce platform? Here's what most people don't know about. We rolled back the brand. We didn't roll back the system changes. Architecturally, DVD and streaming had already started separating. They were becoming two distinct subscription businesses, different entitlements, different fulfillments, different economics.
Even though the consumer saw one Netflix bill, underneath it, we were actually operating two subscriptions that got combined at the billing layer. The consumer experience at one company, the architecture said, two business reconcile together. That reconciliation layer, the stitching together of two subscriptions into a single bill, added complexity that we carried for years. We shut down DVD in September of 2023. That pattern, where the business makes a decision, the system implements it. The decision reverses, but the system does not. You will see this a lot in long-lived architectures, including this one.
Assumption: The World Behaves like the U.S.
At this point, we have only barely started to expand internationally. The first country we expanded to was our neighbor to the north, Canada. Very similar economy to the U.S. It's still incredibly credit card-centric. We didn't really learn much from that launch, except the stereotype about Canadians being nice people. It's true, it's deserved. Then we expanded south, Latin America, Brazil, Argentina, Mexico, and day one, great signups. We had free trial at the time, so we were not actually taking money from people yet, but we were authorizing the payment methods to store them for future use. It all looked so good, until day two. Then things started to look a lot less healthy. We started seeing a large number of those payments we authorized coming back as failures. That's when we learned something pretty important. Turns out that in Brazil at the time, this is 2011, customers were not using credit cards.
They were mostly using debit cards. At that time, most debit cards in Brazil were not actually enabled online. While our system thought we were getting all these great signups, what we were really getting was confirmation that, yes, this card looks valid, but we couldn't actually charge the customers. Architecturally, this broke a very core assumption in our system, and that's the payment validation happens at signup. Our system was designed that once a payment method passes authorization, we could safely treat it as valid and then grant entitlement, which is streaming, and the assumption no longer held. The real online payment method in Brazil at the time, again, 2011, was essentially direct debit, bank transfers. This is where a consumer transfers money from their bank account directly to the merchant. This introduced something we never really had to do before, which is delayed payment confirmations. With credit cards, as everyone knows, you authorize in real time, and you know immediately if the payment method is valid.
With direct debit, you submit transactions to the bank in a file, like a batch file, and then you wait. Sometimes it's a day. Some banks, two, three days. Carnival time in Brazil, the whole country shuts down for a week. Now we had customers signing up with bank accounts, and we wouldn't even know for days whether those accounts were actually valid. This forced the first meaningful architectural system to the commerce system. Up until that point, billing and payments operated almost entirely in real time. Authorization happened immediately. Enabled decisions could happen immediately as well. Pay stream. Direct debit changed that.
Now we needed a way to send transactions to banks, wait for a response, and then reconcile those responses back into the system. We introduced our first batch processing, and the system handled generating the bank files, sending them out, receiving the asynchronous responses, feeding that back into the billing and entitlements. Architecturally, we shifted from a purely real-time platform to a hybrid model, where we had to support both real time and these delayed payment signals. Of course, once you introduce a delayed payment signal into a system that was built for real time, you're going to start seeing a lot of assumptions break. Previously, entitlement model, very simple. Payments succeed, grant access. Payment fails, stop access. Now we had to make entitlement decisions before we actually knew that the payment would ever succeed. If we granted streaming immediately, we could risk giving out free access to accounts that might never pay us.
If we delayed access until validation came back, then we were giving a pretty terrible customer experience. We had to introduce some new concepts: validation windows, temporary holds, conditional entitlements. These are all things that we never really needed before. Of course, it also fundamentally changed how we looked at fraud. With credit cards, you can filter out probably 90% of your fraud at authorization. With direct debit, without that real-time rejection, meant that bad actors could create multiple accounts with invalid bank data, and it would take us days to detect this. Fraud detection moved from real-time authentication to behavioral detection over time, which required not just new signals and new operational responses, but also new vendor integrations. LATAM became our first real lesson that the world is not going to behave like the U.S. Brazil broke all of our assumptions. Once you see that happen in one market, you start to realize, this is going to keep happening. Because if every market behaves differently, then our global launch is not just a rollout. It becomes an architectural problem.
Enter Global Launch
Going to 2015 now, we decided to expand slower a little bit and launch in a bunch of countries. We've had Latin America alive and doing well. We've launched across parts of Europe, and we're learning how different people like to pay in different countries. Some of the things we learned, like in Italy, people like to pay with prepaid cards. In Germany, direct debit. Fortunately, by then, we had already gone through the Brazil experience. We had learned. We had learned how to support payment methods that did not behave like credit cards. We've also launched in parts of Asia. We launched in Japan. That was a great experience about how to build UIs, but not so much about payments. By the end of 2015, our portfolio of payment methods included wallets and partner billing integrations with companies like Vodafone, Comcast, Orange. This is where the partner handles the actual billing and they just feed us back the signal if the customer has paid.
We learned this by expanding market by market. Then came the big one, the global launch. This was roughly 130 countries all at the same time. By this point, we knew something very clearly, every market behaves differently. There's no realistic way you can perfectly localize 130 countries simultaneously, not unless you're willing to spend years doing it. We were not willing to do that. We made some very deliberate decisions to take architectural and product shortcuts. Instead of fully localizing currencies and payment methods and pricing in every market, we simplified. We divided the world into a small number of billing zones. If you're in parts of Europe, say Poland, you get billed in euros, not the złoty. If you are in parts of Asia, like Thailand, we'd bill you in U.S. dollars, not the baht. Was this ideal? Of course not. Did we know it was not ideal? Absolutely. This is what made our global launch possible.
One quick lesson from the simplified model was that geography and payment reality don't always line up the way you'd expect them to. There's this small island off the coast of Canada called Saint Pierre and Miquelon. It looks pretty Canadian on the map. It's actually a French territory, which means they use the euro. The way we discovered this, of course, was by watching every single transaction in that market fail because we're treating it like Canada because geographically, so Canadian. We're using Canadian dollars, and the payment systems were not happy about that. When you launch globally with these simplified assumptions, the reality will find those edge cases really fast, which is, of course, why we were watching things really closely. That's how we found this. That island has a population of 6,000 people. I have sincere doubts that even half of them tried to sign up for Netflix in 2016.
We were watching things. Architecturally, the shortcut meant that our billing and subscription systems had to support customers being billed in currency they didn't always match their local market with pricing that wasn't optimized for the local economics. All of this while we're running our global subscription platform. It also meant that we were deliberately launching in markets where we knew that our payment method coverage was less than ideal. We accepted all of that. Our goal at that time was not perfect optimization. Our goal was to launch quickly, observe real demand, learn from real usage, and then prioritize deeper localization in those markets where it seemed like it would have the most value. Instead of trying to get everything right before the launch, we launched with a simplified model and then we learned in production, which if you work on a large system is often a pretty good way to learn.
One of the ways we accelerated that learning was by running experiments to understand how customers actually wanted to pay in each market because we didn't always know. We did something that's called the painted door test, which is a fairly low effort way to measure interest. The way this worked is if customers came on a Netflix signup screen, we'd show them the payment methods supported in the market, which for the most part was credit card and PayPal. Then we showed them the ones that we thought might be interesting to that particular market, except we didn't actually implement them. People would click on that payment method and they'd get a message that said something like, thank you for your interest, not yet available. Then behind the scenes, we logged those clicks. That gave us the real data on which payment method people actually wanted to use. This became our primary signal of what to localize next. This was a very deliberate use of technical debt. This wasn't accidental and this wasn't sloppy architecture. This was deliberate debt that we took on to accelerate our learning.
Assumption: We Can Store Payment Credentials
We're going to 2020 now. Everyone remembers March, 2020, something very serious happened. I'm talking about regulation in India. The Reserve Bank of India issued a directive prohibiting merchants from storing credit card credentials. They used a great time to do it, March, 2020. That meant not storing credit card numbers, not storing expiration dates, and not even payment processors or acquirers could store that information, only networks and card issuers. If you operate a recurring subscription business that relies on credit card data, it becomes a problem. Of course, there's an alternative, which is tokenization. This is where you exchange a card number with a network, they give you back a token, then you can use that token to do the charges monthly. Of course, RBI did not even issue the card on file tokenization framework for India until September of 2021, which was a full year after the original directive.
The ecosystem in India was not ready to support this. We had this deadline, and we had these primitives that didn't exist yet. The deadline kept shifting, as one does. December 2021 became June 2022. Then, finally, they enforced a regulation in October of 2022. The regulation, it didn't just affect storage, it fundamentally changed how recurring charges work. See, most of the world, our billing system was fairly synchronous. At renewal time, we would send a charge to a processor, and then we get the result. Sometimes it will be a few days, but we get a result. Update state, simple. India made that model illegal. Instead, we had to adopt mandates with pre-approved amounts. A mandate is a fairly straightforward billing concept. It's just a user granting you ability to charge them on a recurring basis. The India mandate has amount ceilings. Also, it has to come with 3DS.
3DS is the Verified by Visa or Mastercard SecureCode, you might have seen if you use credit cards online. It's an additional authentication step. It typically uses a one-time passcode sent to your phone or maybe redirect to a banking app. It proves that the actual cardholder is present for the transaction. In India, the step is mandatory for any recurring billing agreement. Since mandate also introduced the amount ceilings, that meant if we had to do a price change, new mandate. Upgrade, new mandate. Every change that required customer interaction. The charge flow itself also became asynchronous. Before every renewal, we are required to send a pre-debit notification at least 24 hours in advance. It gives the consumer a chance to opt out of the recurring billing, which makes sense. Only after the notification passes can we actually submit the charge against the mandate reference. This changed everything about how our system operated. It introduced a whole bunch of new failure modes, like mandate revocation, amount ceiling exceeded, issuer side mandate expiry, ask me how I know. Or the customer simply opting out.
As we were trying to get a solution working for this directive about card data storage, UPI, India's real-time bank transfer system, was rapidly growing. We adopted UPI Autopay. It was an additional payment method. Of course, UPI has the exact same mandate requirements, explicit consent, and also the pre-debit notification. Same architectural constraints, just a different payment rail, but also comes with the additional complexity that UPI does not have operability between acquirers. If you create a mandate with a single acquirer, you are now forever stuck processing against that acquirer. Now we're dealing with mandates, pre-debit workflows, asynchronous confirmation flows, and tokenization infrastructure that was still being built. Remember, we're not doing this in a new product. We already had millions of customers in India recurring on credit cards. This wasn't, build something new. This was, migrate existing users, wipe their payment credentials, change billing and payment primitives, and then, of course, avoid renewal failures, all at the same time.
This was probably the least relaxing time I had working at Netflix. Of course, it forced us to rethink our architecture. We had built our systems around the assumption that stored credentials were available, a fairly reasonable assumption, and that we could do our billing in real time. Now we had to operate in a market where credentials could not be stored. Recurring charges required explicit mandate with amount ceilings. Every renewal required this pre-debit notification. We had to do all of that inside our single global commerce system because, of course, we did not fork India into a separate architecture. Regulation, it doesn't just add complexity. It can invalidate foundational assumptions about how the system works. When those assumptions break, everything built on top of them needs to be reconsidered.
Highly Optimized Systems are Hard to Change
We're going up to 2021 now, I think. Pandemic. During the pandemic, a lot of things were happening around the industry. I'm pretty sure it's because we're all working from home, had way too much time on our hands. Of course, inside Netflix commerce as well. We had this one big initiative that internally we called MHU, multi-household usage. The press referred to it as the password sharing crackdown. As part of that effort, we introduced a new capability, the ability for a subscriber to purchase an additional member slot on their account. If you were, say, a parent with children at school, you could pay a small additional subscription fee and have them use a separate login into the same account. Now this sounds very simple? Yes, of course it's not simple. The challenge this time, not the payments flow. Extra member still uses the same recurring charges in the same billing cycle as the base subscription.
The challenge was the relationship model. For the first time, we needed to start tracking billing relationship between two different Netflix accounts. The payer who owns the account and the beneficiary who consumes the slot. Up until this point, one account meant one subscription meant one set of entitlements, clean and simple. Now we have derived entitlements. The beneficiary's access to Netflix depends on their own account state, but also on the billing state of another account. This meant introducing a whole new construct, the add-on subscription, is what we called it. A separate billing entity that lives alongside the primary subscription, but can be managed independently. Of course, that independence is what makes this complex. In the original model, the payment state and access state were very tightly coupled. I kind of covered this a lot. If payment fails, access stops. This logic is incredibly straightforward. With extra member, now we have this dependency graph.
If the payer's base subscription fails, the entire household loses access, including the beneficiary. If just the add-on subscription fails, only the beneficiary is affected, the payer keeps on streaming. What if the payer downgrades their plan? Our premium plan allows for two extra member slots, our standard plan just one. A planned downgrade might force involuntary churn on the beneficiary subscription. These are all these scenarios that we did not have yet. I had to introduce new entitlement state, lifecycle events, and new logic to handle these multi-account dependencies.
Of course, we are an incredibly data-driven company. This affected how we measure and monitor our system. The surprising thing, for me at least, was that the changes to our data were actually larger than the changes to the commerce system. Our metrics and dashboards were all built around single account subscription behavior. Suddenly we had linked accounts where one person's payment could trigger state changes on a completely different account. It changed how we track churn, how we attribute revenue, and of course, how customer service can handle escalations. Of course, we had to evolve our fraud models as well. We now had to consider abuse patterns across linked accounts, not just within a single subscription. It wasn't just adding a new feature. It would introduce multi-account commerce relationship. It went through a system that was designed and optimized for single account subscriptions. This looks familiar. We didn't build a separate system for this.
We layered it into the existing subscription platform, which meant all of these new behaviors had to coexist with a system that was already operating at a global scale. We were never designing from scratch. We were evolving an existing system to support a fundamentally different model. That layering of new complexity onto existing flows, this is where long-term architectural debt really starts to accumulate. Remember this system? By 2023, it had turned into five teams, each of which owned several domains and operated multiple services. I was going to make a system diagram for how all these systems operate together, but I'm pretty sure that's something they teach you not to do in presentations. You can imagine it looks something like this. There comes a point when a complexity of a system outgrows the engineering teams that support it. It really took us a long time, way too long, to realize that this has already happened with our payments team.
At the time, we had about 14 engineers all reporting to a single manager, and that in itself is already a huge signal. There was a stronger one, and this was the on-call rotation. The payments on-call rotation has become so stressful that the engineers decided to shorten it from the usual seven days rotation into three days and then four days just to make it survivable. That's not a good signal. That's not a great sign at all. When that happens, you have to ask yourself, is the system so unstable? Is it so full of bugs, or is it just too complex for the team that supports it? In our case, it wasn't that the system was full of bugs. It had its issues as any system does. It was that no single engineer could hold the full complexity of it in their heads anymore. Our first response was organizational.
We split into squads. We did a product focused squad, a tooling/operability one, and then a platform squad working on the underlying architecture. In hindsight, as usual, hindsight is 20-20, I realized that this is an antipattern. When you have a platform team supporting a single product team in this one-to-one relationship, two things will happen. One, the platform engineers lose touch with how their platform is actually used. Two, the product engineers lose ownership of the underlying architecture. It does create a separation, but it doesn't reduce complexity. While this helped temporarily, it didn't actually solve the root issue.
We started decomposing the monolith. Some of the split was pretty straightforward. For example, our method of payment storage, MOP storage, became its own service. This allowed us to scale it differently for different workloads. Signup and real-time user flows have very different scaling characteristics comparing to batch processing and renewals. Serving those from the same runtime meant that they were competing for the same resources. Separating them reduced that operational contention. Another extraction that we did was our business rules engine. Payment routing, over the years, has grown very complex. We tend to have multiple processors for a single payment method in every market, just mostly for redundancy. We introduced a UI so operations team could modify these routing rules without involving engineers. Eventually, we layered in machine learning to simplify the routing decisions. This routing layer could now combine static rules with A/B experimentation and machine learning. That worked well, technically.
It didn't actually reduce the cognitive load on the team. Engineers still had to understand how all of this fits together. That's when we realized the problem was not service boundaries, it was domain boundaries. Even if you split out pieces of the monolith, if one team still owns the whole, every engineer still has to understand most of it. We asked a different question. How do we reduce the amount of system complexity so that any single engineer can hold most of it in their head? That's when we decided to split the system and the team by domain. We separated it into two major areas. One focused on consumer-facing workflows like signups, adding and managing payment methods, entitlement-driven payment triggers. The other one focused purely on payment processing, routing logic, vendor integrations, optimizations, and the lifecycle of money movement. Two very different problem spaces. One is about user interactions and state transitions, and the other is about financial infrastructure, vendor ecosystem, and money movements.
Splitting along that boundary allowed each team to own a coherent domain. This wasn't a clean split. In some cases, onboarding a payment method and then processing that payment method requires integration with the same vendor. Take PayPal, for example. Creating a billing agreement happens during onboarding. That's a different API than the charge API you use during subscription renewals, but it is the same vendor integration. Rather than create a shared integration component between the two new systems, we made a deliberate decision to duplicate integration code. It wasn't because it was architecturally elegant, but because it allows each team to move independently, and it allows us to deliver the split much faster. It also gave each team end-to-end ownership of the workflow that they have to support. We introduced this duplication, and we might actually unwind some of that later. There's already talks. At the moment, this cost of shared coupling was higher than the cost of duplication.
Of course, again, all of this was happening while we were onboarding new payment methods, integrating new processors, launching whole new commerce models. Not a clean rewrite. It was, as people like to say, changing the wheels on a bus as it's going down the highway at 60 miles per hour. The key lesson for me here was this. Sometimes complexity does not require better abstractions. It requires clear ownership boundaries. Sometimes duplication is cheaper than coordination. This was another deliberate tradeoff.
Everything, Everywhere, All at Once
Just as we're going through this multi-year re-architecture of the payment system, in fact, quite in the middle of it, in 2024, something else happens. Netflix decides to get into live events. We're planning to stream the NFL Christmas games, which we'd expect will be one of the largest events we ever have to support. Suddenly, the entire company goes into this all-hands-on-deck mode. For commerce, it's because this introduces a completely new signup pattern than anything we had to deal with before. Up until this point, subscription growth has been incredibly predictable. We knew where major content would launch because, of course, we were launching it. We knew how signups could ramp up across time zones. Globally, traffic behaved like this rolling wave, where regions wake up and then regions go to sleep. That predictability makes everything a lot easier. We have here our capacity planning, autoscaling, fraud detection, processor routing.
Everything is based on this predictable model. Our fraud models, as well, were trained for what normal looked like. The systems were very well-tuned for the cyclical demand. Of course, live events break that. With a large live event, you can get a massive spike of legitimate signups concentrated in a very short period of time in a single region. To a fraud system that was trained on our cyclical usage, that looks exactly like fraud. Now we have to ask ourselves, what do we do? What happens if our fraud systems start blocking legitimate customers right as the event is starting? We also have to think differently about our capacity planning. Do we have enough processing capacity for signup spikes? Do our payment processors have enough headroom? What happens if one entire region gets overwhelmed? One of the key changes we did was make splitting processing capacity between batch workflows and real-time signup traffic.
Batch processing like renewals, settlements, it has flexibility. If we bill our customers an hour later than usual, nobody's going to complain. If someone cannot sign up in the middle of a busy live event, people notice. We designed the system so that signup traffic gets priority, scaling policies are significantly more aggressive, and if it's necessary, we can steal capacity from our batch processing systems. We also introduced some explicit fail open strategies. We examined our signup path and looked for failure points. Then a fraud system becomes overloaded or overly aggressive, which is something that we expect that will happen, we can continue collecting signals, but we allow signups to happen and then just review the outcomes later. If a payment processor in a region becomes overwhelmed, we can always route to another processor. In the extreme scenario, their whole payment system goes down, we can always allow people in and then try to collect the money later. Because during a live event, the priority is letting legitimate customers in. Financial reconciliation can always happen later. You cannot replay a live moment if somebody couldn't watch it in real-time.
After all this preparation, we thought what we had would be a pretty good test event. The Jake Paul - Mike Tyson fight. Anybody here watch it? This was really big in the U.S. Originally, this was scheduled for July. We treated it internally as our full-scale rehearsal for the NFL games. Run a big event, see what breaks, fix it before Christmas. Of course, it did not happen in July. It got delayed until November, which, of course, gave us more time to prepare, which is nice, but it also meant we had very little time to fix anything big that we discovered before the NFL games in Christmas time. The stress levels went up a little bit. The date of the fight arrives, and we're all watching dashboards, just tracking things in the morning, see how things are going. By noon, we had already exceeded our predicted signups for the whole event, and the event was at 8 p.m.
At that point, we looked at each other, and like, yes, this is going to be big. It became a full hands on deck. Everyone was watching dashboards. Everyone ready to react. I will say this, the signup and commerce systems worked really well. This is the actual signup graph from that event. The colors are just different parts of the boxing match. Scaling behaved exactly the way we hoped. There were no unexpected bottlenecks. For reference, the little tail end at the end, that's our typical signup traffic. I'm not sure I can say the same thing about the streaming side of things. There were some rebuffers, and this would get a lot more laughs if you all actually watched the boxing match. We had a lot of rebuffers. It was a record-breaking event in more ways than one. From a commerce and signup perspective, it was exactly the stress test we needed.
It validated the architectural decisions which were separating capacity, our aggressive scaling policies, and the clear fail open strategies, which, thankfully, we did not have to use, but we had in our back pocket. By the time NFL games arrived, we had already seen what happens when a truly global audience tries to sign up all at the same time. Once again, the system did not get there through a clean redesign. It got there through a long series of deliberate tradeoffs. For a decade, correct at Netflix, meant charge before we grant access, flag unusual patterns. During the live event, correct now becomes, get legitimate customers in first, reconcile later, fraud review can wait.
Takeaways
These stories all look different on the surface. Different markets, different constraints, different business models, but, to me, they all reveal the same thing. Architecture encodes assumptions, and reality will eventually break them. The LATAM expansion really taught us that we built our architecture around the behavior of our first market, and the world did not behave the same way. Global launch was an illustration in how not all debt is accidental. Some of it is strategic. Because of the India regulatory changes, we learned that stored credentials were not a guarantee, they were a privilege. When it came to implementing MHU, multi-household usage, and adding the extra member ability, we discovered that the highly optimized, streamlined subscription platform our system has become really worked against us. We realized that systems don't just hit technical limits. They hit cognitive ones too. Sometimes architecture has to change in order to scale the teams, not just the system itself. Finally, when usage patterns change, the definition of correct can change as well. In closing, I tried to find a good quote to put on the screen to capture some of this, but I couldn't, so this has to do. "Great systems don't survive because they were perfectly designed. They survive because they keep evolving as reality changes."
Questions and Answers
Participant 1: In Netflix, has it now become a practice to keep your system open for extension, because there are so many learnings that you have over the decade? Is that a practice, is that a principle, or do you now see these things when designing a system?
Kasia Trapszo: The interesting thing about Netflix is that, when I started there, which was 11 years ago, we're more like a teenager when it came to systems. We're building things fast. We're acting like a startup. We just needed to launch, launch, launch. We were not designing things for the long term. Now we're more of an awkward adult years where we have to start maintaining these systems that we've built over the years. Now we're more focused on good architecture, scaling, making sure that things actually make sense. We're doing actually quite a few rewrites of these older systems.
Participant 2: You talked about how you distributed the load so that the batch systems could be used by the real-time systems. Obviously, they're very different processing patterns. How did you manage to do that? Was that just a CPU resource, or was there more to it than that?
Kasia Trapszo: This is really just about cloud capacity. During these big live events, we pre-scale our systems, when a big load of signups comes in or a big load of streaming comes in, we want to make sure that everything is pre-scaled. We're not just taking time to scale everything up and launch applications. We actually tend to run into capacity issues sometimes in some regions. In order to resolve that, if the load is really that large, we have this whole backup plan of, these are the systems that are safe to shut down, and we can take that capacity and give it to the other systems.
Participant 3: How did you make the decision on scaling globally and using vendors versus building in-house? How did the team take the decision of distributing the application to teams, and how did you execute that?
Kasia Trapszo: On the first part, it's about my pay grade, because I'm an engineer. A lot of these decisions are made at the leadership level. Especially at the time of the global launch, I was just a senior engineer on the payments team. I'm more of a leadership position now, but I would not have been involved in those decisions. I was more involved in the architecture, on the team side of things. I was deeply involved in that aspect of it. We just realized that the engineers had just such a huge cognitive load, and it was so incredibly stressful for them to support the system. It just becomes too much. Especially the on-call rotation, that was just such a strong signal. When engineers come to you and they say, "Listen, I've been on call for three days, and I've been paged every single night. I'm getting no sleep.
I am so stressed out. I don't want to be on call again." That's a time to reexamine things. The first thing you should look at is, why is the system stability so poor? If the system stability is actually fairly decent, it's just it's a very complex system, then that's when you have to start thinking about how do you divide this complex beast into something smaller and more manageable?
Participant 4: You mentioned that you split teams into product-focused teams and more platform teams. When it comes to supporting new payment methods, and I'm sure you probably support hundreds right across the globe. How did you have teams, for example, who are in a certain market, how does payment method basic integration ownership work? Let's say a certain payment method has new capabilities that the platform has never thought of, how do you go about having these teams interact? Does a platform team become a bottleneck to integrate that new functionality into the entire API surface? Really curious to get your thoughts on how you've organized that.
Kasia Trapszo: That was the pattern we had at first, which was having a platform team and then a product team. That was one of the things that we ran into is it slows you down. Also, people just don't have the same level of ownership that they had before, but they still have to support the same complexity. We moved away from that model fairly quickly. I think it took us a couple years to realize this is a bad model. Then we did the domain-based division where payment processing is one domain, and onboarding and user-based workflows is another domain. If we're onboarding a new payment method that has components owned by both of those teams, we would have both teams involved in the project. At Netflix, we are a fairly large vertical team when it comes to payments, so we also have UI teams and a mid-tier that serves between the UI and the backend, so we also have those teams involved. We tend to have these vertical project teams that get involved in the thing end-to-end.
Participant 5: How do you balance tech debt versus new features needed to expand the territory? Do you typically implement a tactical solution first and then think about a more strategic solution? I'm surprised you never mentioned about AI in the whole session. This is the old-fashioned software engineer architecture design we're looking for. I want to ask you, do you think AI is going to solve this architecture complexity, tech debt, all those issues?
Kasia Trapszo: The AI thing, that was actually my goal, except I don't know how many of you noticed that one of the slides had a really awkward-looking long arm. I could not get ChatGPT to fix that. AI was used to make this presentation, but it was not mentioned.
Tech debt versus new features, yes, the everlasting problem. There's no good answer to that question. It really depends on what you prioritize. What is most important to you right now? Is it to launch or is it to have clean architecture that you're going to have to maintain for several years? You have to prioritize what's important to you at a time, because it's never the same answer.
Participant 5: Do you think AI is going to solve those problems for us?
Kasia Trapszo: Not for a while. I use AI fairly heavily. We all do at Netflix. It's very well supported. It's not there yet to think about these large, complex problems. Maybe someday in the future, but yes, I think we need those humans.
Participant 6: You said when you started, you were 14 people working in payments. How many people are working now?
Kasia Trapszo: Right now, the payments team is 35 engineers, I think. When I started, it was 7 engineers. It was 14 when we decided to split the teams.
See more presentations with transcripts