Transcript
Ian Nowland: My name is Ian Nowland. I'm co-founder of Junction Labs. A little bit about my background as it applies to the talk. I was at Amazon 10 years, AWS 8 years, 2008 - 2016, mostly on EC2. EC2 was basically a stack of platforms that I got to learn a lot from. I left there in 2016, I moved to New York. I had a 50-person compute platform team at a FinTech, that's where I met Camille, my co-author. Then I jumped to Datadog, pre-IPO, and just saw the rocket ship there, where again, I just saw these platform teams just challenged by the assumptions that what they should be doing is what all the product teams were doing. I wrote the book, of course. I'm now co-founder of a startup called Junction Labs.
Terminology
I want to start just with a couple of definitions. There's this foundational platform, which is a little bit of a weasel, and I just want to say why I have that. To me, the big difference between a foundational platform and a non-foundational platform is just if you break, does the business go down? That's different to maybe a developer experience platform you deploy, where some engineers are very frustrated with you, but very different to, we're actually no longer making trades, if we're a financial company. If we're an observability company, we're no longer ingesting data. I think it just changes the way that you build platforms. That's the purpose of the talk. Another definition, I actually did not understand this before I started preparing this talk, and I actually don't think it matters too much. If you go back, I think about 5 to 10 years ago, there's one CD, it's a deployment versus delivery.
There's this idea that deployment is fully automated all the way to production. I haven't seen much of that in my career. I don't want to say it doesn't exist at all. Then there's continuous delivery, which is the much more, ok, you get through CI and then a human presses a button, and everything is perfectly automated on either side. I do think for application teams, product teams, that's possible. I think it's platform teams' jobs to make that possible. It's much harder for platform teams itself. A lot of this talk is why that is and what you should do about it.
Outline
My outline is first answering why a typical CD doesn't work for foundational platforms. Then I'll go back, like a lot of what I think I learned works and doesn't work, was just seeing EC2, and it was both of them. We tried a lot of things that should have worked, and they just didn't work for a foundational platform. Then I tried to apply those rules at Datadog and got some more lessons. Then finally just wrapping it up.
1. Why Typical CD Doesn't Work for Foundational Platforms
Why a typical CD doesn't work for foundational platforms. I want to start with, again, you're on a platform team. This is feedback that's been given to me, probably feedback I've given to my teams, and even sometimes feedback that our teams give themselves. The first one is, you join a company and you look at the platforms, and then they're just always built on such old technology. It's like, what are you doing wrong that you're maybe using a Java Spring framework for your UI, as opposed to using React, to use a 2025 example? There's just always this thing with platforms. We're always just using older tech. There's this other thing, especially if you're a fast-moving product organization, there's this constant frustration like, we're doing 10 times as much work as you in the product teams, and you're just doing a tiny bit of work. Why are release timelines bottlenecking on you?
What are you doing wrong that you guys just can't move as fast as we can move? Classic feedback. Another one, which is definitely true in my career, like, why do you cause the company's worst outages? Datadog's two worst outages in my timeline for both the compute platform team, both took down everything. You just look at OpenAI's biggest outage late last year, they were doing an observability rollout, took it down. It's just always true, financial platforms, for reasons I'll talk about, we do seem to cause the company's worst outages. Why can't you test properly? This is like where you start getting a little bit of judgment in the feedback. Again, if you're a product team and CI seems to work pretty well for you, it is like, what's going wrong? Just do some performance testing, some load testing. What is it about platforms that makes it very hard to test properly?
I'll come up with an answer to that. This is the same thing. Why don't you just do normal CI/CD? Again, a lot of what's been documented is great for product teams, just not so great for platform teams. Then there's this other judgment. Like, I do say this, I remember when I joked, so I was at Amazon for 18 months, and I moved to AWS in 2008. I hear my colleagues in Amazon say, those leaders don't understand operational excellence. You go like, why? What is it they're not understanding? They'd be very hand-wavy. They don't build software properly. It's like, no, there has to be specifics. Yes, there is this judgment. I don't think we're crappy engineers. We do maybe have to explain why we're different a bit better.
Why the disconnect? I'm not critiquing this book. I think it's a great book. Of course, DORA is incredibly important. When you write a book, you're writing for the 80%. You're not writing for the 20%. I think "Accelerate" and a lot of the stuff that came out about continuous delivery post Accelerate is really written for the 80%. It's just not that relevant. In particular, there's a bunch of stuff in the book that I remember when I read it, really frustrated me. I just want to have one table. It's just wrong, I think, for most platform teams, definitely foundational platform teams. This is chapter 2. You'll see versions of this in my blog post today, even. The thing that really stood out to me, coming out of EC2, but then also being at the FinTech and just seeing my compute platform was like, I did not see them as not high performers, because they weren't deploying multiple times per day.
If they tried to deploy multiple times per day, which then I would have seen them as low performers. There's just this mismatch between what you really want. If you're a fast-moving frontend team, actually for a company like Datadog, yes, like multiple times per day is great. It's just not great for platforms. There's this other thing here, down the bottom, like, as a platform team, if 15% of our deployments actually caused outages, we'd be fired. That's very different when you're a frontend, where maybe some of your functionality is a little bit unstable and users don't notice. Generally, in foundational platforms, when we have a bug, stuff goes down. Fifty percent failure rate is too high. You see how these two link to each other, is that if 50% outage rate is way too high, then you probably have to think about how often you deploy.
This is just summarizing what I just said, which is the common CI/CD models and practices that they assume at teams. That's not a problem. Most people are app developers. We build platforms for a reason. The wrong assumptions often lead to the wrong practices being recommended. I wanted to justify that first. This is a picture from the book. This is a world with like minimal platforms. You'll see the platforms are in the bottom right corner, the infrastructure is in the bottom left corner. What this picture is meant to represent, I think we call this the over-general swamp. It's just a mess. What you need is more platforms. This is an example of just building a bunch of platforms to make things simple. The architecture has been made more simple, but that complexity doesn't totally disappear. I think we all know that if you're building a good API platform or a good storage platform, for you to be able to provide a good platform to a bunch of different applications, you've internalized the complexity, you've abstracted it.
We repeat that throughout the book. That's the complexity that causes us to be different in terms of CD. Platforms abstract complexity by bringing that complexity into one place, that place is the platform. What does this lead to? This is a little bit hand-wavy. I'll come back to a bunch of these points later in the talk. How do applications differ from foundational platforms in just a bunch of ways? Most foundational platforms are stateful in a way, and we're stateful to make the applications be able to be stateless. This isn't just storage platforms, you'll see API platforms. You'll see a bunch of metadata that we have to maintain to make the platform work that we do on behalf of the applications. That statefulness makes us a lot harder to change it in production. Second one is just like the whole point of platforms is to build on top of the infrastructure.
We all know whether it be modern, whether it be cloud abstractions, whether it be all, whether it be like computers in data centers, we directly use them. We have to think about how they're going to fail in an abstracted process. Reversible changes versus irreversible changes really just comes back, I think, to those two. The more that you are stateful and you're talking to hardware, the harder that it stings, the harder it is to reason about whether you can do a fast rollback and not hurt anything. Fourth one is just the nature of being a platform is you're not just your one business segment when you have an outage, you're affecting large parts of the business. Then you can argue like, taking down one business line isn't that much different. Of course, like for most companies, when everything goes down at once, it's really bad. It just bleeds into it.
The final thing, and this all just comes back. In some ways, this is just saying complexity again. The less complex you are and how much of the system you have to reason about, the easier it is to test locally. Our whole point is to abstract the complexity that comes from the other systems. That's very hard to test outside production. You can't really fake the experience of production in a test setup.
This is just revisiting the earlier feedback I showed. I just put a bunch of categories down the left of what they're saying: bad tech, bad delivery, bad outages, bad testing, bad CI/CD. I said that bad assumptions leads to bad practices. I wanted to say for each, these are all practices I've heard recommended, and I'll come back to later, that I think are wrong. The first one, and you hear this more from the engineers than you do from management. It's often like this reaction to, good engineers will come into a platform team, look at the state of things and go, ok, this is way too much accidental complexity here. We can fix that by rewriting everything. I think everyone knows why rewriting everything is not always the wrong thing to do, but definitely a very difficult thing to do right. Bad delivery. A lot of the CI/CD people love feature flags.
My wife is a product manager. I say this with all due respect. I always say feature flags were invented by engineers and co-opted by product managers. Product managers love them to launch shit quick. Of course, feature flags are great if it's a small segment of the codebase, it's very easy to reason about like on versus off. It's not too long-lived. The longer it takes to deploy, the longer the feature flag's going to live. Platforms in general, you find that the feature flag has the sprawl throughout the codebase. It just becomes this completely second and third order thing that you have to reason about in terms of testing, observability, and whatever. Again, feature flags are great for application teams. I'd say they're good with discretion for the platform team. More process. Bad outages, you get your leader, they get very frustrated. All they want to do, like every outage is another thing that everyone in the platform all has to do in terms of process.
Process, I'll come back to it. Some amount of process is good. More and more process you end up with process for processes sake. You have people doing it, not because they understand what the process is supposed to help with, but just to get that bit of paperwork out of the way to do their real work. Bad testing. You see a lot of people really want a staging environment and more integration tests. This is the thing, again, works great for products. With tests, actually, you can reason about what they do. I'll talk a little bit later about why staging I think is just so doomed for platform teams. Bad CI/CD. You get this obsession over DORA. This is when the book came out, and Camille was my manager, actually. It's like, we need to measure DORA everywhere. I'm like, we should measure DORA for our customers because we should be improving that across the organization.
Don't judge my team by the fact that our deployments only happen once a month. It's just this naive use of DORA that I think is a problem. I actually think DORA, if you use all four of them, is great. I think what you see with DORA is a little bit too much focus on the tool about delivery and a little bit too little focus on the tool about reliability. Otherwise, I'll say DORA's a good thing. Bad people. I think we've all seen this. You bring in a new VP from a FAANG, and they go, I just need better engineers. Sometimes you do need better engineers, but as you lose existing engineers, what you're often missing is the amount of internal knowledge of a system that you're losing. Sometimes, of course, I think we all know, and I think I have this later, is you get the ex-FAANG person who just wants to replicate exactly the system they had at this company, but had a hundred times engineers than you have. Trying to replace people is not the right thing.
Answers for feedback, as I talked about, given that platforms are different. This is not solutions. I think this is just things that we all know. For bad tech, I think what we all know is stability. Even for our customers, even though they'll complain with us, the stability is more important than novelty. This is a reason why old open source, especially, just continues to persist, is that it's going to take us a long time to replace it in a stable form. Bad delivery, as I said, tiny asks. If I choose stateful systems, it's risky to change it. Sometimes what looks to our customers like a small feature is actually very risky for us to deploy at scale. Bad outages, they're a fact of life. Outages hit harder because everyone depends on us. Bad testing, you can't fake production for infra-level systems. CI/CD, I said this already, like everything that's published just assumes that you can automate your way into where anything is a small risk to put in production, small risk to roll back. It's just not true for complex platforms. This is just bad people. This isn't about engineering quality. It's just about, we're doing a different type of software engineering.
This is leading into the next part of the talk. We're different, but what I didn't answer there is what practices should we use? Now the answer to this, and actually Charity Majors from Honeycomb, I think she was at QCon 6 years ago, where she talked about testing in production. I believe in testing in production, but if you've talked to many executives, particularly non-technical executives, you know that testing in production is not a thing that they want to hear you ever say in front of a customer. This is someone else at Honeycomb, 2019. Your executives hate testing in production. I've been an executive now for about five years of my career. I used to be an engineer for a long time. It's like, no, we are testing in production. Like, who's fooling who here? The problem with saying testing in production is that rather than acknowledging, there's just unknown unknowns.
That's what we mean when we say we're testing in production. There's unknown unknowns. We're just not going to know until we actually put shit in production. It sounds like we're being lazy. I think that's a big problem with saying testing in production. Talking to engineers, you can say it all the time. If I was trying to sell progressive deployments to my manager or to my leader, I wouldn't use testing in production. I just got to say progressive deployments. Really, you'll see in the rest of this talk is that it really is just good techniques around what it means that you are testing in production.
2. 2010 - 2016 EC2: Safe Progressive Deployments as a Survival Strategy
I was a customer of EC2 between 2008 and 2010. I joined EC2 in 2010. I just have some numbers up here. 2010 was a long time ago. When I talk about my lessons at EC2, people think about what EC2 is today. EC2 in 2010, like we were a bunch of cowboys. We were learning what it meant to build a cloud as people were using the cloud. Not only were we cowboys, I come to the main thing at the moment. EC2 was super flaky. I knew this because I'd been a customer and got really frustrated. Why do 1 in 1000 instance launches just not come up? No one in EC2 could answer that. It turns out it was about DHCP. My team discovered it about two years later. The cloud was flaky. That was a big thing in 2010. It's why Netflix so emphasized chaos engineering at the time.
Of course, Netflix was a startup at the time as well. Their bar wasn't super high back then either. The other two things from this slide to emphasize, and this comes with the growth. EC2 had a really strong, you build it, you own it culture. There was no test teams, no ops teams, no SRE teams. It was the platform software engineers who were doing everything to make things work. You see our deployment cadence was about every month. We kept that as we scaled up. We had to launch a lot of features to make people want to move on to the cloud. We had to find a way to keep a good cadence. What you don't see is we've managed to reduce that cadence. In fact, within it actually our deployment phasing got a bit slower than we had to do. The other big thing, and I don't want to overemphasize it here, because I'll come back to what it is, this deployment's written up and reviewed.
It's a very controversial thing, but it is a thing that I introduced at Datadog. I think it's very important for platform teams, just given the facts, which we do. On the right, you see six years later, what had happened. Fortune 100 companies were now on us. The bottom thing, because AWS goes down and the world freezes. Every outage should not happen. It's remarkable to me that given the complexity that the big AWS Dynamo, EC2, S3 have, they really only have a bad outage like every two or three years. I think they're doing a lot right to be able to do that.
Now I'm going to go through area by area of what didn't work and then what did work. I talked about deployment planning, this idea that before you do any release, you plan out what's going to change and why. I can tell you, every engineer hates deployment planning. I hate it. Amazon didn't do it. I moved to AWS and we did it. I'm like, why am I doing this bullshit paperwork as opposed to doing real engineering? What didn't work though, like in this story, I think every engineering team, all we need is more automation, better automation. We'll finally make this go away. We'll hit the CD dream of just being able to push a button. It's a good dream and you should be working towards it. The dream did not actually happen. I wanted to talk about, for people who haven't done it, what's the core of a good deployment plan?
The first thing is just what has changed. That meant, especially my team in EC2, having someone review in the last month, every single change that had gone in. Talking to the engineer who made it, if necessary, just making sure that the person who was on point for making the deployment happen, just had a good understanding about what was going to change and why. Second thing is, what is the timing and the phasing? That gets really important. I'll come back to it in a lot of the talk. What AZ, you're going to go where, what percentage of the AZ? What exact commands will be run and have you actually tested them in a staging-like environment? That's an important thing, like making sure no one is inventing the commands that they run on the fly. Of course, if you know S3 had a bad outage, I think in 2017, it wasn't done under change management.
It was a big lesson that even just like one little flag on the command line can totally screw a big system. Any communication steps that will be done, this is a little bit in the weeds, but it's generally important to remind teams who are doing execution, if you're going to have to do communication. What observability is in place? I'll come back to this one. Then a big thing. I think you talk to anyone who's at AWS, we will emphasize rollback and testing that out as being like the crucial thing that has made what could have been a really bad outage into just a minor blip. A big part of a good deployment plan is that you've specified it and you've actually tested it.
What didn't work was removing deployment planning through automation. It really just comes down to the complexity. That you have a massive surface, you have complexity below infrastructure or other platforms, you have complexity in terms of all the different ways that applications use you. Every deployment, they change in what details that they affect. CI is just never going to be perfect. What that means is having someone who is thinking about, is this deployment special, sometimes or not? I think why this deployment might be special and what I should be doing to mitigate any risks to do with that. Engineers document, no, this is just exactly what I wrote up. Peer review is really important. Peer review is another one of those things, it's like code review. Engineers get really frustrated because now I have to go chase someone who doesn't want to help me. Peer review is really important though, because you're constantly bringing in new people to the team and they don't understand the whole team.
Peer review, it's a little bit the finer detail, but mostly it's about to help train the new people. It's actually what's acceptable in terms of process. Then, as I said, actually practicing the rollback path. Charlie Bell was the SVP of AWS for a really long time. If you had a post-mortem and he said like, why didn't you roll back earlier? Your answer was, we didn't practice it. You're going to get yelled at because it was just so essential to do that every single time, just to mitigate damage, to minimize the blast radius. The takeaway, and again, this isn't a pleasant thing. The engineering may sort of recoils. People were very angry about this. I do think on core platforms, definitely on applications on financial platforms, if you touch prod manually, it needs review. That also just means like one-offs we found over time. Like just ask someone on Slack to review the command line you're about to type in. To me, it's just a fact of life. It goes against the CD, ok, this should just be a button. I don't think that's true. I want a human who understands what's going to happen and to have thought through, and roll out.
Staging, I mentioned earlier, I was going to talk about staging. What didn't work, and we tried about three times to create a staging environment. Then later at Datadog, staging came up again. They tried it a few times. This isn't just platform. Every time I've seen staging, I had this dynamic, which is, you get one team saying I need staging to be really stable. Then you go, it makes sense. Why? It's like, but why do you really need it to be so stable? It's like, I've got some stuff that I need to test because I don't actually know if it operates yet. If it's just two people at the whiteboard, this actually is fine. Two people can work out timing. Once you get to about 10 teams all trying to coordinate about staging, you basically just end up with a staging that is perpetually unstable and everyone just perpetually complains about.
The problem with platforms especially is they're the ones who everyone else wants to be stable, which makes them absolutely useless for them to actually test any risky stuff. What didn't work was an EC2-wide staging environment shared env. This is the one staging type thing. Everyone wants it to be stable while they test their own stable stuff. Just N squared doesn't work at all. EC2, we also sunk about five debuts into a different thing that you might as well try where every team gets like a snapshot of the last stable version of everyone else's thing. It's not a bad idea, again, at small scale. At big scale though, it's expensive, for one. Number two, if someone has issues in their environment that's actually caused by interaction between their new unstable stuff and the older system, no one wants to debug. This becomes this sort of like, am I on-call for some other team not configuring their staging environment directly?
It's just one of these things it just doesn't scale as you have more and more teams. It's a tragedy of the commons. Generally, if you could get four or five teams, the social dynamic worked. I was on EC2 data plane. My team by the end was about 50 engineers. We had a staging environment that actually worked really well. It was limited to those teams. What we understood was everything else had to be mocked out. We weren't actually testing against their production characteristics. We're testing against their mocks. That was the compromise that we came to, but it actually worked. Like, you can't really test against all your dependencies before prod, because you just end up with too many people who need to test against each other.
This is just what we have. This is the pipeline stuff before production. We have the image taken from the CI pipeline. You test on team staging. You include rollbacks by the steps. You write up the change and you review it. This is all the pre-deployment work that happens. This is maybe different types of tragedy of the commons, so what didn't work. I saw, I think it was only three. I saw three attempts to rebuild a EC2-wide integration testing platform. Each of the attempts was after like some type of outage where it's like, we didn't run all the tests because they're too flaky. The problem is the test framework is flaky in certain ways. We need to reinvent the test framework to do it. Mostly having seen a lot of integration tests, integration tests are inherently flaky. I think we all know this. The more units you stack and try and test end to end, they're just inherently flaky.
Maybe for frontend, if you have a flaky test, you say, don't run that one anymore. The whole point in foundational software, which is risky, if we want every test to run, but of course every test is flaky. What happens? The other thing you found is anytime we tried to have a centralized team, this was the three times we tried to have a centralized team. I spent all my time writing a new test framework, no time spent fixing tests. This fixing test was this distributed cost that everyone had to pay. I always had to say, like when my test breaks on your flake, do you debug that or do I? We never managed to solve. Whereas within a team, you can solve that. What did work, and this is probably the second thing I think that you find, I find people who are in AWS really see differently to people who haven't just seen that scale or whatever, is you end up with what they call synthetic testing, which is actually just running in production with the detail.
If you're synthetic testing production flake, you page someone. That changes the characteristics of what you test and how you test. This did work. It worked really well for two reasons. Number one, it had to be stable. The flaky problem was well learned. The second thing is it gave you a nice place to deploy the initial version of all your software that's actually in production, running against your dependencies. I'd say our investments in synthetic tests were far more successful than our investments in two-version testing. Here's this initial canary. I just did a table just comparing the two. First thing is I admit the coverage is less. Part of flakiness is often when you get into detailed semantics and synthetic monitoring. You don't want to wake people up at 2:00 in the morning just for a flake. You don't cover as much of the known code. The big thing you get, what I call realism here though, is you are testing against all your dependencies as they are in production.
The realism of what you're testing as a system is actually far realer. That actually makes them far better. Reliability, one tends to be flaky and a bit noisy. The other is stable signal, but again, you lose fidelity. Scalability, and this is a little bit of a lag, so it's changing. I think it's definitely different every day. At least at a scale of 50%, it's very easy to say to each team, you own your own synthetic test. Now you do use more computer resources when you do that, but the flip side of it is you get really clear ownership. You don't get this mixed ownership of who debugs the flake. Then finally, debuggability, again, you don't get this cross-team pointing fingers at each other. You generally get to simple scenarios that are faster to triangulate.
What didn't work is a large end-to-end integration test platform. What did work is high coverage synthetic tests. Basically, the takeaway is just prioritizing synthetic monitoring. That was at both the FinTech, Two Sigma, and Datadog, but something that I took my role as heavily emphasizing to teams. Deployment order, and this always comes up, probably most people have a friend of yours where everyone thinks like us-east-1 is where all the new software goes first, and for some reason why it fails. That's actually not true. us-east-1 is a problem. I think, at least when I was there, and I mentioned this, it was 10 times bigger than the second biggest region. That means any type of horizontal scaling system has 10x the load that it's thinking about. It's not just that. It also means you have 10x the diversity of customer use that you see anywhere else. There's this tendency to think we should deploy to us-east-1 last then.
That didn't work, because the largest diversity of use means you don't want to get 98% pre-deployed, finally go to us-east-1, and then finally learn, ok, we have this issue. I think I'll talk a little bit about one of those issues soon. What did work was to deploy to the lowest sensitive region first. In EC2, in my timeline, that was the Brazil region. I don't know what they call that today. It still has to be one. Someone has to go first, try and find out what that is and avoid it, with that region and come back. Then the second thing, maybe the key point of this slide, is going to us-east-1 second. Not just leaving all the risk until late in your deployment pipeline, trying to bring some of that risk forward and trying to mitigate that risk. Yes, get exposure to diversity of use early.
Deployment speed. What didn't work? I've said this a few times already. It was driving up the velocity for smaller releases more often. It's not like we didn't want to. It's not like we weren't investing a lot in automation. We just found that with the risk that we were taking, the phasing of a deployment was taken up more often. Then there's the second point. This is Marc Brooker, who's a distinguished engineer now. Some problems just take time to happen and take time to be noticed. I'll talk about this 2015 outage that we had. This is a very simplified version of EC2 VPC. You have a global control plane that the APIs talk to. You have this mapping service that sends out all the stuff that's happening in the network. Then down the bottom, where my team operated was the EC2 host. What you see is the mapping service team only had about 10 mapping service hosts per AZ.
We're in the hundreds of thousands at that point. There's this sense that, ok, phasing and going slower only is for us at the bottom because we have so many. Actually, as I'll talk about, it was just as important for the mapping service. This was in 2015. This was an outage. Mapping service, they do a deployment, it had some new features. I think it was maybe IPv6, deploys to first AZ in us-east. Next day, almost 24 hours later, because the full deployment took a few hours, at peak load, the EC2 host fleet, actually, one of my teams started DDoSing the mapping service. We hadn't changed anything. As we're debugging it, it's like, this does look like us creating new connections and DoSing you. We think it's you who changed something because you just did a deployment. What does that mean in EC2? It was just all launches were blocked in that AZ, I think for a few hours to recover.
Root cause. This is a classic when you start doing the 5 whys and getting down to what it was. My team, we tried in the agent. We did have a bug in exponential backoff. We actually weren't doing truncated exponential backoff properly. We were retrying too fast. That was a DoS. That did cause all the customer impact. That actually wasn't a root cause of the outage. That code hadn't been changed in years. Before all the connections had to get reestablished, what we worked out was that the mapping service was just sending a bunch of timeouts for existing connections. That turned out to be triggered by the mapping service having longer GC. It was written in Java. It had GC pauses longer than timeout. GC pauses longer than timeout, you drop that connection. Actually, the team wanted to stop there, and just say, fix the exponential backoff and we're good.
Again, it was the SVP who really pushed the team to go one step deeper. This is where in complex systems, it's like these tiny details as a group. They had changed the SSL certificate to a certificate that, for whatever reason, the library generated more garbage as it was validating that it was a certificate. That was just like one little thing. Most of the change that went out was new functionality. They did this because the security team had been nagging them. That was the thing that actually caused garbage collection to flip over this one-second outage. That's what caused the DDoS. One of the things is there was no way, I think, of any human testing this or understanding this. All you're doing is mitigating the risks. Even for this control plane, because we at the bottom had 100,000 now talking to us, they had to be careful themselves.
Yes, what didn't work was driving up release velocity more often. We ended up at three weeks per deployment as we kept adding regions, and kept having to think about how do we do AZ independence and regional independence, about three weeks per deployment. We kept AZs totally independent because we stressed that to our customers. The customers want regional independence as well, so we had to think very carefully about how to lay those out. Hours per AZ with the classic you start on 10 hosts. You wait maybe an hour. You scale up. I think time finds problems, people fix them. That's Marc's law. I linked to his blog down below. Really good blog post. Like, Marc was there at the same time I was. You get a lot of the rhetoric. Move faster. If you move faster, and you just test harder, again, for some systems, that works great. For a lot of systems, it doesn't work well.
Observability. I think with SLOs as an idea that it's some contract between you and your customers. I've never seen it work and I don't think it works. Now, an SLO, though, as a contract that you have on yourself, I think is really important. I don't hate the concept in general. I just hate the concept that you can use it as a communication mechanism, particularly because all these platforms are very complex. You're now having to aggregate the complexity and try to explain that to a customer, you're explaining every blip on a graph to a customer, even as the graph is synthetic because you have all these complexities. The problems of a large surface area. This gets back to like, why do you need fine-grained SLOs? This was my team's highest impact issue. This was in 2015. The deployment goal was literally just tuning about five TCP parameters that the data plane used to communicate with the EBS service.
This was done by a TCP expert, and exhaustively tested, like performance testing and everything. This guy was very confident that it was going to be right. We passed all the tests. We did the special case load testing. We got it fully deployed to 35% of AZs. Then, why was this my highest impact production issue? It was about 1% of the fleet. I think maybe 10% of who we deployed to in this one AZ was not hitting its IOPS targets. IOPS is if you're running instances barely using its disk, that's fine. If you're doing a database and EBS is used by RDS, slow queries and then the box just becoming throughput limited is actually very bad. I think we're at the point where some customer complaints have come in, but this had the potential to be really bad. This just happened to millions of instances, I think, at that point in time.
The root cause turns out to be in the largest AZ, which is the one we happened to deploy to, and by largest, I mean physically largest. Highest latency between the service of EC2 and EBS, on racks with the worst packet loss. Packet loss was just other customers on the rack doing things. The numbers are tight, like 1 in 100,000, but it was enough packet loss to start spiking it. Then the TCP tunings actually made things worse. To this again, it just comes back to when you have a really large surface area of complexity that you're abstracting over to, tiny little things that are impossible to test against are going to make things worse. The big thing with this one, we did have a metric that we actually weren't watching at the time in our own little data plane software. That's fine. If we'd been watching this one metric, we would have seen this very quickly before we'd even run out to 10% of the fleet and pulled it back.
What didn't work was big broad customer facing SLOs getting into these big broad error budgets. Systems are way too complex. I can tell you each of my teams had between 20 and probably by the end close to 50, they have a dashboard with all these graphs. They weren't looking at it constantly, but they were looking at it constantly around a deployment. Like, I know I went out to 5% of the fleet, did anything just spike? It was very hard to automate it. We did get better over time, worked out how to automate this graph watching. What we needed, especially early on to manage risk was just people looking at the graphs. The takeaway is just roll out slowly and observe deeply.
Putting it all together, this is just a nice pipeline view. Up to the pre-CI stuff, prod canaries first, where your synthetic tests are, go there first. That's the way you discover actually in production, do you work against your dependencies? Small region, something where your customers have lower standards is the ideal, I think, for a general platform. Deploy to a subset of your largest, riskiest AZ early. Then slow rollout hours per AZ. At Datadog we were faster than this, but we did end up using a very similar pipeline.
3. 2019 - 2023 Datadog: Scaling Safe Testing in Production in a CI/CD Culture
Next thing, this is Datadog, 2019 to 2023. This is me taking what I learned at Amazon, coming in as a VP and trying to help a platform team that was having problems. Datadog was 350 engineers. My org was initially 50, within about a year was about a third of this. The big thing with Datadog is we were hiring aggressively. The cultural issues were a big problem. Compared to EC2 we weren't big, but we weren't tiny. We had about 10,000 VMs in two regions. Product teams like Datadog was very proud of its success versus things like CloudWatch. They did have a great product culture. Things like feature flags, things like being able to get something in front of a customer against their real production data and getting feedback from that was a big part of how Datadog grew as fast as it did. That was the product team.
The platform teams were forced to keep up. That was failing as things were getting bigger and bigger. Feature flag-heavy culture, good product. Datadog at the time was acceptably flaky. The customers were customers like Airbnb and Shopify. Generally, with observability, it's not the end of the world if it goes down for startups. As companies get bigger and bigger and it blocks your deployments, it becomes an important thing. That was happening at Datadog. You had these enterprise customers come in with much higher standards about outages. The problem that I had was how to apply safe, progressive deployment without slowing the customer culture.
Introduction to practices. What didn't work was me as platform SVP dictating what worked at his last company. I say this through experience, because Camille, my co-author, based on our time together at Two Sigma, the FinTech, sent me these cups. These are literally cups she sent me. They say at Amazon We. Camille was suggesting I use that phrase a little bit too often. What you learn is processes dictated by an SVP will be done poorly because people don't understand why they're doing it. They're only doing it because someone up on top is yelling at them to do them. Processes done badly are worse than no processes. Generally, like you have to find a way to motivate people within a culture to start doing better than you. What worked for me was to get very regular. I ended up doing an Amazon thing. We did weekly post-mortem reviews.
I made sure I attended. I didn't really yell like Charlie Bell, but I did make sure my voice was heard at those meetings about making sure the team who was just impacted, they were where I was going to introduce the best practices. I didn't go broad with it. It was like starting with the effective teams, find the champions. That's how I managed to scale it. To introduce practices, start small, and expand.
Deployment planning. Now this is on the flip side though. You start doing this, you start seeing success, but you still see particularly when you're hiring fast and then also acquiring that every team almost has to get bitten before they want to start doing the better thing. Of course, that's really bad when you're launching product after product after product, because customers get very frustrated. Deployment planning, when I talked about it earlier, actually making people plan out the deployments. This ended up being product teams as well as the platform team. What didn't work was just saying, here's some best practices. Go talk to the champions, and then go, let you do what you want. What happened was we had one bad outage, but we also just had this trend of increasing number of outages. Rapid hiring plus acquisitions, no consistent norms. Product mindset, again, it's a good thing, but it does tend to mean that reliability is just reactive to current trends.
You see these in the product managers. They don't give a shit about reliability until you have an outage. They're like, why don't you do better, engineers? It's like you're yelling at us to do faster. What did work was after a major outage, my manager, Alexis, and the CTO and myself, I mandated phased deploys. I didn't want to enforce it on everyone without getting some amount of buy-in. I created a group of staff engineers to standardize and we rolled it out to their orgs. This is basically the same takeaway as earlier, if you touch prod manually, it needs review. It's a fact of life for platforms. Feature flags. What didn't work well for us was feature flags. The code was just getting way too complex. The blast radius was too big. Infra rollouts were slow. What did work is this technique of sandboxes and shadow deployments.
I hadn't seen this used this much before, but parts of data were really good at this. A lot of my job was just making sure I was advocated to a larger role. It's basically, find a way to iterate on code without affecting customers. I wanted to do a picture, like this is a very simplified picture of Datadog. This is before doing any shadow deployments. What you see is two things here. The first thing is your data engineering team, this is why they love Kafka. Kafka is this really nice thing because you can just like put this new code in production and make sure it doesn't affect anything downstream, pull from Kafka, and test it against real data. That's a great thing. What you see is this combination with sandboxing on the right. If you can pass through headers saying that you're basically a test account, now they can get down to where they are in the stack and actually test against the storage code as well. You see on the left, the async with Kafka, on the right, the synchronous. The technique is called sandboxing.
Finally, deployment tools. I had the deployment tool team rolling up to me for almost the complete time I was at Datadog. There was this strong drive, like, we just need one great CD tool for all teams. It just never worked. It was just one of those things that, especially as a leader, they're like, I couldn't keep throwing engineers trying to build this very bespoke, but great deployment tool. What was happening was the platform teams, because they had all that complexity, because of the statefulness we talked about earlier, they just needed complex features that the app teams didn't. It didn't make any sense for them to prioritize the platform teams ahead of the app teams, because the app teams were the ones we wanted to move faster. We wanted them to use the proper CD. It did work with the platform teams hiring their own SREs, build platform-specific automation. This is something where it's counter maybe to the Amazon list. It's a place where we just needed teams building automation. Platform teams often need to invest more in automation.
4. Core Takeaways
Core takeaways. As I said, I'm a strong believer in this. I know it's painful. I just haven't seen anything else work for foundational platforms at scale. Prioritize synthetic monitoring. Again, it's the biggest thing, having left AWS that I saw missing in the teams that I inherited. Find a way to iterate on code in prod without affecting customers. Look if feature flags work for you and it's great. Otherwise look at some other techniques. Time finds problems, people fix them. This, I think, is the core thing we know about platforms. It's a nice thing to send Marc's blog post to people when they're critiquing why you were so slow. Roll out slowly and observe deeply. Yes, platform teams, we often need to just do this or everything that I've said above, ourselves. We can't rely on the company platform because we're one and there's a thousand other teams who want something else. This is the final summary, which is just for foundational platforms, continuous delivery requires safe progressive deployments.
Questions and Answers
Participant 1: Can you provide your definition of synthetic monitoring?
Ian Nowland: What it looked like at Amazon was something that called the APIs from the top as if it was a customer focused on one vertical segment of functionality. It wasn't like the dumb just hit a website with HTTP. It was complete workflows. Probably the easiest one to answer was when I was on Elastic MapReduce, the synthetic monitoring I built there would actually start a cluster, run some jobs on the cluster, terminate the cluster. I think of it as a mini user simulation for a vertical, but using the API is just like a gut. I think that's one of the key things. I talk about it in the book. It really forces your internal engineering team to go through the pain. Like if you're flaky, then your customers are going to experience this as well. It's end-to-end workflows, pretending you're a customer, essentially.
See more presentations with transcripts