Transcript
Vanessa Huerta Granda: Welcome to when incidents refuse to end. Let's talk about incidents. If you get trained on incidents, you might see something like this. This is how all incidents happen. No, that's not. If you do an incident review, you might want to do a timeline like this. Something breaks, we fix it, and then we go back to normal. If you've been part of incidents, you know that that's not true. It's not like we have a bug, and then we revert it, and then we all go and celebrate. I wish things were like that. Usually, it's a little bit more like spaghetti. Incidents in real life, they're more like, wait, what's happening? What is that? What is happening? How is that connected to that thing, like that picture right there? Who owns that thing? That's a really good one when you don't know who actually owns that.
How is that connected? Wait, we fixed it. Wait, no, we didn't actually fix it. There's a lot of chaos. It's in that swirling chaos where we see the truth both in our systems and in our teams. Today, I'm going to talk about a special category of incidents, the really long ones, the marathons, the ones where maybe you think, do I still live in my home, or do I live on Zoom or in Google Meets? I think that those incidents really shape us. They show us the difference between work as imagined versus work as done. This is where I show my incident nerd.
First, a brief introduction about me. I don't call myself the incident whisperer, I just really like this picture. I do spend an unhealthy portion of my life thinking about incidents. I lead our resiliency engineering team at Enova. I'm part of the Resilience in Software Foundation. For the past decade, my world has been incident response. I've been the sole incident commander on call for a good four years in my 20s. I have built and I have scaled escalation programs and reporting programs and analysis programs. I've trained responders and I've led retrospectives. I have seen incident management evolve across many organizations. I did do a Rumspringa from Enova. During that time, I did get to work with teams across the industry. I did get a front row seat to all the creative ways and sometimes really chaotic ways that companies respond when systems fail. I also have twin toddlers.
Even at home, I'm deep in incident management. At first, I was thinking that it involves more snacks at home but actually it may not. We might have more snacks during incidents. Something that I have noticed over the past 5 to 8 years is that something has changed. Organizations, they don't just want quick fixes anymore. They want lasting resilience. This resilience can be informed by the hard lessons that we get from these longer lasting, more complex incidents, and what they can teach us. Today we're going to be exploring those and what we can learn from those long-running incidents, the things that stretch from like sprints over to marathons and what they reveal about our systems, our people, our organizations. Before that though, I do want to highlight that I'm an industrial engineer by training. I'm not super technical. I haven't coded in years. The way that I like to describe myself is that my superpower is that I can speak tech to non-tech people.
I think that's a pretty useful skill. It turns out that the best way to get technical knowledge is to just be told to go on call and to be told good luck. I've been in a lot of incidents. I've been in tiny ones and big ones, and ones where I was at the beach and ones were at 3 a.m., I have learned more about my company and my systems, and human behavior from incidents than from any amazing, crazy architectural diagram that any of you can give me.
The Nature of Long-Running Incidents
Let's talk about what happens to us during long-running incidents. I used to joke that I loved it when incidents went through lunchtime, because it meant that I could expense my lunch. I just paid $4 for a banana here at the hotel, so I feel like it's worth getting to expense your food. I do think that nothing bonds a group of engineers more than just sitting at a conference table and eating a bunch of greasy sandwiches while trying to solve a problem. You get to feel like you're superheroes. Other things that I have had to do as an incident commander has been to remind people to take bathroom breaks, rotate responders before they lose brain function. I've had to open up doors to our war room because the war room maybe was built for 10 people and we had like 30 people in there, and we needed to get oxygen in.
Pass my laptop around for people to put in their lunch order. I've had to tell executives, "Please, I will let you know any updates that you need to know. You can go back to your desks." That's a nice way to tell your executives that it's more helpful if we can solve this within our own. I also mentioned that I'm now a mother. As a mom, I know that my job is to make sure that my children are fed, and that they're not too hot and not too cold, and that they remember to go to the potty. It's a lot similar. When you think about it, some of these things that we do as incident commanders are pretty similar. When you're hiring for an incident commander, you don't put in the job description to parent the responders, but it is part of it. Because engineers cannot troubleshoot if they're dehydrated or if they're emotionally crumbling.
Part of the incident commander role is to make sure that responders are best prepared to bring the incident to resolution. It all goes back together. When I talk about long-running incidents, I'm not here to define long-running with like, it has to last these many hours or this SLA. The incidents that refuse to end, it's whatever feels too long. The first talk mentioned that incidents, a lot of it is like vibes. If you're awake at 3:00 in the morning, waiting for a restart to finish, congratulations, that's a long-running incident. If you missed your kid's soccer game, or if you're at the beach, or if you're just really annoyed, that can be a long-running incident. What I truly believe though is that these incidents reveal a lot of realities. They reveal the technical realities of our systems, organizational constraints, and human needs.
All incidents, but particularly long-running ones, amplify the pressures that we feel as responders and commanders: time pressure, coordination pressure, cognitive pressure, and visibility pressure. Time pressure because we know that every single incident means that money is on fire. It's important, we need to get back up and running ASAP. Coordination pressure because we're swapping people in and out, we have a lot of moving parts. Cognitive pressure because of the complexity overload. Then this is the good one, the visibility pressure. Especially with long-running incidents, everyone's watching you. Especially if they're running even longer, and you might have the big wigs, your C-suite executives. It is a lot of pressure. They unlock huge values. The perks of these incidents is that you can learn a lot. You can learn technical learning. Incidents is where engineers are made. It's where you can learn the difference between what the architecture diagram says versus what your architecture actually is built off of.
How your data flows, all of that fun actual engineering stuff. Organizational learning. You learn who actually owns what. You learn how you actually prioritize things. You learn any systemic blockers that have existed for years and years even before you joined that organization. You learn about people. It was during incidents that I learned that I wanted to be a people manager. How do people handle stress? Who becomes a leader? How do we actually care for each other? I genuinely do believe that incidents make us better engineers and better teammates, especially the ones that are longer running.
The Tuesday Issue or the "D" Word
What I'm going to do next is I'm going to go through three examples where I'm going to highlight the pressures and the learnings that we just went through. Story number one is the Tuesday Issue or the D word. In this case, the D word is degradation. I really hate the word degradation. I think it's like bad juju. You can't really quantify it. It just means that the vibes are bad. I just don't want to be here. This story takes place in 2015. To set the scene, it's like baby Vanessa SRE. I had been hired to run an incident program. Remember, I mentioned that I was an industrial engineer in college. I did not have SRE experience at the time. I had worked at a startup where they just needed people to do things, and so I did all of the things. I think now we would call that glue work, but back then, that just was not a word that existed.
I called it Vanessa work. At this point, I had handled a few incidents. I didn't really fully understand how every system connected, but I did understand how our business worked. I knew all of the big things. I knew we had a site, and the site needed to be up. That's a big one. I knew we had a contact center, and they needed to be working. It needed to be functional. A big part of our business was bank files, and I know that they needed to go out. Yes, those were the pillars. Everything else is just details. It is Tuesday. It is 10 a.m., and we start getting reports. First is the contact center's running slow. Things are just slow. What does that mean? I don't know. Things are just slow. Sometimes customer information is not loading. Some applications are struggling. Things are just not on fire. They're just smoking. It's just like that Chrissy Teigen phase. Things are just not very happy.
We start troubleshooting. The database is looking sad. We correlate this with the time that people started reporting issues. Happened around the same time. We bring in more people. We bring in DBAs. We bring in frontend engineers. We even bring in some people who are at the actual contact center to let us know, ok, how do you actually usually do your work? That's actually something that's really important to do during incidents. We don't know why. We don't know how. We just want things to go back to normal. We decide to kill a bunch of things. We wait a little longer. That doesn't work. We're just turning things on and off again. We're in the cycle. Kill things, turn off, restart things, and then pray, do this rain dance. By 1:15, things are looking normal. Then at 2:37, things are not looking normal. We see a spike in errors.
We go ahead. We start killing things again. Just like we're in this cycle. I learned a few fun facts during this incident. Remember, this was my first few months at this job. A fun thing that I learned were the acronyms that people used and what they meant. I think that sometimes people use really silly acronyms. Like our frontend, I thought everyone was saying Snoopy, but they were saying CNUFE. It was spelled C-N-U-F-E. I'm like, that should be pronounced Canoofy. Those are the things that you learn during an incident. You learn as a new engineer that older people might be talking about some things that you have no idea, and there's no place on the wiki where you can learn this. Another thing is that I really love remote work now, but when it comes to incidents, I think that there's something really super special about all of us being in the same room together.
In this case, we were in this one room, all of us together. We were writing things on the whiteboard. We were diagramming, grabbing each other snacks. It was really fun. All the while knowing that every single minute that this incident went longer, we were losing money. People could not use our site, that people in operations couldn't do their jobs. This was a big deal. Incident response is really weird because it's serious. In order to be good at incident response, you can't take it too seriously. You have to know that you need to meet your end of the bargain, and that's bring services back up. If you have very high pressure, that's not very helpful.
We continued troubleshooting. I found my incident report for this, so I have some timelines on this. By 3:11, things are looking good, and so we start bringing some services back up little by little. Some batches, some call center things. Then by 6:30, everyone's really tired, and we're like, "We're good to go. Let's just go home." I don't know what fixed it. Was it magic? Was it brujeria? I don't know. I don't mess with that stuff. We literally wrote in the incident report, which I found and which I wrote like 10 years ago, two children, one husband, four houses ago, I wrote, find root cause. That day was really fun. It was not my longest incident, but I had caught the bug. I had met a ton of people, new tools, new knowledge. I learned what queries take too long to run. Most importantly, I learned how to find that information.
It was the first time I had used Splunk, actually. I learned how all of that then impacts performance in separate areas. I learned how our frontend that our end user sees also impacts what our contact center sees. I learned that we had a lot of unnecessary work happening behind the scenes that we should really be addressing. Then we had our postmortem, and in the postmortem, we talked about the hardware improvements that could take place, but also, we learned that it's up to the IT engineers to schedule those improvements. It's very difficult for them, and they need to collaborate with product owners to figure out the best timing for these things. I learned that as an org, we might think that we have competing agendas, but in reality, we all want the same thing. We want a well-performing system. I learned that our DB team had been pushing to improve some things, like our indexing strategy, but they didn't know how to get this seen by the people who make these decisions.
It turns out, during this postmortem we were talking, we're like, there's this meeting that we can get in, and we can get all these things changed. This incident went from 10 a.m. to 6:30 p.m., and I learned a lot. I learned how our systems are interconnected. How I personally can use our observability tools. Who is in charge of what part of the system. How we don't know what's eating our resources, and how prioritization actually gets done. To me, this is a really great example of what I mean when I say the difference between work as done versus work as imagined.
What I didn't mention in my story is that throughout that Tuesday, people kept saying something. People were like, is this the same thing as last time? No, not that thing again. Then some people were like, this always happens on Tuesdays. Here, I posted the screenshot of like to investigate the root cause from the incident report. You know when you're in an incident and you resolve it, and then you have the postmortem, you're like, let's figure out the root cause. You don't actually investigate the root cause because the incident has been resolved. Then it kept happening. Not every month, but many of them. We figured out that it was always the first Tuesday of the month. Something that we do in my company is that quarterly and yearly, and whatever cadence you want, we take a look at all of our incidents and we make recommendations to different parts of the organization.
We were able to create a tiger team. We analyzed the data. We managed to get these people to take time away from their full-time work so that we could understand what was actually happening within our system, so finally we could investigate the root cause. It wasn't just one thing. It was a combination of multiple issues within our work. Here's a fun one. It was also related to the fact that many people get paid on the last Friday of the month, which means that the first Monday of the month we have really big files and our systems are stressed out on the first Tuesday of the month. Knowing that, we optimized a bunch of things. We optimized our queries, our hardware, background jobs, team structure. The real work, the real fix was not in debugging code. It was in understanding our socio-technical system. If you're an incident nerd, you know how much we love saying socio-technical systems. That was the Tuesday incident.
The Fire Marshal Incident
Now let's go to story number two. This one is the one with the fire marshal. I'm actually stealing credit from a lot of people here. I was not even on call for this incident. At this point, it was 2020, and I had been promoted to manager. I had been handling incidents for about five years at this point. I had pretty much drank the Kool-Aid. I was really into incidents. I really liked them. It's pretty late. It's COVID era, 2020. I'm in my pajamas. My hair is up in a bun, no makeup. I just decide to look at my phone, and there's a lot of activity happening. Everything is down, critical, a data center issue. I'm like, let's see what's happening. I jump on the call, camera off, because again, it's pajama time. The rumor going around is that our data center is on fire. People are really dramatic on these things.
I'm not even on call. I'm not the SRE manager on call even. That's my colleague. I'm like, "I oversee this function. I've got this, go rest. We'll handle it." She's now my boss, so I think that was a really great career strategy, if I say so myself. Something that you should know about my company is that we have lifers. People who really know our systems. Again, rumor is that the data center is on fire, so we have two engineers get in the car and drive to the suburbs where our data center is. I don't know, they thought they were firefighters. They're driving. Meanwhile, on the call, people keep joining because everything is down. Every single alert that PagerDuty could send out, it has sent out. Every single thing is down. Tons of people are on the call. Everyone's camera off because it's Sunday COVID late at night.
Then the executives start to join. I don't know who called the executives, but somebody did. They like having their cameras on. I'm like, ok, Vanessa, splash some water. This now became a cameras on meeting. Google Meet didn't even have the beauty filter back on, so that was like a huge thing. We have a ton of people on the call, and the two people who are actually doing something are driving to the suburbs. Our onsite engineers tell us that the fire marshal is not letting anyone inside, that there is smoke coming from a battery. That's where people were dramatic, that it was on fire. It was not on fire, but there was some smoke coming from a battery, and that people were evacuating. There's progress happening, but there's no timeline. Then we're like, everything is down. What do we do? Do we fail over? Do we fail over to our disaster recovery strategy?
We have this disaster recovery strategy, but failovers are not super easy to do. I don't want to do them if we can help it. We reorganized. We kindly kick everyone out who's not actively working. We kindly tell them, we've got people at the data center, and we will need you all in the morning, so how about you go get some rest? Then we split into two focus teams. They report up to me. The two focus teams, one of them is focusing on prioritizing what needs to be brought back up, and then the other team is focusing on getting things set up if we do decide to DR. Then I'm on the call with the people who are at the data center, and we say, by midnight, we're going to make the decision, go or no go, if we DR.
Then it becomes the IT show, as it should. We don't make the call by midnight. We make it at 12:20 a.m. It's really dramatic, because we're on this call, and we have kicked a lot of people out, but we hadn't kicked all of the executives out, because I don't get paid enough to do that. We have to make this decision, and my executives must have really paid attention to the incident training, because they were actually letting us make the decisions. They were actually letting us talk about what was happening and following our process and all of that. I can tell you that in the year 2020, at midnight, I really wanted somebody else to make that decision. I did not want to be the one making that decision. I did. We decided to no failover. We decided that we were going to wait. Our folks at the data center were like, things seem to be moving along.
Let's just wait. There had also been some crazy things happening earlier in the year. We had acquired another company, and so we just didn't want to mess up with doing a failover. The people at the data center were like, you guys should go to bed. There's nothing we can do here. When the data center is back up and running, we'll call you, we'll page you. We did. At 2 a.m., that call came. They didn't call me, because remember, I was not on call. That was really fun for me. A lot of people did go on and really helped out. We had networking, security, really so many teams working overnight just making sure that things were good. They were verifying systems, connectivity. There were some VMs that needed to be replaced or restarted. By 4:30 a.m., things were coming back up, and then at 7:00 is when I joined again, and I said to the incident commander, who was not my boss, I was like, "Go to bed, I've got it." No one was well-rested during this, but we were not zombies. We were able to make decisions. We were able to send out the communications, the executive summaries, all the things that we needed to do. That was good.
Our incident report literally said, while we were not operational for several hours, this was a win due to our production incident process and our subject matter experts who were able to share their insights and help make decisions. That gave me the feelings when I was going back and reading my incident reports. Why is this a win? Because this incident taught us our true business continuity capabilities. It taught us where our documentation was real versus a fantasy. It taught us how well our teams can coordinate under real pressure, under pressure when everything is on fire, how decision authority really needs to be crystal clear, all the things that cannot be optional. We were able to formalize it. This idea of removing lurkers from the call, the idea of splitting efforts into parallel teams, rotating engineers, writing down assumptions and rules, all these were strategies that we had improvised that time, and maybe a few times before, but then became part of our runbooks. That is something now that I teach others how to do. That was the fire marshal incident.
The Holiday Downtime
Now I have my last story and it's a more recent one. Conveniently, it's a holiday story. For years, I was always on call during Thanksgiving break. I was not born in America, so my family, we never celebrated Thanksgiving until I got married. I just really cared about the sales. For the first few years, we always had an incident during Thanksgiving. Like I said, I didn't care. It was just like, ok, this is fun. Let me ask you, how many of you have had a code freeze policy in your lives? How many of you had still had incidents during that time? It happens to all of us. It's universal, happens to a ton of companies. Don't feel bad, it's a thing. Because before the freeze, there's a big rush to get things out the door. Then after the freeze, there's a big rush to release all the backlog from the freeze.
I do think that it's important to keep stability top of mind during the busiest period. If the holidays is the busiest period for your company, then it's important to keep stability top of mind. Incidents matter because customers cannot use our products. Incidents do not care about Kubernetes. Incidents do not care about the cloud. They do not care. They care that they cannot buy tickets to see Taylor Swift. Holidays are when customers need to fly home or when they need to buy gifts online. That's important. Outages do hit harder during those moments. At my company, we did do away with the strict code freeze. Now we are something like more chill. I call it a code chill. It's just like don't do anything when everyone's on PTO. Don't do a huge version upgrade the day before Christmas, which works most of the time.
Now this incident, this is the third one. This is late November, and we're back from Thanksgiving break. We see degradation again. This is one of our business units only. We don't know why. We think it might be a vendor problem because we had been having issues with one of our vendors. We chased them down. It's difficult because I think that the relationship manager was still on vacation. We had to figure out, like find out where we keep this information, a lot of stuff. Turns out that it's not that vendor. We had already wasted a few hours on that. We do have a manual workaround to get customers checked out, but it's time consuming, it's not sustainable, and the queue is piling up. At this point, it's like an extra four hours. We were like, let's just restart one of these upstream apps. It works. The queue drains and we all go home and we all feel really smart.
I feel like King Charles. Until the next day, it happens again. Then the next day, and then the next day, and then the next day. It was not a one-off. It was just like this little turkey trot that none of us signed up for. Something to note is that this part of the business had come from an acquisition and many of the original engineers had left. I don't know if you've ever been in this position. Some interesting things happen when something like this happens. Number one is that institutional knowledge sometimes is gone. The dependency maps, all this understanding, like the architecture, how all that works, is just not complete. It might be fuzzy. Then some people had become single points of failure. We have single points of failure in December. Chicago public schools are about to go on break, people just want to use their PTO.
We're like, we need to do something about this. We were restarting things constantly. It was like a whack-a-mole on rotation. Here I'm sitting with the people on my team. I'm like, if it happens every day, is this still even an incident, or is this just like how things work now? Do we just have a rotation where we just restart things? It was causing a big impact. It was causing that business impact that we care about. By restarting things quickly, we were not capturing the data that we needed to actually solve what was happening and what the real problem was. We got anxious. We got really twitchy. Like, any single thing that we saw, we were like, incident, incident. We were just paging things. It was a little bit madhouse. We were overreacting because we had under-understood the problem.
What do we do? Right before everyone disappeared for the holidays, we decided it's time to regroup. We assembled a deeper investigative group. We had people from site reliability engineering. We had people from the product teams, business owners, infrastructure. We got all together trying to understand what was happening. Just like my first example, it was not just one issue. It was a constellation of issues. We had, of course, capacity being squeezed by seasonal traffic. We had unnecessary background work. We had a health check that was causing more harm than good. We just also didn't have the alerting that we needed. We didn't know about this issue until things were down. Then we also had some hardware that needed upgrading ASAP. We removed that toxic health check. We added alerting. We prioritized that big IT upgrade that we needed to do. We moved it to Q1 because people were going to go on vacation.
We identified where we needed to be doing some cross-functional changes. All of that. Then we had the postmortem. During the postmortem, we talk about what goes well, what goes wrong, all of that. Then we said, what about the next holiday season? Months later, actually, in Q3, we reviewed what could go wrong this holiday season. We reviewed expected traffic patterns. We got that from business owners and marketing. We reviewed any upcoming changes and any new changes that we had made in that time in between. Marketing campaigns, like I mentioned, because promotions drive traffic, and traffic drives incidents. The next holiday, we were more prepared, and we were able to remain stable.
As you can see from all the incidents, it wasn't just technical. We learned a lot of things. We learned to not assume stability. We learned that systems are going to drift, that context is going to change, that we can't remain complacent. We learned that decisions matter, that architectural shortcuts and staffing choices are going to linger for years. We learned to take people into account when thinking about capacity planning. Because if your subject matter expert is your single point of failure, and that person is on holiday break, we need to take them into account. We learned that a restart is not a solution. We need to learn. We need to adapt. We need to figure out what to do next. The restart is really just a delay. We learned that we need to include people outside of just engineering. I think that's a really big thing that I would love to see our industry change, where we include people outside of engineering in our incidents.
We've got to include marketing, product, operations, they all influence the risk that we encounter. This was not our flashiest outage. It wasn't the data center on fire. It did teach us what I think is an important truth with resiliency, long-running incidents, they don't just reveal system fragility, they expose that organizational fragility. It's important to address both, because that's where resiliency actually lives.
Long Incidents Expose Our Fragility
We have gone through three very different incidents. We had a literal data center fire, an internal systems complexity marathon, and this recurring holiday instability saga. They all have different shapes and timelines and stress levels. What I'm able to see about these is that long-running incidents are not just technical problems. They're more like organizational scans. They show you where your real fragility actually lives. They reveal how fast we communicate, who we depend on, what we actually prioritize when it's needed, where our blind spots are, where we can exhaust our people to keep systems up. We can learn this from all of our incidents, but we learn them more and more easily from the long-running ones. When things break fast, we can react fast. When things break slowly, that's when we see those system interdependencies, that fatigue and cognitive load, how staffing, how attrition can shape resilience.
You can see how decision quality degrades, the difference between hour 1 and hour 10. The way that your brain works, works differently. The business pressure that's creeping in, and how when recovery becomes ritual, instead of understanding and fixing the underlying contributing factors, we need to think, are we doing things right, or is there a better way to do this?
Develop Processes for Endurance
Long-running incidents, they take a toll. When you have a long-running incident, that incident is asking people to give time that they don't have, sleep that they don't have, and patience that they don't have, and childcare that they don't have. Incidents don't care if it's Christmas Eve, or Thanksgiving, or flu season, or your birthday. When you think about resilience, it's not just about the system staying up, it's about the people, and how the people stay up right. My recommendation is that we build systems and processes that are built for endurance. Think about incident role rotations. Think about multi-stream work coordination. Think about smart paging. Think about, do you want to page everyone or only whoever is needed? Think about escalations that make sense, that are humane. Think about knowledge distribution across the entire organization. Do away with those single points of failure. Decisions that consider both the system, and the business, and the people, because we don't just want to survive these incidents, we want to thrive.
My ask to you all is that the next time that you're in an incident that is going on for like six hours, or three days, or two weeks, or however every other month, don't just think about fighting the fire. Think about studying what keeps feeding that fire. Think about the conditions, the people load, the decision flow, and try to figure out how you can improve that. I think many of us know how to respond fast. Then the next step is learning how to respond for as long as it takes. I think that's much more of a people problem than a technical problem.
Summary
In short, long-running incidents expose what actually happens versus what we imagine happens. Success is not just technical, it's the people. It's coordination. It's that decision flow. Pressure reveals brittle processes, unclear ownership, and hidden dependencies. It's important to invest in structure in case of a marathon incident, not just a sprint incident. That means thinking about rotation, stream splits, emotional care on breaks. Finally, each messy and exhausting incident makes us a more resilient organization.
Questions and Answers
Participant 1: You talked about experts being single points of failure, and you talked about knowledge distribution. Do you have any examples of what worked in terms of, in your organization, we want to have better knowledge distribution. Did you have some successes there?
Vanessa Huerta Granda: Yes.
Participant 1: Maybe you could share how that worked.
Vanessa Huerta Granda: I think distributing knowledge can happen in many ways. I actually just had a conversation with some leaders who, it's hard to teach somebody how to troubleshoot because it's more like you just have to do it. Something that I really like to do is invite more people to your incident reviews, and so when you're having your incident reviews, they can see what's working and what's not working and how your systems connect and all of that. If you want to take it a step further, it's thinking whatever artifact comes out of your incident reviews, give that to people who are onboarding. That can be a way to document things. That's something that I've seen some success in. Of course, your regular documentation, like your READMEs and all of that. I personally haven't had a great experience keeping those up to date. I would love to.
I think so far it's getting people to actually do the work and actually be part of it. I think also a lot of the times when you're in an incident, as a subject matter expert, it's like, "I got this. I will do this." That's really great. We definitely appreciate it. Especially when every minute is like money that's being wasted, we want to solve the issue as soon as possible. Think about which incidents you can bring somebody in along so that they can learn, so the next time they can help fix it. Then the other option is, of course, to have them at the review because that's not as high-pressure situation.
Participant 2: For many incidents I've noticed that it's also going to be the first time that the responding engineers are going to have to make some tough decisions. What can I do to foster that decision-making ability and to reassure them that it's ok that they are making it?
Vanessa Huerta Granda: It depends. That's the best SRE answer ever. Different organizations do these things differently. That's why I always like having an incident commander because you're taking away that guilt that the engineering responder is feeling. You're letting them know like, I trust you as a subject matter expert. I trust what you're telling me. This is why this decision makes sense. If you don't have an incident commander, then I think maybe this is where a game base can be really helpful where you're giving them that little bit of an experience in a more controlled scenario where they can try out making these decisions.
Participant 3: You talked about bringing folks into incidents that are outside of engineering. I have a two-part question on that. One, have you seen that those folks are interested? Then, two, what benefits have you seen from bringing those folks in?
Vanessa Huerta Granda: I have seen them be very interested. I think you need to strike a balance of knowing this is going to get very techy and letting them know, we need you because of X, Y, and Z. Let them know very specifically this is why we brought you into the room. Then at the point where their expertise is no longer needed, let them know, feel free to drop out of the call. I think it's very important. A lot of my expertise comes from a place where we do have an incident commander who has the flexibility to make those calls. That's really helpful.
Participant 3: Are those folks open to coming in?
Vanessa Huerta Granda: Yes, I think so. I think as long as you let them know, this is what's happening, this is why. A lot of the time these folks are experiencing the incident themselves, so they want to be part of it. Sometimes you have people, like maybe marketing, they will know that something from their end has caused an incident, but maybe they're not seeing it. That's when you have to give them a little bit of the explanation. That's why it's helpful to have either an incident commander or a comms person, maybe someone who's not actively troubleshooting, have that conversation with somebody else. Another good way to getting other people involved is maybe not even bringing them into the incident, but to bring them into the incident review.
Participant 4: This quote from your talk really struck me, and I wrote it down. Long-running incidents don't just reveal system fragility, they expose organizational fragility, and it's important to address both because that's where resilience actually lives. My reaction to that was that organizational fragility can be a hard thing to address as an engineer, like it's above our pay grades in some sense. Have you had any success in framing a narrative around a certain kind of organizational fragility that gets that to maybe be repaired? I'm assuming sometimes it's an executive level of decision that requires repairing that.
Vanessa Huerta Granda: There's two ways to go around this. There's the ad hoc way where you have these conversations, and you say, this is what we experienced during this one incident, and you try to get into people's brain that way. Then there's the way where I have seen actual change happen that a lot of organizations don't do, but I do recommend, which is sit down every month, every quarter, every year, and then go through a review of your incidents and try to highlight the themes that you're seeing. Maybe one of them will be something related to that organizational fragility, but get that to be seen by somebody, a manager, a senior manager, a director level person. Because I understand that it's very hard to get change happening after one single incident, but if you have a body of work of maybe 5, 10 incidents, then that can get people's attention.
Participant 5: As you talked about, the incident commander's role is really crucial, and knowing what kinds of things they should do and shouldn't do can really make or break the incident. How do you recommend making sure that there always is an incident commander and that they have those critical skills?
Vanessa Huerta Granda: That's a question for the ages. I think that being an incident commander is not the same as being a software engineer, and we are making a huge disservice to our industry by trying to get them to do the same thing. You can be a really great software engineer and not make a good incident commander. If you want to improve your incident response program — I've been to places where every single engineer manager has to be an incident commander — just think outside the box. That shouldn't be it. Think about what are the skills that you need for an incident commander and start training the people that you think would fit those roles. Give people a chance to try it out, and then give people a chance to not do it if they realize that they're not a good fit for it.
Participant 5: How do you make sure they have the skills they need?
Vanessa Huerta Granda: I train incident commanders, and I have a slideshow, and I present the slideshow. Then really what I do is like, watch me for a week, and then the next week I'll watch you, and then you're good to go. I think some of these things you just learn by doing.
See more presentations with transcripts