BT

Facilitating the Spread of Knowledge and Innovation in Professional Software Development

Write for InfoQ

Topics

Choose your language

InfoQ Homepage Presentations Building Resilient Platforms: Insights from 20+ Years in Mission-Critical Infrastructure

Building Resilient Platforms: Insights from 20+ Years in Mission-Critical Infrastructure

47:33

Summary

Matthew Liste explains the 12 core principles of building scalable, reliable platforms for enterprise tech teams. From avoiding technical debt and managing deprecation to building a strong engineering culture, Liste discusses how to design infrastructure platforms that reduce complexity, streamline developer workflows, and deliver long-term value for senior architects and engineering leaders.

Bio

Matthew Liste is the Executive Vice President and Global Head of Infrastructure for American Express. He was previously at JP Morgan Chase, where he most recently served as Head of Platform Services for Global Technology Infrastructure. Prior to joining JP Morgan Chase, Matthew spent nearly a decade at Goldman Sachs, and was named Managing Director and Technology Fellow in 2011.

About the conference

Software is changing the world. QCon San Francisco empowers software development by facilitating the spread of knowledge and innovation in the developer community. A practitioner-driven conference, QCon is designed for technical team leads, architects, engineering directors, and project managers who influence innovation in their teams.

Transcript

Matthew Liste: I'm Matthew. I'm at American Express where I run infrastructure. I've been there a bit over two years. I was at JPMorganChase for nine years before that, similar roles, platform engineering roles. Then I started my career in financial services at Goldman Sachs where I also did platform-y infrastructure roles. Came out of telecom and networks and started my career in oil and gas, so offshore embedded systems doing seismic data acquisition. I had a pretty brokered career. I went in up in financial services as a technologist, mainly because my wife is from New York. She wanted to move to New York. That's where the jobs were 20 years ago. There was no tech in New York back then, so you either worked at Bell Labs or you'd work at a bank if you were a hardcore engineer. That's where I ended up. It's been really good to me because guess what banks do? They do a lot of mission critical software, which means the infrastructure that they run on has to work. This talk is really about building infrastructure and platforms for that. This has come from an infrastructure background, the platform principles. I believe these principles really apply to anyone building platforms for reliability. That's really my experience.

What is a Platform?

Let's start with the definition of a platform, because I think it's useful. This is a platform that I'm standing on, for example. Dictionary definition, a raised platform that can hold people or things. This is one. If you think about the technical definition of a platform in context of what we're going to talk about is really a set of integrated technologies that form a foundation to build applications on top of. What are some of the best examples of that today? Cloud providers, AWS, Azure, GCP, all build these foundational platforms, really infrastructure platforms and different things that sew together that allow you as a developer to write a software on top of. This talk is really about the principles of building that. How many of you are platform builders? We're all consumers. Everyone consumes platforms of some shape or form, myself included. Everyone consumes platforms. If you're not a builder of platforms, and I hope this talk will give you an appreciation for the complexity and everything that goes on under the surface of the iceberg of what a platform is.

Why you should appreciate what others do, and the fact that you don't want to do it yourself. You want to leverage other people. Let me use an analogy for that. Think about a dirty analogy, but it's really a fair one, plumbing. All of the work that goes into making sure that everything that we do every day gets swept away, and that we get water back, and it just works. That's a platform. It's a sewage platform, a water delivery platform, an underlying infrastructure platform that's completely opaque and transparent to most people, except for when it breaks. When it works, you don't even think about it. You don't appreciate it. You just take it totally for granted. I think that's a great analogy for a platform provider. When you do your job well, no one knows you're there. It just works, and it's intuitive, and you take it completely for granted as a consumer. That's the focus.

The Principles of Infrastructure Platforms

I pooled this into the 12 principles that I'm going to talk through this talk. They're in no particular order. There's no reason for number one being number one and 12 being number 12. I've tried to group them based upon making this presentation flow. You shouldn't interpret these as more or less important. These are the principles that I have used and used to build platforms in a repeatable way with my teams. Just to give you some context, the platforms that we build, and we know we're at the cloud scale. I built platforms at Goldman, JPMorgan, now at Amex. Our consumers are internal developers. To give you context, at American Express, that's about 20,000 people that consume the things we build. At JPMorganChase, it was more like 60,000. That's our target audience that we build for. Of course, your cloud provider, you're talking about millions of customers. It's a different scale, but the same principles apply.

1. Deliver an Intuitive Experience

Let's get started talking about what these principles are. The first principle, number one, is about the platforms being intuitive and delivering an intuitive experience. Intuitive also equates to obvious, also equates to transparent just works. Think about platforms that are great and that you could think about in your life. They're a common platform. They're integrated. They're intuitive. They hide complexity. Meaning whatever goes on like sewage or delivering power to this room and this stage is incredibly complex. That complex is totally hidden from us. They appear magical. Every time, and I'm sure all of you can remember back in your life when you saw a new platform, and it was completely magical for the first few days. Then you started taking it for granted. Many examples. Mobile phone, internet, first computer for those that are old enough to not having had computers in your life and then had one.

Myself, old enough to be in that bucket. Or GenAI now. My kids who are two teenagers, they have completely embedded GenAI into everything they do, and couldn't imagine life without it. When I first saw it, I've had friends who have been in AI space and researchers for a very long time. The fact is a change that went from impossible to magic to take it for granted. Those cycles go very quickly. The point about great platforms is that they are magical, because they hide that complexity away. I love this quote from Arthur C. Clarke, "Any sufficiently advanced technology is indistinguishable from magic." It's true. You think about that yourself, you can't tell the difference. It just seems incredible that this complexity just works. The beauty of that is that complexity, another way to think about it is a platform provider is the one that makes sure they manage that complexity so you can manage other complexity.

None of us have the cognitive depth, although there are some incredible renaissance people who can go deep in very many fields. Most people end up specializing. The reason we can specialize in what we do is because someone else has specialized in what they do and built platforms for us. That's the beauty of platform and the layers around it. Steve Jobs also had a very early quote on this, "Simplicity is the ultimate sophistication. It takes a lot of hard work to make something unique, simple, to truly understand the underlying challenges and come up with elegant solutions." Again, I think that holds true for any great platforms you see. The complexity is hidden from you. You can enjoy the systems without having to worry about that yourself.

2. Build Common and Interchangeable Components

The second principle is around platforms, and if we think about the definition I spoke about, are a set of components that allow people to build on top of, which means it's not one thing. It's not a singular thing a platform delivers. Think about like cloud providers, deliver hundreds of different services or thousands for that matter. They all work in a consistent way. They snap together, like Lego. I like Lego as a metaphor. Anyone here, I'm sure, has played with Lego, have kids who played with Lego, have used Lego. Think about the infinite number of shapes you can build with Lego: a dinosaur, a car, the Eiffel Tower. You're dealing with a very finite set of building blocks that have a common API, how they lock together. Everything is built in the same way so they can lock together in certain shapes. We can build anything out of that.

It's really important that platforms have this sense of commonality and set of what I think was freedom from choice. For example, make sure every platform component uses the same set of observability basics or foundations. Make sure they use common identity semantics. Make sure that they use common namespace. Again, think of the cloud providers. You put yourself into one of the clouds. You can observe your entire components from that cloud from a common set of observability components. It's not all different. It allows you to observe your application and the underlying components to it. That platform aspect of interchangeable and interlocking components is super important for platforms to be successful. Otherwise, they're just different parts that you offer, but they don't work as a platform. They work as just a set of components. That aspect is very important when you think about platform building. To let alone, I was just thinking about with cloud providers also, common bill. You get billed for all of this together. The financial aspects of a platform is super important when it comes to that.

3. Use the 3 S's

Thirdly, I like to think of these as the three S's. It works very well with the stool metaphor too, which also has a platform to sit on. Think about a three-legged chair, only works if all three legs work. In my field, building for banking systems, credit card systems, and so on, you have to have platforms that are stable, they have to be secure, and they have to be scalable. Let's start with stability, kind of goes without saying. If it fails frequently, if you're not adhering to a certain high degree of nines to your consumers, they cannot rely upon your platforms to be there for them, and they won't use them. How many nines, of course, that's a commercial decision. Do I go for three nines, four nines, five nines, six nines? Depends on the platform. Also, at American Express, our credit card authorization systems, they run at six nines, seven nines.

They don't tolerate downtime. There are other systems that tolerate more downtime, like a mobile app. We vary based upon, of course, the commercial value of platforms. You can imagine that's incredibly important. Secondly, security, of course. People entrust us with their money. You better make sure that you manage that from a security perspective. Then scalability is really important. I've often built platforms that have worked until they didn't. They worked, and then users started using them, and we had runaway consumption. We found bottlenecks in the system, and they just started failing. I've learned to anticipate scalability. If you build products that people love, people will run at them, which means your volumes will increase. If you haven't anticipated that and built that in, you will then impact stability because you can't deal with scalability. It's super important that you think about those. These three are non-negotiable, again, for mission-critical software.

Now, to be clear, early days of building a platform, you will compromise on stability. You'll compromise on scalability. As you start getting clients who depend upon you, you have to live up to this. You have to be formal about that. Think about putting in place SLOs with your customers and managing closely to that. It has to be true not just day one, but day n. Consistently, you have to deliver to that promise with your customers. I did speak about cost because my view is that cost will fluctuate up and down. If you decide to get cheap on any one of these, you'll ultimately fail for this kind of software. You can dial up and down how much you spend based on how much features you put in. That's really the equation you do. Once you've decided to build a feature or build a product that you use for the software, you have to then be willing to spend the money it takes to do so properly with these three S's.

4. Be Evergreen

This ties into the next point, which is principle four, be evergreen. When you build platforms, your customers take for granted that it's going to continue to work. Unlike, let's say, a sewage system, where you put in the pipes, and hopefully at least for a few years, you don't need to think about those pipes again. They're static. When it comes to software, it's ever-changing. You have to be upgrading and managing the platforms and the upgrade cycle on a consistent and ongoing basis. If I think about one of the things that's been the hardest to do over time has been this. Because it is easy to compromise on this and say, I'm going to wait with that upgrade. I don't have time for it. I'm going to put up the feature functions. I am going to wait because I got other more important things to do. Before you know it, you get caught behind.

Now your maybe major version's behind and you can no longer upgrade your platform easily without major client impact. Because the other principle of this is, if you don't stay on top of this, you will impact your clients more. Ideally, you do this and your clients don't know you're upgrading. We all have hundreds, if not sometimes thousands of different changes ongoing in the environment at any given time, so we're constantly upgrading the environment. The client should ideally not know. Now, it depends on the software, depends on the platform, depends on if you have horizontal scaling and so on. Broadly speaking, the goal should be to do this without client impact. Now, where you do have client impact, client-side libraries, in a certain world where they have to change their API spec, SDK, and so on, then you have to have a contract with the client saying, I will only support n and n-1.

My software cycles, my platform cycles every, let's say, three months, I will support current and last version. I will reserve the right to have you upgrade your integration point at some point, because it's the only way I can stay current. The cloud providers do a really good job of this, of explaining this and being explicit about it. I think it's super important to do that, because otherwise you'll disappoint your clients. Think about it, when you upgrade, because you have to make it current. It is not as if the clients have any value driven typically from upgrading themselves. Meaning, they don't have incentive to do this, so you have to make it explicit that you need them to do it to maintain the platform at scale, and have that dialogue, and have that formal term set up front. I like to also use a moniker with the teams, something I call 0114, about managing this.

I'll explain what that stands for. When you manage a platform, how you should think about maintaining the currency of it. You should have zero people involved in maintaining currency, so it should be fully automated. You should be able to upgrade every single widget, so that you're maintaining thousands, hundreds of thousands, millions of widgets. You should be able to upgrade the entire fleet in less than a day, so that's a 1. You should be able to cycle at least every 14 days. Ideally more often, but at least every 14 days. If you set that parameter and say that's what good looks like, you have to be able to work towards that, it helps a team understand what they need to build and all the scaffolding that they need to put in place to maintain platforms to keep them evergreen. It's really important to set that expectation up front.

5. Avoid Undifferentiated Heavy Lifting

The fifth one here is to avoid undifferentiated heavy lifting. There are way too many times where engineers, myself included, want to chase a shiny thing. They want to go build something because it's attractive, because it's interesting, because it's intellectually stimulating, but not because it adds good value. I think what's very important is to be disciplined around that question of, is this necessary? Why are you doing this? What benefit does it bring? Has someone else done it already? Can you iterate on this other thing that someone else built? Rather than, go off and do it. Maybe some of you work for companies that have infinite resource and depths of pockets. I never worked for such a place. There's always been too much work, too little time. Applying this principle is really important to focus on. Using an example of what I have done. Database is a good example.

All the places I worked have database as a service, Postgres, for example. We wouldn't rewrite a new Postgres engine, perfectly great open-source Postgres engines out there. What we wrote was a control plane to do daily backup because that was required from the regulators. That was the only thing we wrote. How do you integrate the fact that every time someone deploys a Postgres database, it does a daily snapshot of the database to an object store so we have that snapshot and we can document it. Only what was necessary and no more, because the Postgres engine itself, we're grateful we needed just that feature. I think, again, it's really important to focus on the value to the business, the stakeholders, the clients, not in what is interesting from an engineering perspective.

6. Be Opinionated

Principle six, which ties into others, is be opinionated. It ties into the prior one. Back to the finite resources. We don't have enough time ever to do everything. Inherent in this is a couple of principles. You need to listen to your clients intently. I think we do a really poor job of this in enterprise IT. Those that work for customer product companies do a way better job because you have to, because if you don't do this well, customers don't pay your bills. They walk away. You're intuitively trained to this. Those of us in enterprise IT are not as intuitively trained because our customers are captive and they're stuck with us, like it or not. We don't listen as well as we should. That's super important. Listen to your clients and then deliver what has the most value. Deliver to the 80%, 70%, not the 100%.

It's impossible. The other part of this is to retire technical debt. I think this is where we tend to fall over the most is that boat anchor of version after version after version after feature after feature. We want to innovate and do the new things. If we don't shut old things off, we fail. Because again, finite money, finite ability to do things. If you don't shut anything off, you lose your ability to do new things. That retirement of technical debt is often not popular. You go to a client and say, I'm not going to do that thing anymore. The thing that you love, we're going to stop. Because no one else uses it anymore. Very hard conversation, but super important. Then the final part of this is, you're the experts at your platform, your clients aren't. That might sound counterintuitive to listen to your clients.

Yes, you listen to all of them, but, ultimately, you have to make the decision as to what you build on behalf of what gives the value to most clients, not all clients. That means that you have to be willing to say, this is what we're doing. You have to have an opinion. You have to say no. You have to shut things off. If you do all these things well, and it takes time, it takes backbone, it takes fortitude. It's probably one of the most difficult things to do for platform managers to manage this. Because everyone's desire is to help clients and do the right thing. This can be very difficult because you have to say no quite a bit to get this right.

7. Be Long Term Greedy

Number seven, which is, be long term greedy. Building platforms is a long-term commitment. Meaning, you were successful. Clients use what you built. They depend upon you. You can't take the platform away. You can't yank it away from them, because they depend upon it. Just like someone yanked this platform away from me, I'd fall hard and would hurt myself. Your clients get hurt when platforms get taken away, easily. You have to think about, these are long-term commitments you make, multi-year. They don't last weeks. They don't last months. They last years. Every platform I built has been years, even when I want to get rid of it after year one, it took me years to get rid of it. It's really important to be judicious about, should I really do this? Do I really want to put the feature teams around it? The maintenance, the cost, the long-term application driving this?

Or, maybe not. I like to use the analogy of puppies and dogs. Any of you that have kids, I'm sure most of you have had this conversation at one point with your kids around, "I'd love to have a puppy. My friend had a puppy. He's so cute. I really want a puppy." Then you remind your kid that this puppy, assuming that we bring this puppy in, will be with us for 10 years, 12 years. They're going to walk it. Are you going to feed it? Then they say, "Of course I will. I love the puppy." None of the people I know that adopted puppies, especially those during COVID with kids are having the kids still walk the dog. The parents are walking the dog. Of course, they are. The child doesn't understand that, but the parent does. This is no different. Like you have to be the parent and say, am I really ready to bring this dog in, this puppy in?

Because that's what is going to happen, you're going to have a dog in the house for 10 years, 15 years, I'm ok with that. If I already have five dogs, am I ready for another five? Maybe not. That's back to the technical debt from earlier, but this is incredibly important to think about. This comes with time and wisdom because I have built platform components that I really regretted after. Say, "I wish I really didn't do that," but I'm stuck with it. I'm stuck with maintaining it. I'm stuck with the clients on it. It takes up time. It takes up resource. It takes up cognitive bandwidth I'd rather use on other things, and I can't. I regret those decisions. Now I really think hard about if we should do these before we go off and do it because we know that it sticks with us for the long term.

8. Fail Quick, Fail Often

Number eight, which, again, might sound a bit counterintuitive to the prior one, but this is about, once you decided to build something, how you build. This is no different than with software. You should fail quick and you should fail often, because building platforms, you iterate upon them, and they're living objects, organic. You have to iterate your way there rather than thinking that you know beforehand. If you don't do that, you will make some big mistakes as well because you will wait too long to adopt something and it will take you in the wrong path. I'm going to use an example of what we did at one of the places I worked. In the mid-2000s, we got enamored with this thing called container constructs in Linux. We said that would be really powerful, especially in financial services, we could be able to partition customers on the same machine properly with namespaces and so on.

That was pretty attractive to us. We said, let's go do that. We ended up using the primitives, but without orchestration. There was no orchestration out there in the industry, so we built it ourselves. Then about a year later, we were actually out here in San Francisco, we visited Twitter, and they showed us Mesosphere. That's really cool. We can use their orchestration system. Let's go use that. We adopted Mesosphere. Then a year later, we spoke to Google, which I'd spoke to Google already, we knew about Borg and about Kubernetes. Kubernetes, when we spoke to them initially was a science experiment. Then very quickly that became mainstream. Then we shifted then from Mesosphere to Kubernetes. Those learnings allowed us then to get really good at container orchestration very early. If we had waited for Kubernetes, we would have been two, three years later with adopting containers than we were.

That iteration was fine. Then it maybe sounds counterintuitive because it's about building platforms and migrating, but we maintain our client fidelity throughout. Our clients are using containers. We help them migrate from our own orchestration. Their primitives was containers. It could actually be migrated relatively straightforward from our proprietary one to Mesosphere to Kubernetes. You have to think about then when you have to do these migrations, how you own that migration on behalf of your clients and don't leave them saddled with that because then it becomes really hard to move forward. If you think about this mindset of, I'm going to continue fail forward. Failure is not a bad word. Failure is a learning. I'm a big believer in trying things on a consistent basis. Take the lessons from it and continue to iterate your way forward.

9. Share Responsibility

Principle nine, which is on shared responsibility. As a platform owner, and I spoke about this a bit before, you have to be very clear on what you do and what you don't do, what you do well, what you won't do at all. As much as possible be formal about it, because the more formal you are, the less you're going to disappoint your clients. I'm talking about formality that is easy to understand. I'm not talking about 100-page EULAs that no one reads that is there written by lawyers to save your legal ass. These are well-understood contracts that you have with your customers. They could be contracts in software like APIs. They could be contracts in process. They could be contracts in SLOs, SLAs. Having well-defined boundaries that says, this is what I do. This is what you need to do. These are my SLOs. This is what you can expect from us.

Amazon do a really good job of this. If you go to this or just Google, or ChatGPT, whatever you use nowadays to do your search, but look at shared responsibility model, they have a really good definition of what they do and what they expect the customers to do. This has evolved over time, but that formality has been very helpful when you build on top of it because you understand what happens. For anyone who dealt with this outage, if you had read that shared responsibility model, you had followed the Well-Architected Framework, which is a different framework they have, to build resilient, that Amazon issue in NA east-1 would have been non-impactful. Very few of our banks were impacted. Why? Because they had built multi-region, they had built around it, they had understood what it took to build this, and were relatively non-impacted. It's important then to define that so customers understand what they're getting themselves into.

The other aspect is also, as a platform owner, as I said, you only have finite resource, finite time, finite ability. You can only vend ice cubes, not snowflakes. You cannot build artisanal products for every single customer. You have to think about building like Lego shapes, ice cubes, whatever analogy you want to use, but you can only build fixed shapes. Again, you have to make it clear to the customer, here are the shapes I can give you, these other shapes that you want, I cannot give you. You have to know that as you consume this. Make customer choices. Maybe I'll use this platform, it has the shapes I need, or maybe I won't, it doesn't have the shapes I need. That should be an explicit and eyes wide open decision.

10. Abstract, Don't Obfuscate

Principle 10, which is about abstract, don't obfuscate. When you build platforms, you can have different levels that customers interact with. They can interact with an API, SDK. They can interact with Infrastructure as Code, like Terraform. They can interact with a UI. Different customers need different interaction points, and different use cases require different complexity. It's not your job as a platform owner to determine which ones your customers should use. You should meet them where they are. At least at a big place like us, they're everywhere. They're all across that spectrum. A couple of things out of that. First of all, you have to have different abstraction layers. As I mentioned, for example, in our case, we tend to have, again, APIs. We have Infrastructure as Code, typically Terraform. We have UIs that can spin Terraform or can call the APIs directly. What we make sure of is that we also give customers insights into what happens.

We don't obfuscate it away. We try and expose the Terraform, the API calls, the environment. For two reasons, A, if something breaks, you need to know what happened under the hood, because if you don't know, you have no idea. If you hit a UI and try to spin up something, it didn't happen. There's no good telemetry, no observability, no underlying understanding what happened, that's not a great place to be as a client. Secondly, when you want to try and do something subtly different than the platform did. Let's say we vend a certain database pattern, we generate Terraform for it. You want to do something slightly different, now you can take our Terraform and change what you need yourself because we didn't obfuscate that away from you. It also allows customers to have a lot richer interaction with the platform than if you obfuscated it away.

You can imagine the extreme version of this is you could write machine code. You could write assembly. You could write Java. You could have Copilot generate your Java, or you could vibe code and have no clue what happens. Ultimately, it's machine code that runs at the bottom of this. These are all different levels of abstraction across the stack. I started my career writing machine code, which was awful. I never want to do that again, with a hex editor. Then I went to assembly, which was slightly less awful. These layers of abstraction are actually valuable, because I never want to do those things again. I'm very happy that I had a higher order of abstraction that I could start using, and in my case was C, not that high, but at least better than those. There is real value in this.

11. Stand on the Shoulders of Giants

Principle 11, which is about standing on the shoulders of giants. Really the metaphor for giants here is open source and open standards. There is no way that the cloud providers, myself as a platform builder, could have done what I've done, or they could do what they did without having open source, and all the open-source components available to build on top of. Because now you can focus on integrations, you can focus on your value add, but you don't need to focus on building every single bit across the stack. You get enormous mind share. You get great engineering scale, of course. You can stay current because there's consistently innovation happening. Also, it can be portable. Think about every cloud. Every single cloud provider offers a Postgres variant. They offer a Kubernetes variant. They offer a Kafka variant, and so on, which means you can port your applications because you're using the same underlying semantics.

The platforms are subtly different, but they're relatively consistent in some of these components. Hugely valuable. The quote from Isaac Newton, this is not a new thought, like science builds upon science and innovation upon innovation. "If I have seen further, it is by standing on the shoulders of giants." This was all about how his innovations, he didn't take credit for it. He said, the reason I was able to do what I did was because of the people before me that told me what they did, and so on. This principle is incredibly important. We started using Unix for a long time, but we started using Linux heavily in financial services about 20 years ago. Since then, I would say the vast majority of our critical systems are today built on top of open source. I'm sure the same holds true for most of you. It has really allowed us as a computing industry to innovate at a scale we would completely be unable to without this innovation. There's nothing wrong with closed source. I'm sure a lot of you write software that you sell. There's nothing wrong with that. It's that plus open source will have gotten us to where we are. I would say that generally speaking, platform building, open source is particularly important, because it allows you then to assemble upon those components.

12. Build Culture, the Rest Takes Care of Itself

Then the final principle, which is my favorite one, I didn't say any one was more important or less important, but this is my favorite one, which is about culture. If you build culture, the rest takes care of itself. Remember when I spoke about platforms? Platforms are many components, many teams involved in building those components, which means that if you're going to do so consistently, you have to have a culture that assembles those teams together in a way that allows them to build a platform that is assembled with all those components. Great culture builds great teams, and great teams build great products. If you get the culture wrong, you might get lucky occasionally and build a great product, but you won't get lucky consistently. The focus on culture is incredibly important. Where I spend most of my time is on culture. There are a couple of things that you should think about when you think about building culture.

First of all is empowerment. Empower your teams, your feature teams, your product teams, to make as many decisions as they can themselves. Give them freedom from choice. For example, you'll use this observability stack, you'll use OTel. You don't get to choose. Don't worry about that. That's choices made for you. You'll use this identity SDK. You're going to use it. You don't get to innovate there. What you need to innovate is in, if you're building a database, build the best database out there. Put all your innovation in there. There you have autonomy to make your decisions. Secondly, having teams that have diversity of thought. Every time I built monochromatic teams that all think and come from the same background, we've had worse outcomes than where we had teams with people of many different backgrounds. Because they challenge each other more naturally. They think about things in a different way.

We get better outcome. It might be more dynamic, to use that word kindly, more noisy, because you've got people that argue a bit more and that have quite different positions. It tends to build better outcomes over time. Managing that is really important. Then, thirdly, no different than a sports team, you have to be willing to manage the players in the team and know which players you need. Managing the dynamics of who's in the team and being intentional with it is important. You cannot be hands-off for that. Yes, the culture is really important, but teams don't fully self-select. Because if they do, then mostly you get teams where everyone are friends and they all go work together. You don't end up with diversity of thought. You don't end up with the right balance. I think it's very important to think about how you manage the teams to do this as well to get the best possible outcomes.

Key Insights

To bring this full circle, I spoke about platforms today. They are pervasive in our lives. I'm standing on one right now. We consume them all day long in our day-to-day life. We build our software and our platforms on top of other platforms. Think of this as a layer cake of platforms. Like I build my platforms on top of open source. They're built on top of other platforms and so on. It's also a circle layer cake. I probably should have illustrated this as some big space circle with platforms all on top of each other. You get the point. Everyone does this. We take the platform beneath us for granted. We don't want to think about it. It just works. If we do our job well, then the people that use us have the same experience. These principles are not meant to be exhaustive. These are the principles that I have derived from my work. I hope some of these resonate with you, and they cue some of this.

Questions and Answers

Participant 1: Often, platforms start with someone building for their own particular use case, and then it starts to expand to more and more people. There's a decision point, I feel sometimes like, should we invest in making this a platform? Any thoughts about how does that transition happen? How do you prioritize?

Matthew Liste: Most of the time platforms, to your point, ideas start in teams that build it for themselves. For example, a few years ago, at the place I worked, there were probably seven different teams doing Kafka in their own way, assembling it together. They'd added different stuff to Kafka to make it enterprise ready, ready for basic financial services. Somebody looked once at that, it's not particularly commercial having all these teams do it. None of them really want to do it for anyone else, want to do it for themselves. We decided to centralize that and create a common managed Kafka service for the enterprise. It often starts like that, where people have built internal software that becomes a platform. It depends on the culture of the company, is the best way to put it. If there's a pot of money to fund the people to build platforms for others, it's a lot easier.

When there is no money and everyone has their own budget, and it's constrained by I don't have money to support you, I only have money to support me, it becomes a lot harder. What we try and do is make that a bit flexible, because we want the right people to build. If you built a great platform, and you are willing to support the company to give you money, that's the best outcome. We at least try and find funds to let the team that wants to do it, and has a bound to do it without being too parochial about it. Often politics get in the way of that.

Participant 2: I have a two-part question. One, how do you balance simplicity and speed? You emphasized simplifying the architecture, simplifying things. Oftentimes the tradeoff, I can keep going, or I can invest in simplicity. The second part of it, you mentioned creating these paved roads, these central organizations that are funded that can create this meaningful way of doing Kafka right, for example. How do you champion the adoption of these initiatives, and how do you make sure that people are using the arc that you have built?

Matthew Liste: First of all, complexity versus speed. I think it's one of those it depends question, because sometimes you have to optimize for speed, because you're trying to get to market, and you're ok managing complexity after the fact. Typically, we do that. We opt for speed first, and then we have to slow down and remove complexity. Because as we're going quickly, we're leaving technical debt in our wake. It's everywhere. It builds up. If you don't then take a step back and start stripping it away, you end up with this behemoth that is unmanageable. That happens every time you build a platform, it starts with speed. You're racing to get clients. You're trying this out. You're putting it out into the marketplace. In our case, a virtual marketplace internally, but still seeing it get adoption. Then you have to take a step back at some point and say, now that we have adoption at scale, again, think about these three as, I have to scale this platform now.

I have a lot of clients on it. I now have to start stripping away complexity. That's a hard place to be too, because I have to tell people, I'm going to slow down feature velocity, because I'm going to start focusing on removing technical debt. I've had to do that every time. No one likes it. It's a real important thing to do, because otherwise I cannot keep running. It's like going through the water, I'm just building more barnacles on my boat. At some point, that boat can no longer move, because it's so encrusted. You become encrusted with technical debt. Certainly, on that one as well just the central teams will decentralize this. You can imagine any bank or place I work, I have a lot of highly opinionated development teams who think they know better how to do the job that I'm doing than I do.

They will tell me, why should I use your crap? I can do better, faster, cheaper. There's a bit of like carrot and stick. The carrot is, if you use ours, it will be less work for you. You can spend your time innovating where it's really important. You have to do that. Typically, over time, people come around and say, I no longer want to maintain my own version of Postgres. I don't want to write my own Linux distro. I don't want to write my own database. We've had that. We've had people even write their own compilers. They've done everything. At some point, it comes around to writing my own compiler might not be the best way to spend my money building better trading algorithms, and so on. At some point, people come to that. The other thing, the stick is, in a bank, there's such a thing as regulators, and satisfying them.

At some point, you have to say, you have no choice here. This thing you're doing is dangerous, is going to get us into trouble. Therefore, there's certain rules that you cannot violate. We often tell people, listen, if you want to build a highly compliant system, knock yourself out. Most of the time, I don't want to deal with that. It's just too much work, too much headache. You do that for me. I think that's a bit of the balance we try and achieve.

Participant 3: This is another balance question, because it seems like there's a lot of this. In terms of when you've talked about deprecating things, and especially getting rid of technical debt, it makes me wonder how you do that with clients. Because in my experience, the clients who get the most invested, if you really want serious customers who are using your platform heavily, then they invest in it to such a degree that when you then need to deprecate something or change something, it takes multiple years for them to get off it. Do you have any strategies for helping stay agile in your development, or any strategies for how you manage clients and help them to move or encourage them to move?

Matthew Liste: It's probably the biggest challenge, is, how do you help clients get off the thing that you built when you have to do a forklift upgrade? First of all is, ideally, you avoid ever having to do that to begin with. Because if you do that, it's really an antipattern. Think hard about, if ever I have to make my client shut the thing off and migrate, I failed, because I should have never had to. Again, realistically, it's happened. I've done it. It does happen. Then you have to just think about, how do you make it as easy as possible to reduce that bar? How do you send as much as possible? It's super hard, because again, it's not in their interest. Unless there's a feature they need, a value add, there is no real reason for them to want to move. Sometimes you have to beg, you have to use goodwill, you have to bribe them.

Internally, it's a little easier. We can put money aside. Internal strategy with enterprises, we sometimes put aside a pot of money that we give to them and say, here's money to help your developers do that, so it won't take away from your future velocity what you're building. You can get from our money, we'll pay you to do this, because you recognize you don't want to do it. There's no value to your business by doing it, so we're going to help fund it. Clearly, it's more difficult when you have outside customers, but you can think about other mechanisms to ease the pain a bit.

Participant 4: You want to strive for the ice cubes, but then you got these snowflakes that you built, so then, how do you get rid of those snowflakes?

Matthew Liste: Just to state the point even more specifically, your incentives are not aligned with the client's incentives, usually not as a platform provider. Their incentives, once they use your stuff and it works and they don't need anything new, they just stay on it, never talk to you again. It just works. Your incentive, of course, is to maintain the platform on behalf of all clients, which does means occasionally you need to shut things off. Your incentive is often not aligned and you have to find the best way to manage it. Again, different techniques, but it is one of the hardest things I do. That's why you want a finite number of dogs in your house. Because if you have too many of them, it becomes unmanageable. Now you have a zoo. Often, that feels like my everyday managing a zoo, but hopefully it's well structured. That's a really important thing for all of you to think through.

 

See more presentations with transcripts

 

Recorded at:

Software is changing the world. QCon San Francisco empowers software development by facilitating the spread of knowledge and innovation in the developer community. A practitioner-driven conference, QCon is designed for technical team leads, architects, engineering directors, and project managers who influence innovation in their teams.

Oct 06, 2026

BT