BT

Facilitating the Spread of Knowledge and Innovation in Professional Software Development

Write for InfoQ

Topics

Choose your language

InfoQ Homepage Presentations The Five Stages of AI Maturity in Engineering Organizations - Where and Why Teams Get Stuck

The Five Stages of AI Maturity in Engineering Organizations - Where and Why Teams Get Stuck

38:54

Summary

Quotient CEO Lizzie Matusov explains why soaring AI spend often fails to improve software delivery. She presents a research-backed AI maturity framework designed to help engineering leaders move beyond vanity metrics like token usage, align organizational AI adoption, and address critical bottlenecks across the software development life cycle to deliver measurable business outcomes.

Bio

Lizzie Matusov is the co-founder and CEO of Quotient. She also co-authors Research-Driven Engineering Leadership, a newsletter that translates academic research into practical insights for engineering leaders. Before founding Quotient, she worked in engineering roles at Red Hat and Invitae.

About the conference

QCon AI is a practitioner-led event focused entirely on the engineering discipline required to scale these workloads safely. It provides direct access to the architectural playbooks and failure metrics that peer organizations use in production.

Transcript

Lizzie Matusov: My name is Lizzie Matusov. I am the co-founder and CEO of Quotient. We spend a lot of our time studying bottlenecks in software engineering, especially as those bottlenecks change with AI. I think one truth that all of us have felt is that innovation is moving very fast. If you're feeling like this movement is almost unprecedented in its nature, know that you're not alone in that feeling, because AI spend is moving faster than really anybody forecasted. Just in 2026, we've seen a 47% year-over-year growth in global AI spending. We are now projected to spend $2.5 trillion on AI, and 72% of enterprises are running at least one AI workload in production. That pace is exceptional. I think it goes to show that at this point, AI is a necessary part of how we build software. The thing is, when AI spend is moving faster than anyone has forecasted, that also means that we're blowing past all of our enterprise's expectations of that spend.

Let's take Uber, for example. They recently announced that they had burned through their entire 2026 AI budget in just four months. That's one-third of the amount of time. Of course, monthly adoptions skyrocketed, especially moving from December over to March. Curiously, when asked about the return, the CTO said that, frankly, they could not find a stable equivalent relationship between the expenditure and the productivity output. Even more recently, Microsoft just announced that they're going to be dropping Claude Code by the end of this month due to the budget and the way that they blew right past it. Again, if you read the transcripts, they're not really able to tell you their ability to say what is the ROI for the AI use. The thing is, we know we're getting value out of AI. All of us here can talk about the way that we've become more efficient, how it's become indispensable in the way that we develop software. Then, why is it that individual usage is not translating to organizational outcomes?

Theory of Constraints

My goal is to answer that question for us today. I think to help illustrate this problem, I want to help us think about it through a different lens. Let's take a plant factory. The basic concept of a plant factory is pretty straightforward. Raw materials go in, it moves through a series of steps in production until it hits the final product. There's this famous story in the "Nineteen Eighty-Four" book called The Goal, where there's a plant that's in trouble, and the factory manager really wants to find a way to save the plant, so she comes up with an idea to maximize efficiency at every station. Her idea is, run every machine at capacity. We hope that this might increase production, will keep everyone busy, and that will save the plant. What happens is things don't get faster. In fact, things get slower. They start to break.

What ends up happening is that we've discovered there's a bottleneck. The bottleneck in one of their processes was the rate limiting step. Even though they tried increasing throughput, they couldn't make it past the limitations of that single bottleneck. This is the theory of constraints. Every system is limited by a single bottleneck. If you're trying to speed up everything else, it's not going to increase the output if you don't actually address that bottleneck. Three things will happen. One, your work in progress inventory will explode, oftentimes piling up right behind that bottleneck. Two, your lead times get longer, so things will enter into the system and take longer to get through to the final product. Three, quality problems will compound. Now you're increasing strain on the system and different steps of the process.

If this feels a little bit familiar, it should, because it's also how the software development life cycle works. We've designed the software development life cycle to go through numerous gates to make sure that we're developing high quality software to the customers that we serve. We make sure that code goes through testing, reviewing, different deployment gates, and production support to make sure that we deliver excellent software. The theory of constraints holds here, too. Let's say that you have a bottleneck in the code review process. Then just increasing the number of tokens we spend or code that goes through the system might not necessarily translate into higher throughput if we're not also addressing the areas where the software gets stuck. The thing is that bottlenecks can look different for every organization. I'm sure a ton of us are experiencing challenges in the code review process, but it could be something else.

It could be how we test. It could be how we deploy. It could be any number of things. Even across different teams within the same organization, your bottleneck might be different. Actually, the DORA report in 2025 also found this. They did a study of about 5,000 technology professionals, looking at how they're leveraging AI. They found that while individual effectiveness went up at the top, software delivery throughput did not meaningfully change. In their analysis, they posit that this is because the systems surrounding our individual usage of AI has bottlenecks, and that we haven't properly addressed those to help make sure we're getting end-to-end gains.

Roadmap

I think I've sufficiently convinced you that the gap does exist. Now the question is, how do we solve for it? In this talk, I'm going to convince you of two things. I want you to think about how the most effective engineering organizations do two things. First thing they do is thoughtfully improve AI usage across the software development life cycle. The second thing is that they resolve the bottlenecks that limit their outcomes. We'll talk about these a good bit. Let me give you a brief roadmap for our discussion. The first thing that we're going to do is I'll introduce the AI maturity model. This is a structured, research-backed framework for how engineering organizations advance their AI usage. All of you are going to recognize your team somewhere along this maturity curve. Second thing I'll do is put the framework into practice and walk you through how do you identify your stage and set the right near-term goal for your organization. Then the third thing we'll do is measure outcomes. We'll talk a little bit about how we move away from looking at just usage and look more towards how we measure outcomes so that we can identify those gaps and the bottlenecks that might exist.

The Five Stages of AI Maturity

The five stages of AI maturity is a framework for how organizations advance their AI usage thoughtfully. Before I get into the details, I do want to spend just a moment talking about our model design. The reason why is because there is a lot of guidance and best practices, and they are changing very quickly. Oftentimes, I encourage people to think about the incentives of who's generating that guidance. To give an example, if Jensen Huang is saying that he would be deeply alarmed if his $500k engineer didn't consume at least $250k of tokens, you might want to ask yourself, why is Jensen Huang giving that advice? The reason is, his company is going to do exceptionally well if everybody takes that advice. We might see in a few years that it actually did make sense for each engineer to spend half of their engineering payroll on tokens.

I would argue that today, that's likely not the case. We worked and built a model using a first principles approach that would allow us to minimize bias and focus on the goal of improving engineering outcomes. What did we do? We first did a synthesis of about 15 different research papers coming from universities, research arms of organizations like Microsoft and Google, and independent but peer validated research. We looked at papers focused on AI's impact on developer productivity. We also were careful not to go back too far. Anything more than six months feels a little bit outdated, especially with the inflection point that came at the beginning of this year. Second thing we did is we did researcher interviews. We worked with the researchers who designed many of the productivity frameworks you maybe heard of, like DORA and SPACE, to get their input and also to see what are they thinking about today, and where are their research directions taking them.

The third thing we did was field discovery. We spoke to over 100 organizations that are prioritizing AI as a top goal for their organization this year. As you can imagine, there was no shortage of companies to talk to and collect data from. The model describes five stages. Moving from stage one, which is ad hoc experimentation, all the way through to stage five, which is end-to-end autonomy. As you're thinking about each of these, I'll give you a bit of a guide. Each of those five stages is defined by six characteristics. First one is enablement. That's learning and skill growth. Policy and governance, so organizational guidance and guardrails. Validation and testing. What is the quality of AI-generated work? Embedding and workflows. How well is AI integrated into how we work? Workflow automation, which is the triggering of deployment workflows. Finally, data context and access. The internal data that's available to AI systems. Five stages, six characteristics. Let's get into the details.

Stage one is what we would call ad hoc adoption. This is the earliest stage of AI maturity, where usage is really driven by our individual experimentation without any formal guidance. Think of this as the window of time when AI tools were first hitting the market. We were all just trying out different things to see what worked. We would paste in whatever context into chat windows or maybe early usages of Copilot, and really just trying to figure out, what is this tool? As far as enablement, again, that individual experimentation is what's guiding enablement. For policy and governance, there's really nothing there because we don't yet know how AI is going to impact our organization. As far as validation and testing, it's very informal. Developers are self-reviewing. They're getting outputs from chats and looking at that and saying, LGTM, throwing it into their system, and then leveraging that.

As far as embedding and workflows, AI is living really outside of those normal workflows. For workflow automation, it's manually triggered. As far as data context and access, AI can't see anything other than what you are providing within maybe a context window. All of us started here. Maybe for a small handful of people, your organization might still be here, especially if you're in a very regulated environment or you have very secure systems. For most of us, we then moved on to stage two. Stage two is what we would call assisted development. AI is beginning to spread across teams as the org is enabling that early adoption. There's likely some amount of tool sprawl as we're figuring out, do we try this tool or that tool? Which workflow works better for us? Engineers are using AI to assist them with their own personal tasks. As I'm writing code or I'm writing documentation, I'm leveraging AI to help me with my personal tasks.

I think the way stage two feels is that engineers are really empowered to use AI and support them, but it's really individually focused. We also see the bifurcation of usage start to happen in stage two, where you have some engineers, probably the folks that we would call early adopters, who are really leveraging AI, pushing boundaries, figuring out the right use cases, and working through it if it isn't perfect the first time. Then you also start to see the emergence of skeptics who try it once, twice, three times, and say, "It's not for me. We'll try something else." In enablement, you see informal learning and experimentation is really driving how we work. There's early emergence of best practices being formed, but again, we're still trying to figure out, how is this best suited for our organization? As far as policy and guidance, there's a basic policy or maybe a limited rollout, but overall, we still have it in a draft state because we're refining the way that we use it.

For validation and testing, we've got basic checklists for AI generated work, but for the most part, things haven't really changed. For embedding and workflows, developers will manually trigger AI to assist them. Again, workflow automation is pretty limited. It's helping us with our individual tasks, but it's mostly manual when we're asking for AI. Again, data context, mostly user supplied and a little bit limited. I think what's important about stage two is that the harnesses or the systems around the code that we write mostly haven't changed or adjusted for the increased volume that AI is going to generate across the SDLC. Once we do start thinking about those harnesses, that's where we move to stage three.

Stage three is what we would call standardized workflows. This is where AI is becoming more embedded in development workflows across the teams. It's supported by shared organizational best practices. I think the key word for this stage is the word shared. Teams are starting to adopt consistent expectations. You might still see some variance from team to team. Overall, there's evangelist teams who have figured things out that are now sharing these best practices to the other teams. We're starting to see real standards form across the whole engineering organization. What's important about stage three is that the shift really starts to happen from the individual to the team and the systems around how those teams work. What does this look like across our characteristics? As far as enablement, you start to see formal training really come to light. If you've done the hackathons, the knowledge transfer sessions, the learning groups, those are really key traits of stage three.

For policy and governance, there is a clear or wide policy. Engineers have a good understanding of where they can and can't use AI. It helps guide their own autonomy and exploring. For validation and testing, you'll see that automated checks will start to validate AI generated code. The systems are really starting to adapt to, again, the increased volume. Embedding and workflows, you'll start to see that AI is embedded with some consistent expectations. It's not just supporting the individual. It might be part of the development process, or linting, or the code review process. It's not something that always has to be manually triggered. In workflow automation, you'll start to see that it does get auto-triggered for defined tasks. Finally, for data context, it can now start to access parts of code and docs beyond just what you supply to it. There's a pretty big jump from stage two to stage three, because this is where we're starting to focus on the development harnesses around our code. That's a lot of investment that starts to happen, especially as AI is starting to become automated across the software development life cycle. Stage three is where those bottlenecks start to become very visible. As you start thinking about those bottlenecks, that's really where you want to unblock them to get to the next stage.

I am many minutes into my talk, and I did not say the word agent once. That is because it is really important to think about the infrastructure from stages one to three, which will enable you to get to stage four, which is what we would call supervised automation. In stage four, agents really start coming into the scene. To start, engineers are configuring agents to execute bounded tasks across the software development life cycle. They'll handle lower complexity tasks with high oversight. A great example of a stage four automation would be if you've configured an agent to pick up a P3 support ticket, look at it, maybe generate the PR. You still have a human that's reviewing it, making sure the tests pass, and then ultimately deploying it. In order to be confidently getting the benefits of stage four, you want to focus on the infrastructure that you've built in stages one through three to help support you, because you're handing off control to agents, which as we've seen, it is extremely powerful.

It's also pretty risky. When we think about enablement, AI literacy is widespread. This is really a time where you've got lots of engineers thinking about different ways to solve problems and automate our tasks. Policy and governance is not only defined, but it is now aligned with the company's broader strategy. We're not just using AI for AI's sake. We understand how leveraging AI allows us to support our business goals. For validation and testing, you've got dedicated agents that can validate AI work before release. Heard some great strategies here, too. For embedding and workflows, agents are now embedded in delivery workflows. Again, they'll be quite bounded. There isn't necessarily a limit on where they can exist within the SDLC, and we're starting to play around with different areas. For workflow automation, we find that AI is now executing bounded tasks, but the human is still approving. Then for data context and access, AI retrieves relevant context automatically. I think that the most important thing to call out with stage four is that you really start to see the control of the software development life cycle shift away from the engineers to the agents. Especially as engineers are building more trust with agents, it's becoming more pervasive throughout the SDLC.

As that trust and autonomy of agents grow, that's what allows us to eventually get to stage five. This is what we would call end-to-end autonomy. AI orchestrates end-to-end multi-system, multi-agent workflows across engineering systems. Humans are now focused on oversight, high leverage decisions, and managing exceptions. This is that truly AI native feeling. The way that we describe that feeling of stage five and how you know you might have it is, think about the ownership and autonomy that each of you has as an engineer to develop software, test it, approve PRs, merge production incidents. When you've given agents the exact amount of ownership and autonomy you'd give to a senior engineer, that's when you've really hit stage five. The thing is, it is at the very tip of innovation, which means that the pace of change is really quite high. When we think about the characteristics, the theme here is continuous iteration.

Enablement looks like continuous learning that is embedded into the culture. If you're not constantly learning about new ways to use agents and the way that they can support your systems, you might actually miss the opportunity to better secure your software or better leverage those gains. Policy and governance. You're now seeing organizations continuously evolve their policies as systems evolve. Validation and testing. There's continuous production-linked feedback loops. Embedding and workflows. You're seeing that AI is now coordinating complex workflows end-to-end with autonomous execution. In order for it to have the same amount of ownership and autonomy that we do as engineers, there's a structured org knowledge graph that allows it to basically access anything that an engineer could.

How to Identify Your Stage, and Set the Right Goal

I want to talk a little bit about how to actually identify your stage. I'm sure as we went through these stages, you could feel yourself pulling in the direction of a stage, or maybe there's one characteristic where you're like, we're definitely here. In a different stage, you're like, but we're definitely here. I want to talk a little bit about how you can evaluate what stage your team is at and how that compares to your own company. There are three steps to diagnose your stage. The first one you want to do is score each capability. For each of those six capabilities, you'll want to identify the stage description that best describes your team. I'll show you an example, but just to finish that thought, your organization's maturity might be a little bit different than your own team's maturity. The reason why is because, again, we're in a room of people who are listening to and trying to learn the most cutting-edge information about AI.

Your teams might be a little bit ahead. Your organization's AI maturity is representative of how an average team would respond. Let me show you what that looks like. Here's a really basic table that shows the different stages and capabilities. Let's say you're going through this. You're like, let's start with enablement. We've got formal AI learning. I would say we've still got a little bit of bifurcation of usage, but we've got those hackathons, those knowledge transfer sessions. We've got that. For policy and governance, we have pretty clear expectations. I wouldn't say they're maybe aligned to our company's strategy, but I know when to and not to use AI. For validation and testing, we have basic checklists. If I think about it, we haven't quite yet designed the testing harnesses around our software to be representative of the increased volume we're seeing with AI. For embedding and workflows, we are starting to play around with agents.

They're quite bounded across the SDLC with a lot of oversight. That's ok. Again, we're manually approving the things that our agents do. Then for data context and access, maybe we work in a more regulated environment, and our IT teams have been a little uneasy here, so we've limited the access to be partial. If this looks like your organization, you would be at stage three. Again, there's parts that are at stage four and parts that are at stage two. Overall, you would rate yourself as a stage three. Again, you could do the same thing for the average team in your company. That's what would give you your company's stage. The second thing is to resolve your weakest capabilities. You want to choose the one to two characteristics where your stage is the lowest, and make investments to improve them. If we're looking at this same graph, we're much more interested in improving our validation and testing characteristics than we are about expanding our embedding and workflows and our workflow automation.

This is really important because when those gaps start to form, that is a perfect place for bottlenecks to live. Also, as you move into stages four and five, it becomes really important for you to have that necessary foundation set so that you can give agents the autonomy and ownership that we're all hoping for. The final thing is to pick that next-stage unlock. Find the one or two capabilities that if leveled up would really move your organization up to the next stage. I'd venture to say that most folks are thinking about those next stage unlocks. I don't have to tell you guys how to do that one.

What is the right stage for your org? I think if we looked at the headlines, we would see that every organization is moving to level five, end-to-end autonomy within the next two months. I think we can do it. We're in a room of the folks that are actually building these systems. We know that the truth is a little bit more nuanced than that. In our work with dozens of companies, we actually found that the average big enterprise is somewhere between stage two and three. The average growth stage organization is somewhere between stage three and four. This makes sense. If you think about a big enterprise, they've got a lot of things. They've got increased complexity. They have more people, more processes, and oftentimes more regulation, especially if they're publicly traded or have some certifications like FedRAMP. If we talk about a growth stage company, they tend to move a little bit quicker.

They have a little bit lower process, a little bit lower complexity code. This is why we start to see that shift happen. What's interesting is that we don't often find people at the polar opposite poles of stage one and then also stage five. That's because the risk is not monotonic in maturity. When we first designed this model, we were very nervous about calling it a stage model because we thought that everybody would just go straight from one to five as fast as they can get there. Actually, on both ends, there's risk to consider. At stage one, there's two issues. One, there is a shadow IT risk. You're pasting company context into a chat window that lives outside what the company knows about. That can get a little bit dicey at times. There's also an innovation risk. Most organizations are at or well past stage two.

If you're seeing this and thinking, my company is at a stage one, I would consider the competitive risk of folks in your industry that can move much faster. On the other side, stage five is pretty risky as well. The reason why is because agents should be able to do everything that humans do, and probably at a rate that's much faster than we were doing it before. If you're going to relinquish your control to agents in this manner, you better be confident that the systems in place are going to actually maintain the quality and integrity of the software you build so you can maintain trust to your customers. Because humans can't manually support this load. We're probably all doing this right now and finding that there's a lot of bottlenecks that we're struggling through. That's why for most organizations we recommend targeting somewhere between stages three and four, depending on where you are right now. It's a good 2026 goal.

Measuring Outcomes, Not Usage

We've talked a little bit about the model. We've also talked a little bit about how to set your own goals within the model. I want to bring us back to our original question. How can we bridge that gap between individual effectiveness and organizational delivery? How can we measure and see it? I want to start with a brief story, if you'll humor me. There's an old adage, some of you may have heard of this, of a police officer who's on night patrol and he finds a guy near a streetlight on his hands and knees digging around. The officer comes up and says, "Sir, what are you doing?" The man says, "I lost my keys." He goes, "Quiet night. I'll come help out." They spend a few minutes digging around on their hands and knees in front of the streetlight. The officer goes, "I don't really see it here.

Are you sure you dropped your keys here?" The man says, "No, I dropped them over there." He points to a dark alleyway. The officer goes, "What are we doing here?" "This is where the light is." This is the streetlight effect. We search for answers where it's easy to find, not necessarily where we know the truth lies. We might as well start here, because, again, that's where the light is. Even if we know that it's not the best place to look and it's not the likely answer, it's at least a good starting point. That's always been the story with developer productivity and measuring outcomes. Long ago, we used lines of code because it was an excellent proxy for how much work is being done in engineering. Today, we're using tokens. Measuring tokens are easy because they are right there. You can spin up a dashboard and quickly see who is spending tokens.

It can be a directional signal to show how increasing usage is happening across the company. It is most certainly not the best way to measure outcomes. I would argue that token maxing or using tokens as a proxy for AI's impact on productivity is a much more dangerous measure than lines of code. Let me tell you why. Lines of code had an indirect but very real cost on the software development life cycle because of the higher complexities on the system, the technical debt, and the cost of managing it. The thing is that each token has a literal dollar value associated to it. When you're incentivizing that behavior, it's going to translate to money being spent right there. To show another example of why tokens are quite difficult to look at as a measure of productivity, recent study at Concordia University looked at token usage across 30 development tasks running with agents.

They found some really interesting things. Like for example, code reviews took up about 60% of token usage. That was due to the back and forth between developers and agents as they were tweaking. Then, tokens were disproportionately spent on input rather than output or reasoning. Researchers again pointed to the communication tax where a lot of our tokens are actually spent on communicating context rather than generating output. We can see how tokens provide a pretty weak signal of outcomes.

There's another principle at play here, too. I know we all know this one. It's called Goodhart's law. "When a measure becomes a target, it ceases to become a good measure." We see this all the time in the world of productivity where when it's defined by a single metric, you know that we'll get really good at optimizing that single metric to look good. I'm going to give you an amazing example of this today in the world of tokens. In a recent Substack called the Pragmatic Engineer, the author interviewed engineers at different companies who have token leaderboards. He asked them what they think. Here's the response he got from one senior engineer at a well-known FAANG company. "We have internal dashboards and metrics tracking AI usage, token usage, percentage of code written by AI, and hand-written code. I am conscious of not wanting to be seen as 'uses too little AI.' I'm not ashamed to say that I do token maxing here.

Things I do to inflate my token usage metrics. One, I ask AI questions about the code already in the documentation. The AI pulls up the documentation, processes it, and gives me the results ten times slower, all while burning lots of tokens. I could use an internal product, but then my token numbers would be lower. Two, I ask the AI to prototype a feature I have no interest and no intention of working on. Prompt it a few more times, and then throw the whole thing away. Finally, I default to always using the agent even when I know I could do the work by hand so much faster. Then, I watch it fail." It's a pretty tough example, but I think the industry is starting to catch on to the fact that tokens really aren't the best way of measuring productivity. In fact, so much so that over the weekend, Amazon shut down its token leaderboard.

In their quotes, the internal memo, the Amazon senior VP said, "Please don't use AI just for the sake of using AI. Use it to help you solve customer problems, to help you solve business problems." We still gravitate towards activity metrics. The reason why is because, like we said with the streetlight effect, it is just easy to measure. It's probably the first thing we look at. We all know that this is not a good signal of our outcomes. As we increase our AI maturity, we need to start looking at the outcomes it has on our software delivery. This will tell you if AI usage is actually translating to the end result that we want, because at the end of the day, our goal as engineers is to write software that delivers value for our customers. If we can do that faster, that's great. We also want to maintain the bar that we set so that we can achieve our business goals.

I want to show you a couple of different frameworks for measuring outcome signals. The key to this is to manage tension. Again, we want to avoid Goodhart's law by making sure that we're incentivizing the right behaviors here. The first one is very simple: speed, quality, ease. This is actually the framework that the Google research team uses to understand productivity. It powers a lot of their productivity research. It does a good job of balancing that tension of software development that we're always thinking about. At its core, speed is about getting value to customers as quickly as possible. Quality is about maintaining excellence to build and retain customer trust. Ease is about making the product development process efficient, friction free, and predictable. This is one very simplistic one that you can use. The other one, which I'm sure some of you have heard of before, is the SPACE framework.

This is a very popular framework for measuring developer productivity. It was designed by some of the most famous researchers in the productivity space, like Nicole Forsgren from Accelerate, and the Microsoft Research team. They designed this to help find friction in the software development process. Unlike the DORA 4 or 5, that really focuses on software stability versus the SPACE framework really focuses on team friction and outcomes. The five dimensions are satisfaction and well-being, performance, activity, communication and collaboration, and efficiency and flow. Recently, the researchers came together for a summit. Even though this was published in 2021 before AI exploded, with all of the research that these researchers are doing, they still came back to the fact that a holistic outcomes-based and tension-oriented framework is the best way to understand AI's impact. They all recommended the SPACE framework again. To tie things back, when usage grows but delivery outcomes aren't improving, this is what can signal new bottlenecks are happening.

These productivity frameworks can help you diagnose where those bottlenecks might be so you can lift them. This is a continuous process. You'll find as you advance in one characteristic of AI maturity, that you might identify a different bottleneck in the system. That's ok. That's a normal part of improving our software processes to manage the increased volume and ownership that we're giving to AI.

Putting This All into Practice

We have covered a lot of grounds today. I want to give you some tools to put this into practice. First thing we talked about is the maturity model. Each stage and organization is going to fall somewhere in the five stages of AI maturity. We talked a little bit about how stage one is probably not super common and definitely not for most folks. Also stage five is not necessarily the best near-term goal for your whole organization. We recommend somewhere between standardized workflows and supervised automation. Each stage is defined by these six characteristics, through enablement, policy and governance, validation and testing, embedding and workflows, workflow automation, and data context and access.

Actionable Takeaways

Four actions that I want you to take with you. One, think about how you would diagnose your stage. I've got some links at the end that will make this super easy for you, but you can use the maturity model to identify which stage your organization falls in, where your team falls. Again, what characteristic you might want to invest in to improve, and then what your next stage unlock is going to be. Second thing is set that near-term goal for stage three to four. Ask your team or your organization, what would we need to do or how would we need to solve bottlenecks so that we can actually unlock that next stage of maturity? Measure impact, not activity. We know that tokens and seats are easy to count, but outcomes are what prove that we've increased productivity. Especially as the cost of tokens and the cost of leveraging AI goes up, it becomes really important to understand what is the outcome of all of that work.

Number four, watch the gap. When usage is not translating to outcomes, when you see your token spend is going up faster than you could imagine, and you look at your outcome metrics and say, "I don't see what's happening here. We're not getting that end-to-end delivery," that is your signal that there are bottlenecks in the system. Again, the most effective organizations do two things. One, they thoughtfully improve AI usage across the software development life cycle. Two, they resolve the bottlenecks that limit their outcomes.

Resources

If you're feeling very inspired to take the next step, this is one QR code that will give you access to three different things. There's an AI maturity model white paper that's an easy shareable that you can send to your team or your colleagues or your boss to help talk about this and form a language around it. There's a really quick 12-question assessment that you can use to actually identify what stage is my team on without having to run through the chart and poking dots. Also, there's a link to the latest research. Our team publishes every week a three-minute digest of the latest research that's happened in developer productivity. I will tell you sometimes a paper from three months from now will negate something we read about two weeks ago. It's a really great way to stay at the cutting edge.

 

See more presentations with transcripts

 

Recorded at:

Aug 04, 2026

BT