BT

Facilitating the Spread of Knowledge and Innovation in Professional Software Development

Write for InfoQ

Topics

Choose your language

InfoQ Homepage Presentations Context Is the New Code

Context Is the New Code

50:23

Summary

Patrick Debois discusses how to manage, evaluate, distribute, and observe context using proven software engineering practices. He shares how treating context like code - complete with testing, CI/CD, package managers, and security scanning - enables engineering leaders to reliably scale AI coding agents, maintain control over non-deterministic outputs, and build long-term organizational knowledge.

Bio

Patrick Debois is a practitioner and researcher exploring how AI agents reshape software development for teams. As Product DevRel lead at Tessl, he studies AI-native dev patterns, context engineering, and organizational knowledge loops. Organizer of AI Native DevCon and pioneer from DevOps to AI-native dev, he focuses on one question: how do teams get better, together?

About the conference

Software is changing the world. QCon London empowers software development by facilitating the spread of knowledge and innovation in the developer community. A practitioner-driven conference, QCon is designed for technical team leads, architects, engineering directors, and project managers who influence innovation in their teams.

Transcript

Patrick Debois: Context is code. Who believes that is true? There's two ways that I've seen this manifest. One is the like very vibe coding, Andrej Karpathy who says, we're just giving it prompts. We're giving it context. It's writing the code for us. Definitely context is code there. The second piece is something we experienced in our own company. We were writing a piece about onboarding people to our tool. You cannot imagine the discussions we had and then the number of pieces of code we had to write for all the edge cases of doing the onboarding. Then in the end, we said, let's write a skill in natural language, give it to the agent, it will cater to whatever the person in front of it needs to have, all the scenarios. So much code compressed to a couple of words. That's also code. It's not just building it, it's like a piece of the product as well.

I personally love to think in parallels, and in 2009, I was thinking like, what if ops would be more like dev? We had DevOps. This talk is about what if context is like code. That's the mindset. The Software Development Life Cycle is the Context Development Life Cycle. I'm not talking about context engineering within your coding agent, but think of it as context. How do we deploy? How do we manage that across our whole setup? We remember the infinity loop, nothing new there. I've put a few different words on that. We're generating context. We're evaluating whether that context is good. We're distributing that to team members, the organization to agents, and then we'll observe whether that actually is working, yes or no, and based on the feedback we get, we improve or generate and come back. Simple loop, same as we try to do with coding. I'm going to walk you through it in the parallels between those four things, and in the end, we'll have a bonus, the context flywheel.

Generate - Create and Curate Content

First part, generate. I guess you've all been generating. You're the human context generator. You're typing into your Claude Code. You're prompting. You're giving it context. No, it's like this. No, please listen to me. Here's my CLAUDE.md. This is what I want. The humans as a context engine. It doesn't scale well. Context engineering became a thing, what we put in the context window of the LLM, and we're deciding what to do. I'm not going to talk about context engineering, because that's typically about what the agent will pull in the context, yes or no. That's not the talk. I'm thinking about all the context that we need to prepare that goes in that thing. It's a different cycle. It's not context engineering, but it's more the context as an artifact. People have been putting rules, instructions. Luckily, we now have a little bit of a standard. It's called AGENTS.md that we can reuse across.

Funny thing is, Claude exists on being CLAUDE.md still. All the rest of the industry has caught on. There's a whole PR on this. There is a little bit of standardization happening, which is good because now it is something that we can reuse across different pieces. There's other pieces of context, which is about your code and your docs and libraries that you're using. They have docs. They have code. These tools, like Context7 and Ref, they look at your code dependencies and say, I'm using the Express library. Let me pull in the context that is really good at coding with that library. That's another piece of context that people have been catering for. Then you can say, I'm on version 1, 2, 3 of that library, and it pulls in the right context. They're producing this for you to consume. It's not something you write. You could write it from your own libraries. Maybe you do that for your microservices. Microservices team 1, they publish the context, how to use the microservice, and then another one can use that part.

Then you have connectors. What's weird here, like one of the vendors was unblocked. They have an MCP agent, they look on Slack, they use all your sources that you typically, as an engineer and as a person, will look for context. Where was that decided? Where was that chatted about? Where do we find that information? We can now do a lot with MCP and find that information. Then there's a new style of context that is not like in your CLAUDE.md, but it's your specs. "Please implement this," and it's not a prompt, but it's more like a product requirements document in a very loose form. I want you to perform this. There are different ways that different tools think about this. Kiro thinks of this as a task specification. SpecKit actually thinks about this as more of the product features that I keep in my project all the time.

There are different ways of people that do it. Then Baruch also put in intent integrity, where you can verify the specs while you're running these as well. It's a different kind of context that you're now providing to the agent to do the job. A little bit more sophisticated as your prompt. That's the rough generate. There's a lot of things that we're starting to generate with this, whether that's your documentation, your code, your guidelines, your rules. That's all good. You become the curator of context. The agent will decide what goes in, but you'll make sure it has access to up-to-date, the right information, and to go from there.

Evaluate - Test and Measure Context Quality

Second piece is we have code, and in our SDLC, we build tests. It's very common. Nobody likes to write tests. We have things like test driven development, write the test first. All good, but we're doing this to measure the quality of those pieces. What's the equivalent for context? What are tests for context? Most people will have heard about something called SWE-bench. It's a bunch of problems that says, here's an input of a task, here's a model. Take this PR and solve this. How good does it solve it? This is a way to like reflect and make a case, a test case for testing things. Most of the things you see around that are about benchmarks. I'm training a model. I want to see how efficient the model is. I would suspect that maybe 0.01% of you are actually building a model, if not less. You're building context.

This is useful as a way to think about test cases and how they do things. Then there became tools that allow you to write your own benchmarks. This is starting a Docker container. It's running the coding agent. It's input, output. You run the tests. Did you succeed? Yes and no. That's how you would test context by saying, this is without the context, this is with the context, this is with the old context, this is with the new context. That is the baseline of testing that people are going to. This has not yet been into the mainstream because it was thought of like too complex. This is for the LLM providers. You can actually do this yourself. Simple way of testing. Here's that skill. A skill has typically a metadata portion or a frontend matter in the YAML, I can check this similar to a linter.

Does it have the lines that it needs to have? Does it have the correct syntax? That's very much like a linting of context. Maybe if we have some structured context, we can run a linter on that and do some verifications. Similar to linting of code. Then we can see, maybe the description field is too long, this is not according to the spec. Second way, think of this as, we write a skill or a piece of context and the LLM is giving us feedback. I often compare this to Grammarly, like what's your writing style? Is the writing style of your context effective for this model? We can have AI, LLM-as-a-judge, what's your context, and give us feedback on how well this is written, where this is well-worded, where this is too short, where this is mixed, and so on. Grammarly for context.

Let's take this up a notch. So far, we've just been looking at a linter, the syntax, what it can do. A skill often has a purpose, a task it wants to complete. Before, I just looked at like the wording and where it was spec compliant as well. Because I know the goal of the skill it wants to achieve, I can now put this to the test. I have this thing that I want to get done, I run the skill, it gives me of all the test cases, this kind of result. I change it, and then I get better results. It is different from testing because some tests that will pass when you do this start failing when you change it. It's a little bit different. Think of this as I'm testing the functionality of the context that you've provided. A different notch. Much as I told you that vendors are providing docs for their libraries, here's an example of Vercel testing their own versions of models, how well they actually work on their library.

They're testing what information, what best practices they need to give the LLM, and maybe the model of today will not succeed but the model of tomorrow does. They have to provide you as a test harness to trust their models to use the right thing in their library. They're optimizing also for agent use and not just for human use. This is now getting more and more traction. As a library owner, I want to provide that testing and not just the context in there. The extra notch is you take your piece of context, you have a skill, it needs to complete a task, but the rubber meets the road if it's within your repo. Otherwise, it's like a synthetic check. It does that on its own. We can check it no matter what, but the real case, it needs to work in your codebase. Much like BDT, we can come up with scenarios, and say, if you start from this commit, you give it that prompt, we want to achieve those tasks and then we run that set of tests whether that task is completed.

That is the end-to-end test in my opinion for context. Then you even start building a library of testing and go from there. If you know what's wrong, you can start optimizing this. This is like, please enhance, and it gives you hints on how to do this. Similar to, you're in your VS code and it gives you, like you need to do something about this, it's not working, it's not well. It's doing things. You could say things like, please remove all the stuff that the LLM already knows. Why are you putting this in? It doesn't make an impact. It just takes away of your context window. That's one piece of feedback. Conflicting things there as well. The big thing is when you have the test, you can optimize. It's the one thing that we talk a lot about like vibe coding, but the vibe is also like, how do I test this? This was the piece of the puzzle that actually makes that loop. Like, I write things, that's working, but I install a new agent, it's not working. I rerun my test, I now see what's failing, and I get that feedback loop. That's what we're trying to achieve much as we were with writing our own tests.

If we have this whole piece, we can run the tests, the C, whatever you do, the optimize, and then we run that in CI/CD. Unfortunately, I already said it, it's a game of Whack-A-Mole. If you change context here, some tests will succeed, some will not. Compare it to when you have an insanely big codebase, and you've been in a big enterprise, you know that some pieces will fail, even if you're going to a deploy. It's a choice. Is it risky? Is it not risky? There's another piece of that, which is when I run this, it's non-deterministic. I'm going to run the test, let's say, five times. I'm going to say it works most of the time, so we're good. You can run this locally as well. Similar to when I write my tests, I pick two tests, I skip the rest. I have my fast feedback loop.

Then I open the test to more, so that is the same loop. One great way of dealing with that, if you see while you're typing in Claude Code, it's not doing the right thing, take that as a test case. That is something you can bring in. You write the test case without the context, with the improved context, and then you got it working. That is giving you a level of certainty. In the beginning I said it's non-deterministic, and you run those tests, and what you'll start seeing is that somebody has to take a decision whether we're going to go live with this. It is different from going live with code, because you say, test suite works, put it up. Like, 98% of my evals work, what do I do? Who decides? Often, that means actually the product owner has to start deciding whether you go live, yes or no.

The way you think about it, it's a risk game. There's evals you care about that must pass, and other pieces can be a little bit more erratic. The metaphor is, think about error budgets. I give a certain set of functionality an error budget that's very limited, and then I have another one that I can relax. That's how you deal with the non-deterministic part and the Whack-A-Mole to still get something in production. There's a lot of vendors who will try to sell you out-of-the-box metrics. I don't care about whether it works in your environment that's very clean room. I need to work it in my code, in my environment. The evals that are generated, you have to make sure they make sense. You can have an AI generate the evals, and the more context you give it, the likelier it will also generate the best test for that. Don't just go like, I pick a vendor, it has evals, we're done. It doesn't work like that. You've got a whole stack of things you go through. I'll come back to the security scan later in another phase, but that is also pieces that you can put in as well.

Distribute - Package and Share Context

We've got generate. We have the equivalent of testing with evals. Now, how do I distribute that stuff? Most people, if they do share it, could be that they copy and paste it on Slack. I'm not talking about that. That's not sharing. We share it in GitHub. We check the file in, and so somebody else who checks the project out, they can get the context that they need. People have been using Git submodules, but in the end, nobody likes it because you always forget that you need to pull the submodules. It's annoying, and then they start putting everything in one repo. Then one project and then the other project, ok, how do we get the context out? We have a solution for that. It is exactly libraries and package managers. If I have my codebase and I had my library that works only for this project, that's fine, but if I want to reuse that, I create a library.

I package this into an npm package or whatever your poison. I push it to a registry, and somebody else can pull the same thing. In this case, we're just having registries of context that we pull. A simple install this context in this project, version blah. That's a package manager that I want to do. It also helps if you have libraries that provide that context that can just say, from that vendor, from that library, install the version that matches. I have my team's rules that I'm going to install. I'm going to install the frontend package because that is good. We have pieces that you can put together, and it's an artifact, so we version that because you can have the updates and pulling from there. What you'll see is that there's a bunch of marketplaces who have a bunch of skills. Problem is, were they even written for reuse?

A lot of people put their personal skills up, "I built a thing," but they don't have any testing. You don't know. Then you have to look at the whole context whether that's useful, yes or no. There's a maintenance aspect of this that also gives you a credibility if you start adding tests to this, much like you would look at a GitHub repo and say, they have well documentation, that's good, they have good tests. I'm going to reuse that. If you just share a SKILL.md or a markdown file, I know what happens, and you don't even know what happens in your codebase. That's challenging. Registries alone are not the solution. Luckily, I feel that skills is a little bit on its way to become the standard JAR, npm package, or whatever of context. Much like we standardized AGENTS.md, we now have a way, ok, all the major vendors of AI coding will support skills. It becomes interchangeable. It's not like you have instructions, like you have rules just in skills, and we're reusing that same format as well.

When I distribute this stuff, much like security scanning, I want to make sure that there's no malicious skills in there, because skills also contain code, which is like everything's now melting together. This got a lot of attention with OpenClaw and people doing weird stuff. It's weird. It needs like the one thing for the whole industry to wake up. If you think in parallel, you know this was coming. It's so simple to predict. It's one of the things I have with my talks is like, I give the overview, and then I just wait until a vendor implements some of these pieces through the whole year. In this case, Snyk known for its security scanning, also had a context scanner in there. Then, obviously, we have Bill of Materials. How was the context created? Who created the context? Like, why? When was it created? Similar to a Bill of Material that we implemented for security on code, we can do the same thing.

An AI Bill of Material, and here's a tool, git blame, that actually shows like this line of code was written by this person, but they know whether it was AI or not because they captured that in GitHub annotations, and they didn't just look at who checked it in. Part of that provenance of where the context came from also gave you a certain level of credibility, much like it would do your libraries for code. Then, if you take this a step further, instead of like doing this only in annotations, why don't we have almost like a transaction log? Everything that's been done on this code, and we record that. It has an advantage, not just for tracking provenance, but it's also context. Remember five steps ago, we did this, and that person did that, and you've got those sources. The agents love this because then they can read this almost like a backstory of why this happened. Hope this gave you some idea that it's not that dissimilar as code, but we have just a bunch of new tooling that will be spawned around managing context there as well.

Observe - Monitor and Improve in Production

Finally in the cycle, the observe. What if Claude could self-improve its own thing? It is creating a loop that it's looking at the failures of the things it saw in its logs and say, let me rewrite that better. It's another step of optimize that we did with the skill, but it's, in its form, a little bit of an observability thing. It looks at what it did and then improves that piece as well. Hugging Face did similar things on their codebase, and they call it a retrospective. They got the agent to talk about, we did this thing. Was it a good thing? Was it a bad thing? Reflecting actually almost like a post-mortem on what the agent did. They're collecting that and then they're feeding this back into the agent as well. Often, that's still in a very closed loop on an individual system, but you can think about if that optimization gets repackaged, pushed into a registry, the next iteration of all the other agents will get the new context and the optimized context.

That is the feedback loop of actually running this thing in production, not eval as the synthetic checking, but on actual use cases as well. Luckily, the industry is starting to settle on another standard. How do we not only get our CLAUDE.mds to standardize, but how do we standardize our logs? Because they're all starting to tap into that observability layer. Ideally, one standard and then all the tools could work similar to OpenTelemetry on this, but on the logs as well. No surprise, this is the company, Devin, that also does a lot with the agents and autonomous agents. It's building that self-reflection loop as well.

What's another way of looking at feedback? Every commit is a story. You've chatted with this, but imagine I put my agent and said, please auto-push to git. The review there is capturing context. The review is an observability and feedback layer to whether it actually worked. It doesn't have to be a commit. It could be on Slack that you capture this. Like, we deployed whole thread, feedback. That is another loop of observability that we might not think about. We think about like, logs and servers and agents, but this is also feedback of whether your context worked in the wild, yes or no. A whole business, like a company just dedicating to that. Then there was Hud, and I liked the idea. They instrument your code, and when it's deployed in production, they actually capture when things happen. These were the lines of codes that were impacted.

Now, traditional operations would then say, "Great, you had a failure. Here's the problem. Let me raise a ticket with the team." What they could do now is they can say, agent, pull out all that information. Use that information as context so you're not repeating the same mistake again. Save this as a lesson learned, put that into your context, and push that up. It could be as simple as is shown here. The system is being used, like that's the standard deviation of usage, but it could also be like, this is the failure rate on this one, or this is the risk rate, or these are like 10 million people using that function, so probably take care if you're writing this code that there might be an issue and that you're pushing this out. All contextual, but not siloed up.

Gathering information and running these things requires you to do some sandbox, typically. It makes sense, and this is a great article. Like, we want to control what the agent is doing, but imagine it's trying something it wasn't intended. If I get it into my logs, that's great, but what if it's reaching out to an endpoint that we don't know? What if it's doing something crazy? I try to sandbox Claude myself, in a way, and then it says like, "No worries. I can look at your environment variables, even if you scrub me all kinds of secrets." Ok, remove them. "I can still look at your memory." Ok, let's remove that. "No worries, I can just send it over DNS IP exfiltration of the data." That's the kind of thing that you want to know. I know I'm exaggerating a little bit, but this is also part of observability, like what is going wrong and what it's not supposed to do, and not just in your context.

When we deal with context, we used to have the web application firewall, and it's always a pain writing those rules. You get a lot of noise. People question like, "Is this useful or not?" Then it still gives you a signal where something's going on, makes sense. In this proof of concept, I've created a different kind of firewall. Imagine you're in a sandbox, and you say, we're safe. Good. Ok. It's downloading stuff, but the problem is, whenever I start, immediately my coding agent CLAUDE.md is loaded. I have zero control on what it's loading there. Looking at the network, whatever, it is loading. What I did is I said, why don't I create a shared library that will intercept file read calls and actually will check what it's reading from disk. I learned you can do that with eBPF, but not on the Mac. The prototype is, are you safe?

It's not just about network access, it's also like, what is the context? Most people think about prompt injection, but it could be different ways that it's overruling things. Think of this as your context firewall inside of your thing. It has the same problems. You have to write the rules and that is tedious, but it's another way of getting feedback. Even if I'm just logging the things and I'm not blocking the things, I can get that information. What you'll see is that even for training the models, OpenAI is using the app logs, or whatever it's doing to actually correlate that to problems in Codex. They're using it not just in production, but they're using it also to train the model. This is a feedback loop that makes sense, but we can do the same thing with our context and not training a model as well.

Then we get into dark factory land. Who has heard about the term dark factories? There's this futuristic idea, and I hate to break it to you. There are people that just say, give me a spec, I'll pull it through a factory of coding agents, and software comes out. I'm not saying this can't work, but what's interesting is that they're also relying on digital twins to simulate Google, Okta, Jira, and so on. They're basically building that observability also for the agent to see whether it's actually working, whether they're doing the right thing, and so on. I personally think that autonomy is not the end goal. The way that I explain it is, there's different style of managing coding agents. There's the micromanager, and likely, you're doing a lot of micromanagement, "No. Yes. Ok." It doesn't scale. Next step is, I'm going to write you a spec, you generate, and then I do the review.

Then the team lead is just saying, I have a bunch of agents, they do, and one is criticizing the others, and I'm just going to take the call if somebody doesn't agree on that. I'm going to run this nightly, and maybe the agent keeps going, and they just call me when I need it. I'm not even in the loop anymore, just when there's a conflict they can't solve. All different styles. Is one better than the other? If you're a manager, and that's what you're actually becoming, of the agents, the point is that you know which management style to apply when. If you truly, deeply care about something that cannot go wrong, you get into the room, and you're micromanaging the situation. If the risk is low, and you mitigate it like you probably mitigated maybe deploys or similar things, you can say, go at it.

It's a known situation, we got guardrails, everything around it, and if we hit it, you can call me. That's why I'm saying autonomy is not the goal, it's just knowing when to apply autonomy in that full spectrum. That's my take on this. There are obviously people who believe that it's the ultimate autonomy, we can run billion companies with two people and it's done. As Hannah showed us, they're still on-call, there are still solutions, we still have to make decisions. Good luck if this is your money, and you're just blindly trusting the agents. I hope that gave you an idea on where the similarities are, managing context, and managing code, and how the tool space is adapting, and we're getting there.

Context Flywheel

Final piece. This slide is what my initial talk was, but I decided to rewrite my talk. It talks about you as a developer changing and shifting roles. As a developer, you're managing agents, so you're basically operations. You do not know anything anymore about the code. They throw it over the wall, and you have to make a decision. Great. If you write your specifications, you actually become the QA architect that writes context, and kind of gives them all the specifications, how the agents should be working. Great. Then you become the product owner because you decide what the agents need to do, what they need to build. You get closer to that field and working out. Maybe you build that proof of concept, you build five proofs of concepts. Why not? Why limit it to one, and you make the call, what is actually good. In the end, you become the data manager who collects all that data and turns that into knowledge. Senior engineers know that their role was never only about the development. They were the ones who went out into other parts, and now AI is pushing us to do the same thing in there. End this slide, but I think that still holds, even though this diagram was made two years ago.

I'm an engineer. What's in it for me? There's a funny thing happening. Engineers did not like to write tests, but all of a sudden, they like tests because the agent is writing the code, and without the test, they can't tell whether it's good. Engineers didn't like to write documentation, but now they write context because it's very selfish of them to write context because the agent does a better job, and they don't have to do that much of review. For the first time, there is an alignment of people doing jobs and things that they weren't liking before, but I see it almost like they start doing good engineering practices, and no matter how hard me as a VP engineering, somewhere like, "Please don't do this. Please do it that way." If I could just distribute the context and say, my agents are going to do the job for me, and say, "No, no, not that way.

Please do it that way." That makes me scale as well. For some time, we try to get the alignment of dev and ops, kind of like you care about the production, you care about these things. For some reason, this is working as an auto-alignment across people doing that stuff automatically. We've been talking a lot about doing this as a solo person, which is great, a lot of people do it that way. I really said, what if I share my best practices and my context to my teammates or to my teammates' agents? Then they don't hit the same problem again. This is also a perfect place to get all the naysayers on board, and it's like, no, AI can't do this. Can you write the context so they can do that? Great. The thing is, the more you share this and not just on Slack, but the more you share this, there is a benefit effect of distributing that. Why one team, why not a set of teams? You maybe as a platform team say, let's write the guidelines, push that also into the context, and you're scaling this out as well. Everybody can contribute pieces of the context, and eventually, hopefully, your organization collect benefits from that flywheel as well.

People have been thinking about maturity levels. Do I manually create context? Do I persist it? That was the distribution. Do I make it adaptive? That's the optimize and the learning. Do I look at observability? That's my proactive context looping back. Then, who owns context? It's good. Maybe you own the context or the code within your repo, but if you start sharing, and that's one of the puzzles that people have like, but who owns the CLAUDE.md? Who owns my skills? Who owns that? It's an artifact. We have to start thinking like, that team is best positioned for that piece of context. They own it much like a library or a microservice, kind of that is ownership. Here's an example of a product that says, everybody can submit rules, but there's only a few in the team that can say, yes, no, before the distribution happens to the whole company.

There's an ownership and handling PRs of context. If you think about maturity, our agents are still very much optimized for coding. How much will they get optimized for creating context, downloading the right context, doing that stuff. There's an agent maturity around context that you have to think about, like the technical tooling, how well does it actually deal with context? How mature are your team or your organization context practices? Are you putting in the tests? Are you pushing them just like YOLOing out? That's a maturity level on its own. Then, as a manager, how do you enable other teams to share their context? What are you doing for them to make it available? Are you putting that platform where they can push things out, where they can share things? Do you give them education on how to write good context? Do you have the right tools in place and the right practices? You start focusing people like, every time somebody says the I did a bad thing, the reflex is like, did you update the context? That kind of reflex is enablement, going around that.

In dev and ops, we bust it and try to bust quite some silos. What's funny right now is, even if the people sometimes are hoarding context, as long as I can get access to the data, my agent will figure it out. There's a silo busting happening like, typically if I'm not in engineering, I can create a PR and then I have to see whether that's working. What happens right now is, as long as I have access to the repo, I can send the team pull requests. I'm not blocked by a PR. Even like one of the things we're doing in our company is that I report things from the outside, and we get into a conversational Slack, and we just say, @Claude, can you fix it? Can you create a PR? We're having just a joint conversation, and then it just does the thing, and then only one person has to say approve, yes, no.

Before that was like, create a ticket, wait until the triage was happening, and it just goes a lot faster. There is definitely some silo busting going on there. It's not all perfect, and I don't want you to believe that. Context will change. The model will change, whether it's still effective with your context. Is it still relevant? Is it the thing that we did a year ago? Now, we have to remove this. Obviously, AI can help you and see, and picking up the signals, whether somebody said no, and five developers said no, and then you say, let's have a discussion about that. Context drift is a thing, much, and as I told you, distributing the context. If I'm still using the old version of the context, you're using the new version, it is still an unsolved problem.

A couple of unfinished thoughts. What is chaos engineering for context going to look? Are we just injecting context and see what happens, and we'll survive, and the coding keeps working, and our whole system detects bad things? Context analytics, maybe, is that the number of downloads of context? Is that the effective use of context that we want to flag? I don't know what that's going to look like. What about A/B testing? Can I push out new context, and try it on two teams, and see whether that succeeds, and then roll it out to the rest of the team? I didn't find those products. Who will be the next GitHub? Who will be the next Kubernetes for context? Then the question comes, why are we still doing this? All the work, can't the agents do it? Maybe they can, but we're still responsible. We still have to assess the risk, and they can all assist us with that. In the end, some things are just a human call. Then maybe they'll just invent their own language, and we can't see what context they're sharing within each other, and then the end is near.

I didn't dwell too much on knowledge, but I think this is where we're heading. We code. We're sharing all the context. The ultimate goal is that your context actually become knowledge that is specific, because it's not a bunch of markdowns, it's like what you put in there that actually counts. Your moat as a company might not be the coding, might not be the first thing, but it's providing the right context, and maybe about your business, the right context, and this is the differentiator. I've done things, like, I have my own solution. I see there's like five GitHub repos that do the same thing. Can you just extract what features they have, and we don't, and can you incorporate them into my project? That's the copying in the end, if you have that world, but I will not have the business understanding of what to build and what to go for. I want you to think about that. There's still a moat to be had in the industry. A lot of people focus on the coding agents, but I think context is the fuel, and coding agents are your engine. If you're not giving it the right fuel, they choke and underperform in that way.

Questions and Answers

Participant 1: I have a question about language of the context. What language is the best for the context? Maybe it's a code for some type of context, or English, or another abstract language?

Patrick Debois: Language models are typically trained on examples. Most of the LLMs that we're using out there are a lot trained on English. With embeddings, you can have similarities between English and other languages, but you sometimes would struggle. We had one case where somebody made a skill in Spanish. Our review said, you're using Spanish. Is that bad or not? There is definitely a challenge there. Like, is that working for all the languages, how you describe? I think in reality, from what I can tell, some languages are better than others, and that's just because of the training of the model.

Participant 2: My question is around confidence. Let's assume that you apply all the things that you describe in this talk. It seems to be like quite a rock-solid strategy. It seems like that has some logic. In enterprise software that is basically ruling the world, I'm thinking like banking, insurance, health, hospitality. How confident are you with this approach that the output will be good enough for running in production, considering that usually you don't have like a greenfield, you are working in brownfield, maybe with a lot of intricate logic and stuff like that.

Patrick Debois: There's a couple of pieces. How confident I am? Let's say you're doing legacy and the language has the same problem as that language. Has it been trained on that? Will it do good? Some AI coding vendors, they train on more niche languages. That's already one step you can take. The providing context in most of the cases gives a lift to do that. I talked about the management styles, and it's the risk. Would I say straight to production in a bank without it? No. Can I come up with cases where I would do it if I have the right tests and the guardrails in place? Yes. It's a little bit like the story of within DevOps, it says like, can the junior come in and push to production on day one? How confident are you? It was how far, not only about your tests, but also about your guardrails. You can put a lot of guardrails in place. I still feel a little bit hesitant about like the auto going to production. Yes, there's people doing that, but they accept the risk for their business. If you're not willing to accept, don't.

 

See more presentations with transcripts

 

Recorded at:

Software is changing the world. QCon London empowers software development by facilitating the spread of knowledge and innovation in the developer community. A practitioner-driven conference, QCon is designed for technical team leads, architects, engineering directors, and project managers who influence innovation in their teams.

Sep 30, 2026

BT