The Operational Context Gap in AI-Native Service Management
by Dave Rosenlund, Chrissy Clements, and Sven Peters
About the Authors
Dave Rosenlund
Global Director of Software and Solutions at Trundl, an Atlassian Platinum Solution Partner Enterprise, and an Atlassian Community Champion. He works at the intersection of enterprise service management, AI adoption, and Atlassian platform strategy, and writes the "AI That Works" series on the Atlassian Community, one of the platform's most-read practitioner perspectives on AI in service operations.
Chrissy Clements
Global AI-native Service Reinvention Lead at Accenture, an ITIL Master, and an Atlassian Community Champion. She advises enterprise organizations on AI-enabled service management, governance, and the operational changes required to make AI a trusted participant in service delivery.
Sven Peters
AI Evangelist at Atlassian, where he focuses on how teams adopt AI to improve velocity, reduce friction, and change how software gets built and shipped. A longtime developer advocate and keynote speaker, he has spent more than a decade studying what makes engineering teams work, and what gets in the way.
Contents
Executive Summary
The Gap Behind the Numbers
AI is arriving in service management, but it's early for most organizations.
From what we've seen, most organizations are experimenting: piloting agents on incident triage, letting copilots surface knowledge, automating a slice of request fulfillment.
Where it's deployed, the early numbers look encouraging: tickets deflected, resolution times down. Those results are real. They're also narrow, and they measure what AI produces, not what it requires.
What AI requires is context: the current, accurate picture of services, dependencies, ownership, and decision rights that lets it produce decisions grounded in reality, not just reasonable guesses.
In most organizations that picture is incomplete, and skilled people quietly compensate for the difference. That compensation has always worked. However, it doesn't scale as AI expands, and it doesn't transfer when the people who carry it leave.
This paper names that gap, traces how it forms and why it persists, and advances a single argument: closing it is not an IT cleanup task. It's a larger operating-model decision that determines whether AI investments become durable enterprise capability or a growing source of friction that erodes customer satisfaction.
The Central Argument
The organizations that treat it that way now, early, will operate differently from the ones that wait for the compensation to fail.
01
Why the Stakes Are Higher Than They Look
How AI adoption metrics leave the context question unasked
What the Numbers Measure
AI is moving into service management faster than the operational models underneath it. That's not a criticism. It's a description of where most organizations are right now. Agents are triaging incidents. Copilots are surfacing knowledge. Automation is handling request fulfillment that used to require human routing. By the metrics that get reported, it's working. Tickets deflected. Mean time to resolution down. Practitioner capacity freed for higher-complexity work.
The problem isn't that those outcomes are wrong. They're real. The problem is that they're incomplete. They tell you what the AI delivered. They tell you nothing about the quality of the context it drew on to get there.
Every AI-assisted workflow operates on context. Not the surface record of what a service is, but the operational picture of how it works and who decides. That's the difference between the AI doing something useful and something merely approximate. That context exists in your systems. What most deployment assessments don't ask is whether the context your systems contain is the context your AI needs. Whether it's current, complete, and structured in a way that makes it executable. Whether what the AI sees matches what the work requires.
The Gap That Stays Invisible
In most organizations, the answer is that it mostly does. And practitioners handle the rest. This isn't a failure condition. It's how complex operational systems function. Skilled people compensate for the gap between what documentation says and what operational reality is. They carry that knowledge and they apply it. The system works because they make it work. The gap stays invisible for as long as the compensation holds.
When the Compensation Starts to Thin
What changes the stakes is when the compensation starts to thin. AI deployment at scale means more autonomous action and less human review per decision. Organizational change means the practitioners who carry the compensating knowledge leave, and it doesn't transfer cleanly into documented context. The consequences of acting on incomplete context grow, because the human catch is no longer in the loop where it would have mattered.
None of this shows up in a ticket deflection rate. None of it registers on an AI adoption dashboard. It accumulates quietly until a threshold is crossed, and then it doesn't look like a context problem. It looks like an AI reliability problem, or a process problem, or a people problem. The structural cause stays buried.
That gap isn't IT housekeeping. It's becoming an enterprise problem. The operational knowledge that makes AI reliable is the same knowledge that leaves when experienced people do, and it was never written down anywhere the AI can reach. The more of the business that runs on shared context, the more the quality of that context decides whether AI spending turns into durable capability or into risk that stays quiet until it's costly. That's not a service-desk metric. It's a decision about how the organization operates, and most are making it now, by default.
A Competitive Question
This paper is about that structure. It's not a catalog of AI risks, a maturity model framework, or a buyer's guide to platforms. It's an argument. The argument is that there's a specific and nameable gap between what AI-assisted service management requires and what most organizations have built. That the gap is being managed through human compensation that can't scale and doesn't transfer. That the path to an AI-native service model requires closing that gap deliberately, not waiting for the compensation layer to fail.
The stakes are competitive, not just operational. As models become increasingly interchangeable, durable advantage shifts to the one layer that doesn't commoditize: the quality of the context an organization can feed them.
The AI-native service model is a direction of travel. It's not where most organizations are now, and it's not an immediate destination. It's the model toward which the platform decisions, process investments, and capability choices of this period are pointing. What organizations build now either compounds toward that model or makes it harder to reach. The context gap described in this paper is the difference between those two outcomes.
The first thing to understand is why the gap exists in the first place, and why the systems that are supposed to prevent it don't.
02
What Your Service Model Doesn’t Know
The practitioner knowledge AI cannot find in any system
A Harder Problem Than Data Quality
Much has been written about the data quality problem in AI-enabled operations, and rightly so. In earlier work, all three of us have argued that AI can only act on the context it’s given, and that weak context produces weak outcomes. That observation stands.
This paper is about something harder.
The challenge isn’t only that service management data is often stale, incomplete, or inconsistently owned. The challenge is that even when the documented record is reasonably accurate, it still doesn’t capture what experienced practitioners know and do. There’s a layer of operational reality that has never been in the system. Not because nobody bothered to document it, but because it was never designed to be captured there.
We call this the Human Compensation Layer.
Key concept
The Human Compensation Layer
Every service management environment depends on knowledge that lives in people rather than systems. This is not a documentation failure. It is a structural property of how complex organizations work. The Human Compensation Layer is the operational knowledge practitioners carry and apply continuously to bridge the gap between documented process and operational reality.
What lives in the system
What lives in people
Ticket records: what was requested, when it was resolved, which SLA was breached
Which CMDB dependency maps stopped being updated after a migration
CMDB entries: what the asset is, who owns it, where it lives
Which SLA clocks effectively reset at escalation, regardless of documentation
Runbooks: the steps a practitioner is supposed to follow
Which approval paths require a specific VP regardless of risk classification
Service catalog: documented owners, SLAs, escalation paths
Which escalation routes only work because one person knows who to call
The Human Compensation Layer
Every service management environment runs on it. It’s the knowledge that lives in people rather than systems: the analyst who knows the CMDB’s dependency map is accurate everywhere except the legacy cluster that Finance runs, which the system stopped tracking after a migration. The agent who understands that the SLA clock on a specific service type effectively resets at escalation, regardless of what the documentation says. The manager who knows that any change touching the payment gateway needs a particular VP’s approval, regardless of what the risk classification assigns.
None of that’s in the ticket. None of it’s in the runbook. It works because people carry it.
This isn’t a documentation failure in the conventional sense. Organizational knowledge researchers have understood for decades that tacit knowledge (the kind embedded in practice, judgment, and experience) resists formal capture by nature. Michael Polanyi named the concept.1 Nonaka and Takeuchi built on it: converting tacit knowledge into explicit, documented form is a deliberate organizational capability, not an automatic byproduct of documentation effort.2 Most organizations never fully develop it. They don’t need to, because human practitioners bridge the gap continuously and invisibly. That arrangement has worked. Until now.
None of that is in the ticket. None of it is in the runbook. It works because people carry it.
What Changes With AI
AI systems encounter the formal, documented service model. They can’t encounter the Human Compensation Layer. It was never encoded anywhere AI can reach. The on-call contact who is on sabbatical. The legacy system with the undocumented quirk. The escalation path that only works because a specific person knows who to call. None of that transfers.
What changes with AI isn’t the existence of the gap. That gap has always been there.
What changes is the consequence of it.
Human practitioners noticed when something felt wrong. They recognized when the documented reality diverged from the operational one, and they compensated. Often without being asked, often without knowing they were compensating. An AI system operating against the same documentation doesn’t notice the divergence. It reasons forward.
It routes the incident, drafts the response, closes the request. All against context that was always incomplete. Without the judgment that made incomplete context workable.
The gap did not get worse. The catch did.
This isn’t the same problem that a better knowledge base solves. The Human Compensation Layer was never in the knowledge base. It was in the people.
A New Design Requirement
A service model becoming executable context isn’t a metaphor for a technology shift. It’s a description of what the service model is now required to do: not record how service is intended to work but provide the operational surface AI reasons over and acts against.
Everything underneath that surface just became more consequential.
03
The Map Most Organizations Are Lost On
Five stages of operational context, and where most organizations actually stand
Data versus Context
The distinction matters architecturally. Data describes what happened. Context explains what it means. In practice, that difference determines whether AI can do useful work or just produce confident summaries of things you already know.
The service management systems most organizations run are built for the first job. A ticket in Jira Service Management captures what was requested, who requested it, when it was resolved, and which SLA was breached. A configuration item in an Assets record reveals what the asset is, who owns it, and where it lives. A runbook in Confluence documents the steps a practitioner is supposed to follow. All of that’s data. It’s accurate, or approximately accurate, or accurate as of eighteen months ago. Even a full history of resolved tickets is data of this kind: a record of what was done, rarely of why it worked or the judgment that made it work.
None of it was organized to answer the questions AI needs to answer. Which services depend on this component? Who makes decisions when two teams disagree about scope? What does “critical” mean for this customer, specifically? Why does this incident pattern keep recurring even after three remediation efforts? Those answers exist in the human part of the organization. They’re not part of the data.
Why Most Organizations Overestimate Their Context
Organizations don’t move from poor context to AI-ready operational context in one step. They move through recognizable stages, each capturing more of the relationships AI depends on to reason effectively. Most organizations are somewhere in the middle. Most also overestimate how much operational context their current service model actually provides.
Five Stages of Operational Context
01.Asset Register
You know what you own: what it cost, who is accountable for it, and where it sits in its lifecycle. You don’t know what it does, what depends on it, or what breaks if it changes. An asset register is a financial and contractual record, not an operational one. AI at this stage can answer ownership and inventory questions. Nothing operationally useful follows from that.
02.CMDB
You have mapped relationships between configuration items, but the map was built for change management and is maintained by a small team against a large, moving organization. The dependencies are documented. They’re rarely current. Engineers verify the CMDB against reality before significant changes because experience has taught them to. AI reasoning from a CMDB that requires human verification before use isn’t a reliable operational actor.
03.Service Catalog
You have a catalog of services, anchored on the business services users request, with documented owners, SLAs, and escalation paths. AI can triage and route reasonably well for standard scenarios. The tell is the edge cases: anything with operational nuance gets routed to the wrong team, escalated unnecessarily, or resolved against criteria that no longer match how the service works. It was designed for human consumption: routing requests, satisfying auditors, onboarding new staff. It wasn’t designed to answer the questions AI needs to answer.
04.Knowledge Graph
Relationships are first-class. Not just what assets exist but how they connect: which services depend on which components, which teams own which decisions, which architectural choices drove which constraints, which unresolved issues are blocking downstream work. AI reasoning from a knowledge graph can explain why something is the way it is, not just what it is. This is where the performance gap between AI grounded in structured context and AI operating on unstructured data becomes measurable. Most organizations haven’t crossed this line. The ones that have know it immediately because the failure rate on AI-generated recommendations drops sharply.
05.Executable Context
The service model is designed from the start for AI participation: governed relationships, explicit decision rights, stewardship built into how work gets done rather than assigned as a maintenance task. This isn’t a state most organizations have reached. It’s the design requirement this paper is working toward.
The Architectural Shift
The asset register, the CMDB, the service catalog, the runbook library: each was built for a legitimate purpose. None of those purposes was “provide the operational surface AI reasons over.” That’s a new design requirement. Most of what organizations have built doesn’t meet it.
AI has no equivalent to the Human Compensation Layer. It reads what’s available. If the service model is organized for operational convenience rather than contextual completeness, AI operates on a partial map. The gap between today’s service model and AI-ready operational context is the critical one. For most organizations, the transition from service catalog to knowledge graph is where that gap first becomes visible.
A knowledge graph requires a different intent behind how information is organized: relationships, not records; ownership, not assignment; dependency, not proximity. The question changes from ‘where is this information stored?’ to ‘how does this piece of information connect to everything else?’
Once operational context becomes structured as a knowledge graph, a second possibility emerges: that context becomes portable. It’s no longer confined to the application where it was created. It can move across systems and become available to any trusted AI capable of reasoning against it. That’s the architectural shift. Application integration moves information. Portable Context moves operational meaning.
Portable Context and the Architecture Beneath It
The clearest production example of this shift today is Atlassian’s Teamwork Graph. Rather than simply connecting applications, it models the relationships between work, knowledge, people, and services that AI needs to reason across organizational boundaries.
It demonstrates what becomes possible when operational context is treated as a connected asset rather than a collection of records.
44%
More accurate AI responses when grounded in structured context vs. unstructured organizational data. Atlassian Teamwork Graph benchmarks, May 2026.
48%
Fewer tokens needed for AI operating against structured context vs. unstructured data. Same benchmark. The implication: organizations should audit their context, not compare models.3
That finding inverts the way most organizations are thinking about AI performance. They’re comparing models. They should be auditing their context.
At Team ’26, Atlassian made a structural decision that clarified where the ecosystem is heading. The Teamwork Graph was opened to any MCP-compliant AI agent, not just Rovo. Claude, Cursor, GitHub Copilot, and others can now consume the same context layer that powers Rovo. That reveals a three-layer architecture for AI-native service management:
Governed operational context (with Teamwork Graph as today’s clearest production example), the foundation everything else depends on
The model layer (the underlying LLMs, increasingly interchangeable)
The interaction and agent layer (Rovo, Claude Code, and other MCP-capable clients), where work gets done
The layers aren’t equal. The model layer matters, but it’s becoming easier to swap. The interaction layer matters, but it’s only as capable as the context beneath it. The more autonomy you expect from AI, the more it depends on deep, governed operational context.
That’s why the foundation carries the weight. Application integration moves information. Portable context moves operational meaning: not data moving between systems, but relationship-aware context any compliant AI can reason against. The agents and models above it are increasingly interchangeable. The context layer isn’t.
Why the Architecture Applies Beyond Atlassian
The architecture isn’t Atlassian-specific. Forrester analyst Charles Betz, writing after observing both major service management platforms in May 2026, concluded that context graphs have become “the center of gravity of the IT management platform” broadly.4 His warning about what happens when organizations don’t make that investment matches the argument this paper makes: AI, he wrote, “will make the cost of not doing data quality and governance show up faster, in the form of agents producing confident, plausible-sounding nonsense at machine speed.” The failure mode is universal. The Atlassian examples in this paper reflect where the authors’ experience runs deepest, not the limits of the argument.
The strategic asset in an AI-native service environment isn’t the model or the agent. It’s the organization’s ability to produce, govern, and share operational context that trusted AI can use. Models will improve. Agents will evolve. Operational context is the foundation they depend on.
04
What Breaks When the Context Is Wrong
How incomplete context produces plausible failures at scale, and why they get misdiagnosed
Plausible, Confident, and Wrong
AI doesn’t fail dramatically. It produces a plausible answer. That’s what makes this hard to diagnose.
An incident gets routed to the wrong team. A recommended action overlooks the dependency everyone in the room knew about but no one documented. An escalation misses the person carrying the institutional knowledge because the service model still points to someone who left months ago. None of these failures are new. What changes is that AI executes them at greater speed, greater scale, and with greater confidence.
We described this in earlier work as fast, confident, and operationally wrong. That captures the symptom. This section examines the structure beneath it.
The Failure That Does Not Look Like a Failure
The Human Compensation Layer quietly prevented many of these failures from becoming incidents. Practitioners noticed when something looked wrong. They recognized when documented context diverged from operational reality and adjusted accordingly, often without realizing they were compensating.
AI doesn’t do that.
It reasons forward using the context available. If that context is incomplete, the reasoning remains internally consistent while becoming operationally wrong. The output still looks plausible. The failure appears later, dressed in familiar clothing, and it gets diagnosed the way failures always have: process failure, communication failure, team failure. More runbooks. Tighter SLA definitions. Better onboarding.
But the service model wasn’t wrong. And AI wasn’t broken. The context was incomplete, and nobody knew to look for that.
In a pre-AI service environment, context gaps usually produced localized failures. A stale CMDB entry affected the change that touched it. A missing ownership record delayed the ticket that depended on it. The consequences generally remained close to the gap.
Portable context changes that arithmetic.
How a Context Failure Gets Misdiagnosed
Context gap exists
Stale dependency.
Missing ownership.
Outdated escalation path.
AI reasons forward
Internally consistent.
Plausible output.
No visible error.
Failure appears
Wrong team.
Overlooked dependency.
Missed escalation.
Diagnosis
Process failure.
Communication failure.
Team failure.
Fix applied
More runbooks.
Tighter SLA definitions.
Better onboarding.
When the Blast Radius Scales
Most AI initiatives don’t fail dramatically, they just fade away. ‘We’re using AI’ isn’t the same as ‘our engineering teams are shipping better and faster because of it.’5
Sven Peters, AI Evangelist, Atlassian
Sven Peters described the organizational version of this problem before this paper. Individual AI use can improve while organizational performance remains flat. The difference is operational context. When context becomes portable, every agent drawing from the same context graph inherits the same strengths and the same weaknesses.
The Teamwork Graph is open to any MCP-compliant agent, which is the right design. But it also means a relationship gap in the service model propagates to every agent consuming that shared graph at once. The blast radius of poor context scales with portability.
$161B
Annual coordination losses across Fortune 500 companies from the fragmentation tax. Atlassian State of Teams 2026, survey of 12,035 knowledge workers and 173 Fortune 1000 executives.6
87%
of knowledge workers reported that with everyone in execution mode, they lack the time or capacity to coordinate across teams: sharing context across handoffs, connecting work happening in parallel, and aligning on decisions before they diverge.
Those numbers describe organizations where individual tasks become faster while coordination remains fragmented. AI accelerates execution. It doesn’t repair missing operational context.
The gains evaporate where the context runs out.
Making the Right Diagnosis
Organizations that diagnose context failures as process failures will keep fixing the process, and the underlying problem will persist, wearing different clothes each time.
Organizations that recognize them as context failures make a different diagnosis: the service model was never designed to be the operational surface that AI reasons over and acts against. Filling that gap isn’t a documentation project. It’s not a cleanup effort. It’s an architectural decision. And it requires treating operational context as the design problem it is.
That’s what the rest of this paper develops.
05
Context Doesn’t Maintain Itself
Why operational context must be a property of work, not a maintenance task
How Operational Context Drifts
The previous section argued that the service model was never designed to be the operational surface AI reasons over and acts against. That diagnosis leads naturally to the next question.
What happens once you’ve built better operational context?
The answer is the question most organizations haven’t asked yet.
Context isn’t an asset you build. It’s a condition the organization either sustains through how it works, or loses.
Operational context reflects the living state of the organization. Services evolve. Teams change. Exceptions become norms. Norms become exceptions. A process that worked six months ago may have been quietly modified by the team responsible for it. Not in any system. In how they work. A service owner who understood a particular dependency may have moved to a different team. A workaround that papered over a structural gap may have become the default path, invisible to every formal record that still describes the original intent.
For most of the history of service management, that drift was manageable. Not because it wasn’t happening, but because practitioners compensated for it continuously. The Human Compensation Layer kept the gap closed.
AI doesn’t work that way. It acts on what the context says. When the context has drifted, AI acts on a picture of the organization that no longer matches operational reality. The gap that skilled humans close through judgment stays open. It doesn’t generate a visible error. It generates a plausible response to a situation the context no longer accurately describes.
The natural instinct is to treat this as a governance problem. Assign ownership. Build a review cycle. Define metadata standards and audit schedules. This instinct comes from treating context as an asset to be stored, versioned, and managed. It’s the same mental model organizations use for data.
That frame is too narrow. The Human Compensation Layer isn’t a data quality problem. It’s operational knowledge that was never designed to be captured. The context that matters most is the operational knowledge that shapes how work gets done. That context isn’t produced by a documentation process. It’s produced by the organization doing its work.
Why Governance Programs Alone Fall Short
There’s a legitimate objection to raise here. Formalizing tacit knowledge isn’t free. Knowledge management has a long track record of expensive initiatives that produced documentation no one read, captured knowledge in forms that degraded on extraction, and created maintenance burdens that outpaced the teams assigned to carry them. That objection is worth taking seriously.
But it’s aimed at the wrong target. This paper isn’t arguing for a formalization program. It’s arguing for something different: designing work so that doing it keeps the operational picture current. The context this paper is concerned with can’t be extracted from how people work. It can only be built into it.
The better question isn’t, ‘Who is responsible for keeping the context accurate?’ It’s, ‘What does the organization have to do, routinely, for the operational picture to stay accurate on its own?’
Context as a Byproduct of Work
The answer looks different from a governance committee:
Three Behaviors That Sustain Context
Service owners resolving incidents in ways that update the service model, not just close the ticket.
Engineers capturing root cause as operational context, not just as a record of what broke.
Leaders making decisions that are visible as decision rights, documented as authority rather than as outcomes.
In each case, the context is a byproduct of work done deliberately. Not a separate maintenance task. A property of how the work gets done.
Stewardship as Organizational Capability
This is why context stewardship isn’t a new IT function. It’s a new organizational capability.
The distinction matters. A function is something you add to the org chart. A capability is something you build into how the organization operates. Organizations that treat stewardship as a function will assign it, staff it, and discover that the people assigned to it can’t keep pace with an organization they’re separate from. The context that requires stewardship is produced continuously, across every team, every incident, every decision. No function maintains that. Only the organization does.
In February 2026, ITIL 5 was released, an evolution of ITIL 4 that updated the qualification scheme broadly, not just at Foundation level.7 Among its changes: it named stewardship of service and product data as a distinct operating-model responsibility. Not an administrative task. Not a documentation requirement. A named, ownable organizational capability, with its own governance qualification and explicit connection to AI-enabled service management.
That’s notable not because ITIL defines best practice, but because it names something the industry had previously left unnamed. Stewardship has always existed, carried informally by the people closest to the work. ITIL 5 corroborates that shift. It doesn’t create it. What changes is the consequence of doing it poorly. When AI acts on operational context, the quality of that context determines the quality of what AI does.
Organizations that sustain operational context well have made a specific design choice. They have built work so that doing it keeps the operational picture current. That’s an operating-model decision before it’s a governance decision.
It also makes the next question unavoidable. If operational context is produced across the organization, who has the authority to define it when different teams disagree? Stewardship keeps the context current. It doesn’t decide whose understanding becomes authoritative.
That’s the question the next section answers.
06
Who Decides What the AI Can Do
The authority question most AI deployments have not confronted yet
Authority the Documentation Never Encoded
The previous section ended with the question this one has to answer: if operational context is produced across the organization, who has the authority to define it when different teams disagree?
Every service management environment has always had implicit answers to that question. The CAB knew what was risky enough to review. The service owner knew what the SLA meant in practice. The incident commander knew when to escalate. None of that was usually documented as decision rights. It didn’t need to be, because the people doing the work understood the boundaries. AI doesn’t inherit those boundaries.
When an agent triages an incident, routes a request, recommends a change window, or closes a ticket without escalation, it’s exercising judgment. That judgment used to belong to a person who understood not only the task, but the limits of their authority. If the context doesn’t encode those limits, the agent still acts. Not maliciously. There’s just nothing in its operating surface that tells it where authority stops.
This is the governance problem most AI deployments haven’t confronted yet. Not whether AI should be governed. Everyone agrees it should. The operational question is narrower and harder: how do you express decision rights in a form AI can respect?
The Spectrum from Tool to Autonomous Agent
There’s a spectrum. At one end, the agent is a tool. It surfaces information, suggests options, drafts responses. A human reviews and acts. Safe, but also where much of the value stays constrained by human review. At the other end, the agent acts autonomously: resolving known incident patterns, executing standard changes, and closing requests that match documented criteria. That’s where the transformative value lives. It’s also where the authority question becomes unavoidable.
The space between those two ends is where most organizations will operate, and it’s ungoverned territory for almost all of them.
AI inherits permissions. It does not inherit authority.
An agent that reassigns a ticket to the wrong team creates a delay. An agent that closes a problem record prematurely frustrates the teams watching it. An agent that modifies a firewall rule based on incomplete context creates a security exposure. An agent that rolls back a deployment without understanding the data migration between versions can cause damage that’s not cleanly reversible. The difference isn’t just severity. It’s consequence.
And these are still the easier cases.
The Agentic Infrastructure Problem
Agentic AI is moving beyond ticket operations into infrastructure. Scaffolding environments, provisioning cloud resources, modifying network configurations, scaling services, triggering deployment pipelines. Not recommending changes. Executing them.
In that world, the governance question isn’t whether the agent triaged correctly. It’s whether the agent is authorized to change the operational environment itself.
Can it provision a new environment in response to a capacity alert, and if so, in which regions, under which cost envelope, and with what security baseline? Can it modify firewall rules to resolve a connectivity incident it has diagnosed, and does it understand that the rule it’s about to change was put there deliberately by a security team that’s not visible in its context? Can it scale down infrastructure during a quiet window, and does it know the apparently quiet service is running a batch process for a regulatory filing due in four hours?
Some of those actions are reversible. Some aren’t.
That’s where decision rights have to account for consequences, not just capability. It’s not enough to define what the agent can do. The organization has to define how reversible each kind of action is, what confidence threshold is required, and when the agent has reached the edge of its authority. A ticket level action that’s easy to undo can run with more autonomy. An infrastructure action that’s not easy to undo needs stronger gates until the organization has earned trust in both the context and the agent’s judgment.
What the Human Compensation Layer Carried
The Human Compensation Layer didn’t just carry operational knowledge. It carried authority knowledge.
At what threshold does an incident require human review? Who authorizes an exception to a standard process? What constitutes a decision ambiguous enough to require approval versus one the agent can execute within defined parameters? Those answers existed in practice. They were rarely encoded in the service model.
This introduces a capability most organizations don’t have yet.
Not a technology role. Not an AI administrator or a model tuner. A capability that sits at the intersection of service design, governance, and AI operations, responsible for defining and maintaining what agents can do, under what conditions, and how that changes as context quality and organizational trust evolve.
Some organizations will see this as an extension of existing roles: the service owner, the change manager, the process owner. But none of them fully owns this question today. Service owners define how services work for humans. Change managers assess risk for human-initiated changes. Process owners design workflows that humans execute. None of those roles, as typically defined, owns what an AI agent is sanctioned to do within those services, changes, and workflows.
And this capability will be contested.
Each of us has sat in rooms where IT governance, security, and service management all claimed this was theirs. That’s exactly why no one owns it. It sits across all three. No single existing power structure wants to cede ground to create it, and no one has the mandate to force the question.
The organizations that resolve this tension early will have a structural advantage. The ones that let it sit unresolved will find their AI governance fragmented across teams, each governing a slice, none governing the whole.
Adding a “Head of AI Governance” to the org chart doesn’t solve this, for the same reasons context doesn’t maintain itself.
A Governance Model at Three Levels
What this capability requires is a governance model that operates at three levels.
1.Agent-level decision rights
Define what a specific agent can do autonomously, what requires human confirmation, and what is outside its scope entirely. These rights should be explicit, auditable, and version-controlled. Not buried in configuration, but visible as governance artifacts service owners and risk teams can review.
2.Service-level trust boundaries
Define the envelope within which an agent operates for a given service. An agent triaging password reset requests operates in a different risk environment than one routing incidents affecting payment processing. The boundary has to reflect the consequence profile of the service, not the technical capability of the agent alone.
3.Organizational escalation
Defines what happens when an agent reaches the edge of its decision rights. The answer can’t be simply that it stops. It has to escalate to a defined human authority with the context required to decide. Otherwise the failure mode isn’t that the agent acts without authority. It’s that the agent defers, and no one owns the decision that follows.
The organizations that get this right treat governance and the confidence to use it as things that grow together. They start narrow, where mistakes are cheap and reversible. Clear escalation paths. Measured outcomes. An envelope that widens from what they learn, not from what the policy hoped. The ones that get it wrong won’t always discover it through a dramatic failure. They will discover it through the slow accumulation of individually defensible, collectively ungoverned decisions, each one eroding trust in the system they were trying to build.
Beyond IT: The Consequence Problem Across ESM
This section has focused on IT service management, but the argument doesn’t stop there, and neither does the platform.
Atlassian’s service collection, like most enterprise service management platforms, already spans HR, legal, facilities, and customer operations. IT may administer the platform. It doesn’t administer every consequence profile that platform touches.
An agent handling HR service requests operates under employment law. Discrimination protections, grievance procedures, privacy obligations: those don’t appear in the IT risk model. A customer-facing agent that acts outside its decision rights doesn’t produce a bad ticket. It produces a customer commitment. A legal service request carries privilege concerns that postmortem governance can’t undo.
The platform is unified. The consequence profiles aren’t. Domain-aware decision rights aren’t an ESM add-on. They’re a condition of operating a unified service platform responsibly.
The organizations that navigate this best won’t create a new governance structure for every domain. They will expand the mandate of the function that already owns the service model. Service management owns the definition of what services do and how they’re governed for human practitioners. Extending that mandate to cover what agents are sanctioned to do is a logical expansion, not a new org design problem. Security and risk functions become stakeholders in that governance model, not owners of it. They set the risk thresholds. Service management operationalizes them.
ITIL 5 points in this direction. Naming stewardship of service and product data as a service management responsibility, not an IT task, positions the function that knows the service model best as the natural home for the governance decisions that follow from it.
The question should not be whether someone needs to decide what the AI can do. It’s whether your organization has decided who that someone is, and whether the service model is explicit enough for that decision to hold when AI starts acting against it.
07
What an AI-Native Service Model Requires
Four design properties of a service model built for AI participation from the start
The Design Intent Question
An AI-native service model isn’t an existing service model with AI added to it. The distinction matters. A model with AI layered on top inherits the gaps this paper has been describing: incomplete operational context, informal authority, human stewardship that lives outside the system, and relationships that remain accurate only because practitioners know where the record is wrong. An AI-native model starts from a different assumption. AI participation is part of the design requirement, so the model has to make context, relationships, authority, and stewardship explicit enough for trusted agents to reason against.
The design intent question is the right place to start.
Most organizations ask: what AI capabilities should we add to our service model? That’s a reasonable question. It’s also the wrong one. It treats the service model as fixed and AI as the variable. An AI-native approach inverts that. It asks: if AI participation were assumed from the start, what would we build differently?
The answer isn’t a different technology stack. It’s a different set of design decisions about how work is organized, how operational context is produced, how authority is encoded, and how the model stays current as the organization changes.
The Old Question
What AI capabilities should we add to our service model?
Treats the service model as fixed and AI as the variable
Measures success by AI features deployed
The New Question
If AI participation were assumed from the start, what would we build differently?
Treats the design assumptions as the variable
Measures success by quality of operational context AI can use
The Four Properties
Those design decisions map to four properties.
Structured Context
The service has a dependency graph that reflects how it works now, with decision records that explain not only what was chosen but why.
Governed Relationships
The dependency graph is a live operational artifact. When a service changes, the relationship model changes with it. Relationships are explicit, maintained, and authoritative.
Explicit Decision Rights
Decision rights are first-class properties of the service model. Who decides, under what conditions, and with what accountability. Encoded, not inferred.
Human Stewardship Built In
Stewardship is a design property of the operating model. Incident resolution updates context. Service changes propagate through the relationship model.
1. Structured context
In a human-operated service model, a ticket captures what was requested. A CMDB entry records what an asset is and who owns it. A runbook documents the steps to follow. This is useful information. It’s not enough. Operational context answers different questions: What does this service depend on? What breaks if this component goes offline? Why was this architectural decision made three years ago, and is the constraint it was responding to still real?
A service model designed for AI participation makes those answers structural. The service doesn’t just have an owner and an SLA. It has a dependency graph that reflects how the service works now. It has decision records that explain not only what was chosen, but why. The context is organized around the questions AI needs to answer, not the questions a human would ask a colleague after sensing that the record was incomplete.
2. Governed relationships
In a human-operated environment, dependencies are often managed through coordination. Practitioners know who to call, which upstream service causes recurring incidents, and which dependency is missing because the team that built it left two years ago. That knowledge works while those practitioners are present. It doesn’t become operational context until the relationships are explicit, maintained, and authoritative.
Governed relationships mean the dependency graph is a live operational artifact, not a snapshot from the last migration. When a service changes, the relationship model changes with it. When a team takes ownership, the context reflects that. This is where operational context becomes more than data quality. The organization has designed the work so the relationships AI depends on are kept current by the work itself.
3. Explicit decision rights
The previous section covered the authority question in depth. The point here is narrower: decision rights have to be a first-class property of the service model. They can’t be inferred from org charts, approvals, or institutional memory if an agent is expected to act against them.
An AI-native service model knows which team has authority to approve a change to a service. It knows what threshold triggers escalation. It knows which decisions can be made within the service and which require cross-team coordination. This isn’t a permissions list. Permissions control access. Decision rights encode authority: who decides, under what conditions, and with what accountability.
When decision rights are explicit and structural, an agent operating against the service model doesn’t have to guess where its authority ends. That’s not just better governance. It’s what makes autonomous action safe enough to expand.
4. Human stewardship built in
The corollary to the previous section is simple: an AI-native service model can’t depend on a separate maintenance effort to stay accurate. If context is produced by how the organization works, the model has to be designed so that doing the work keeps the context current.
In most environments, context maintenance is separate from the work. Someone updates the CMDB. Someone reviews the runbooks. Someone checks whether the service catalog still reflects reality. That separation is why context drifts. No team has enough spare capacity to do the work and maintain a parallel record of the work with the precision AI requires.
An AI-native service model reduces that separation. Incident resolution updates operational context. Decisions are recorded as authority, not just as outcomes. Service changes propagate through the relationship model because maintaining the model is part of making the change. Stewardship isn’t an assignment added after the fact. It’s a design property of the operating model.
The Direction of Travel
Most organizations aren’t operating this way now. They’re operating human-designed service models with AI added, and they’re still relying on practitioners to bridge the gap.
That’s not failure. It’s the current state. The useful question is whether each design choice moves the organization toward operational context AI can trust, or locks the compensation layer more deeply into the model.
That makes the transition practical rather than abstract. An organization at the CMDB stage isn’t failing to be AI-native. It’s at a specific point in a known progression. The next move isn’t to redesign everything. It’s to choose one domain where the relationships matter, make them explicit, define the authority around them, and redesign the work so the context stays current.
Each property has an entry point before full maturity. Structured context begins with asking what questions AI needs to answer and whether the service model is organized to answer them. Governed relationships begin with one domain where dependencies are explicit and maintained. Explicit decision rights begin with one service where authority is encoded rather than inferred. Stewardship begins with one workflow redesigned so that doing it keeps the model current.
The organizations that make meaningful progress won’t be the ones that launched the most AI initiatives. They will be the ones that made different design choices about how operational context is produced, how relationships are governed, and how authority is encoded. Those choices compound into a service model AI can reason against reliably.
The Payoff
When the model works, the effects are organizational, not just operational. Incident resolution improves because the context the agent reasons against reflects how the service works today. Governance scales because authority is structural rather than stored in the heads of senior practitioners.
The service model becomes an organizational asset in a way it hasn’t been before. Not because it’s better documented, but because it’s trusted by the people and agents acting against it. That trust is earned through the accuracy and currency of the operational context. Once earned, the organization can expand what it asks agents to do without treating every expansion as a new leap of faith.
It also changes what you commit to. A service level can read green while the experience underneath runs red: the metric met, the person still poorly served. What used to keep that gap honest was the Human Compensation Layer, someone who noticed and made it right, and that someone is what AI removes.
An agent can hit every service level and still leave the person worse off. So the commitment that matters moves from the service level to the experience itself, what the field is beginning to formalize as an experience-level agreement, offered alongside the SLA rather than instead of it. It’s a statement of accountability for whether the service actually worked for the person, not just whether it hit its numbers, and it’s where the quality of your operational context finally shows.
This is where the service model stops being an IT concern.
When the context AI depends on is produced and governed by how the organization works, its quality becomes an operating-model property, not a service-desk metric. It decides whether AI investment compounds into enterprise capability or accumulates as risk that stays invisible until it’s expensive.
That’s the decision your CIO is already making, whether the organization frames it that way or not. Not which AI to buy. Whether the operating model produces operational context good enough to trust AI with more.
That’s why the final question is practical, not theoretical. If this is the service model AI requires, what decisions can’t wait?
08
The Decisions That Can’t Wait
Three decisions that compound, and a diagnostic question every organization can answer now
From Argument to Action
The argument this paper makes is no longer hypothetical. AI is already operating against service models that depend on people to supply operational context the systems don’t contain. The consequence is straightforward: as AI participation expands, that hidden dependence becomes an operating-model liability. The decisions in this section follow directly from that conclusion.
They’re consequences of the operating-model argument developed throughout this paper, not a maturity checklist. Each decision increases the organization’s ability to produce and sustain trusted operational context before the service model reaches full maturity.
Three Decisions That Can’t Wait
Name Ownership
A named owner. An explicit mandate. A regular forum where context quality is visible as a metric, not just as a feeling.
Fix One Service End to End
A mapped set of real dependencies. Encoded decision rights. Workflows redesigned so that doing them updates context as a byproduct. A template for every service that follows.
Define Decision Rights Before Deployment
A one-page decision rights statement: what the agent can do autonomously, what it escalates, and who it escalates to. Version-controlled alongside the agent configuration.
The Honest Assessment
The one decision that unlocks everything else is an honest read of where you are on the maturity progression. Not where your documentation suggests. Not where your last implementation project claimed. Where your operational reality is.
Most organizations overestimate. Not because they’re careless, but because the Human Compensation Layer works. The system appears to function. Requests get resolved, incidents get closed, services get delivered. The gap between what the context says and what the context is stays invisible because skilled practitioners bridge it continuously.
The diagnostic question is this: if the people currently bridging that gap weren’t here tomorrow, what would your AI do?
If you can answer that specifically, naming the gaps, estimating their scale, and understanding which services are most exposed, you have the information you need. If the question itself is uncomfortable, or if the honest answer is “we’re not sure,” that’s your baseline. Everything else starts there.
The Three Decisions
Once you know where you are, three decisions compound.
1. Naming ownership
Context stewardship needs a home before it becomes improvable. This isn’t a hiring decision. It’s a mandate given to someone who already understands the service model: you’re responsible for what the context says, whether it’s current, and how work is designed to keep it accurate.
The person for this role is rarely hard to identify. They’re usually already doing a version of it informally: the service manager who knows which CMDB entries are wrong, the process owner who understands which runbooks don’t reflect how work gets done. Making it explicit changes what is visible and what is accountable. Context quality stops being everyone’s background concern and becomes one person’s measurable responsibility.
A named owner, an explicit mandate, and a regular forum where context quality is visible as a metric, not just as a feeling.
2. Fixing one service end to end
The full context estate isn’t the target. Picking one service and making it right is. The criteria: high AI interaction volume, meaningful operational complexity, and a service owner willing to be part of the work.
Map the service’s real dependencies, not from the CMDB, but from the practitioners who work with it. Identify every category of decision that gets made about it and encode those as explicit decision rights: who approves changes, who escalates incidents, what threshold triggers cross-team coordination. Redesign the core workflows so that doing them updates the context as a byproduct rather than as a separate task.
Then instrument it. Track how often AI recommendations get overridden and why. Run it for ninety days. The override patterns tell you where the remaining context gaps are. Fix those.
What you have at the end isn’t just a better-governed service. It’s the template for every service that follows, and a proof point that the work is possible.
3. Defining agent decision rights before deployment
This is the decision most organizations delay until they have a reason to regret it. Before any agent goes into production, someone should be able to answer three questions: what can this agent do autonomously, what does it escalate, and who does it escalate to?
The artifact isn’t a lengthy governance document. It’s a one-page decision rights statement, reviewed by the service owner and whoever owns risk for that domain, version-controlled alongside the agent configuration. When the agent’s autonomy changes, the document changes. When the service context changes materially, the rights get reviewed.
The organizations that do this find that the discipline of writing it down surfaces assumptions they had not examined. The agent that “handles standard incident patterns” turns out to handle patterns that aren’t as standard as assumed. The escalation path that “goes to the on-call engineer” turns out to route to someone with no context on why the escalation happened. Defining rights before deployment catches these before they become incidents.
The question
The defining question isn’t whether your AI is capable. It’s whether your organization produces the operational context your AI depends on.
Do you know what your AI doesn’t know? Which gaps are still being bridged by practitioners, and what happens when they’re unavailable, leave, or scale outpaces their ability to compensate?
If you can answer that with specificity, you know where to start and what these decisions are worth.
If you can’t, the work described in this paper isn’t a future consideration. It’s the present condition of your AI deployment, running quietly, compensated for by people who are doing two jobs and won’t do so indefinitely.
AI didn’t create the Operational Context Gap. It exposed it. The organizations that deliberately produce and sustain operational context will gain an enduring advantage, not because they chose a better model, agent, or platform, but because they built an operating model that continuously produces trusted operational context.
Appendix
Sources and References
[1]Polanyi, M. (1966). The Tacit Dimension. Doubleday. See also Polanyi, M. (1958). Personal Knowledge: Towards a Post-Critical Philosophy. University of Chicago Press. [Page 5]
[2]Nonaka, I. and Takeuchi, H. (1995). The Knowledge-Creating Company: How Japanese Companies Create the Dynamics of Innovation. Oxford University Press. [Page 5]
[4]Betz, Charles. “Atlassian And ServiceNow: The Dominant AI-Enabled IT Management Platforms Lean Into Context Graphs.” Forrester Blogs, May 1, 2026. [Page 10]
[7]ITIL 5. AXELOS/PeopleCert. Released February 12, 2026. [Page 16]
Trademark attribution and copyright notice
Atlassian, Jira, Confluence, Jira Service Management, and Rovo are trademarks or registered trademarks of Atlassian Pty Ltd. Accenture and the Accenture logo are trademarks of Accenture. Trundl and the Trundl logo are trademarks of Trundl Inc. All other trademarks are the property of their respective owners.