...
Itential Platform Pricing Explore flexible plans and options for your team
Itential logo
Blog

The Autonomy Paradox in Network Operations

Headshot of Karan Munalingal, SVP of AI Strategy and Innovation at Itential, driving AI-driven automation strategy that helps global customers modernize and scale network and infrastructure operations.
Karan Munalingal
SVP of AI Strategy & Innovation
Last updated September 24, 2026
AI agent icon connected to a network server through a guarded barrier with a shield checkmark, showing governed changes

Key Takeaways

    • 82% of network leaders are comfortable letting AI make at least some production changes without prior approval, and 69% require detailed explainability for agent-driven actions. In production operations, those answers describe the same requirement.
    • The bottleneck is broader than alert handling. It is the capacity to turn an operational signal into a justified, completed, verified action.
    • Autonomy is a set of permissions, not an on-off switch: observe, recommend, execute with approval, and execute within policy. An organization can use all four levels at once.
    • Every action an agent might take needs defined answers for scope, evidence, authority, execution, verification, failure, and record.
    • Let agents reason and choose among authorized capabilities. Let governed workflows carry out the actions those capabilities represent.
    • Start with one recurring incident, measure the whole loop, and expand authority on the basis of results.

Would you let an AI agent change your production network without asking first?

For some actions, a surprising number of network leaders would. In new research from Cisco and Omdia, 82% of respondents said they were comfortable allowing AI to make at least some categories of production network changes without prior human approval.

The same report found that 69% require detailed explainability for agent-driven actions. Respondents also identified approval mechanisms, policy limits, access controls, overrides, and audit trails as essential safeguards.

At first, those answers seem to pull in opposite directions. Give the agent more freedom. Put more controls around it.

  • 💡 In production operations, they describe the same requirement.

    A network team can preauthorize an action only when it knows exactly what the agent can do, what information it must check, how the change will execute, and what will happen if the result is wrong.

Autonomy depends on making those answers explicit. That is the challenge the industry now has to solve.

The Workload Has Outgrown the Manual Model

The Cisco and Omdia report puts useful numbers behind a problem network teams already feel. Organizations surveyed generate roughly 4,100 monitoring alerts and events a day, with about half related to the network. The report estimates that clearing the daily network alert volume manually would take around 100 specialists.

Of course, every alert does not deserve a full investigation. That is precisely the difficulty: someone or something has to distinguish noise from an emerging incident, gather context, decide what action is warranted, and confirm whether it worked. The report says 46% of network alerts are closed without investigation. It also says 45% of investigation time is spent pursuing false positives.

Meanwhile, the network keeps changing. Nearly three-fifths of respondents make production network changes daily or more often, and 57% say their current change processes cannot keep up with the speed required.

The bottleneck is broader than alert handling. It is the capacity to turn an operational signal into a justified, completed, verified action. Adding another alert queue does not create that capacity. Neither does asking the same engineers to work faster.

AI can help at each step, but only if the path from decision to action is designed for production.

A Network Incident Rarely Stays in One Domain

The report’s findings on fragmentation matter as much as its alert numbers. Ninety-two percent of respondents say performance issues commonly span multiple domains. Organizations use an average of nearly ten separate tools to maintain full visibility across those domains.

Consider an application latency complaint. The first symptom may appear in an application monitor. The cause could involve a cloud network setting, a WAN path, a security policy, or a device configuration changed hours earlier. An operator may need to consult several sources, coordinate with another team, make a controlled change, and check the application again.

That sequence explains why the report’s incident resolution figures are so far apart: a 12.5-hour median and an 88-hour mean. The hardest incidents can last far longer than the typical one.

Observability and AIOps tools have a vital job here. They collect signals, reduce noise, and help identify probable causes. But identifying a likely cause does not complete the response. The action may have to cross systems, teams, and approval boundaries. It needs a reliable way to execute and a way to prove the result.

  • 💡 The industry has spent years improving the ability to see what is wrong.

    Agentic operations raise the next question: What happens after the system decides what should be done?

Autonomy Is a Set of Permissions, Not an On-Off Switch

“Let the agent act” is too broad to be an operational policy. Reading interface status and rerouting production traffic are both actions, but they carry different risks.

A useful way to think about agent authority is as a progression:

  1. Observe. Gather live information and assemble the context for an operator.
  2. Recommend. Propose a specific action and explain the evidence behind it.
  3. Execute with approval. Prepare the action, then wait for a person to authorize it.
  4. Execute within policy. Complete a narrowly defined action without prior approval, verify the outcome, and report what happened.

An organization can use all four levels at once. A diagnostic agent might gather evidence autonomously. A low-risk remediation might be preauthorized for a particular device class and time window. A change with a wider blast radius might always need human approval.

This is how to read the report’s finding that 82% are comfortable with at least some production changes occurring without prior approval. It does not imply a mandate for unrestricted agents. It suggests that network leaders are willing to assign authority to specific actions when the boundaries are clear.

Moving an action up that ladder should be an evidence-based decision. Did the checks identify the right condition? Did execution behave consistently? Did post-change validation show the expected result? Could the team recover quickly when it did not?

Those are operational questions. A more capable model cannot answer them on its own.

Give Every Agentic Action a Production Contract

The report says enterprises want explainability and guardrails. To turn those requirements into an operating model, I would ask seven questions of every action an agent might take:

What Is Its Scope?

Which systems, environments, and operations may the agent access? An agent diagnosing a WAN issue should not acquire every available tool simply because those tools might prove useful someday.

What Evidence Does It Need?

Decisions should use current network and service state, with the source and freshness of that information visible to operators.

Who Granted Authority?

Is this action covered by a policy, or does it require approval from a person? That answer should be determined before the agent reaches the decision point.

How Will It Execute?

Prerequisite checks, change steps, and constraints must run consistently. An agent should not be able to skip a safety check because it believes the answer is obvious.

How Will Success Be Verified?

A successful API response does not prove that the service recovered. The workflow needs post-change checks tied to the intended outcome.

What Happens If It Fails?

The team needs a defined stop, escalation, or recovery path. Where rollback is appropriate and available, it should be part of the procedure.

What Record Remains?

Operators need to reconstruct the trigger, evidence, decision, approval, action, and result. That record matters during an incident and long afterward.

This contract is what makes greater autonomy possible.

  • 💡 Without a production contract, “human in the loop” can become a person approving an opaque recommendation.

    With it, a team can make an informed decision about when a human approval adds value and when a well-bounded action can proceed automatically.

Let Agents Adapt & Keep Execution Dependable

Network incidents do not follow a single script. An agent may need to investigate several plausible causes, compare evidence, and change its plan as new information arrives. That flexibility is useful.

Production changes need a different quality. A pre-check must run when required. A policy must apply every time. The system must know who initiated the action and whether it completed successfully.

The answer is to connect adaptive reasoning to dependable execution. Let an agent assess the situation and choose among authorized capabilities. Let governed workflows carry out the actions those capabilities represent. Connect both to the monitoring, inventory, ITSM, security, and network systems that provide context and receive updates.

That separation gives an agent room to reason without making every production step improvisational. It also gives operators a consistent control model whether the work starts with a person, a scheduled process, an event, or an agent.

Where Itential Fits

This is the operating model Itential helps infrastructure teams put into practice.

FlowAI, the agent harness of the Itential Platform, lets teams build agents that reason through operational goals and use tools explicitly selected for them. Those tools can include existing Itential workflows, API integrations, existing automation scripts, external MCP servers, compliance actions, and other agents. The agent’s ability to act is therefore tied to capabilities the team has chosen to expose.

That matters because most enterprises already have useful automation. They have workflows for diagnostics, configuration changes, approvals, service provisioning, and validation. The path to agentic operations should make those investments more valuable. A proven workflow can become an authorized action an agent invokes when the situation calls for it.

Itential also helps connect actions across domains. An incident response may start with an observability signal, consult inventory, open or update an ITSM record, run network checks, apply a change, and return the outcome to the systems and people following the incident. Orchestrating that sequence is how an AI recommendation becomes an operational result.

Governance runs through that execution path. Access controls determine which capabilities are available. Workflows can incorporate approval and validation steps. Agent Sessions give operators a chronological view of an agent’s reasoning and tool calls, with links to workflows it started and the ability to intervene while a session is running.

Itential’s role is specific: it connects AI reasoning to governed, cross-system action. Monitoring and AIOps tools still provide much of the signal and diagnostic context. Good inventory and source-of-truth data still matter. No execution platform can compensate for a remediation procedure that has never been tested.

That is why the value proposition goes beyond adding an AI interface. It is about giving an agent a way to do useful work through the controls an infrastructure team needs in production.

A Customer Example: Building the Loop Before Removing the Human

Lumen’s network automation journey shows how these pieces can fit together.

Lumen uses Selector to consolidate operational signals and surface likely issues, Itential to orchestrate governed actions, and ServiceNow as a point of access for operational processes. Its published case study describes more than 350 live workflows and an AIOps effort that reduced over one billion raw alerts to 57,000 actionable incidents.

The distinction between those achievements is important. Reducing alert noise helps a team find the work that matters. Reusable, governed workflows provide a way to act on it. Lumen’s stated direction is to expand machine-to-machine operations as particular actions prove reliable in particular contexts.

That progression is more useful than announcing a target percentage of “autonomous operations” with no explanation of what earns an action that status. A team can examine the device family, environment, time window, quality of the signal, failure history, and recovery path. It can remove an approval for a defined slice of work while retaining visibility, audit, and the ability to intervene.

  • 💡 The control does not disappear.

    Its location changes: from an operator deciding every step in real time to an operating model that defines, tests, and monitors which steps may run without that operator.

Start With One Action You Can Measure

The report says 84% of respondents expect an AI-led operating model within twelve months. That is a statement of intent, not evidence that every organization will reach the same destination on the same schedule. It does, however, make the near-term planning question urgent.

I would start with one recurring incident that has a recognizable signal and a known response. Give the agent read-only access to the context it needs. Ask it to assemble the evidence and recommend a specific action. Then connect that recommendation to an existing workflow with human approval, pre-checks, execution, and post-checks.

Measure the whole loop: time from alert to diagnosis, time to resolution, false-positive effort, approval delays, failed changes, recovery events, and whether the action record is complete. Once the team has enough evidence, decide whether that particular action can run without prior approval under a narrower policy.

That is a concrete path from AI assistance to governed autonomy. It also gives executives and practitioners a shared way to discuss progress: actions completed safely, time returned to engineers, and authority expanded on the basis of results.

The Cisco and Omdia research makes clear why network teams want agents to do more than advise. The opportunity now is to make agentic action dependable when it reaches production.

The question I would put to any network operations team is simple: Which action would you trust an agent to complete tomorrow, and what evidence would you need before giving it more authority?

👉 Read the full report from Cisco and Omdia

👉 Learn more about FlowAI’s Agent Harness

Headshot of Karan Munalingal, SVP of AI Strategy and Innovation at Itential, driving AI-driven automation strategy that helps global customers modernize and scale network and infrastructure operations.
Karan Munalingal is the SVP of AI Strategy & Innovation at Itential. Previously, Karan ran systems engineering at Ciena, focusing on carrier ethernet and core switching platforms. At Itential, Karan drives AI strategy enabling global customers to adopt AI-driven automation journeys that modernize and scale network and infrastructure operations.
Keep Learning

The Latest in Agentic Operations

Get Started

Agentic infrastructure operations starts here.

See how Itential connects AI reasoning to governed execution across your entire infrastructure.