The fifth agent in this lab finally gets write access to the network, but only through a gate that always stops for a human, composed almost entirely from pieces this series already built.
Home Lab Series · Part 6 of 6
1. Closed-Loop Config Backup2. First FlowAI Agents3. Continuous Compliance4. Ticket-Driven Diagnostics5. Platform Observability6. Governed Remediation
That agent is built now, and it’s the fifth and last one in this particular lab’s arc. Every prior agent stopped at the same boundary on purpose. NetBox Source of Truth just answers questions. Arista Device Ops reads live state. Network Compliance and Network Diagnostics can both tell you exactly what’s wrong and why, but neither one can touch a device. Network Remediation is the one that finally can, and the design question I sat with before writing a line of workflow logic for it wasn’t whether that could be automated. It obviously could. It was how much of that decision the agent should actually get to own.
Every prior agent in this lab stopped at the same line: NetBox Source of Truth, Arista Device Ops, Network Compliance, Network Diagnostics, four different jobs, and not one of them could change anything on a device. Here is that lineup as it actually sits in the platform.
Catching drift only matters if something eventually fixes it, so the fifth agent couldn’t stay read-only. That’s the part I sat with before writing any workflow logic, not whether closing that loop could be automated, it obviously could, but how much of that decision I was willing to let the agent own.
I landed on the same answer twice. Once for the workflow’s own write step, and again for whether the agent could trigger it without asking: never autonomous. The agent gets to look at a device, compare it against NetBox, and decide for itself whether a fix is even safe to propose, drift confined to one field, isolated, nothing broader going on. But the moment an actual config push is on the table, a real person has to approve it. No exceptions, no trusted devices, no autonomous mode to flip on later once the agent’s earned some notional trust.
What surprised me was how little new engineering that guardrail actually took. The Itential Platform already had the pieces sitting there. Building Network Remediation meant composing them, not building an approval system from zero.
The NetBox-authority drift check, the same logic that compares a device’s hostname, primary IP, and Ethernet1 IP against NetBox’s declared record, already existed from the Network Compliance agent back in Part 3. Network Remediation just reuses it. Work Center already had a real queue for paused, human-facing tasks, so a form asking approve this change, yes or no, drops straight into it, no new interface to build. On the FlowAI side, turning that whole workflow into something the agent can call was just a matter of pointing it at one tool: Pre-Change Validation Gate. The platform handles the part where the agent’s session sits and waits, however long a human actually takes to respond, then picks the result back up and reports on it honestly.
gate_result values count as a real, verified push.That last constraint is written directly into the prompt, not left to hope: NEVER claim a device was changed unless the tool's own gate_result for that device is exactly 'pushed_and_verified'. Everything else, rejected, no_action_needed, blocked_broader_drift, pushed_but_still_drifted, has to be reported plainly, with no implication that anything succeeded.
The run below is a real session, not staged. I gave the agent one instruction and let it work.
One tool call, and it’s deceptively simple looking in that log. Behind it, the gate checked all five devices, found two with drifted hostnames, and then just waited. Two separate approval cards landed in Work Center, one for ceos2 and one for ceos1, each one blocking that same tool call until a person looked at it.
hostname ceos2 line via send-config, nothing else. Reject, and the device is left untouched.The second card, for ceos1, read almost the same, except this one named the actual drifted value directly: core1 -> ceos1. Readers of Part 2 might recognize that name. Back then, Arista Device Ops pulled ceos1’s running config and noticed its configured hostname was actually core1, and called the mismatch normal, expected drift in a lab where an inventory label doesn’t always match a device’s own hostname setting. That was true then. It stopped being true the moment Network Compliance made NetBox the fleet’s declared authority in Part 3. What counted as a harmless quirk two entries ago is exactly the kind of narrow, single-field drift Network Remediation exists to catch and fix now.
Both approvals landed, and the session picked the tool call back up right where it left off. What came back wasn’t a vague confirmation, it was a fact-by-fact table, one row per device, with the gate’s own verdict for each.
pushed_and_verified, the only gate_result the agent’s own prompt allows it to describe as a successful change.dist1, access1, and isp1 had no drift on any of the three tracked facts, hostname, primary IP, Ethernet1 IP, so the gate took no action on any of them. ceos1 and ceos2 each had exactly one field out of sync, and only that field, isolated enough to count as safe. The gate proposed a fix, waited on a human for each, and independently re-verified the change had actually landed before reporting pushed_and_verified for both. Nothing here came from the agent’s own say-so. Every line in that table traces back to the tool’s real output.
Governance
Work CenterPre-Change Validation Gategate_result
Itential Platform
FlowAgentsAgent ProjectsNetBox-Authority Drift CheckItential Gateway
Devices
Arista cEOS-lab (arm64)Containerlab
Agentic Ops
Claude Sonnet 5Itential MCP Server
Five agents, each one building on the last: a source of truth, then read access, then compliance, then diagnosis, and now a gated write path. That’s the arc I set out to build, and it’s done. But the real point was never whether an agent could reason its way to a good answer, any capable model can do that with enough context. It was whether the platform underneath it could let that agent act safely once it had one, scoped tools, a full audit trail, a human standing at the one step that actually mattered. That’s not a policy I added to make this lab look responsible. It’s the default Itential already builds for, on a five-node home lab or on infrastructure at a scale I’ll never personally touch. What’s next is bigger, a multi-vendor lab in EVE-NG, closer to what a real production network actually looks like. The same pattern comes with it, propose first, always gate the write behind a human. The near-term focus shifts from building new agents to getting that environment stood up and watching how this approach holds up somewhere a lot less tidy than a five-node lab.
Ready to see what I build next? Connect with me on LinkedIn.
See how Itential connects AI reasoning to governed execution across your entire infrastructure.