Itential shipped its own Grafana dashboards for Platform monitoring, and the lab finally gets a way to see whether four agents and a growing pile of workflows are actually healthy.
Home Lab Series · Part 5 of 6
1. Closed-Loop Config Backup2. First FlowAI Agents3. Continuous Compliance4. Ticket-Driven Diagnostics5. Platform Observability6. Coming soon
Four agents deep into this lab, and I hadn’t built a single way to check on any of them beyond opening a terminal and reading logs by hand. That gap sat there through the last several entries, mostly ignored, until Itential shipped the answer for me: official Grafana dashboards for Itential Platform Monitoring. The first thing I did once I saw it was wire it up and see what it actually tracks.
Put simply: Itential Platform Monitoring is a pre-built Grafana dashboard, published straight to the Grafana Marketplace under dashboard ID 25527. Point it at a Prometheus data source already scraping the stack, and five tabs, Overview, Workflow Engine, Platform Process, Redis, and MongoDB, cover the health of the whole deployment. No Loki, no Elasticsearch, no custom PromQL to write before the first panel lights up.
Every prior agent in this lab had something to point at, NetBox, a live Arista switch, a ticketing API. This entry doesn’t add an agent at all. It adds a way to see the platform all of them run on. Itential Platform Monitoring showed up in the Grafana Marketplace as a finished, official dashboard, not something to hand roll from a blank panel. Search for it by name, or grab it directly by dashboard ID 25527, point it at a Prometheus data source, and the import is done. Setup was a same-day task, not a project.
The screenshots below are pulled from a home production stack I run on Proxmox, separate from the Arista devstack this series has been following, a deployment I’ll get into in more detail down the road. I chose it here because it gives a better picture of what these tabs can actually show than the quiet Arista lab would: real cluster topology, real job history, real resource numbers, instead of a handful of idle containers with nothing much to report.
Overview rolls up the four core components, the platform application, the workflow engine, Redis, and MongoDB, into a single UP or DOWN tile each, alongside active jobs, job error rate, active sessions, total task throughput, and per-node CPU, RAM, and disk.
Workflow Engine breaks the same deployment down by job activity: job start, completion, and cancellation rates, a job success rate gauge, jobs currently in progress, job records by status pulled straight out of MongoDB, and a further layer of task-level detail, task start versus complete, task distribution by server, total tasks, task error rate, and watcher reconnects.
That distinction matters more than it looks. The gauge isn’t a running lifetime average, it’s a trailing 5-minute completion rate, sampled at the latest point inside whatever range you’ve got selected. The MongoDB table underneath it is cumulative. Read them together, and five quiet minutes doesn’t get mistaken for a broken pipeline.
Platform Process is the tab that doesn’t stop at “the platform is up.” It breaks the underlying Node.js process apart: how many applications and adapters are running, how much memory and V8 heap each is using, thread and open file descriptor counts, and CPU and memory broken out per individual application and adapter.
V8 Heap Used % holds steady near 95%, which looks alarming at a glance, but Heap Used vs Total and Process CPU (User/System) tell a calmer story once you look at the absolute numbers behind it, nothing here reads as a leak or a runaway process. Look at the legend on Top Applications by CPU and Memory, too: every component is named Itential, the same internal engine tag that’s been showing up in every agent’s raw session trace since Part 3.
Scroll further down the same tab, and Platform Process breaks out a full component resource table, every Itential-named application and its adapter counterpart, one row per component, each with its own CPU %, memory %, resident set size, thread count, open file descriptors, and uptime.
The Redis tab doesn’t stop at a single UP tile. Cluster Members lists every node individually, one MASTER and two REPLICAs in this deployment, each with its own health status and master link state, alongside clients connected and blocked, memory used against the configured max, queue operation throughput and latency, commands per second, and total keys stored.
MongoDB gets the same treatment. Cluster Members lists the primary and both secondaries by hostname, replica set name, and health, next to a plain Replica Set Status tile, replication lag, operation rates, connection counts, cache utilization, and page faults.
The health roll-up logic underneath all of this is more careful than a simple ping check. Redis only reports DOWN when no primary is reachable at all, a downed replica, like the two shown above staying healthy on their own, shows DEGRADED instead. MongoDB works the same way: DOWN only when no primary is elected, not when a single secondary drops out of the replica set. That’s the difference between an alert that gets a team’s attention and one that trains them to ignore the channel entirely, and it’s built into the dashboard itself, not something bolted on for this deployment.
Observability
GrafanaPrometheusNode ExporterItential Platform Monitoring (ID 25527)
Itential Platform
FlowAgentsAgent ProjectsStudio WorkflowsItential Gateway
Devices
Arista cEOS-lab (arm64)Containerlab
Agentic Ops
Claude Sonnet 5Itential MCP Server
One entry left in this series, and it’s the one every prior agent has been circling without crossing. Every FlowAgent so far has stopped short of the device, on purpose, propose a fix, flag a drift, never touch a config. The final entry is where that boundary finally moves, a governed agent that can make a real change to the network, with a human still standing in the approval path.
Want to follow along with me? Connect with me on LinkedIn.
See how Itential connects AI reasoning to governed execution across your entire infrastructure.