...
Itential Platform Pricing Explore flexible plans and options for your team
Itential logo
Home Lab Series: Part 5

Building Platform Observability With Grafana & Itential

Itential shipped its own Grafana dashboards for Platform monitoring, and the lab finally gets a way to see whether four agents and a growing pile of workflows are actually healthy.

Headshot of author and solutions engineer elliot conner
Elliot Conner
Solutions Engineer
Last updated September 30, 2026

Home Lab Series · Part 5 of 6

1. Closed-Loop Config Backup2. First FlowAI Agents3. Continuous Compliance4. Ticket-Driven Diagnostics5. Platform Observability6. Coming soon

Key Takeaways

    • Itential Platform Monitoring is now a standing Grafana dashboard (Marketplace ID 25527), imported directly rather than hand built, running against a live Prometheus-scraped deployment.
    • Five tabs cover the deployment end-to-end, Overview, Workflow Engine, Platform Process, Redis, and MongoDB, all on Prometheus alone, no Loki, no Elasticsearch, no custom PromQL required to see the first panel.
    • Redis and MongoDB tabs surface real cluster topology, not just a single UP or DOWN tile: one master plus two replicas, one primary plus two secondaries, each individually tracked.
    • Platform Process breaks the deployment down to the Node.js process itself, 23 applications, 4 adapters, memory and V8 heap per component, CPU and open file descriptors down to individual Itential-named applications and adapters.
    • Health roll-ups are cluster aware by design: Redis only reports DOWN when no primary is reachable, MongoDB only reports DOWN when no primary is elected, so a single downed replica shows DEGRADED instead of a false outage.

Four agents deep into this lab, and I hadn’t built a single way to check on any of them beyond opening a terminal and reading logs by hand. That gap sat there through the last several entries, mostly ignored, until Itential shipped the answer for me: official Grafana dashboards for Itential Platform Monitoring. The first thing I did once I saw it was wire it up and see what it actually tracks.

Put simply: Itential Platform Monitoring is a pre-built Grafana dashboard, published straight to the Grafana Marketplace under dashboard ID 25527. Point it at a Prometheus data source already scraping the stack, and five tabs, Overview, Workflow Engine, Platform Process, Redis, and MongoDB, cover the health of the whole deployment. No Loki, no Elasticsearch, no custom PromQL to write before the first panel lights up.

Terms Worth Knowing First

  • Prometheus
    A time-series metrics database that scrapes numeric data from exporters on a fixed schedule. The only data source this dashboard needs.
  • Exporter
    A small process that exposes a system’s internal metrics, CPU, memory, job counts, in a format Prometheus can scrape. Node Exporter covers host-level stats here, alongside Itential Platform’s own built-in metrics endpoint.
  • Cluster-Aware Health Roll-Up
    Status logic that only reports DOWN when an entire redundant system has actually failed, no reachable primary, so a single downed replica shows DEGRADED instead of triggering a false outage.
  • Dashboard ID
    Grafana Marketplace’s identifier for a published, vendor-provided dashboard. Import by number instead of hand building panels from scratch.

Five Tabs, One Import

Every prior agent in this lab had something to point at, NetBox, a live Arista switch, a ticketing API. This entry doesn’t add an agent at all. It adds a way to see the platform all of them run on. Itential Platform Monitoring showed up in the Grafana Marketplace as a finished, official dashboard, not something to hand roll from a blank panel. Search for it by name, or grab it directly by dashboard ID 25527, point it at a Prometheus data source, and the import is done. Setup was a same-day task, not a project.

The screenshots below are pulled from a home production stack I run on Proxmox, separate from the Arista devstack this series has been following, a deployment I’ll get into in more detail down the road. I chose it here because it gives a better picture of what these tabs can actually show than the quiet Arista lab would: real cluster topology, real job history, real resource numbers, instead of a handful of idle containers with nothing much to report.

Everything Green, At A Glance

Overview rolls up the four core components, the platform application, the workflow engine, Redis, and MongoDB, into a single UP or DOWN tile each, alongside active jobs, job error rate, active sessions, total task throughput, and per-node CPU, RAM, and disk.

Fig. 1 – Four green tiles, four active sessions, and per-node CPU and RAM tracked across two separate nodes in this deployment.

Real Numbers, From Real Runs

Workflow Engine breaks the same deployment down by job activity: job start, completion, and cancellation rates, a job success rate gauge, jobs currently in progress, job records by status pulled straight out of MongoDB, and a further layer of task-level detail, task start versus complete, task distribution by server, total tasks, task error rate, and watcher reconnects.

Fig. 2 – Job Success Rate reads 0% because nothing finished inside this particular window, not because anything’s broken. Job Records by Status pulls a longer history straight from MongoDB: 480 complete, 46 errored, 41 canceled.

That distinction matters more than it looks. The gauge isn’t a running lifetime average, it’s a trailing 5-minute completion rate, sampled at the latest point inside whatever range you’ve got selected. The MongoDB table underneath it is cumulative. Read them together, and five quiet minutes doesn’t get mistaken for a broken pipeline.

Down to the Node.js Process Itself

Platform Process is the tab that doesn’t stop at “the platform is up.” It breaks the underlying Node.js process apart: how many applications and adapters are running, how much memory and V8 heap each is using, thread and open file descriptor counts, and CPU and memory broken out per individual application and adapter.

23
Applications
4
Adapters
3.23 GiB
App Memory
309
Threads
1.74K
Open FDs
Fig. 3 – Twenty-three applications, four adapters, down to individual CPU and memory lines. V8 Heap Used %, Heap Used vs Total, and Process CPU are all populated here, this deployment has had time to warm up.

V8 Heap Used % holds steady near 95%, which looks alarming at a glance, but Heap Used vs Total and Process CPU (User/System) tell a calmer story once you look at the absolute numbers behind it, nothing here reads as a leak or a runaway process. Look at the legend on Top Applications by CPU and Memory, too: every component is named Itential, the same internal engine tag that’s been showing up in every agent’s raw session trace since Part 3.

Scroll further down the same tab, and Platform Process breaks out a full component resource table, every Itential-named application and its adapter counterpart, one row per component, each with its own CPU %, memory %, resident set size, thread count, open file descriptors, and uptime.

Fig. 4 – Twelve of the tracked components, Itential core down to individual services like WorkFlowEngine and ToolRegistry, each pinned to a node and port, each running for the same 1.73 weeks.

Redis, Down to the Replica

The Redis tab doesn’t stop at a single UP tile. Cluster Members lists every node individually, one MASTER and two REPLICAs in this deployment, each with its own health status and master link state, alongside clients connected and blocked, memory used against the configured max, queue operation throughput and latency, commands per second, and total keys stored.

Fig. 5 – One master, two replicas, all reporting UP, alongside live queue operation and command throughput.

MongoDB, Down to the Replica Set

MongoDB gets the same treatment. Cluster Members lists the primary and both secondaries by hostname, replica set name, and health, next to a plain Replica Set Status tile, replication lag, operation rates, connection counts, cache utilization, and page faults.

Fig. 6 – One primary, two secondaries, replica set status reading Healthy in plain text, not just a green tile.

Cluster-Aware, On Purpose

The health roll-up logic underneath all of this is more careful than a simple ping check. Redis only reports DOWN when no primary is reachable at all, a downed replica, like the two shown above staying healthy on their own, shows DEGRADED instead. MongoDB works the same way: DOWN only when no primary is elected, not when a single secondary drops out of the replica set. That’s the difference between an alert that gets a team’s attention and one that trains them to ignore the channel entirely, and it’s built into the dashboard itself, not something bolted on for this deployment.

What’s Actually Watching the Watchers

Observability

GrafanaPrometheusNode ExporterItential Platform Monitoring (ID 25527)

Itential Platform

FlowAgentsAgent ProjectsStudio WorkflowsItential Gateway

Devices

Arista cEOS-lab (arm64)Containerlab

Agentic Ops

Claude Sonnet 5Itential MCP Server

Where the Lab Goes Next

One entry left in this series, and it’s the one every prior agent has been circling without crossing. Every FlowAgent so far has stopped short of the device, on purpose, propose a fix, flag a drift, never touch a config. The final entry is where that boundary finally moves, a governed agent that can make a real change to the network, with a human still standing in the approval path.

Want to follow along with me? Connect with me on LinkedIn.

Headshot of author and solutions engineer elliot conner
Elliot Conner is a Solutions Engineer at Itential, where he helps enterprises turn fragmented network automation into governed orchestration. A CCNP Enterprise certified network automation engineer, he has built more than 70 automation tools for multi-vendor networks using Python, pyATS, Ansible, and Nornir. He documents his own Itential Platform lab builds in public, testing the platform the way customers actually deploy it.
Keep Learning

The Latest in Agentic Operations

Get Started

Agentic infrastructure operations starts here.

See how Itential connects AI reasoning to governed execution across your entire infrastructure.