...
Itential Platform Pricing Explore flexible plans and options for your team
Itential logo
Demo

Choose Your Model, Not Your Vendor: Running FlowAgents on Any LLM, Anywhere

A demo of FlowAgents running on any model, open weight, proprietary, or something your team built in-house, so your compliance team stops being the reason your AI roadmap stalls.

 
Headshot of John Capobianco, Head of AI and Developer Relations at Itential, helping organizations adopt AI safely in network automation with deep experience across enterprise, government, and cloud networking.

John Capobianco

Head of AI & Developer Relations
Headshot of Joksan Flores, Principal Solutions Engineer at Itential, advancing infrastructure automation through AI-driven orchestration with 10+ years of networking architecture experience at Cisco.

Joksan Flores

Principal Solutions Engineer

How to Adopt AI When Your Data Can’t Leave the Building

For a lot of teams building agentic AI right now, the model is the platform. Pick a provider, send your infrastructure data to their cloud, and everything downstream, every agent, every workflow, is locked to that one choice. That works fine until it doesn’t: a European data residency requirement, a regulated industry review, a compliance team that needs to know exactly where a model runs and what it can see, or simply a better model shipping from a different provider six months from now.

This isn’t a fringe concern. Cisco’s 2026 Data & Privacy Benchmark Study found that outright bans on GenAI tools dropped from 28% to just 7% in a single year, not because the risk disappeared, but because blanket bans don’t work and organizations shifted to governing AI at the point of use instead. A 2026 Cloudian survey of enterprise IT decision-makers found 93% had already repatriated AI workloads from public cloud, were in the process of doing so, or were actively evaluating it, with data sovereignty as the leading driver. Nobody’s banning AI anymore. They’re demanding control over where it runs.

That’s a compliance question about one model on one cloud, not a rule about AI itself. Once the model can run entirely inside your own infrastructure, the answer changes from “we can’t” to “not with that model, but yes with this one.”

Powered by FlowAI, the agentic harness of the Itential Platform, FlowAgents run on any LLM: open weight models like NVIDIA Nemotron, Kimi K2, Llama, or OpenAI’s gpt-oss running on infrastructure you control, proprietary models like Anthropic’s Claude, OpenAI’s ChatGPT, or Google’s Gemini, or a fully custom model your team built and fine-tuned in-house. The model is a choice you make per agent, not a commitment you make for the whole platform. Every agent executes through the same governed path: scripts and automations running through Itential Gateway, direct API calls, or platform workflows, with the same RBAC, approvals, and audit trail underneath, regardless of which model did the reasoning.

What You’ll See

    • A FlowAgent handling a task where data has to stay put entirely, deployed against an open weight model (Nemotron, Kimi K2, or Llama) running fully inside the environment, with nothing leaving the building.
    • A different FlowAgent in the same platform reasoning through a more complex task against a proprietary model, side by side with the first, same guardrails underneath.
    • Swapping the model behind an existing agent, no rebuild, no re-architecture, just a configuration change.
    • The audit trail: RBAC, approvals, and logging applied identically no matter which model made the call.

Why You Should Watch

If “we can’t do AI here” is the answer your compliance or security team has been giving, or you’re an AI or platform team tired of your agent roadmap being hostage to one provider’s pricing and release schedule, this session is for you. You’ll leave understanding what exactly changes that answer, and how to make the case to the people who need convincing.

Why Itential

Most agentic platforms pick a model and build the whole product around it. Itential was built the other way: FlowAI is the harness, the model is a plug-in. Use Itential’s own default out of the box, or bring your own, open weight, proprietary, or built and fine-tuned in-house. That’s not a bolt-on feature, it’s how the platform is architected. No re-platforming when a better model ships. No exception process to run one agent differently than the rest. Just a setting.

Why It Matters

Betting an entire agentic platform on one model provider creates two problems at once: a compliance problem, if that provider’s cloud isn’t where your data is allowed to go, and a strategic problem, if a better or cheaper model ships somewhere else next quarter. Choosing per agent instead of per platform means the compliance team gets a model that never leaves your infrastructure where that’s required, and the AI team gets the best available model everywhere else, without anyone re-platforming to get either one.

+

John Capobianco • 00:04

Welcome, everybody. This is going to be a really exciting episode and podcast and webinar, whatever you want to call it. I’m joined again by Joksan Flores, our principal solution engineer. I’m the head of AI in Dev Rel, John Capobianco. And today, we’re going to address probably the number one thing that we encounter from prospects, from customers, from the industry. I read about it all the time. It comes up.

John Capobianco • 00:25

It came up at my talk in Autocon, one of the questions from the audience. What if we can’t use a cloud provider? What if our industry prohibits it, like other cloud activities, especially with AI? We can’t rely on a cloud provider. What if we don’t have those agreements in place yet with the hyperscalers, maybe to keep our stuff private, or we don’t have a strategy around this yet? Is this going to slow us down? And it’s slowing down a lot of industry.

John Capobianco • 00:54

We at Itential have a sort of bring-your-own model approach, and we’re going to expose that today, right? And choose your model, not your vendor. So, you don’t have to get locked into a vendor. You can start right away, especially with, and we’re going to have a demo today, things like read-only activities, things like ticket enrichment, things like basic troubleshooting. The model from the open weight community and the open source community has really gotten pretty incredible. I wouldn’t say it parody at Joxon with you, but they’re not far behind those frontier and foundational models.

Joksan Flores • 01:29

Yeah, I think the commercials, the commercial models are still winning that battle by a tiny bit, but the open weights have caught on greatly. I think, John, when the 1st time you and I were talking about this year and a half ago, I think Lama, what, 3.1 or something like that? And it was okay. It used to do some stuff, but it wasn’t quite at the reasoning depth that you need when you have semi-complex use cases. But, man, I’m going to tell you, like the last month or so, I’ve been very impressed, and there’s more coming out every day.

John Capobianco • 01:58

And they have things like tool calling capability. So, that to me has been a gap. Even with modern open source models, like I’m not going to pick on one, but let’s say Gemma 4. Gemma 4 is remarkable at reasoning and at inference, but it doesn’t have the tool calling capability, right? So, in our world, the ability to call MCPs or call CLI or interact with skills and tools, it’s very important. And those other models, some of these other models have that capability, right? So, let’s, I’m going to start sharing the screen and we’re going to do a little slide wearing out to maybe ease people into this.

John Capobianco • 02:31

If open weight and open source and frontier and foundational, there’s a lot of discussion around this. We’re going to try to go slow. And we’re actually going to show you in practice how, with the Flow AI builder, that you can bring your own model at build time, right? So, another thing on a lot of people’s minds, Joxon, is tokenomics. And I don’t want to maybe gloss over that, but. Some of the tokens, right? We thought tokens were going to go down, and we also used to have an open buffet.

John Capobianco • 03:02

Where you could just a la carte go to each buffet and fill up your plate and eat as much as you want for a certain fixed fee. Now it’s more like you have to weigh your plate, and depending on what you’ve put on your plate per serving, you have to pay for that now, right? So the whole model has shifted and tokenomics have become extremely important. Obviously, because of the scale of human usage, right? A lot of people are using this across enterprises now. And is it a per-monthly token budget? Is it what is that budget?

John Capobianco • 03:34

What if the model changes and there’s different pricing? It’s so fluid and in flux all the time, Joksan. Whereas you can take control of your own destiny in terms of zero cost or somewhere in the middle, maybe you don’t have the hardware on-prem, but you find a middle ground where someone in the cloud is offering that hardware and you put the model you want on it, right?

Joksan Flores • 03:52

Yeah, there’s a lot of reasons, right? And tokenomics is a big one, right? So one of them is people want to control their destiny, which you mentioned before, right? We talk a lot to financials, energy companies and things like that. They don’t want to deal with commercial or cloud providers of LLMs. So they want to control their destiny, but tokenomics is a huge one. And you get a lot of bang for your bucks, especially.

Joksan Flores • 04:13

And one of the things that, John, we can talk about later that that’s not probably scoped in, so I might be getting myself a little bit in trouble. But we can talk about how in the Flow AI platform, we allow you to actually segment agents out and you can actually mix and match different providers, right? So you can use really good stuff for the fancy things that you need to do, and you can use open source models for the other more basic things. But yeah, tokenomics is a big deal, and especially things are so expensive nowadays. It’s crazy.

John Capobianco • 04:43

Yeah, and I’ve turned to it during my dev cycles, even as I’m developing my flow agents. I’m using open source models during my entire dev cycle. And then I sort of shift gears. It’s like, is this ready for production? Maybe now I can switch to a higher capable model. So let me start sharing the screen and we’ll get into some of these slides to bring everyone up to speed. And let me make sure.

John Capobianco • 05:05

Sorry about that. And let me hide this. Okay. All right, perfect. So choose your model, not your vendor, running flow agents on any LLM anywhere. I’ve written a blog about it two years ago now, actually, Joxon, about fine-tuning Llama 3.1. It’s funny you mentioned that, but I was into a fine-tuning phase with open source models.

John Capobianco • 05:25

It’s quite remarkable what you can do on local hardware. To remind people of a reasoning and action agent, you hear a lot about agents. We like to clarify and actually be very specific that it’s a React agent. And there’s Arvix papers about this. But the idea is that the agent has reasoning capability and the ability to call tools. Joxin, maybe do you want to break this down for everyone?

Joksan Flores • 05:50

Yeah, and I think it’s super important. You mentioned this earlier, right? There are some models that don’t have the capability, especially in the open weight world, to do to perform tool calling. And that’s very important, right? Because those are models that are only suited for generative AI, right? For chat applications, building images or code or what have you. And here, we’re specifically talking about React agents.

Joksan Flores • 06:09

So we need a reasoning loop cycle that can do tool calling. And this is what Atinsil does, right? What we implement in Flow AI is exclusively React agents. The whole idea there is that we leverage an external LLM to the platform. We expose from the platform perspective the tooling that the platform provides. That could be either an API call that comes in through an integration. That could be a script or Ansible playbook or an OpenTOFU plan.

Joksan Flores • 06:37

Or it could be a workflow that you build on the platform that becomes a deterministic tool that you expose to that agent. You give it a prompt, you tell the agent what it needs to do, right? Its goals and its reasoning goals. And then you also pass in: we have parameters, we can pass variables and things like that. And then you let the agent rip. And essentially, the agent goes and does its thing, right? I can give it a series of steps to go troubleshoot an interface or an IP.

Joksan Flores • 06:59

That’s kind of the example that we’re going to pick doing today. Or we could do, hey, go and provision me a DNS record or an EC2 instance, or go and enrich this ticket with some data from the environment, right? There’s all sorts of things that you could do that leverages the reasoning aspect of the agent, but then the tooling itself resides on the attentional platform. And then we just expose it and making it a React agent.

John Capobianco • 07:20

What I love about the platform is the ability to scope and choose tools at build time. Some of us have even been building these sort of agents before MCP was the standard. Some of us have been actually building agents and trying to do this the hard way with different frameworks and different harnesses. But you always came up against, well, I imported an MCP server, it’s got. 74 tools, and now I’m having a tool selection problem, and it’s hallucinating, and I’m not getting good outcomes. And now I’ve tried to add four MCPs, and each of them has 30 or 40 tools. All those tools are cataloged, and there’s a point-and-click exercise.

John Capobianco • 07:57

And I made some videos about importing the Catalyst Center MCP server into the gateway and exposing the tools to this build. And that tools box, you literally check the tools you want and bind them to your agent. This entire box. Has RBAC and AAA around it. So the agents I build here at Itentral are exposed to the solution engineers group, but the other groups don’t have access to my agents, Joxon. Like all of this is secured and behind the appropriate enterprise-level guardrails, right? Yep.

John Capobianco • 08:27

Yep. So we mentioned this about the deterministic execution, right? So what I love about my favorite thing about the combination of AI with classic network automation, and it’s almost classic network automation now, is that we have a workflow and you’ve got a starting, a start, and an end point. You understand your logic, everything goes great, and then you start hitting errors or edge cases or things you maybe have missed. Or maybe there’s something like the API you’re using has a rate limit that you didn’t realize. And now you’re starting to have to handle rate limits, maybe pagination, on and on, right? And that’s just a simple API doing crud activities.

John Capobianco • 09:06

Well, you can keep those workflows. And I’m not saying throw away all that stuff, but the reasoning is where it steps in, right? Run this workflow. Here’s my A to B to Z steps. Oh, you hit an error at step J. Here, the reasoning can step in and work around that, right? The fix, the patch, the back off timer, whatever it is dynamically with the reasoning, right?

John Capobianco • 09:33

And it’s not one or either or both. It is a spectrum here, right? So we have totally perfect workflows that do not need reasoning, maybe to augment them a little bit, maybe with ticket enrichment or communications enhancements, or like I said, error handling, or purely reasoning situations, like maybe a triage where we don’t have anything that’s deterministic because we don’t know what’s coming in the ticket. So we can use that reasoning to suss out root cause or probable cause or what systems to test or what logs to collect for the human. And you can see that there is some tokenomics associated with this as well, right, Joksan?

Joksan Flores • 10:11

Yeah, and 100%, John. You said that the tokenomics and this is a bunch of aspects and also the learning curve, right? The build a learning curve, right? We have customers that with workflows, it takes them a little longer to get used to the canvas and do the data manipulation bits and all that stuff, right? There’s work, right? I think our CTO calls it. We’re still building code, even though it’s slow code.

Joksan Flores • 10:31

But then with the agents, it’s just natural language. I’m just spelling steps, right? If I know the mop to do something, to troubleshoot something, or if I have the mop in a Word document, I can just paste it in. And now I did some work. And then, like you said, John, then we start marrying the workflows, which are super deterministic and cheap to run with the reasoning. And now you can make super awesome, super powerful use cases.

John Capobianco • 10:54

And the other thing that we’re going to show today is when you see Flow AI agents’ cost to run high, we have tokens in brackets there on purpose because they can be free to run. And we’re going to show you that today or absolutely little, as little cost as possible. If you have OLAMA, e.g. , right? There’s a couple of things I want to mention about open source and some of this cost to run being high. We have an asterisk. That means tokens. It doesn’t necessarily equate to dollars.

John Capobianco • 11:20

So don’t think that if you’re getting into the agent building business, that it’s going to be expensive necessarily.

Joksan Flores • 11:27

Yeah, on this one, we could go really quick, John. I think this is just a recap of the entire architecture. And whenever we’re talking about building agents or building platforming, right, all the things that we’re including inside of the, you know, the offering on the platform, these are things that you have to build some way, one way or the other. If you look at the integration layer at the bottom, right, you talked about talking to APIs or be able to execute scripts at scale or talking to network devices over SSH or executing MCPs. All those things we include. Imagine having to build this stuff on your own, right? That’s a lot of work.

Joksan Flores • 12:01

Even if you vibe code, you still have to own the lifecycle. So we also always do a refresher on all the capability we have to offer. But then there’s also the bottom bits there, right? We talked about secrets, single sign-on, RVAC, all those kinds of things, which belong to this whole thing. So we’re going to be very focused on agents and models and tool call today. But this is the entire framework that makes that possible.

John Capobianco • 12:23

Let’s get into the demo. And let me stop sharing my screen. But yeah, it’s quite remarkable. As soon as I heard, I got wind quite quickly that you had got this, you had some really solid demos with open weight models. And right inside the platform, everyone, I’m going to highlight this. This is the new skin. And you can see that we have a model registry.

John Capobianco • 12:45

And this is when we say BYOM, this is what we mean. And these translate into drop-down options at build time in the flow. Or if you’re using our anthropic skill, our builder skill, you can actually have the builder skill reason about what model is available in your model registry to suitable for the job that the flow AI agent is being built, right? So I just want to mention that I love our builder skill at the CLI. Yes, you can use the GUI, and our GUI is incredible, not to take away from it. But that experience of having an LLM actually help you and say, I want to build this age/slash/flow agent is the DNS agent I want to build. Pick the most suitable model given the task, right?

John Capobianco • 13:30

But yeah, let’s take a look. Let’s take a look at what you got. That’s a great point, John, before we go into this, too.

Joksan Flores • 13:36

Your mileage may vary, right? This is a very much a thing for each different persona, right? Some people love the GUI, some people love the terminal. I live in the terminal for the most of the day, but then there’s also the UI to do things for our users. Some of our users, like I said, they just come with a mob. They don’t have a lot of AI experience, they don’t have API experience and so forth. And we try to make it as easy as possible for them.

Joksan Flores • 13:56

Like you said, tick boxes for the tool selection. You get to type in your prompt, you get to make it pretty well marked on if you want to. The model registry is all right here. We try to make it as straightforward as we can. Okay, so let’s get into this. So, this is a model registry. And as you can tell here, it probably doesn’t mean a whole lot to a lot of people, but once I start breaking it down, then you’d start understanding, right?

Joksan Flores • 14:16

Like, we have gone to town here, John, quite a bit with model selection. So, we have a series of models at the top here that are essentially what we use for our lab. So, we have an attention-managed offering because this is in our cloud. So, for some of our customers that consume the attention provided LLM, this is what they would see. And then we also have our own lab offering. So, I have one that’s direct through our cloud, and then I have one that goes through our lab gateway. So, this will connect into Anthropic today only.

Joksan Flores • 14:46

But then the interesting stuff starts happening down here. So, when you start looking at everything that I got down here, John, configured, we got bedrock onboarded. We got Open Router and laptop Olama. So, these two, of course, is very much me going to town as the nerd and the focus of I think today’s session because I’ve done quite a bit of testing. And when you start looking at these two, actually, specifically, these are going through my laptop. And I have a gateway that’s running on my laptop, has an active gateway server process, and it’s connecting into our cloud environment. Now, that all happens through a mutual TLS channel.

Joksan Flores • 15:23

And the reason why I want to bring that up is because we have a lot of customers in the energy sector, in the finance sector, in the enterprise sector that want to host their own Olama LM Studio, or they have a Microsoft Foundry local setup, or perhaps they’re using the NVIDIA NIM platform when they have their own GPUs on-prem, and they don’t want to expose that API to the internet. So, we actually have a gateway, and the communication comes from the gateway. The gateway is hosted on-prem. And by the way, in addition to this, I actually have my own home PC here that has a 4080 RTX 4080 in it. And I have connected it through my gateway and running on my Mac. So, my gateway running on my Mac, talking to my PC 4080. Running an attention clock, right?

Joksan Flores • 16:08

So you get all the connectivity through a mutual TLS channel. It’s all very secure, certificates provisioned, and that interaction happens all that way. Okay, so these two are the ones that we’re going to focus on today. I have Open Router set up, and Open Router is a great way of getting started for people, right? They actually have work using a commercial paid-for offering here, right? But in the case of you having your own GPUs, you can actually run this for free. And they have a variety of models that they run.

Joksan Flores • 16:35

If you look at here, I have a bunch selected, John. I’m actually going to go and edit this here and go fetch models. They have a lot of stuff that they expose today, right? From the they let you use things like OpenAI all the way to using OpenWave models. And I have, I think, one OpenAI model selected. Everything else that I have is Muse Glimmer just came out, I think, today, John. Meta Muse Glimmer.

Joksan Flores • 16:55

Okay. So they have already a crow there. I did some testing with it. I need to do more. They got Kimi, they got so.

John Capobianco • 17:03

I just want to pause here. It’s literally like you’re going model shopping and you just click the knob and away you go, right?

Joksan Flores • 17:09

Yeah. So this is exactly from the model registry perspective. If you look at, I actually logged in, I glossed over it, but this is already provided credentials and it’s got all the secrets. So the key and the URL get set up for this. So I pointed to Open Routers API. I provide my API key and I clicked on models and it gives me everything and I get to enable them. And John, one of the things also talk about the platform Ebits and the other stuff, right?

Joksan Flores • 17:33

If I scroll all the way to the bottom, this is hundreds of more, I actually have access group control. For the models as well. On key, and I am with the searching boon, and the firewall team brings their own key. They get to have their own model registry instance separated. Nobody can reuse. There’s a lot of ways you can skin this whole thing, but there’s just a lot of levels of control and protection, and everybody can set up. Really cool, really cool.

Joksan Flores • 17:59

Yeah. Any other highlights here, Don, on the model registry?

John Capobianco • 18:02

No, I love that it shows the number of agents being used. It’s great for audit. Maybe you start seeing models that have less and less agents and you start to clean up and drop connections or maybe you start to notice trends here, right? There’s a lot of metadata available here. And all of this has a REST API too, right? So if you wanted to call this through an API and give you a nice dashboard and whatever, Grafana or something of your model usage and all that stuff, it’s all behind APIs. The whole platform is REST API driven, right?

Joksan Flores • 18:31

Yep. That’s a great point for model deprecation, especially, right? Things like Sonnet 4.6, I think, goes and does support at some point in October or November, something like that. So model deprecation, you could do the same thing for the OpenWay models and all that stuff. And we didn’t talk a lot about fine-tuning. You talk about fine-tuning and you have a blog, John. We should probably link those.

Joksan Flores • 18:49

But when you start fine-tuning, you can actually take some of these open weights. The base of a definition behind OpenWait is that you can customize it, right? You can make it your own. So if you choose to fine-tune it, you can host it in here as well. And like I said, we have all the provider examples. So you can create all these. Most of them fit in the OpenAI compliant one, right?

Joksan Flores • 19:06

Olama fits NVNIM, all those open router as well.

John Capobianco • 19:10

I think that’s worth repeating for people that because there are industries, there are places, and I think it’s going to happen more and more over time, is that they’re actually fine-tuning their own model. But those are hard to. Let me just say that sometimes fine-tuning a model is hard to expose through some of the other harnesses. It’s not as easy as just loading up your fine-tuned model sometimes, but we can bring them into this platform as well. So, that because maybe you have your guardrails, your documentation, your change management approvals, your basic ACLs, everything should have whatever you want to bake into a model. So, the model knows that without doing rag or external inference, right? So, these show up in the agent project in the builder, right?

John Capobianco • 19:53

So, once we have our model registry, we put in our prompt, we pick our tools like I suggested, and then we actually drop down and pick our model.

Joksan Flores • 20:04

Yep. So, we have a simple agent that we’re going to use for this, John. I kept it very tactical, didn’t expose it to any other pieces of the platform. Of course, a lot of these variables on the prompt can be templatized, and that gets exposed via our operations manager, which you get a pretty form and all that from an operator perspective. We are builders here. John and I are builders today. So, we’re going to assume that everything is provided for testing purposes.

Joksan Flores • 20:25

So, I’m giving a lot of static data here. I got a decently complex workflow, right? I have five steps that the agent has to do. It’s got to ping inventory, it’s got to grab inventory, it’s got to check connectivity to a particular route on some devices, and then it’s got to present me with a load report and then send an email as well. From a tooling perspective, I have selected a variety of tools, John. So, I have Python scripts, a couple of Python scripts in here. This is an API call to the platform.

Joksan Flores • 20:49

This is a human-in-the-loop task, and then I also have a workflow. So all this stuff that we’re putting the agent through, right, this model through, it’s actually something that it’ll have to do everything that the platform will do. So I thought baseline, John, and I don’t know what you think here, but I think for baseline, we can run this on GPT-56 Terra, which is the middle-of-the-line kind of Sonnet flash Gemini flash competitor, right? It’s a foundational, so we can have a comparison, right? So we run that guy and let it run for a minute and see what happens, right? So this is, of course, making API calls. Like I said, it’s set up for open router today.

Joksan Flores • 21:25

And we’re just going through their proxy to, in this case, to the GPT API. But we’re going to go ahead and switch it once this is done. So let’s go and this is kind of our baseline of the experiment.

John Capobianco • 21:39

This is so cool. This is so cool. Yeah, the agent. And it’s calling tools. And it’s calling tools. Yeah.

Joksan Flores • 21:45

Yeah. Yeah. So this is running GPT. It’s running, what’s calling tools. It’s supposed to be troubleshooting a route. It’s got a couple devices I have access to. It’s got a Linux host and it’s got a CE and it’s got a workshop in a human and loop activity.

Joksan Flores • 21:57

So I’m going to go ahead and open that. And here we go. Let’s see. Okay. So we got a simple review the completed connectivity assessment. Site A host successfully reach the route. Right.

Joksan Flores • 22:08

One at one. Okay, perfect. So no packet loss. Everything is great, John. So we should get a similar result out of our open weight test. Okay. So let’s go ahead and complete that guy.

Joksan Flores • 22:18

Our agent execution will finish here. So we’re going to ignore that one and we’re going to close it. So now, and this is John, how complicated it is when you did Lang Chin and ADK and all those things? How tricky was it to switch a model?

John Capobianco • 22:33

No, yeah, it’s not fun. It’s not fun at all. Some of them have different schemas entirely.

Joksan Flores • 22:38

A lot of code changes. I remember early on, things have gotten better, but a lot of them, if I didn’t take models for the API and all that kind of stuff, in here, we’re trying to bring everything into the harness, right? So I just said edit agent. I scroll all the way to the bottom. I got my tools selected here. Nothing changes. My prompt, my tools, nothing changes.

Joksan Flores • 22:55

And I’m going to go ahead here. And in open router, I got all my lists that comes from my model registry that I can see because I’m in solutions engineering. And I think let’s pick, let’s pick on Nemotron Ultra. Nemotron Ultra, even though it’s a high model, we can use any of these. I have tested with this specific agent as well. I have tested GPT-OSS 120, Quen3, 235. I’ve tested Nemotron Super 120, which they all seem to do real well with this use case.

Joksan Flores • 23:25

And your mileage may vary with some of these, John, right? Like I’ve done testing on running models that are 35 billion quantis running on my laptop. And it works. The speed suffers a little bit, of course, because you’re sharing memory and things like that. And there are some situations where you may want to use one agent versus the other or combine them, right?

John Capobianco • 23:44

But there’s also an aspect sometimes. It’s funny when we say when we use the right code, sometimes you don’t care about the outcome. Do you know what I mean? Sometimes you’re just testing a workflow. And does the agent get from me clicking run agent to a report of some kind? Do all the tools get called? Did I miss anything?

John Capobianco • 24:02

So to that point, sometimes you’re not interested in what the, I’m not going to say slot, but whatever the outcome is, if it’s not a high grade, it’ll get better when I change models. What I’m trying to prove is that my workflow works, right?

Joksan Flores • 24:16

Yeah, that’s true, John. That’s a good point. Actually, I hadn’t thought about it that way. This is actually, I like this. It’s very organic, right? Because I hadn’t thought about the fact that you can use that for testing and actually get an almost like an inverse positive effect where if your prompt works with a weaker model, it’ll be amazing with a really great model.

John Capobianco • 24:37

Absolutely. Absolutely. Yeah. And maybe it doesn’t have markdown or this or that, but did you get through your entire workflow, especially if you’re tying to something that’s deterministic? And maybe it’s your 1st time running the agent. You just want to get a green light at the end.

Joksan Flores • 24:50

Yeah, of course. All right, let’s go ahead and run it. Yeah, let’s check this out. And we could run it however many times, because John, you saw, like, I just clicked a couple of times and I just switched the model. So if we want to run it multiple times, we can run it.

John Capobianco • 25:02

And there’s the tool calling. So that’s important that this model is capable of calling tools. We can see this in the logs.

Joksan Flores • 25:08

Yeah. And also the visibility, John. I think we highlighted this last time, right? The visibility that you get at each step of the way. You get to see the tool calling in detail every time.

John Capobianco • 25:17

And I know that we’re enhancing this more and more all the time, right? And tokenomics is going to become part of our dashboards and some of our information. What might be neat, Jox, when we can report back to the developers is a dollar saved or emissions saved, right? How many, how much energy or how much money did we save using an open source model, right?

Joksan Flores • 25:36

Absolutely. And these are actually pretty cheap to run. Good thing that you say, because I’ve been experimenting a fair bit to identify, hey, what’s the best way to showcase this, right? And I think I spent John between, I’ve probably done 100 runs in the last couple of weeks or so of these different variety of agents that we have. And they’re demo agents, right? The majority of them are not creating new ones. I’m just taking them and switching them over from Sonnet 5 or whatever.

Joksan Flores • 25:59

And I think I’ve spent 60 cents in or so because I’m using a variety of them, but some of these 30 billion parameter models are like, gosh, like fractions of a cent. They’re like 0.00028. Cents per run. So, very cheap. Like you said, I think for development and even from production for a simple here, we’re running Nematron 550, which is like top of the line from an open way perspective. And I think this is what most people want to run. It’s a great example that, hey, this will succeed at some of these use cases that are fairly complex, right?

Joksan Flores • 26:36

They’re getting string data back, like some of this command outputs, right? They’re getting the entire route entry command from V, but it’s got to reason through and so forth.

John Capobianco • 26:45

And I love that we’re doing this live. It’s not edited. We’re not going to cut this down. We’re going to really see how long this runs and we can just keep talking over it. We can see that we’re at about 1200 output tokens right now. I think all the tools have been called because they get called in parallel, which I love that we do parallel calls and we’re not just, and they’re all asynchronous, right? So you saw some of these tools finish before other tools.

John Capobianco • 27:06

The other thing, and I don’t want to blow people’s hoo away or go too fast, but let me just maybe paint a picture to you where, there we go. So it’s waiting for tool output. We’re going to open it in Work Center. Look at that. More way richer output, too. It’s even better output. Yeah.

John Capobianco • 27:25

And this took minutes. We were just talking through it the whole time. Look at this. And look at that. It’s the same info, right? In terms of the Pepsi challenge, it did get the same info, but even more rich data. Same tool calling, different model, right?

John Capobianco • 27:38

Yeah. So this is now that costs, that costs a nickel, something like that, right?

Joksan Flores • 27:44

Probably something like that or less. And mind you, this is one of the, like I said, flagship and it gives you a likely cost of what’s happening. It’s actually interesting, right? Because I didn’t tell it a whole lot about complaining about the CE not being able to reach the internet. You and I know, John, we live in the networking world. The CEs shouldn’t be able to reach the internet. So this is expected.

Joksan Flores • 28:03

But it’s actually got, because I didn’t tell it, it’s like, hey, your host can reach, which sits behind. So it starts going. Oh, hey, the trace route says we’re going through the CE. Look, it identified it. I didn’t tell it anything. I gave you information. It’s like, hey, I identified this.

Joksan Flores • 28:18

It’s going through the CE, but the CE can’t ping. This is the likely cost. You need a default route. It starts going through all these crazy things.

John Capobianco • 28:24

Unbelievable. What a great demo. What a great demo. And would you mind? So that’s under agent sessions, everyone. And if we go back to the main route of agent sessions. You’re going to see that it’s all there.

John Capobianco • 28:37

And all this is, again, behind APIs. So you can write some or have Claude write some additional scripts to poll our platform and Grafana this. How many agent runs? How many failed? What ones are failing? Why are they failing? What was the reasoning?

John Capobianco • 28:51

Who’s using the most tokens? It’s all there, everyone. Joxon, maybe do you mind stop sharing your screen for a sec? I’ve got some thoughts about maybe where people could go with this as well. We’ll do. We can, that agent you just built could be part of a larger deterministic workflow, meaning we can have agents call other agents and have agents interact, right? Correct.

John Capobianco • 29:12

Yep. So there’s no reason why you couldn’t have some, just like you would organize humans, your triage agent, your junior agent, they do the best they can with free or very inexpensive models. And maybe a supervisor agent at the end that really looks at this stuff that is more expensive. That is, you have the option to mix and match and give agent different capabilities and then stitch the agents together where maybe the agent punts, I’m not capable enough. I’m going to send this off to a more capable agent in the chain, right?

Joksan Flores • 29:46

Yeah. And that’s something that we’re seeing actually. The pattern has been emerging, which is very encouraging for me. And every time I talk to customers about these patterns, it’s moved from less of, oh, I want to use AI for doing chat body stuff to now people actually want to do those things. So, John, so they want to during our pod planning. I can’t say too much because a lot of this is under NDA, but a lot of our customers come to us and say, hey, I want to do X, Y, and Z. And if we can’t find a particular root cost where we can’t get to a point where we’re giving, we give it a runbook, right?

Joksan Flores • 30:16

We give the typically the agent a mob. And if it can’t get to the point where there is enough information that it can make an assessment of. Either open a ticket very enriched with a lot of information for tier two people or even mount it to a more highly capable model, like you just said. So, ticket enrichment alone, right? I like to call it a tier one looker, a level one engineer. And it’s, it is what it is, right? You talked about digital co-worker, right?

Joksan Flores • 30:39

You use that phrase a lot, and it truly is, right? This is like doing stuff that a CCNT or CCNA or a junior engineer would do, right? These are the kind of things and just make them way more powerful and useful. Even these people can use these agents to even learn from them.

John Capobianco • 30:55

I was reading some great paper talking about how AI is actually going to, so it’s the inverse of the fear and the haters and what they’re, you know. About junior engineers. This paper was actually saying the inverse: that AI is making junior engineers as capable as senior engineers. Without the 25 years of experience, because they’re asking the right questions, and the agents are so capable, and they have so much access to internal data and databases and APIs. And so, juniors are actually elevating much faster and doing much more difficult work than in the past. I also love this because, honestly, a lot of conversations I’ve had, some people are let down or they feel deflated that their industry or their company is never going to be able to embrace this. They’re not going to be able to enrich themselves and augment themselves because they happen to work in healthcare or pharmaceuticals or finance or manufacturing or government or education.

John Capobianco • 31:54

There’s many sectors that want this. They’re clamoring for it. So, as soon as I got wind of this, I said, let’s have a podcast. We’ll show the people how this is done in Flow AI. And, you know. It’s really exciting. It’s an exciting time because you can have distribution of this.

John Capobianco • 32:11

So, all of your network engineers have it local and also centralized in the platform, right? So, you can do local experiments with free, cheap models, get the agents working perfectly, and then promote them up into the platform as production agents, and you switch out the model, right?

Joksan Flores • 32:29

So, Joksan, do you want to leave everyone with anything? No, I think this is great, John. I think that’s a good wrap-up. I think this there’s work to be done for these things, right? You can get the most expensive, let’s say it the other way around, right? Even with the best model, you can still screw things up. And with the cheap model, you can do a lot of stuff if you know how to do it properly.

Joksan Flores • 32:47

So, I think as a learning tool, at the very least, learning how to prompt properly and also using the proper guardrails, as you see, right? We don’t let a lot of room in the attention platform for mishaps, right? Because we have a lot of guardrails and everything is very much purposely done, right? Build time, tool selection, all those kinds of things don’t let my agent run a mock. But that’s the whole idea here, right? It’s like you have the flexibility to appease your security and privacy concerns and the need for executing LLMs on-prem, or the data that you’re sharing to the LLM can’t go to a cloud, all the way to, hey, I just want to play with something that is on my laptop that I got to train myself in. So, this gives you a lot of flexibility.

Joksan Flores • 33:27

We’re having these conversations every day. This is really becoming very apparent, right? Just said that there’s multiple reasons: taxonomics, privacy, data latency, and so forth.

John Capobianco • 33:35

I’m going to put you on the spot here for our next, maybe not our next video, but another one down the road. And what we’re seeing more and more, and I’ll leave this with people to think about. We’re seeing more and more vendors releasing small language models as open weight models. We’ve seen this from Cisco releasing a CVE small language model. I don’t know how many, I think it’s 3 billion parameters, something like that. It’s a small language model. Those, we can bring them in with the techniques that you’ve seen today.

John Capobianco • 34:01

So, if you want to actually start bringing in vendor-released models, those are again bring your own model. We showed you how to bring them into the platform today. So, Joxon, why don’t you and I pull that model from Hugging Face and spin up a Flow AI agent that is powered by the Cisco. CVE model.

Joksan Flores • 34:22

I bet we’re going to save a lot of tokens parsing through pieces. Let me tell you that much. So I’m very excited. I actually like that idea.

John Capobianco • 34:27

All right. I got you. I’ll put you on the hook for next time. Everyone, thanks for joining us. This is really exciting. Please share and circulate this video, especially if you’re in those industries that feel a little bit left out or feel that this is a tool that you want to embrace but have not been able to due to governance or security or privacy concerns. And we’re here to work with you.

John Capobianco • 34:50

Jox and I will do demos. We will meet with you. Please reach out. Please put stuff into our calendars. Please put stuff into our mailboxes. We’re both available on LinkedIn and we have incredible content on the itential.com site. Thanks again.

John Capobianco • 35:02

We’ll see you next time.

Keep Learning

The Latest in Agentic Operations

Frequently Asked Questions

+

Through model hosting, the same way you’d connect to any model provider: Amazon Bedrock, OpenRouter, Hugging Face, or Ollama for fully on-prem deployment. Whichever one fits where that agent needs to run.

+

No. Every agent executes through the same governed path regardless of which model is reasoning behind it: RBAC, approvals, and a full audit trail apply identically either way.

+

Yes. A custom or fine-tuned model built in-house is hosted the same way any other model is. The platform doesn’t care who built the model, only that it’s reachable and governed like everything else.

+

No. The model is a configuration choice on the agent, not something baked into how the agent was built. Swapping it is a setting change, not a rebuild.

+

Yes, that’s the point. One agent can run on a model hosted entirely on your own infrastructure for a task where data can’t leave, while another runs on a different model for a task that needs more reasoning power. Same platform, same governance, different models.

+

No. Regulated industries feel the compliance side of this most acutely, but any AI or platform team that doesn’t want its agent roadmap tied to one model provider’s pricing and release schedule gets the same benefit.