You Bought the Monitoring. It Told You Everything Was Fine.

A WAN link with one wrong digit in its gateway address. The dashboard read Connected. Latency read one millisecond — the number you get when nothing is being measured at all. None of the tools were broken, and the answer they produced together was wrong.

You Bought the Monitoring. It Told You Everything Was Fine.

Earlier this week, a managed site turned out to have a WAN link configured with a single wrong digit in its gateway address. One character. Off-subnet, unreachable, dead.

The dashboard reported that connection as Connected. Latency read one millisecond — which, if you think about it for a second, is the number you get when nothing is being measured at all. A real link on that circuit reads about twelve.

The site kept working, because the backup link silently carried everything. Nothing alarmed. Nobody called. And every inbound sales call to that business rode a double-NATed satellite connection, which does not present as an outage. It presents as the phones are weird sometimes. A lost job looks like a customer who changed their mind.

Here's the part worth being careful about: none of the tools were broken. Every one of them reported exactly what it was designed to report. The monitoring platform checked whether the device answered. It did. The dashboard showed the interface status field. It said connected. The RMM agent was healthy, because the machine it ran on was fine.

Good products had been purchased. They were all working correctly. And the answer they collectively produced was wrong.

If you run an MSP, or an internal IT team, or you are simply the person who gets called when the phones are weird — your stack would very likely have reported the same thing. That is not an insult. It is how the tools are built, and there is a reason for it.

The finding is never the typo

A mistyped octet is a two-second human event. It will happen again, to all of us, next year. That's not interesting.

What's interesting is that a WAN sat in a failed state and nothing in a stack somebody paid real money for was structured to notice. Not because the vendors are lazy — because of arithmetic. A vendor builds one integration that has to serve every customer. So it reads the handful of endpoints the average customer needs, and it stops. The interesting state — the settings that actually break things — tends to live one tier deeper than that.

On one of the controllers at that site there are effectively four levels of access: the current documented API, a second API for a different subsystem, a legacy interface where port forwards and firewall rules actually live, and shell on the box. Every off-the-shelf tool I've worked with stops at the first one. That's not a criticism. That's the only rational place to stop when you're building one thing for everybody.

Their firewall makes the same point from the other direction. Plenty of firewalls log to disk; theirs didn't, and had no analyzer either, so its entire record lived in a volatile memory buffer that's wiped on reboot. Pulling that out over the API gives you log history the appliance itself cannot keep. None of their products were going to do that — not because the products are deficient, but because none of them knew that was what was needed there.

And that's the actual point, which I'd rather state plainly than leave as an implication: whatever the gap turns out to be at a given site, it can be filled. Not a category of gap. The one in front of you.

None of which is a complaint about the tools, and I want to be careful not to let it read as one. They are genuinely useful, and I wouldn't necessarily tell anyone to stop paying for them. But useful and demonstrably valuable are not the same thing. Useful is a product doing what it says on the box, for everybody who bought it. Valuable is that product answering the question you actually have, at your site, on the day it matters. The distance between those two is exactly where mileage varies — and it's the only place worth building.

What actually changed

For twenty years, build-vs-buy for operational tooling was settled, and the answer was buy. Not because bought tools were better, but because the integration surface was where projects went to die. You'd write a connector, the vendor would change their API, and you'd own that forever. Buying was right because building never ended.

That arithmetic changed, and I don't think most people have repriced it.

Not because AI writes code faster — that's the shallow version. The expensive part of integration was never typing the code. It was the discovery: reading undocumented interfaces, working out which of four tiers holds the field you need, handling the auth quirk nobody wrote down. That discovery cost is what collapsed. And separately, vendors started shipping MCP servers — Microsoft ships one for Azure — which means the integration surface is increasingly something the vendor maintains for you rather than something you reverse-engineer from them.

So the line moved. It didn't disappear.

Where buying is still obviously right

I want to be honest about this, because the version of this argument that goes build everything is written by people who haven't operated anything.

Buy the thing where the vendor's scale is the product — threat intelligence, patch catalogs, anything whose value is that thousands of customers feed it. Buy the commodity where your requirements are genuinely average, because they usually are. Buy anything you'd have to staff a team to keep alive.

Build in one specific place: where the gap between what the API permits and what the product surfaces is costing you money. That's a narrow test and most things fail it. The WAN link failed it expensively.

And the answer isn't a bigger tool either

The reasonable objection to all of this is that the category already exists. Ship everything into a SIEM — Splunk, or any of its competitors — correlate across sources, and the contradiction surfaces. That's fair, and worth conceding, because it's true.

Two problems with it. The first is fit. Splunk is built and priced for organizations with a security operations function; entry tiers run into the thousands a year before anyone has logged in, and the pricing turns on ingest volume or on compute units that are famously hard to forecast until you're already committed. For a business with a handful of sites and no SOC, that is a bridge spanning a canyon to cross a creek. You pay for the span, you maintain the span, and the creek is still four feet across.

The second problem is the one that actually matters. A SIEM would have ingested the same wrong answer. The interface status field said Connected — that's the value that would have shipped, faithfully, into the index. Nothing in the platform knows that one millisecond is an impossible reading on that circuit. Somebody has to know that, and then write the rule that says so.

So the expensive version of this still leaves you building the detection. It just sells you a licence first and hands you a query language to build it in. That's the shape of the trade far more often than the category names suggest: the heavier tool rarely removes the build. It relocates it, and prices it.

Reading is the easy half

Everything I've described so far is a read. Ask a device what it thinks, compare that against what's true, raise a hand when the two disagree. It's useful. It's also the timid version, and I'd rather say out loud where this goes.

The same access that lets you pull a firewall's log history out of volatile memory lets you write a rule into it. Two things in that direction already run against my own estate, and both exist for the same reason the WAN story does — something reported a comfortable answer that wasn't true.

The first checks that the login page of the platform a business actually runs on renders — the real page, the real content — rather than checking that a server answered. A SaaS outage that still returns a healthy response is invisible to everything else in a stack. The second writes site documentation from live calls to the equipment, and marks every line as either confirmed on the device or assumed. Documentation that quietly presents a guess as a fact is the same failure as a dashboard reading connected, and it's most of why inherited documentation is so often worse than none.

The rest is conversation rather than product, and I want to be precise about that — none of what follows is built yet:

Call routing that changes on a schedule and puts itself back, with something confirming the change actually took — and the same mechanism moving traffic off a degrading link instead of alarming about it. Firewall rules that carry an expiry, so nobody discovers a temporary any-to-any three years later during an audit. Remote execution on an endpoint, where the half worth paying for isn't running the command but that it was scoped to the right machines, recorded against a named person and a reason, and checked afterward. Drift detection across sites that were built identically and have not been identical since. Joiner-mover-leaver as one run across mail, directory and network access together, ending with the check that the person is genuinely gone everywhere and not just from whichever console someone happened to open. And tickets that arrive with the evidence already inside them — not device down, but device down, the WAN event that preceded it, the config change before that, and the last three times it happened.

None of that is exotic. The access was always there. What changed is that reaching it stopped being the expensive part.

Every one of them has to follow the same shape, though: proposed as a plan, approved by a person, applied, then verified independently — because a success response means the call was accepted, not that the intended thing is true. Which brings me to the part I'd rather people copy than the rest of it.

Building buys you an obligation

I've written before about what I put around agents that can actually do things — scoped credentials, independent verification, tiered approval, checking the outcome rather than trusting the response code. I still believe all of it.

But I've come around to something stronger since. The best guardrail isn't a rule. It's an absent mechanism.

My deployment tool accepts the name of a thing to deploy and nothing else. There is no parameter that takes code. It is therefore not possible to deploy something that was never written down and reviewed — not discouraged, not policy-blocked, not possible. Where a tool was too powerful, I didn't switch it off behind a setting; I removed it, and a check fails the build if it ever comes back.

A permission can be granted in a hurry at midnight by someone tired. A mechanism that doesn't exist can't be.

That's the real cost of building, and it isn't the code. If you're going to hold deeper access than a vendor would give you, you owe the corresponding discipline: verify the plan independently before, verify the implementation actually completed after, and write the decisions down as they're made rather than reconstructing them later from memory.

If you sell software, the question flips

Everything above is the buyer's side. The mirror image is the more interesting one.

If the cost of customers building against your API just collapsed, three things follow. Your API stops being a checkbox and starts being product surface — the tier of access you expose becomes a pricing and positioning decision rather than an engineering afterthought. Partnerships get more valuable, not less, because the integration you don't have to build is worth more than the one you build badly. And the customers most likely to build around you are your best ones, the ones who've outgrown the average configuration — which makes their workarounds some of the highest-signal roadmap input available.

The uncomfortable version: somewhere in most install bases, a customer has already built the thing that should have shipped. Finding them is cheaper than a research program.