A Dashboard Read "Connected." The Link Was Dead.
A WAN link with one wrong digit in its gateway address. The dashboard read Connected. None of the tools were broken — each reported exactly what it was built to report, and the answer they produced together was wrong.
Earlier this week, a site we look at had a WAN link configured with one wrong digit in its gateway address. One character. Off-subnet, unreachable, dead.
The dashboard reported that connection as Connected. Latency read one millisecond — the number you get when nothing is being measured at all. A real link on that circuit reads about twelve.
Nothing alarmed, because the backup link silently carried the traffic. And every inbound sales call to that business rode a double-NATed satellite connection, which does not present as an outage. It presents as the phones are weird sometimes. A lost job looks like a customer who changed their mind.
We want to be careful about the lesson here, because the obvious one is wrong. None of the tools were broken. Every one of them reported exactly what it was designed to report. The monitoring platform checked whether the device answered; it did. The dashboard showed the interface status field; it said connected. The RMM agent was healthy, because the machine it ran on was fine.
Good products, working correctly, producing an answer that was wrong.
Why a bought integration stops where it stops
This is arithmetic, not laziness. A vendor builds one integration that has to serve every customer they have. So it reads the handful of endpoints the average customer needs, and it stops there. That is the only rational place to stop when you are building one thing for everybody.
The problem is that the interesting state — the settings that actually break things — tends to live a tier deeper than the average customer needs. On one controller class we work with there are effectively four levels of access: the current documented API, a second API for a different subsystem, a legacy interface where port forwards and firewall rules actually live, and shell on the box. Every off-the-shelf tool we have worked with stops at the first.
Their firewall makes the same point from the other side. Plenty of firewalls log to disk; theirs did not, and had no analyzer either, so its entire record lived in a volatile memory buffer that is wiped on reboot. Pulling that record out over the API gives you log history the appliance itself cannot keep. None of their products were going to do that — not because the products are deficient, but because none of them knew that was what was needed there.
Which is the actual point, stated plainly: whatever the gap turns out to be at a given site, we can fill it. Not a category of gap we have prebuilt. The one in front of you.
None of this is a complaint about the tools, and we would not want it read as one. They are genuinely useful, and we would not necessarily tell anyone to stop paying for them. But useful and demonstrably valuable are not the same thing. Useful is a product doing what it says on the box, for everybody who bought it. Valuable is that product answering the question you actually have, at your site, on the day it matters. The distance between those two is exactly where mileage varies, and closing it is the work.
What we actually do about it
We connect to the equipment at the tier where the answer lives, and we check the thing the business actually depends on rather than the thing that is easy to poll.
Two examples that run today. The first checks that the login page of the platform a business runs on renders — the real page, the real content — rather than checking that a server answered, because a SaaS outage that still returns a healthy response is invisible to everything else in a stack. The second writes site documentation from live calls to the equipment, and marks every line as either confirmed on the device or assumed. Documentation that quietly presents a guess as a fact is the same failure as a dashboard reading connected.
The answer is not a bigger tool either
The fair objection is that the category already exists: ship everything into a SIEM, correlate across sources, and the contradiction surfaces. That is true, and we would rather concede it than pretend otherwise.
But a platform like Splunk is built and priced for organizations with a security operations function. Entry tiers run into the thousands a year before anyone logs in, and pricing turns on ingest volume or on compute units that are hard to forecast until you are already committed. For a business with a handful of sites and no SOC, that is a bridge spanning a canyon to cross a creek.
And it would not have caught this one for free. A SIEM would have ingested the same wrong answer. The interface status field said Connected; that is the value that ships into the index. Nothing in the platform knows that one millisecond is an impossible reading on that circuit — somebody has to know it, and write the rule that says so.
Which is the cost argument in one line: the heavier tool rarely removes the build. It relocates it, and prices it. Building into the specific gap costs a fraction of licensing a platform that spans every gap, most of which you do not have.
Reading is the easy half
The same access that pulls a firewall's log history out of volatile memory can write a rule into it. That is the direction this goes, and we would rather describe it plainly than imply it already ships. None of the following is built yet:
Call routing that changes on a schedule and puts itself back, with something confirming the change actually took — and the same mechanism moving traffic off a degrading link instead of alarming about it. Firewall rules that carry an expiry, so nobody discovers a temporary any-to-any three years later during an audit. Remote execution on an endpoint, scoped to the right machines, recorded against a named person and a reason, and checked afterward. Drift detection across sites that were built identically and have not been identical since. Joiner, mover and leaver as one run across mail, directory and network access together, ending with the check that the person is genuinely gone everywhere. Tickets that arrive with the evidence already inside them — not device down, but device down, the WAN event that preceded it, the config change before that, and the last three times it happened.
Why we can say that without flinching
Every one of those follows the same shape, and the shape is the product: proposed as a plan, approved by a person, applied, then verified independently — because a success response means the call was accepted, not that the intended thing is true.
That verification is Fidenta, and it sits at the centre of everything we build. Fidenta checks the plan before it runs and confirms the implementation actually completed after. It is not a policy layer that can be argued with; where a capability should not exist, the mechanism for it is not built, so it cannot be granted in a hurry at midnight by someone tired. Fidenta keeps the work accurate. Governance and our AI management system keep it controlled.
Holding deeper access than a vendor would give you is a real obligation, and we would rather be judged on how we discharge it than on how much we can reach.