Before You Blame the ISP, Build the Witness
After a lightning strike, a homeowner was sure his cable provider was the problem. The honest answer was "maybe, but not yet." Here's the order I worked it in, and the monitoring, documentation and password handling I built on Cloudflare so the next answer is evidence instead of opinion.
It started with a lightning strike. The house took the hit, and the network never quite recovered. Devices fell off, the internet felt worse than it should, and the homeowner landed where everybody lands. It's Cox.
He'd already done the reasonable thing. In a house over 6,500 square feet, a single router was never going to be enough, so he'd bought good equipment: six Omada Wi-Fi 7 access points. That was the right instinct. He just didn't go far enough.
It might be. But "it's the ISP" is a claim about the far side of the modem, and it only holds if the near side is healthy. Nobody had looked at the near side. There was nothing there to look with.
The network had no witness
The access points were cabled into an unmanaged switch. Traffic moved, and nothing reported. From the controller's point of view, the network was six radios and five clients.
Five. In a house full of garage door controllers, thermostats, irrigation, and whole-home audio.
The first fix was not configuration. It was instrumentation. I had him buy an Omada Access Pro PoE switch, and I adopted it into his existing Omada cloud site next to the access points. Now there was one pane that could see the switch, the radios, and every client.
The finding was a setting, not a failure
The SSID was set to WPA3-only, with protected management frames required and Wi-Fi 7 multi-link operation on. On paper, that's a modern, secure config. In practice, a lot of smart home devices ship with 2.4 GHz radios from the WPA2 era. They don't fail loudly on a network like that. They just never associate, and so they never show up anywhere as a problem.
I moved the SSID to WPA/WPA2-Personal, turned off MLO and the 6 GHz band, and kept the network name and passphrase identical, because he's used them for years and every device in the house knows them.
Connected clients went from 5 to 33. Then I went back through signal strength until every device sat at a level I'd trust, not just a level that connects.
None of the equipment was broken. Every product was doing exactly what it was configured to do. The config was the problem, and nothing in the stack was built to say so.
MAC addresses are not documentation
A client list full of MAC addresses tells you almost nothing. So I worked every one: OUI lookup for the manufacturer, hostname, which access point or switch port it was on, and then a match to the actual physical thing. Where it mattered, I confirmed it on the device itself. A thermostat's MAC, read off the thermostat's own screen, is a fact. A guess from a hostname is an assumption, and the record says which is which.
Every device got a real name in the controller.
The record is a page, not a file
Documentation handed over as a PDF is out of date the moment something changes, and it tends to get lost. So the source of truth for this house is a private web page: gateway, switch, access points, Wi-Fi settings, every labeled client, open items and next steps.
It sits behind Cloudflare Access. The homeowner enters his email address and receives a one-time passcode. No account to create, no password to forget.
Passwords that are never on the page
He wanted his passwords in the record. I didn't want them stored in the record. Those two requirements turned out not to conflict.
The secrets live in Bitwarden Secrets Manager. The page gets its own read-only machine key, scoped to that one home's entries and nothing else. If that key ever leaked, it could expose only those few entries, and it can't write anything.
When he presses Reveal, the page's server validates his Cloudflare Access token properly (genuine, current, issued for this app, and belonging to an allowed address), and only then pulls that single secret from the vault and returns it. The value is never cached, never in the HTML, never in git. It re-masks after two minutes. The Wi-Fi row also renders a join-by-camera QR code, drawn in his browser, so the passphrase never goes to a third-party QR service. Every reveal is logged: which secret, when, and from what device.
The easiest secret to protect is the one that isn't sitting there.
Then, and only then, watch the ISP
With the inside healthy, "is it Cox?" is a fair question again, and it deserves data instead of a feeling.
His Cox gateway doesn't expose any management ports to the internet, and I turned on WAN ping response so it could be measured from outside. Cloudflare Workers can't send ICMP, so the probe runs on a small cloud server: every two minutes it sends 20 pings to his line and posts latency, jitter and loss to a Worker I call jarvis-extmon. That Worker writes to its own D1 database, rolls samples up into days and outages every ten minutes, and publishes a stripped-down payload to KV. A separate public page reads that payload and nothing else. It shows latency, jitter, loss and outages. It never shows his IP, his name, or how the probe works. His page stays private to him, but a page like this is the one I run for my own Cox and Starlink links.
Two details make it worth trusting.
Vantage matters. The first version probed from a box on my own home connection, which meant every reading included my ISP's behavior too. I moved the probe to a cloud server in Dallas so the numbers describe his line, not mine.
Blind is not down. Every run also pings a reference target. If his line doesn't answer but the reference does, that's an outage on his side. If neither answers, the probe itself was offline, and that sample is marked blind and never counted against his provider. An outage report that can't tell those apart isn't evidence. It's a second opinion from a witness who was out of the room.
Let it bake
Now it runs. Over the next several weeks, the data turns into a baseline. If Cox is the problem, he'll have a dated, timestamped record to put in front of them. If it isn't, he avoids switching providers and still having the same problem.
The next phase is already written down in his record. It covers equipment that can carry a multi-gig plan when he wants one, and a second WAN with automatic failover if uptime ever matters enough to pay for. Neither is urgent. Both are decisions he can make later with a baseline in hand instead of a feeling.
That's the whole method, and it isn't clever. See the network. Fix what's actually broken. Name everything. Write it down somewhere it stays true. Then measure the thing you were going to blame, from a vantage you can defend.
The ISP might still be the answer. But now it'll be an answer.