This week ANYbotics launched Shift, a platform for running fleets of inspection robots. One module takes sensor data from the robots and flags overheating, leaks and misalignment. Another, in the company's words, turns confirmed anomalies into work orders in SAP or IBM Maximo. ANYbotics says one utility saved more than $3.5 million in the first months, and one cement plant ran over 33,000 autonomous inspections in 16 months.
That is where building monitoring is going, at every size. A sensor sees something, software decides it matters, and a task appears in someone's queue without a person in between. I think that's the right direction. But it puts a lot of weight on one word: confirmed. Something has to decide that the anomaly is real and that the fix worked. Small buildings will get the same pipeline, scaled down: a $40 smart plug, a script, an AI that summarizes, and an alert on a phone.
I've spent the last month auditing my own version of that pipeline, a sump pump in my basement watched by a monitor, a backup "guardian" process, an assessment script and an hourly AI health check. I didn't read the dashboards. I read the raw logs on the host, line by line, and published what I found, including the times I was wrong. The system looked healthy almost the whole time. It wasn't.
Here are the seven questions that month taught me to ask before I trust any monitor, mine or anyone else's. Each one comes with the number that taught it.
1. Who else can command this actuator?
Before you analyze what a pump, fan or valve did, list every process that can switch it. I found four for one smart plug: the monitor, the guardian, a pair of dormant schedules stored on the plug itself, and the AI health check, which can clear a lockout and restart the service.
The monitor and the guardian disagreed about the relay over and over. Since June the guardian has cut the pump 73 times for running too long. The monitor treated at least 43 of those cuts as a fault and switched the relay straight back on, usually within a second or two. Each log recorded its own success. Neither log mentioned the other. You only see the argument by joining the two files on timestamp. (The full story is here.)
Ask: Show me every component with write access to this relay, valve or setpoint, and one timeline that interleaves all their commands.
2. Is the verdict downstream of the controller it's judging?
My assessment script classifies inflow. For 22 of 22 daily digests it said the same thing: high inflow. It wasn't measuring groundwater. When the controller is in its top tier, it runs the pump on a fixed timer, two minutes on and ten off, and the assessor counted those timed runs as evidence of water. It had hundreds of "high inflow" verdicts on hours with zero minutes of real demand.
The same mistake shows up in the controller's own repair routine. More on that in question 4.
Ask: For each alert or classification, which actuator is it not allowed to depend on? Does it suspend itself when that actuator takes over?
3. Does "sent" mean someone received it?
A blank line in a configuration file (NOTIFY_EMAIL_LOG=, present but empty) quietly gave one alert tier an empty recipient list. The code checked for a missing setting, not an empty one. In the guardian, an early return meant the next line, "email sent," ran anyway. Twenty-one alerts were logged as sent to nobody.
In the monitor, that tier has now skipped 5,018 sends since August, 125 of them in the last 24 hours. The urgent tier's clean record is why nobody noticed. A healthy loud channel tells you nothing about the quiet one.
Ask: How many alerts were delivered last month, by tier, counted at the receiving end? Not how many fired.
4. Is the recovery test harder than the fault?
When my controller thinks the float switch is stuck, it rapid-cycles the pump and then holds it on for 60 seconds to see whether it stops. The controller's own normal pulse runs the same pump for 120 seconds, and in 15,000 of those pulses it has never drawn less than 400 watts. A one-minute test of a fault that the system fails for two minutes at a time was never going to pass for the right reason.
As of this morning it has tried 954 times and reported success 8 times. All 8 relapsed. The two newest, on September 26 and 27, lasted 8 minutes and 21 minutes before the controller escalated again. The longest of all eight lasted 93 minutes. Each attempt flips the motor on and off seven times, which is wear you pay for in starts, not kilowatt-hours. (Details here.)
The general recovery check is even weaker. After 90 idle seconds the system declares itself "confirmed unstuck" and sends a back-to-normal notice. The median time before it escalates again is 34.5 minutes.
Ask: What does "fixed" mean here, how long does the system wait before believing it, and how does that compare with the measured time to relapse?
5. What happens if the controller disappears mid-command?
On September 27 the smart plug dropped off Wi-Fi 28 seconds after being told to run the pump for two minutes. The "off" command was a separate message that never arrived. The pump ran for about 55 minutes with nothing supervising it. The plug came back reading 65 °C, its hottest ever, which tripped an overtemperature lockout. That lockout left the pump without power for four more hours in the rain. (The correction is here.)
Most networked relays can carry their own countdown, so a timed command ends on the device even if the network doesn't. Mine wasn't using it.
Ask: For every networked relay, if the controller vanishes right now, does the load fail on or fail off, and is that the right answer for this load?
6. Did you read the setting from the device or from the docs?
Three documents in my repository say the plug restores its last state after a power loss. When I finally queried the plug, it said "initial_state": "off". After a power blip on September 16 it came back with the pump off, and only the monitor's normal-mode logic turned it back on. (I wrote it up here.)
Ask: When was each safety setting last read back from the hardware, and does anything check it after a reboot?
7. Does the watchdog share a failure domain with what it watches?
Two of my lost critical alerts failed for the same reason the plug was unreachable: no DNS, so neither email nor push notification could get out. The redundancy was on paper only.
The AI health check that finally cleared a seven-and-a-half-hour lockout on September 14 ran on a subscription with a weekly usage limit. For most of that outage, every run stopped at the limit. The lockout ended when the quota reset at 5 a.m. (Here's that one.) An AI watcher that is smart enough to explain an event can also explain it away, so my rule now is that explaining and alarming are jobs for separate components.
Ask: List what the watchdog needs to run and to send (network, DNS, cloud account, quota, the laptop it lives on). Which of those also fail when the thing it watches fails?
What this means if you're buying, not building
None of these failures showed up in a status light. My daily digest said OVERALL: OK every day for three weeks straight. A field that never changes carries no information, whatever it says.
Nexus Labs surveyed building owners ahead of NexusCon this month. About 70% of the ones just starting to scale building technology named data, infrastructure and interoperability as their main obstacle, ahead of the AI itself. That matches what I found. The model was rarely the weak point. The plumbing around it was: who can command what, whether a message arrived, what a setting really is, and when the system decides something is fixed.
If you're evaluating a monitoring vendor, or a robot platform that will write work orders for you, put these seven questions to them. A good vendor will have answers with numbers. If the answer is a dashboard screenshot, that's an answer too.
The short version: know every commander, keep verdicts independent of the controller, count deliveries at the receiving end, make "fixed" harder to reach than "broken," decide fail-on or fail-off on purpose, read settings from the hardware, and give the watchdog its own failure domain.
Want these seven questions answered for your building?
I build sensor and edge AI monitoring for small buildings, and I audit existing setups from their raw logs. Every finding above came from my own system, with the numbers published. I can do the same review for your pumps, boilers, freezers or rooftop units.
See what I build →