Skip to content

Counting the Wrong Events: 1,148 Sags and 63 Outages

Counting the Wrong Events: 1,148 Sags and 63 Outages hero image
Modified:
Published:

There is a particular kind of argument that happens in buildings where equipment keeps dying. The maintenance side says the supply is bad. The supply side says the outage log shows the supply was fine. Both are reading real records, and both are right, because the record everyone is arguing about does not contain the events doing the damage. #EnergyAndPower #RemoteMonitoring #PowerQuality

The installation

Monitoring at the incoming board of a building on mains with a UPS behind it, reading supply voltage, the minimum voltage inside each window, frequency, harmonic distortion, and how long the UPS battery would last if the mains went now.

The owner installed it after a run of equipment failures that nobody could explain. Their conclusion, once the data arrived, was that it was sags rather than the equipment. This post is the arithmetic behind that sentence.

Every figure in this post is measured over one fixed window, 28 May to 7 August 2026, which is the window the charts show. Naming it matters: these devices keep running, so a claim about “the last ten weeks” stops being true the week after it is written, while a claim about a named window stays checkable against the chart above it forever.

The record everybody already has



Mains availability across thirty days, sitting at 100 per cent for most of the window with occasional drops when the supply actually failed

This is mains availability, the number that goes in the report. For most of the window it reads 100%, dropping only when the supply genuinely failed. Across the full ten weeks the meter recorded

That is a real and irritating number, and it is the number the whole building talks about. It is also the wrong place to look for the cause of equipment failure, because a machine that is switched off by an outage is not being damaged by it. It stops, the generator picks up, it starts again.

The record almost nobody has



Minimum voltage and supply voltage over thirty days. Supply voltage holds near 239 volts with occasional vertical drops to zero at the outages, while minimum voltage dips repeatedly to between 195 and 215 volts without ever reaching zero

Two channels over the same window. The upper trace is supply voltage, sitting at its nominal 239.8 V and dropping vertically to zero at each outage. Those drops to zero are the 63 events everybody knows about.

The lower trace is the minimum voltage seen inside each reporting window, and it is a different picture entirely. It dips constantly. It dips to 210, to 205, to 200, and at its worst to 192.2 V, and it does this without the supply ever reaching zero. Nothing switched off. Nothing was logged as an outage. Availability still reported 100% for that window.

Counted properly over the ten weeks:

Using the conventional definition of a sag, a dip below 90% of nominal,

the minimum voltage fell below that line in 1,308 of 10,366 reporting windows, which is

Not 12.6% of a bad week. 12.6% of the whole ten weeks.

Why a sag hurts more than an outage



Panel (a): an outage, where voltage falls from 240 volts to zero for minutes or hours, the generator starts and somebody records the event. Panel (b): a sag, where voltage falls to 192 volts for a fraction of a second, never reaches zero, is never logged, and the connected drive drops out anyway

The deepest sag here reached

which sounds survivable and is not. A contactor releases below about 80%. A variable speed drive trips on undervoltage and needs a manual reset. A switched-mode supply rides it out by pulling proportionally more current, which is the part that does the quiet damage: the same power through a lower voltage means more amps, more heating in the windings and the capacitors, every time.

That gives the failure the shape monitoring handles best and inspection handles worst:

  • It is invisible without instrumentation. It is over before a person can look up.
  • It is cumulative. No single sag kills anything. A thousand of them age insulation.
  • It is absent from the record everybody trusts. Availability of 100% is true and useless.
  • It misattributes the cost. The failure lands on the equipment budget, and the equipment was never the problem.

What this is worth



The numbers below are a worked example, not a promise. Put your own figures in.

The asymmetry is unusually stark here because the measurement is so cheap. One monitoring point at the incoming board covers everything downstream of it. Against that:

  • Attribution. Knowing that failures cluster with sags stops you replacing the same equipment repeatedly on the theory that it was a bad batch.
  • Leverage with the supplier. A count of 1,148 events with timestamps is a conversation. “The power is bad here” is not.
  • The right fix, rather than a bigger one. Sags concentrated on one feeder point at that feeder. Sags at 06:00 every weekday point at something local starting up. Both are cheaper to solve than a whole-building UPS.
  • Sizing the protection you already own. The battery reading tells you how long you would actually last, rather than how long the nameplate says.

Watching costs one instrument at one board. Not watching costs a recurring equipment bill that is filed under the wrong heading, which is the version of a cost that never gets challenged because nobody has connected it to anything.

What to measure if you run a board



Minimum voltage per window

Not the average. An average over ten minutes erases a 200 millisecond dip completely, which is exactly the event you are hunting.

Sag count against outage count

Two separate tallies. The ratio between them is the argument, and on this installation it is over eighteen to one.

Harmonic distortion

Distortion rising with load points at something on your own side of the meter rather than at the utility.

UPS battery state

How long the protection would actually hold, measured rather than assumed from the rating.



Building power monitoring watches the same class of supply for a multi-tenant block, and its own post takes the other half of this argument: how thirty logged outages can add up to a number that reads like healthy availability. Where there is no usable grid at all, an off-grid tower site runs on solar, battery and a generator, and the fuel reconciliation post covers what goes wrong there instead.

For the equipment on the receiving end of all this, a production line drive motor and an injection moulding machine are both on the public fleet with bearing temperature and vibration, which is where repeated undervoltage eventually shows up as wear.

That everybody counts outages and nobody counts sags is availability bias with a switchboard attached: the events that announce themselves are the events that get recorded. The cognitive biases in engineering lesson works through why the vivid event crowds out the frequent one.

See it running



The device page shows the live readings, the charts above in their interactive form, and the alert history including the outages this post counts. If the device link below ever stops resolving, the public device gallery lists everything currently shared, and devices come and go from it as owners decide.

To put your own board on a chart, you can send readings without creating an account and see them arrive on a live chart in a couple of minutes, or read how to connect a first device.

Open this device Monitor your own equipment

Comments

Loading comments...


© 2021-2026 SiliconWit®. All rights reserved.