The alarm metrics worth reporting monthly
Six numbers describe the health of an alarm system. Most sites report none of them and discuss alarm performance anecdotally.
Alarm system performance is unusually measurable. The data already exists in the alarm and event log, the calculations are simple, and nearly every historian or alarm management package can produce them. Despite that, the common state is that nobody produces the numbers and the topic is discussed in terms of impressions.
Six measures cover it. Reported monthly, on the same basis each time, they show whether the system is improving, degrading or stable.
Average alarm rate per operator
Alarms annunciated per operating position per hour, averaged over the reporting period. This is the headline figure and it must be per operator rather than per unit, because it is the operator's attention that is the constrained resource.
Workforce analytics can add staffing and workload context to operational performance measures. For a related reference, see workforce analytics software.
The wider alarm-management lifecycle is described in the ISA-18 alarm management standards.
The published guidance places a manageable long-term average at around six alarms per hour, with twelve per hour described as maximum manageable and anything above that as increasingly unacceptable. Sites frequently measure ten times the upper figure at the start of an improvement programme.
Report it alongside the operating context. A rate measured during a month containing two unit shutdowns is not comparable with a quiet month, and the comparison is what people will draw unless the context is stated.
A control room where one operator covers three units at night and three operators cover them by day has a different rate at three in the morning, and that is when it matters most.
Peak and flood behaviour
The average conceals the distribution, and the distribution is where the risk sits. Two figures capture it.
The proportion of ten-minute intervals in which more than ten alarms annunciated — the conventional definition of a flood condition. Guidance suggests this should be under one percent of the time; measured values of ten to twenty percent are common on unimproved systems.
The peak count in any ten-minute interval during the period. This single number tells you what the worst case looked like, and it is usually the one that gets attention from people outside the control room.
The top ten contributing alarms
A ranked list of the alarm tags generating the most annunciations, with counts, over the period. This is the most immediately actionable of the six.
The characteristic finding is extreme concentration: it is common for the top ten tags to account for a third or more of all alarm activity, and for the top one to account for several percent on its own. These are almost always chattering alarms fixable with a deadband or delay change rather than genuine process problems.
Because the list changes as fixes are made, it also functions as a progress report. A top-ten list where the leading tag generates two hundred annunciations is a different system from one where it generates fifteen.
Standing alarms
Alarms that are active continuously for an extended period — commonly counted at a daily snapshot, or measured as the number active for more than twenty-four hours.
Standing alarms are corrosive in a specific way: they occupy the alarm list permanently, so the operator's view of current abnormality is cluttered with conditions that have been true for weeks. A list with thirty standing alarms means that a new alarm arrives into a screen that already looks alarmed.
The causes are usually a failed instrument nobody has repaired, a piece of equipment out of service with its alarms still enabled, or a threshold that no longer matches how the plant is run. All three are fixable, and the number should trend toward single figures.
Alarm distribution by priority
The split of annunciated alarms across priority levels, reported alongside the configured split.
This tracks priority inflation, which is otherwise invisible until it is severe. A configured distribution drifting from five percent high toward fifteen percent high over three years is a slow failure that no single change caused.
Operator response and stale alarms
Two related figures describe whether alarms are being acted on.
The proportion of alarms acknowledged within a defined period. A low figure suggests the operator cannot keep up, or that acknowledgement has become a bulk action performed to clear the list rather than a per-alarm decision.
The count of alarms that annunciate and clear without acknowledgement — fleeting alarms. A high count means conditions are appearing and self-correcting, which is a signal that the threshold or deadband is wrong rather than that anything was wrong with the process.
What not to do with the numbers
Alarm metrics describe a system, not the people operating it. Used to assess operator performance, they change behaviour immediately and in the wrong direction: bulk acknowledgement to improve the acknowledgement figure, reluctance to report nuisance alarms, and hostility to the measurement itself.
The measures belong to the engineering team responsible for the alarm system, and they measure the quality of its configuration. The operators are the people who tell you which numbers are misleading.
Making the report actually happen
Alarm metrics are typically produced enthusiastically for a benchmark exercise and then stop. The report has to be automated and owned to survive.
Automated: generated from the alarm log on a schedule, without anyone assembling it by hand. A monthly report that requires four hours of work will not survive a busy quarter.
Owned: a named person who reviews it and, importantly, is expected to do something about the top-ten list. A report circulated to a distribution list with no owner is read for a few months and then filed.
The most effective format is short — the six figures, the trend against the previous six months, and the top-ten table — attached to a meeting where the top contributors are assigned to someone. That meeting is the mechanism; the report is only the input to it.