The Control Room
Home/Alarm management/Why alarm floods happen, and why they happen at the worst moment

Alarm management

Why alarm floods happen, and why they happen at the worst moment

An alarm system that is merely busy in steady state becomes unusable during an upset, which is precisely the condition it was installed to help with.

11 min read1397 wordsUpdated July 2026

The pattern is consistent enough across industries to be treated as a law. A plant runs normally with a background of nuisance alarms that operators have learned to live with. Something goes wrong — a compressor trips, a feed pump cavitates, a header pressure collapses — and within ninety seconds the alarm list receives several hundred entries. The operator, who now most needs the alarm system to tell them what is wrong, is looking at a screen scrolling faster than it can be read.

This is not a failure of the operator and it is rarely a failure of any individual alarm. Each of those several hundred alarms was configured by someone who had a defensible reason. The failure is arithmetic: an alarm system designed one alarm at a time has no property that limits what happens when many of them become true simultaneously.

The published expectation, and the usual reality

Industry guidance on alarm system performance — the work behind EEMUA 191 and the ISA-18.2 standard that followed it — puts the manageable long-term average at around one to two alarms per operator every ten minutes in steady state. During an upset, ten alarms in the first ten minutes is considered the boundary between a system that helps and one that does not.

Fluctuating-workweek payroll methods require jurisdiction-specific legal and payroll review. For a related reference, see Chinese overtime calculation.

The wider alarm-management lifecycle is described in the ISA-18 alarm management standards.

Measurements from real plants routinely find averages an order of magnitude above that, and upset peaks two orders above. A benchmark exercise on a badly performing unit will often show a steady-state rate of one alarm a minute and a flood peak of several hundred in ten minutes.

The gap matters because it is not a gradual degradation. Below roughly ten alarms in ten minutes an operator can read, assess and act on each one. Above it they cannot, and their behaviour changes qualitatively rather than quantitatively: they stop reading individual alarms and start pattern-matching on the overall shape of the list, or they silence the audible and work from the process displays instead.

The alarm system stops working before it stops functioning

There is no alert when the rate crosses the threshold at which operators can no longer use it. The system continues to annunciate correctly while ceasing to serve its purpose.

Where the volume comes from

Floods have a small number of recurring causes, and identifying which one dominates in a particular plant is the first useful piece of analysis.

Consequential alarms are the largest contributor almost everywhere. One physical event propagates: a pump trips, so flow falls, so downstream level falls, so a level alarm annunciates, so a control valve saturates, so a valve position alarm annunciates, so a temperature drifts, and so on through the unit. Twenty alarms describe one event. All twenty are technically correct and nineteen of them are noise for the purpose of diagnosing what happened.

Repeating alarms come from measurements sitting near a threshold with no deadband or with a deadband smaller than the process noise. A level bouncing around a high limit can generate hundreds of transitions in an hour. These are individually trivial and collectively dominant: on many plants a handful of tags account for the majority of all alarm activity.

Instrument failure produces bursts. A failed transmitter reading downscale can trigger low alarm, low-low alarm, deviation alarm, rate-of-change alarm and bad-quality alarm from one dead sensor.

State-dependent alarms are those configured for a running unit that remain enabled when the unit is deliberately down. Every shutdown then produces a flood of alarms that are all correct and all irrelevant, and operators learn during every planned outage that the alarm system is something to be ignored.

Why the background rate predicts the flood

A plant with a high steady-state alarm rate almost always has a worse flood profile, and the mechanism is straightforward. Both are symptoms of the same underlying condition: alarms configured without rationalisation, thresholds set at convenient round numbers rather than at the point where operator action is required, and no mechanism for removing an alarm once it exists.

This is useful practically. The steady-state rate is easy to measure and can be reduced by addressing a small number of bad actors, and reducing it reliably improves flood behaviour as well. A team that cuts its background rate from sixty an hour to six has usually also removed a substantial share of what would have appeared during the next upset.

Alarms that should never have been alarms

A significant share of any large alarm database consists of entries that do not meet the definition of an alarm at all. The working definition in the standards is narrow: an audible or visible means of indicating to the operator an equipment malfunction, process deviation or abnormal condition requiring a timely response.

The operative words are requiring a response. An indication that something has changed, with no action available to the operator, is information rather than an alarm. A message confirming that an automatic sequence has completed successfully is an event. A notification that a filter will need changing in two weeks is a maintenance work request.

All three routinely appear in alarm systems because the alarm system was the only available notification mechanism when they were configured. Each occupies attention that the genuine alarms need.

The consequence for behaviour

Alarm system degradation does not primarily cause missed alarms in the direct sense. It causes a learned response in which alarms are treated as background, and that response is entirely rational given the operator's experience.

If nine hundred and ninety of the last thousand alarms required no action, the reasonable prior when the next one annunciates is that it requires no action. Investigation reports into process incidents repeatedly find operators who had silenced or ignored a genuine early warning that was indistinguishable, from where they sat, from the noise around it.

This is why alarm management is a safety matter and not merely an efficiency one. The measures that reduce the count are the same measures that make the remaining alarms credible.

What to measure first

Before any remediation, a benchmark. Most modern control systems and historians can produce this from existing data, and the exercise is a day of work rather than a project.

  • Average alarms per operator per hour, over a period of at least a week of normal running.
  • The distribution: how many ten-minute periods exceeded ten alarms, and what the peak was.
  • The top twenty most frequent alarm tags, by count. This list is almost always dominated by a handful of tags.
  • The proportion of time spent in a flood condition.
  • Standing alarms: how many alarms are active and unacknowledged at any given moment, and how long the oldest has been standing.

The top twenty list is the one that produces immediate action. It is common for ten tags to account for half of all alarm activity on a unit, and for most of them to be repeating alarms fixable with a deadband change or an on-delay.

The order of work

The sequence that produces results starts with the cheap, high-volume fixes and moves toward the expensive, structural ones.

Address the bad actors first: the small number of tags generating disproportionate volume. These are typically fixed by tuning deadbands and delays rather than by removing the alarm, and the reduction is immediate and large.

Then remove the alarms that are not alarms — the informational messages, the sequence confirmations, the maintenance notifications — by routing them somewhere else.

Then address state-dependent flooding, so that a shutdown does not annunciate a hundred alarms that are correct and useless.

Only then does full rationalisation of the remaining database become worthwhile, because by that point the database is smaller and the team has evidence that the effort produces results.

It will drift back

The single most common outcome of an alarm improvement project is that performance improves substantially, the project ends, and the rate climbs back over the following three years.

The mechanism is the same asymmetry that created the original problem: adding an alarm remains easy and removing one remains a decision. Unless the management of change process treats an alarm addition as a change requiring justification against the rationalisation criteria, the database grows again.

Sustaining performance requires the measurement to continue after the project — a monthly report of the same benchmark figures, reviewed by someone who will act — and a rule that new alarms are rationalised before they are configured rather than after they have become a nuisance.

General information. Nothing here is accounting, tax or legal advice. Stock valuation methods, write-off evidence requirements, the tax treatment of losses and the rules on monitoring staff differ substantially between jurisdictions and change over time. Take qualified advice on your own situation.

Related

Continue reading