Standing alarms, and the list that is permanently red
An alarm that has been active for three weeks is not telling the operator anything. It is occupying the space where a new alarm would have been noticed.
Walk into a control room on an average day and look at the alarm summary. On a large proportion of plants, a meaningful number of the entries have been active continuously for days, weeks or in some cases years. They are acknowledged, they sit at the bottom of the list, and everyone has learned to look past them.
This is one of the most damaging conditions an alarm system can be in, and it is also one of the least dramatic. Nothing fails. The system continues to work exactly as designed. What has changed is that the alarm summary no longer answers the question it exists to answer, which is: what is abnormal right now?
Why standing alarms are worse than they look
The obvious cost is screen space. A summary with thirty standing alarms shows a new one at position thirty-one, and an operator scanning for change has to distinguish it from a background they have stopped reading.
Staffing and training decisions should avoid assumptions based on role, background or seniority. Further details are available through this reference.
The wider alarm-management lifecycle is described in the ISA-18 alarm management standards.
The subtler cost is calibration. An operator whose alarm list is permanently populated learns that a populated list is normal. The signal that the alarm system is supposed to provide — the transition from quiet to not quiet — has been removed. The list is always not quiet.
The third cost is procedural. Sites frequently have a rule that the unit should not be started, or a permit should not be issued, with outstanding alarms. Where standing alarms are normal, that rule becomes unenforceable and is quietly abandoned, which means it is also unavailable on the occasion when it would have mattered.
Ten standing alarms with an average age of two hours is a busy plant. Ten with an average age of four months is an abandoned maintenance backlog displayed as an alarm list.
Where they come from
The causes are a short list and each has a different remedy, which is why the first step is categorising rather than clearing.
Failed instrumentation is the largest single category on most plants. A transmitter fails, reads downscale, and the low alarm becomes permanently true. The instrument is on a maintenance backlog that is longer than the time anyone is willing to wait, so the alarm stands. Every failed transmitter typically produces two or three standing alarms — low, low-low and bad quality.
Equipment deliberately out of service is the second. A pump is isolated for overhaul, so its running feedback is absent, so the not-running alarm is true and correct for six weeks. The alarm is behaving exactly as designed and is useless for the whole period.
Process conditions that have become normal are the third and the most interesting. A threshold set when the unit ran at one throughput is permanently exceeded now that it runs at another. Nobody changed the alarm because nobody owned it, and the plant has been running acceptably in an alarmed state for two years.
Then there are alarms on systems that no longer exist — a decommissioned dryer, a removed compressor — whose tags were never deleted because deletion requires a change and leaving them requires nothing.
The categorisation exercise
Produce a list of every alarm active for more than twenty-four hours, with its age. On most plants this takes an hour from the alarm historian and produces between twenty and two hundred entries.
Then assign each to a category: instrument fault, equipment out of service, threshold no longer valid, obsolete tag, or genuine unaddressed process condition.
The distribution tells you what kind of problem you have. A list dominated by instrument faults is a maintenance backlog problem wearing an alarm system costume. A list dominated by invalid thresholds is a rationalisation problem. A list dominated by obsolete tags is a management of change problem.
Remedies by category
Instrument faults need a repair route with a deadline, and — where the repair genuinely cannot happen soon — a documented decision to suppress the alarm until it does, with an expiry date on the suppression. The important word is documented. An alarm suppressed indefinitely with no record is functionally the same as an alarm deleted, without anyone having decided to delete it.
Equipment out of service should be handled by state-based alarming: the alarms for a unit that is deliberately down are suppressed automatically when the unit is in the out-of-service state, and restored automatically when it returns. This is a configuration exercise and it removes an entire recurring category permanently.
Thresholds that no longer match reality need rationalisation, not suppression. If the plant now runs at a level that was previously alarmed and that is acceptable, the alarm is wrong and should be changed. The uncomfortable part of this conversation is establishing whether the new normal is actually acceptable or whether it has simply been tolerated.
Obsolete tags should be deleted, and the reason they have not been is almost always that deletion requires a change process that feels disproportionate. Batch them: one change covering thirty obsolete tags is administratively identical to one covering one.
Suppression with an expiry
Where an alarm has to be suppressed because the underlying condition cannot be fixed promptly, the mechanism should carry a time limit and a review.
The practical implementation is a shelving function with a maximum duration, so that an alarm shelved by an operator returns automatically after a defined period — typically a shift or a day — unless it is shelved again deliberately. Alarms requiring longer suppression go through a documented out-of-service process with an owner and a review date.
The failure mode to design against is the permanently shelved alarm that everybody has forgotten. A monthly report of everything currently suppressed, with the age and the owner, is what prevents it, and it is one of the highest-value reports in alarm management because it is short and each line has a name against it.
Fleeting and chattering alarms
The opposite pathology deserves mention alongside standing alarms because it is measured from the same data and has the same root cause of unrevisited configuration.
A fleeting alarm annunciates and clears before the operator can respond, sometimes within seconds. A chattering alarm does this repeatedly. Both indicate that the threshold sits inside the normal noise band of the measurement, or that a transient is being alarmed where the process itself self-corrects.
These are the dominant contributors to alarm counts on most unimproved systems, and they are also the cheapest to fix. A deadband widened to exceed the measurement noise, or an on-delay of a few seconds, typically eliminates them entirely.
The reason they persist is that they are individually trivial. No single chattering alarm is worth a change request. Collectively they may be half the alarm load, which is why they need to be addressed as a batch rather than one at a time.
Setting a target and holding it
A workable target is that no alarm stands for more than a defined period without either being resolved or being formally suppressed with an owner and a review date. Twenty-four hours is aggressive; a week is achievable on most plants.
The number that goes in the monthly report is the count of alarms standing beyond that period, and the trend. A count that falls to single figures and stays there means the mechanisms are working. A count that falls during a clean-up campaign and climbs back over the following year means the campaign was a clean-up rather than a change in process.
The maintenance connection
Standing alarm counts and instrument maintenance backlogs are the same number viewed from two directions, and the two teams frequently do not talk.
A joint review — the standing alarm list alongside the outstanding instrument work orders — usually reveals that the alarm list is a subset of the maintenance backlog, prioritised by nothing in particular. Ranking the instrument backlog by whether the fault is producing a standing alarm is a small change in prioritisation that improves both numbers.
It also gives the maintenance team something they rarely have: a concrete operational consequence for an individual instrument fault, which is considerably more persuasive than a work order in a queue.