The Control Room
Home/Alarm management/Auditing an alarm system you have inherited

Alarm management

Auditing an alarm system you have inherited

A structured audit turns an impression that the alarm system is bad into a costed list of specific work.

11 min read1007 wordsUpdated July 2026

Engineers frequently inherit an alarm system that everyone agrees is poor and nobody can characterise. Operators say there are too many alarms, management asks how many is too many, and the discussion runs on anecdote.

An audit converts that into evidence. It is a bounded exercise — a week or two for a unit, depending on the size of the database — and it produces the two things needed to get work funded: a comparison against published expectations, and a ranked list of what to do first.

Scope it before starting

Define the boundary in terms of operating positions rather than equipment. The unit of analysis is the console: what one operator sees, because it is their attention that is being audited.

Time records are only useful when the process accounts for predictable forms of manipulation. Further details are available in the linked article.

The wider alarm-management lifecycle is described in the ISA-18 alarm management standards.

Choose a measurement period long enough to include normal variation — a month is typical — and note what happened during it. An audit period containing a major upset produces different numbers from a quiet one, and both are legitimate as long as the context is recorded.

Part one: the quantitative benchmark

From the alarm and event historian, produce the standard measures.

  • Average alarms per operating position per hour, and per day.
  • Distribution of ten-minute alarm counts, and the proportion exceeding ten.
  • Peak ten-minute count and when it occurred.
  • Top twenty contributing tags by annunciation count.
  • Standing alarms: count at a daily snapshot, and the age distribution.
  • Priority distribution, configured and annunciated.
  • Fleeting alarms: those clearing within a short period, as a proportion of the total.
  • Acknowledgement behaviour: median time to acknowledge, and the proportion acknowledged in bulk.

These are all derivable from data that already exists. The work is in extracting and presenting it rather than in gathering it, and most alarm management packages produce the whole set automatically.

Part two: the database review

Separately from the behavioural data, examine the configuration itself.

How many alarms are configured per operating position? A figure in the thousands per operator is common and is a strong predictor of poor performance regardless of what the rate measurements show, because it represents latent volume waiting for the right upset.

How many carry documented rationalisation? On most inherited systems the answer is none, and establishing that is useful because it defines the scale of the remediation.

How many are configured with default deadbands and no delay? This is a good proxy for how much nuisance volume is available to be removed cheaply.

How many are duplicates — the same physical condition alarmed from two sources, or a condition alarmed at both the field device and the control system?

Part three: talk to the operators

The quantitative work describes the system. The operators describe the experience, and the two frequently disagree in informative ways.

The questions that produce useful answers are specific. Which alarms do you routinely ignore, and why? Which alarms have you never seen? What do you do when the list floods? Is there anything you would like to be alarmed on that is not? What do you do first when you take over the console?

That last question is diagnostic. On a healthy system the answer involves reviewing current conditions. On an unhealthy one it is frequently some version of clearing down the list so that new alarms will be visible, which is the operator manually compensating for a system that does not distinguish current from historical abnormality.

Part four: check the mechanisms

Beyond the alarms themselves, examine the surrounding process.

Is there a documented alarm philosophy — the site-level document defining priority criteria, performance targets, shelving rules and the rationalisation method? Most sites either have none or have one written for a capital project and never used since.

Is there a management of change route that covers alarm additions and priority changes? Where alarms can be added by an engineer without review, the database will grow regardless of any remediation.

Is anyone producing performance metrics? Is anyone reading them?

What is currently shelved, suppressed or out of service, and does anyone know?

Producing the output

The audit report should be short and should lead with the comparison against published expectations, because that is what converts an internal impression into an external standard.

Then the ranked remediation list, in order of effort against benefit.

First, the bad actors: the top twenty tags, with the specific fix for each — deadband, delay, threshold, or removal. This is usually days of work and removes a large share of the volume.

Second, the alarms that are not alarms: informational messages and maintenance notifications to be rerouted.

Third, state-based suppression for the flood scenarios identified in the data.

Fourth, standing alarm clearance with the maintenance team.

Fifth, full rationalisation, scoped by unit and estimated realistically at the rates a real session achieves.

Sixth, the governance work: the philosophy document, the change route, the monthly report and its owner.

Estimating honestly

Rationalisation estimates are the part most often got wrong, and getting them wrong is how projects lose credibility partway through.

A session with a full team covers a limited number of alarms per hour and cannot be sustained all day — the concentration required means three or four hours is a realistic daily maximum. Multiply through for the database size and the resulting figure is usually large enough that someone will ask whether the whole database needs doing.

That is a reasonable question and the answer is usually no. A risk-based scope — full rationalisation for the high-consequence units and bad-actor treatment elsewhere — is a defensible position and a far more likely one to be completed.

Re-audit on a cycle

The audit is worth repeating annually, using the same measurements on the same basis, so that the comparison is meaningful.

The characteristic finding on the second audit is that the metrics improved and then partially regressed, and that the regression is concentrated in alarms added since the first audit. That finding is the argument for the governance work, which is always the least popular item on the remediation list and the one that determines whether any of the rest lasts.

General information. Nothing here is accounting, tax or legal advice. Stock valuation methods, write-off evidence requirements, the tax treatment of losses and the rules on monitoring staff differ substantially between jurisdictions and change over time. Take qualified advice on your own situation.

Related

Continue reading