Archive
Every guide, in one list
40 guides across five sections. Nothing here is sponsored, and no software vendor has paid for a mention.
An alarm system that is merely busy in steady state becomes unusable during an upset, which is precisely the condition it was installed to help with.
Rationalisation is a slow, structured argument about every alarm in the database. It is tedious, it is the only method that works, and most attempts fail for predictable reasons.
Remote workforce management can help coordinate support teams working across sites and time zones. A related reference is available at https://www.monitask.com/remote-workforce-management-software/.
The broader control-system security framework is outlined in the ISA/IEC 62443 series.
Related lifecycle and security responsibilities are also covered by ISA/IEC 62443.
If a third of the alarms are high priority, there is no priority scheme. There is a list with decoration.
Six numbers describe the health of an alarm system. Most sites report none of them and discuss alarm performance anecdotally.
An alarm that has been active for three weeks is not telling the operator anything. It is occupying the space where a new alarm would have been noticed.
Turning an alarm off is sometimes correct engineering and sometimes the removal of a safeguard. The distinction is procedural, not technical.
A rationalised alarm database displayed badly is still an unusable alarm system.
A structured audit turns an impression that the alarm system is bad into a costed list of specific work.
Most process graphics accurately depict the plant and answer none of the questions an operator actually has.
Most control systems have detail displays and unit displays. The overview that tells an operator whether the whole plant is normal frequently does not exist.
Colour decisions made on a desktop monitor in an office are evaluated on a different screen, in different light, by people who have been awake for eleven hours.
A number tells the operator where a variable is. Only a trend tells them whether to act now.
An operator diagnosing an upset should not be searching a menu. Most navigation schemes are designed by people who already know where everything is.
Graphics are designed for the ninety percent of time the plant runs normally, and used most intensively during the other ten.
Without a style guide, every display reflects the preferences of whoever built it, and the operator does the translation.
Display improvement is usually deferred until the next system migration, which may be a decade away.
Storage is cheap enough that the default answer is to historise every tag. The cost has moved from disk to the ability to find anything.
Historian compression is enabled by default with settings nobody chose, and it discards exactly the detail an investigation needs.
Retention policy is usually whatever the system was configured with, discovered during the first request for data from four years ago.
Sequence-of-events analysis depends entirely on clocks agreeing. On most sites they do not, and nobody knows by how much.
A failed transmitter reading downscale produces a number. Nothing about that number announces that it is meaningless.
A tag with a value and a timestamp answers almost no useful question on its own.
Two departments produce two different figures for the same month from the same historian, and the meeting is about the discrepancy rather than the plant.
The data is the asset. The software is replaceable, and the migration is where decades of it are quietly lost.
A loop check performed by two people in a hurry, signing a sheet, is a document rather than a test.
A FAT run as a demonstration finds nothing. A FAT run as an attempt to break the system finds what site commissioning would otherwise find at ten times the cost.
The project is finished when the plant runs. The operations team's problems start there.
The document that explains why the control system does what it does is written once, by someone who is leaving, for an audience they never meet.
A management of change procedure that takes three days is a procedure that gets bypassed during the situations it exists for.
Every site knows its drawings are out of date. Very few know by how much, or which ones.
Every commissioning generates outstanding items. What distinguishes projects is whether they close.
An interlock that has never been tested is a design intention. Testing is what makes it a protective function.
A control system becomes unsupportable years before it stops working, and the warning signs are all commercial rather than technical.
The spare you need is the one that has been sitting untested in a cupboard since 2011.
Most sites have control system backups. Considerably fewer have ever performed a restore, and the difference only becomes apparent once.
Most control networks were flat when they were built and have been connected to progressively more things since, one reasonable request at a time.
Almost every remote access arrangement was set up quickly during an outage and never revisited.
The advice to patch promptly assumes a system that can be restarted. Control systems frequently cannot, and the resulting position needs managing rather than ignoring.
Operators become expert at normal running through daily practice. The abnormal situations are the ones they encounter least and need most.
The five minutes at the start of a shift determine what the next twelve hours are working from.