Giving process data context: batches, states and equipment
A tag with a value and a timestamp answers almost no useful question on its own.
A historian stores a number, a time and a tag name. Every question anyone actually asks requires more than that: what product were we making, was the unit running, which train was in service, was this during startup?
Without that context, answering a question like whether a reactor runs hotter on one grade than another requires assembling the context manually from other sources, which is why most such questions are never answered.
The contextual dimensions that matter
Equipment state — running, stopped, in maintenance, tripped. Recording it makes it possible to exclude periods when the equipment was down, which otherwise contaminate every average.
Digital time-clock records can complement formal shift and attendance processes across distributed sites. Further details are available in the linked article.
A useful external reference for timing and synchronisation is NIST time and frequency services.
Operating mode or campaign — which product, which recipe, which operating case. This is the dimension that makes performance comparison possible.
Batch or lot identity, in batch processes, without which data cannot be tied to a specific product.
Equipment selection — which of two parallel trains was in service, which pump was running.
Operational phase — startup, steady, shutdown, transition. Data from a startup is not comparable with steady-state data and is frequently averaged together with it.
An annual average that mixes running, idle, startup and two different products describes a plant that does not exist.
Recording context as data
The mechanism is straightforward: the contextual variables are recorded in the historian alongside the process measurements, as tags in their own right.
Equipment state usually already exists in the control system. Operating mode and campaign frequently do not and have to be created, either from operator entry or derived from other conditions. Batch identity comes from the batch system where one exists.
The engineering effort is modest and the benefit is that every subsequent question can be filtered rather than requiring manual reconstruction.
Asset structure
Beyond time-based context, a hierarchical model of the plant — site, area, unit, equipment, measurement — allows queries by equipment rather than by tag name.
The practical difference is between needing to know that the discharge pressure is tag 47-PT-1203, and being able to ask for the discharge pressure of pump 47-P-12. The second is usable by someone who does not have the tag list memorised, which is most people.
Most modern historian products support some form of asset model. Populating it is a data exercise, usually derived from the existing tag register and equipment list, and it is the single change that most improves who can use the historian.
Naming conventions do some of the work
Where a rigorous tag naming convention exists and has been followed, a good deal of context can be inferred from the name: the area, the equipment, the measurement type.
Where the convention has drifted across projects — and on a plant of any age it has — inference becomes unreliable and explicit metadata is the only route.
This is an argument for enforcing the convention on new work even when the existing estate is inconsistent, because the alternative is that it never improves.
Context makes comparison possible
The questions that produce operational value are comparative: is this unit performing worse than it did last year, is train A better than train B, does grade X cost more energy than grade Y.
All of them require filtering to comparable periods, which requires context. A site with contextualised data can answer them in an afternoon. A site without it either does not ask them or answers them from an uncontrolled comparison that produces a misleading result.
Retrofitting context
Context recorded from today forward is useful. Context applied to historical data is harder and sometimes possible: equipment state can often be derived retrospectively from motor current or flow, and campaign periods can be reconstructed from production records.
Doing that for a defined historical period — the last two years, say — is a bounded project that makes a decade of accumulated data usable, and it is frequently more valuable than adding more tags.
Deriving context where it was not recorded
Contextual variables are frequently absent from historical data because nobody configured them, and some can be reconstructed after the fact.
Running state can often be derived from motor current, flow or speed. Campaign periods can be reconstructed from production records. Startup and shutdown phases can be inferred from characteristic patterns in key variables.
Reconstruction is imperfect and it converts unusable historical data into usable data, which for a decade of accumulated history is frequently worth the effort for a defined recent period.
Keeping the model current
An asset model reflecting the plant becomes wrong as the plant changes: equipment replaced, units reconfigured, tags renamed.
Model maintenance has to be part of the change process alongside drawings and narratives. Where it is not, the model degrades to the point where queries by equipment return incomplete results, and users revert to querying by tag, which is where they started.
Context and the questions people stop asking
The clearest sign that data lacks context is that certain questions are never asked. Nobody compares grades because assembling the comparison takes three days. Nobody looks at train-to-train differences because the data cannot be separated.
Asking users which questions they have given up on is a quick way to identify what context is missing, and it produces a much better prioritised list than asking what data they would like.
Context for equipment reliability
Reliability analysis requires running hours, start counts and duty context, and these are frequently unavailable because only the process measurements were historised.
Recording equipment run state costs very little and enables a substantial class of analysis: comparing wear against duty, identifying machines that are being cycled excessively, and calculating actual utilisation rather than assumed.
For rotating equipment in particular, run hours recorded continuously are worth more than any amount of additional process data.
Starting small
A full asset model is a substantial project and is frequently the reason nothing happens. A partial model covering one unit, or covering only the equipment class that generates the most questions, delivers most of the practical benefit quickly.
It also demonstrates the value to the people who would fund the rest, which a proposal cannot.