The Control Nobody Argues About

The list of controls required of an industrial operator grows between revisions and does not shorten. Items are added. Items are elaborated. What does not happen is withdrawal because a requirement turned out not to be worth having.

The documents doing the requiring sit at different levels. IEC 62443-3-3 states system security requirements and maps them to capability security levels. It is careful about its own scope: target and achieved levels belong to other parts of the series, and this document states what a system must be capable of rather than what any particular plant should aim at. It is equally careful to say that it specifies functional requirements and leaves how they are met to integrators and suppliers. NIST SP 800-82 Revision 3 wraps risk management guidance around a control overlay and tells the reader explicitly not to treat the document as a checklist, but to assess and tailor. It goes further than that. It contains a section on building an OT cybersecurity business case, with a subsection on the benefits of cybersecurity investment. The NIST Cybersecurity Framework sits higher again, specifying outcomes rather than controls and speaking directly about prioritisation, risk tolerance, available resources, and progression where a cost-benefit analysis supports it.

So the frameworks are not silent about benefit and cost. They discuss both, and one of them gives the subject a section of its own. What none of them does is attach either figure to an individual control in a form that supports a decision about that control. No requirement states what reduction it produces. No requirement states what it will take to put in place. Neither could be stated by a document written for every plant. The reasoning that does exist lives at the front, at the level of the programme, and has disappeared by the time the reader reaches the requirements themselves. What the operator inherits is the requirement without it.

These two look like one problem. They are two, and they have different causes.

Why the benefit figure is absent

Most of the required list sits, by dominant purpose, on one side of a distinction that is rarely drawn. These are likelihood controls. Their reason for existing is to make the event ending in loss less probable. The assignment is testable rather than asserted: remove the control and ask whether what changes is the chance that the loss occurs at all, or the loss once it has occurred. Few controls are purely one thing, and segmentation answers to both. Monitoring belongs chiefly to the first where the required capability includes timely response, rather than merely producing a record of the intrusion. What matters is which effect the control is justified and measured on, and for most of the list that is the likelihood side. That is the axis the missing figure belongs to.

Waiting for better data is not a plan, because the data has no mechanism of production.

The waiting is explicit in the documents. One of them states an intention to move its levels and requirements towards quantitative descriptions and metrics as data and experience with industrial security systems accumulate. In the meantime, it describes its level definitions as deliberately unspecific, so that one vocabulary can serve every document in the series. The generality is the point, and the numbers are pending. That was written over a decade ago.

They are pending because rates are computable when failures are sampled from a stable population, and adversarial events are not. Equipment failure rates are produced the first way. A pressure transmitter fails at a frequency observable across a large installed base, and the observation holds because the transmitter is not choosing when to fail. Adversarial events are selected. An adversary picks the target, the moment and the method, and revises all three in response to what defenders do. That remains true where the individual event is automated and indiscriminate. Selection happens at the campaign rather than at the plant, in the tooling and the class of target at which it is pointed, and it is revised on the same terms.

This removes the comparison. Establishing how much a control reduces risk requires an observed rate of loss with the control and an observed rate without it, drawn from comparable populations over a period in which the adversary was not adapting. That last condition is not one any study design can impose, which is why the pairs on offer rest either on populations that are not comparable or on periods in which the adversary was adapting. The empirical question is not unanswered. It is malformed, and better instrumentation will not help, because the comparison it asks for is not one that anything observes.

The same obstacle appears at the level of a single site, where it is sharper. Measuring what a control bought at this plant would require observing this plant, with this architecture, against this adversary, both with the control in place and without it. That comparison is not available at any price, and there is no physical law to substitute for it, because the initiating event is chosen rather than independently sampled.

Something else is available in its place, and competent practice uses it. Adversarial likelihood is assessed through the adversary’s capability, intent and targeting, scored on an ordinal scale by people who know the plant and the threat. That method is legitimate, and it is what a competent risk assessment does. What it produces is a judgement. A judgement can be documented, defended and revised. It is aggregated and compared across sites routinely, and the aggregate carries no information the underlying judgements did not. What cannot be checked afterwards is whether the assigned likelihood was right, because the outcome does not reveal it. The judgement carries the missing term rather than supplying it.

Quantification does not alter the position. Replacing ordinal scores with calibrated estimates and propagating them through a simulation produces a distribution rather than a rank, and the distribution is assembled from the same judgements. Calibration is trainable and checkable, but it is checked against questions that resolve. This one does not. What transfers is a general estimator skill rather than demonstrated accuracy on the quantity in hand.

The same series requires the operator to go further than assessing likelihood. Existing countermeasures are to be identified and evaluated for their effectiveness in reducing likelihood or impact, and likelihood is then to be reassessed in light of them. That is the missing figure, required as an output. No method for producing it is given, and any risk-assessment methodology is permitted provided the requirements are met. The obligation is to arrive at the reduction. What discharges it is a judgement, which is what was available before the requirement was written.

Defining a local measure does not supply the missing term; narrowing the population leaves it carrying the judgement it was introduced to replace.

There is a second reason the question never becomes well formed, and this one is a property of the requirements rather than of the evidence. A requirement that can be shown complete can be asked what it delivered. Most of the required list cannot be shown complete, because it carries no termination condition. Segment the network: to what depth, against what assumed adversary position, and when is it finished? Achieve visibility: over what fraction of what, at what detection latency, and when is it finished? That last absence is not an oversight. Continuous monitoring is required to be timely, while the time that makes it timely is left as a local matter outside the standard’s scope. The term that decides whether observation became timely response is therefore absent from the requirement. Two competent engineers at the same plant, looking at the same systems, can reach opposite conclusions about whether either requirement is met, and there is no artefact that settles the disagreement. Completion ends up being set downstream by whichever instrument was bought to satisfy the requirement, and the resulting coverage threshold is not derived from anything physical. Ninety-five per cent asset visibility is not a number that came from anywhere, and nobody can say what the missing five per cent is connected to.

What the controls still establish

None of this disputes the controls, because the missing figure is not the only kind of justification available, and not every control is justified the same way. Where a control’s contribution is whether a path exists at all, mechanism is the justification: a severed connection cannot be traversed, and a boundary that passes no inbound sessions limits what is reachable from where. Those are claims about mechanism, established by engineering rather than by an observed rate, and they hold without one. Where a control’s contribution is a matter of degree rather than a testable state, credentials and monitoring among them, the judgement already described is the justification instead, and it needs no mechanism either. Neither kind converts into a figure for the change in the chance of a loss, because that change depends on what else is available to whoever is choosing. What the absent figure withholds is a further thing: how much, compared against what else.

Those claims also have to keep being true. A severed connection stays severed because something governs what gets reconnected, and a boundary passes no inbound sessions only while its rules remain what they were. What holds a mechanism claim in place between the day it was established and the day anyone checks is the control programme. That is not sizing and it is not a figure. It is maintenance, and it is what the programme is for.

Coverage measures are a legitimate and adequate instrument for confirming that a control justified in the first way is in place across the estate and functioning. Metrics discharge that real function, which is different from sizing what the control bought. Coverage is also what remains when the thing worth measuring has no number. It is countable, and a discipline that measures what it can count is behaving reasonably under a constraint it did not create.

None of this waits on the figure. The controls are required, they get put in and they get maintained, and none of that has ever waited for a number. What the missing figure prevents is not the work. It is the claim about how much of it was enough.

The boundary between the two is worth stating, because it is routinely crossed. A recurring complaint in programme reviews is the compensating control that exists in the ticket and not in the plant, and it is usually treated as a fidelity problem: someone recorded a measure that was never put in. Part of it is that. The rest is not. Confirming that a compensating measure is present is a coverage question, and metrics answer it. Confirming that it compensates is a question about what it supplies against the specific gap. Where the answer is a testable mechanism, the divergence between the ticket and the plant is discoverable: run the restore, test the boundary, inspect the severed connection. Where the claimed compensation is a reduction on the likelihood side, closing the gap requires the figure that does not exist. That is why the complaint recurs and never resolves.

Other disciplines in these plants do not work this way, and the comparison does not have to be imported. The standard draws it. Its own discussion of security levels opens by setting them against safety integrity levels and attributes the difference to what can be quantified: the probability of a component failing from random hardware causes can be measured and carried into the required protection calculation. Security, it says, is harder because the consequences and circumstances are broader and the root cause need not be a random failure at all.

Two properties are doing the work in that comparison, and they are worth separating because only one of them transfers.

One is the rate. It is available in functional safety because equipment failure is sampled, and it is unavailable here because adversarial events are selected.

The other is the termination condition. A safety instrumented function is bounded before anything is calculated: this function, these elements, this demand case, this test interval. That boundary makes the question askable in the first place, and it owes nothing to the availability of a rate. It is a drafting property.

Why the termination condition is absent

Segment the network has no termination condition, and its absence has nothing to do with the missing comparison. A requirement could oblige the operator to state what must be unreachable from where, in what as-built configuration, and verified by what means, and to be finished when that state holds. None of the requirements asks for that.

One adjacent requirement does not reach for a boundary at all: it asks for segmentation without saying which networks are critical or when the work is finished. Another does reach for it, asking only for the capability to enforce the boundary the operator has already documented. Neither requirement asks for verification that the partitioning holds in the plant as built. The defect is not the division of labour but the incomplete handoff.

The state that ends the control could be required tomorrow. An operator specifying the control for its own plant need not wait to be asked.

A termination condition is necessary and not sufficient. Give one to a likelihood control and the question of what it delivered becomes askable; it remains unanswerable, because the answer needs the counterfactual pair that does not exist. Give one to a consequence control and the question becomes both askable and answerable. That is what allows one item on the list to be measured against what it delivers, rather than merely against its specification.

Why the cost figure is absent

The cost figure is missing for an unrelated reason, and the difference matters.

Cost is estimable. That is a weaker word than knowable, and it is still far stronger than anything available on the other side. What it takes to deploy network monitoring across a distributed plant estate, in engineering hours split between work that needs plant judgement and work that does not, in outage windows consumed, and in the standing draw after handover, can be worked out in advance from the operator’s own records by someone who has done the work before. The apparatus for doing it already exists and already runs. Turnaround planning, management of change and shutdown scheduling price engineering hours and outage windows for non-security work every year, closely enough for the business to commit capital against the result. It is not pointed at a required security control, because nothing asks it to be. Some of the cost stays uncertain until the architecture is fixed. The number is not missing because it cannot be produced.

It is estimable, but only of something bounded. What it takes to deploy monitoring depends on how much monitoring, and the requirement does not say. The termination condition arrives inside the cost question: with no state at which the control is finished, the figure has no scope to attach to. What gets costed in practice is an increment somebody scoped, and the scoping was theirs rather than the requirement’s.

The reason lies in the format of the requirement rather than in the difficulty of the question. A requirement is stated as an obligation, and an obligation has one column. It says what must be present. It does not say what the presence is worth against what it displaces, because the document is not constructed as a trade. Nothing in the requirement format invites the question, and nothing in the audit format asks it. The auditor establishes whether the control meets the requirement. There is no field on the form for what meeting it consumed.

Where a trade is acknowledged, it is handed back. Physical network segmentation is described as removing a single point of failure at the price of a more complex and more costly design, and evaluation of the trade is assigned to the reader’s own design process. The trade is named. It is not priced, and nothing requires the process it is assigned to to record what the price was.

An obligation with no second column is a decision taken elsewhere by someone who did not have to show their working. The operator inherits the conclusion without the arithmetic, and cannot reopen it on the terms in which it was handed down.

What the cost actually is

The cost is not the licence fee. Licence fees are visible, negotiated and readily assigned to the programme. Their visibility is part of why the real figure goes undiscussed.

The cost that goes unwritten is OT engineering competence, consumed.

Deploying and sustaining a security control on an operating plant requires people who understand both the control system and what the plant is doing with it. That is not process engineering, which knows different things about the same plant, and it is not general IT engineering either. It is specific: what this unit operation is sensitive to, what the interlock was put there for, what happens to the column if the network segment carrying that traffic is interrupted during a transition, which changes can be made online and which cannot. There are few such people at any site. They are already the constraint on the automation backlog, on obsolescence upgrades, on alarm rationalisation, on loop performance work, and on modification support for every project under way. Every hour spent on control deployment is an hour not spent there.

Part of that spend comes back outside security. An asset inventory earns its keep in obsolescence management, spares provisioning and lifecycle planning whether or not anyone ever attacks the plant, and the traffic visibility that comes with a monitoring deployment is useful to a control engineer for reasons of their own. That is an offset against the competence consumed, but it is not what the requirement asked for.

The constraint is also scoped rather than universal. A large share of security work does not require plant judgement at all. Perimeter architecture, infrastructure build, identity and access administration and patch distribution mechanics can largely be separated from plant knowledge and resourced from the general engineering and IT market. What does not separate is the part requiring a decision about what the plant will tolerate. The proportion of a programme falling into that second category determines whether the programme is constrained, and it is higher for controls placed close to the process than for controls placed at the boundary.

The constraint responds to budget, and not in a way that helps. An operator can hire people and build the required capability. It takes years, and the cost it adds is standing rather than retiring with the programme. What it returns is general engineering capacity, which is real, and no figure on the protection axis used to justify the spend. The operator who builds the capability is measurably more expensive, and more secure in a way that no available instrument will register.

Nor does the knowledge transfer at zero cost. The cost of moving that knowledge falls with similarity. Inside a fleet running the same control system, the same unit operations and the same corporate standards, an engineer moves between sites with modest loss of immediate productivity. Across the wider market, the loss is large. Whatever scale exists in this competence therefore accrues to operators and to integrators working within a fleet, not to anyone selling across the market at large.

Time to patch is the clearest instance of the cost being incurred somewhere the measurement does not look. The patch does not exist until the vendor produces and validates it, and application then requires a process window that production controls. The guidance states as much itself: updates have to be tested by the vendor and by the operator before they go in, outages have to be planned weeks ahead, and revalidation may be required as part of the process. The terms that actually move the number are therefore elsewhere: whether the installed base is still supported, whether the architecture permits patching without taking the process down, and how often the plant turns around. The architecture and the original turnaround assumption were set at capital sanction. They have moved since, but only through decisions outside the patch process: production and maintenance planning, vendor support decisions, and projects that changed the architecture without revisiting whether it could be patched online. The requirement prices none of them. The measurement bills current operations for architectural and procurement decisions that current operations did not take.

Why the cost figure stays unwritten

The cost figure stays unwritten for a reason visible in what each regime requires to be recorded.

Requirement sets are explicit that they are to be applied proportionately, and most of them describe a tailoring process for doing so. What that process is constructed to record is applicability: the control does not fit the architecture, something already in place does the same job, or a compensating measure covers the objective. The assessment records which condition applies and why. That is routine, and it carries none of the exposure attached to a rejection on cost. Tailoring on cost is a different artefact. It records that the organisation decided a required control was not worth what it would take, which is precisely the document that gets read aloud if anything happens. The provision is symmetrical on paper and the exposure attached to its two uses is not.

Under the UK regime, the safety discipline in the same plant does write the second kind of document. Measures are considered and recorded as rejected where their cost would be grossly disproportionate to the reduction they would provide. The record goes to the regulator and is as discoverable as anything else on the file. Two conditions make it writable. The reduction carries a figure, so the rejection is a comparison rather than an assertion. And the record is required, so producing it is compliance rather than admission.

Neither condition holds here. The measure a safety case rejects is bounded before it is priced: this valve, this layer, this arrangement. Most required controls are not. An operator rejecting one on cost would need what the control costs, and most of the list does not say how much of it there is. They would have to settle that themselves, price what they settled, and set the result against a reduction that carries no figure. Neither side of the comparison is anchored to anything, and the exercise never gets as far as being uncomfortable. And what the security regime requires the operator to record is applicability, not cost: the tailoring record asks whether the control fits, not what it would have taken. The document stating that a required control was not worth its cost is one nothing asks for.

Where a control is bounded by its own object, the cost is real and only the other side of the comparison is missing. There the operator is best placed to establish what the control would take and is the only party holding the artefacts needed to do it. The operator also has the strongest reason not to produce the figure.

The two figures are not missing in the same way. The benefit figure is missing on the likelihood side specifically, for reasons peculiar to that side. The cost figure is missing everywhere, for every item on the list, including the ones nobody argues about.

The control nobody argues about

One item on the list behaves differently.

The backup and recovery requirements apply at every capability level in the system requirements standard. Backup carries two enhancements, and both concern verifying that the copy is good and automating its production rather than the obligation to hold one. The recovery-and-reconstitution requirement carries none. Most requirements get harder as the level rises; recovery is expected from the start. Backup and restore attract arguments about scope, architecture, cost and timing like everything else. The argument they largely escape is whether the claimed function can be demonstrated at all. Elsewhere the guidance goes as far as naming how quickly a system can be brought back as a major characteristic of a good programme, which is a rare thing in these documents: a quality marker stated as a quantity. Operators can establish whether a restore works, how long it takes, and what state it returns the tested system to, and they can establish all three by running it.

Whether operators have run that test is a separate matter. Two things stand in the way, and neither is indiscipline. The test needs a window and the same people everything else needs. And the test returns a number: a recovery time longer than the one the organisation has been assuming is a discoverable artefact of exactly the kind the requirement format elsewhere avoids producing. What distinguishes restore from everything else on the list is not that operators measure it. It is that they can choose to measure it, without waiting for an event and without appealing to a threat rate.

The reason is that restore is not a likelihood control, and unlike most of the list it is not partly one either. It makes no event less probable. It reduces the consequence of one that has already happened, and its object is an artefact: this system, this configuration set, restored on this plant within this time. The requirement contains an operation capable of termination rather than an indefinitely extendable capability. The operator still has to supply the state and the time, but once supplied they end the claim. Restore this system to this state within this time is a statement that terminates. Either the system can be restored on those terms or it cannot.

The distinction is not that the other controls cannot be tested. A detection rule can be exercised against a specified technique, cheaply and at will, and adversary emulation does exactly that. Repeated across a programme it yields a comparison over time: whether the same objective takes longer to reach than it did before. That is more than a single rule test produces, and it is still not the figure, because the comparison is against the emulation rather than against whoever turns up. The distinction is what the test establishes. Exercising the rule confirms conformance to a specification but does not establish what the firing bought against a real adversary. Running the restore measures what it hands over: the state the system comes back in, and the time it takes to get there. Most of the list can be tested against its specification and not against what it delivers. Restore can be tested against both.

What a restore test sizes is restoration time for the specified scope, under the conditions of the test. Not risk reduction. For the causes restore principally answers, that is the figure. Against an adversarial event it is not a forecast, because the timing is not the plant’s, the concurrent load is not the plant’s, and the actor that caused the event may have reached the copies first. It pays out across causes rather than against any particular one: the same restore that answers a security event answers a corrupted upgrade, a failed disk and an engineering mistake. The rationale accompanying the backup requirement names system failure and misconfiguration rather than attack.

Those causes are not adversarial, and that is the sharper half of the point. Failed hardware generates demands at a frequency observable across an installed base, in the same way the pressure transmitter does. Corrupted changes and human error recur without an adversary selecting their moment. Restore is therefore not merely justifiable without the missing likelihood comparison. Its principal causes are among the few on the list that generate observable demand populations. They produce events with causes that can be assigned, populations that can be observed, and histories the operator already keeps in maintenance and incident records for reasons having nothing to do with security. What that buys restore is not a different relationship to attack. It is a population, generated by causes no adversary selects, in which the thing restore delivers can be observed. A restore known or assumed to be good may also change what an adversary chooses to attempt. That effect turns on what the adversary believes rather than on what the control does, and it sits on the likelihood side with everything else there: asserted, unsized, and not what restore is justified on.

Restore is the clearest instance of a general property rather than a special case. Others on the required list share it. Whether a unidirectional gateway prevents reverse traffic is testable at will, and what it attracts is argument about cost and operational friction rather than argument about whether it does what it claims. Whether a decommissioned connection is actually gone is testable. Whether a physical interlock actuates is testable. In each case the object of the question is an artefact rather than a population, and in each case there is a state at which the work is finished. What those tests reach differs. For the gateway and the connection they establish that the mechanism holds, which is the first kind of justification and not the missing figure. Restore is the case where the test also reaches what the control delivers, and that is what the consequence axis adds to a termination condition. The property that makes these controls uncontroversial on the efficacy question is not an argument against the rest. It explains what the argument about the rest is actually asking: not whether the mechanism works, but how much it buys against what else, and therefore where priority and budget should go.

What restore does not cover

Restoring configuration does not restore the operating state of the plant. It returns the control system to a condition from which operations can begin the established startup sequence, with engineering support where required. On many units that sequence takes far longer than the restore itself. The restore has delivered what it claims, and the operation is still not recovered. The recovery time that matters is not the time to reload the controllers.

Restore also assumes there is something to restore from. A control system’s application is accumulated over decades: code, tuning, interlock configuration, graphics, and modifications that were never recorded anywhere except in the system itself. Where no valid copy of that survives, the outage is not long, it is open ended. Rebuilding it can become a multi-year engineering project, and on some units the honest comparison is against not restarting the plant at all. Nothing in a recovery time figure distinguishes a plant that is four hours from running from a plant whose logic no longer exists.

And recovery works only up to a physical limit. It answers events that leave something to restore to. Past a threshold there is nothing to restore to, because the equipment has been damaged, the containment has been lost, or the consequence has been realised in a form that no configuration file addresses. A plant that comes back in four hours and a plant that cannot come back are different outcomes, and backup performance does not distinguish between them.

Where that threshold sits is an engineering question about the physical plant. The disciplines that own it can establish it in advance, and nothing in the required control set does so. Whether the plant can be driven to it is a separate question, and the required control set does not answer that either.

Close

Three things are missing from the requirements, and they are missing for three different reasons.

The benefit figure on the likelihood axis is not coming. No amount of better instrumentation will produce it, because the comparison required to establish it is not one anything observes. The termination condition is absent because nothing asks for it, and the operator specifying a control for its own plant need not wait to be asked. That is not an evidentiary question. The cost figure is estimable from records the operator already holds, and goes unproduced because nothing asks for it and the obvious way of asking creates a document nobody wants to have written.

None of this stops the work. The controls go in and they are kept, and they always have been. What is absent is the account of them, not the doing of them.

The items nobody argues about are the ones with something to show. For most of them, what shows is that the mechanism holds. On the consequence side, where the claim terminates, what shows is what the control delivered. That is not a fact about the consensus. It is a fact about what can be shown.


Where that axis produces a stopping point, and what obligation it belongs to, are treated in Compliance Has a Working Range.