Independence Cannot Be Discounted

Where a configuration authority can create the initiating demand and defeat the credited safety instrumented function in the same causal sequence, that function is not an independent protection layer against that cause. Its integrity level does not entitle it to partial credit there, and a probability derived from the security measures in place is not a substitute, because it describes a different conditional event. This paper does not establish, for any particular installation, that such an arrangement exists.

The wider proposition is that some dependencies do not discount protection-layer credit but defeat its admission on the affected cause-consequence pair. The shared-authority arrangement is one example: the affected function remains creditable against causes it is independent of, but leaves the product against the cause that defeats it while creating the demand. Other dependencies may produce the same result, and they are outside the scope of this paper.

The scope is one arrangement, one initiating cause and one method. The arrangement is a standing write path reaching the configuration of both the control function and the safety function. The initiating cause is intentional malicious manipulation of that path. An erroneous engineering change is a different mechanism, addressed elsewhere in the lifecycle, and it is not what follows. The distinction is not that an erroneous change cannot affect both functions, since a common download or a common project error can. It is that what follows examines intentional use of the authority, in which the demand and the defeat are selected together as the object of one sequence. Accidental and systematic common-change mechanisms need their own dependency treatment and their own frequency basis, and they are not assumed independent here.

That division between malicious action and erroneous change is the standard’s own: a note to its definition of human error expressly excludes malicious action (3.2.29, Note 2 to entry), while the security risk assessment covers intentional attacks and unintended events separately. It does not by itself settle what protection-layer credit remains admissible against the malicious cause. The method is layer of protection analysis; others are permitted and are treated at the end.

Throughout, $f$ with a subscript denotes a frequency in events per year, and $F$ with a subscript a contribution to the consequence frequency in the same units. $\mathrm{PFD}_{\mathrm{avg}}$ denotes the average probability of dangerous failure on demand assigned to a credited protection layer in this demand-mode analysis. It is dimensionless. $P$ with a subscript denotes a dimensionless probability. $f_T$ denotes the tolerable frequency applicable to the defined consequence under the site or company risk criteria.

No universal value is assigned to $f_T$ here. It belongs to a consequence rather than to an installation, and it may differ between scenarios at the same plant according to severity, risk endpoint and allocation convention. The figure used in the worked cases later is an example, and the result is the equation in $f_T$ rather than the value put into it.

The requirement and the test attached to it

Clause 11.2.9 obliges the designer of a safety instrumented system to account for, in the standard’s words, “all aspects of independence and dependency”, both with the basic process control system and with the other protection layers.1

A separate requirement addresses the initiating cause considered here directly, and it is where the authority for that cause sits. The standard requires a security risk assessment for the safety instrumented system (8.2.4), covering the devices in scope, the threats capable of exploiting vulnerabilities, intentional attacks on hardware and application programs, their consequences and likelihood, and any additional risk reduction required. The design is then required to be resilient against the security risks so identified (11.2.12). Intentional manipulation of a configuration path reaching both systems therefore falls inside requirements the standard expressly makes.

It might be said that intentional causes are discharged by that assessment, and that the layer of protection analysis is unaffected by them. That does not follow from the requirements. The obligation at 11.2.9 is not scoped to non-malicious causes; neither is the requirement that common cause, common mode and dependent failure between protection layers be assessed against the integrity those layers are required to deliver. Crediting a function in a risk calculation is a claim that it is independent of the cause it is credited against. An assessment conducted elsewhere in the lifecycle does not license that claim where the dependency it identifies reaches the credited function. The security risk assessment identifies and assesses the intentional cause and determines the requirements for any additional risk reduction. Its completion does not by itself determine what protection-layer credit remains admissible in a calculation governed by the independence and dependency requirements.

A further requirement bears on the dependency itself. The likelihood of common cause, common mode and dependent failure between protection layers, and between protection layers and the control system, must be assessed and shown to be sufficiently low in comparison to the overall safety integrity requirements of those layers (9.4.1). The comparator matters: the demonstration becomes more demanding as the integrity those layers are required to deliver increases. Above a risk reduction of ten thousand, or an average frequency of dangerous failure below $10^{-8}$ per hour, the standard makes a quantitative methodology mandatory and requires dependency and common cause between the safety system and any other layer whose failure would place a demand on it to be considered (9.2.7). That threshold is high, and the worked examples below sit well beneath it, so the provision does not mandate the treatment developed here at ordinary integrity levels. What it shows is that the drafters recognised the wider class of dependencies in which failure of another layer both places a demand on the safety system and affects its ability to respond, and required that class to be handled quantitatively where the stakes were highest.

One further requirement bears on what independence is taken to depend on. Where the control system is not intended to conform to the IEC 61511 series, clause 9.3.5 requires each control system protection layer to be independent and separate from the initiating source and from the others, to the extent that its claimed risk reduction is not compromised. A note to that clause sets out what the assessment of separation and independence can consider. The list opens with hardware, naming processing units, input and output modules, relays and field devices, and then extends to application programming, networks, the program database, engineering tools, the human machine interface and bypass tools.

Two limits apply to using that here. The requirement is scoped to control system layers rather than to the safety system, and the note is informative, so neither imposes an obligation in the case considered here. What the clause and its note show is that the standard’s treatment of independence is not confined to the components carrying the signal during a demand. A third condition is satisfied rather than limiting: the clause bites only where the control system is not being managed to the series, which is the arrangement this paper concerns.

One clause supplies something different again: not the authority for the cause, but the analytical form the test takes. Clause 11.2.10 restricts sharing. Where one device serves both systems, and its failure could produce the demand while also causing the safety function to fail dangerously, the sharing is prohibited. One exception is supplied: an analysis confirming that the overall risk is acceptable.

Three sources sit behind that sentence and they carry different weight. The prohibition and its exception are normative text in Part 1, and the normative text names no parameter: it requires only that the analysis confirm acceptable overall risk. The dangerous failure rate of the shared device appears in an informative note, offered for the case the note has in view, which is a shared device. The guidance in Part 2 elaborates on that basis and gives operational meaning to sufficiently low (A.11.2.10).

That elaboration is worth reading closely. It combines the dangerous failure rate of the shared equipment with the probability of failure of the other protection layers, expressly other than the safety instrumented function, and assesses the result against corporate risk criteria. The function that shares the equipment is named out of the product rather than entered at a reduced figure.

What the multiplication is licensed by

Layer of protection analysis produces a mitigated event frequency by multiplying an initiating event frequency by the average probability of dangerous failure on demand of each credited independent protection layer. In conventional practice, a safeguard is admitted as an independent protection layer only where it meets the applicable admission criteria, one of which is independence of the initiating cause and of every other layer credited in the same scenario. IEC 61511 separately defines a protection layer as an independent mechanism reducing risk by control, prevention or mitigation (3.2.57), and requires dependencies between layers to be addressed. The product form is a consequence of that admission condition rather than a convention adopted for tractability. The standard’s definition of dependent failure turns on exactly this (3.2.12). What marks a failure as dependent is that multiplying the unconditional probabilities of the events causing it does not give its probability. Where the condition does not hold, the product is no longer justified as a model of the joint probability, whatever the individual figures are worth on their own.

Within conventional layer of protection analysis, admission of a protection layer on a particular scenario is not graded. A safeguard meets the independence criterion or it does not, and the method supplies no convention by which a partially dependent layer contributes a fraction of its risk reduction. Dependence itself can be modelled, in a fault tree or a Markov model or elsewhere, but not by retaining the independent product and substituting a reduced figure for the affected layer. Conventional treatment does the same: an alarm sharing a transmitter with the initiating loop is refused credit against that cause while remaining creditable against causes it is independent of.

What a shared authority adds

One arrangement is taken as given: a write path reaching the configuration of both the control function and the credited safety function, standing open rather than requiring something to open it, and through which one malicious action, or one coordinated sequence of actions with no independent barrier intervening, can reach both. An engineering environment from which writes to both can be authorised or executed is the ordinary form of it. What is assumed is not merely that such a path exists but that a sequence can be completed through it: that the writes are accepted, that the resulting process movement produces the demand, and that no barrier outside the same authority intervenes. Whether that holds at a given installation is a site question and is not settled here.

One set of requirements bears on the premise directly. The standard sets requirements for the maintenance and engineering interface: the interface design is to ensure that its failure does not adversely affect execution of the safety function, and it notes that this may call for engineering interfaces to be disconnected during normal operation; access security is required over the functions that add, delete or modify the application program; and enabling and disabling read and write access is to be carried out only through a configuration management process with authentication (11.7.3).

Those requirements plainly exclude an uncontrolled, unauthenticated write path. They do not expressly address the arrangement considered here, which is an authenticated authority that remains available, operates as designed, and is used maliciously by or through an entity able to exercise it. They govern the act of enabling and disabling access rather than the duration for which access remains enabled, and the interface requirement is drafted around failure of the interface rather than its deliberate misuse. Compliance with them therefore does not by itself settle what credit the function receives against that cause, and that is the question taken up here.

Independence is a property of a layer within a scenario, not a property the layer carries everywhere. The wider guidance to the separation requirements (A.11.2.4) makes the same scenario-specific point for shared devices, stating as an expectation, in informative terms, that failure of a control system device does not both initiate the hazardous event and produce the dangerous failure, defeat or bypass of the function protecting against the specific event under evaluation, subject to a carve-out where a redundant device is able to actuate the safety system. A function coupled to a shared configuration authority remains fully independent of a stuck valve, a blocked outlet or an operator error, and its credit against those causes is untouched.

The shared authority identifies a scenario that has to be represented distinctly. Where the study contains no malicious initiating cause, the scenario is added. Where it carries the cause but credits the coupled function against it, the scenario is present and modelled wrongly, and what follows applies to the credit rather than to the omission. On that scenario a common malicious sequence supplies the demand and removes the response. One authority, or one coordinated use of that authority with no independent barrier intervening, can drive the process toward the hazard and leave the function unable to answer. The contrast that matters is not between one command and several, but between a common causal sequence and two independent events.

A material distinction separates this from the common cause examples the guidance offers. Plugged lead lines, maintenance error and misoperated isolation valves are all failures of something. A shared configuration authority need not fail at all. It can function exactly as designed, authenticate correctly, log correctly and write correctly, while the write it carries out is the one that produces the hazard. There need be no malfunction anywhere in the path, so diagnostics directed at hardware or software malfunction need not detect it. Detection requires a control concerned with authority, command legitimacy or configuration state rather than with component failure. The mechanism does not resemble the equipment failure mechanisms for which the guidance supplies an explicit parameter.

The function is therefore not independent of the selected cause. It may still operate against causes it is independent of, and an attempt that creates the demand while leaving it effective belongs to a different branch. What it cannot do is enter the product as an independent multiplier, because the initiating action as defined already includes its defeat, and any allowance for the action not succeeding belongs upstream in the characterisation of that action rather than downstream as the function’s ordinary figure. The distinction between remaining effective and being creditable is what the arithmetic turns on. The guidance is consistent with that treatment, taking the probability term over the other protection layers and naming the shared function out of it. That wording can also be read as preventing a function already credited elsewhere from being counted twice, and nothing here depends on which reading is preferred. The independence argument reaches the same place on its own.

What the arithmetic becomes

The addition is a line, a further cause-consequence pair, not a correction:

$$ \begin{aligned} F_{\mathrm{ordinary}} &= \sum_i f_i \prod_j \mathrm{PFD}_{\mathrm{avg},j} \\ F_{\mathrm{selected}} &= f_{\mathrm{sel}} \, P_{\mathrm{surv}} \\ F_{\mathrm{total}} &= F_{\mathrm{ordinary}} + F_{\mathrm{selected}} \end{aligned} $$

The index $i$ in the first term runs over the initiating causes the study already contains, and $j$ over the layers credited against each of them, with the function entered at its assessed figure against causes it is independent of. The second term covers the shared-authority cause, on which the function is not credited, and $P_{\mathrm{surv}}$ runs only over the layers that survive that authority.

Consequence modifiers such as occupancy or ignition probability are omitted throughout for clarity. In an applied analysis they belong in the relevant line unless already embedded in the frequencies, and they need not take the same values on the ordinary and selected lines.

An erroneous change made through the same authority is not this term, for the reason given in the scope. Where a single mistake both degrades the function and drives the process, the causal structure is the one described here, but the frequency basis is not, and that case needs its own dependency treatment rather than this one.

Where the criterion applies to the total, the condition on the new line is not that it fall below the tolerable frequency but that it be no greater than whatever the existing causes have left:

$$ F_{\mathrm{selected}} \le f_T - F_{\mathrm{ordinary}} $$

Allocation targets the criterion, so the margin available is whatever the sizing has left over.

A high integrity level does not by itself make a function independent. The same physical function holds two different analytical statuses at once, and both are correct: it is credited at its assessed figure against every cause it is independent of, and it receives no separate credit as a protection layer against a cause that defeats it in the act of creating the demand. An integrity level establishes requirements for safety integrity. It does not establish independence from a cause that reaches the function directly.

Defining the selected frequency

Guidance to the separation requirements (A.11.2.4) states that where failure of common equipment can cause a demand, an analysis can be conducted to establish that the overall average frequency of failure satisfies expectations, and that the analysis can cover control and safety devices generally, including data communications, utilities, operator stations and engineering workstations. Configuration infrastructure is therefore not outside the dependency analysis the standard contemplates. The same guidance states that shared interfaces and devices can be managed as SIS components unless hardware and software configuration provides functional separation, and identifies restriction of writes as a consideration for preventing unauthorised or unintended writes to the safety system.

The first step here is the paper’s own: the borrowing itself. The guidance brings engineering workstations, operator stations and shared interfaces into the dependency analysis in general terms, but it does not say that a shared configuration authority is the common element the shared-device clause contemplates, and that clause and the common cause requirements alongside it are drafted in the language of failure. A configuration authority used deliberately need not fail at all, and the standard’s own separation of malicious action from human error suggests the failure-drafted provisions were not written to carry intentional acts. So the authority for the malicious cause rests on the security requirements set out earlier, and the shared-device clause is borrowed for the shape of its test rather than for its trigger. That is an argument from analogy and it is offered as one.

The second step, selecting the parameter, is also the paper’s own. The normative requirement names none, so selecting one suited to the shared element in front of the analyst is what the normative text leaves open rather than a departure from what it specified. The need for a quantity of this kind is not left open once the method is chosen, however. The security risk assessment is required to result in a determination of the requirements for additional risk reduction (8.2.4), and a requirement for additional risk reduction cannot be determined without some characterisation of the risk it is additional to. Where the site elects to demonstrate compliance against a numerical frequency criterion through a layer of protection analysis, the cause has to be represented in a form that can enter that calculation. For a configuration authority an equipment failure rate is not the right shape. What follows proposes instead the frequency of a completed demand-and-defeat sequence, covering reachability, action and effect together, as the representation most compatible with a layer of protection analysis. It occupies the initiating frequency position on the selected line without being a frequency of access, of compromise, or of configuration change. Other methods represent the same mechanism differently.

Let $f_{\mathrm{sel}}$ denote the frequency of that chain up to the point at which the demand exists and the function is defeated: the authority is reached, an action is taken, and that action both produces the demand and removes the response. It stops there. What the surviving protection layers then do is the probability term that multiplies it, and folding them into $f_{\mathrm{sel}}$ would count them twice.

$P_{\mathrm{surv}}$ is the product of the $\mathrm{PFD}_{\mathrm{avg}}$ values of the surviving layers, so that $F_{\mathrm{selected}}$ is $f_{\mathrm{sel}}$ multiplied by $P_{\mathrm{surv}}$. A layer is surviving for this purpose if the same authority cannot prevent, bypass, reconfigure or disable its required response in the defined sequence, and one that is not drops out of the product on the reasoning already applied to the function itself. Where nothing survives, $P_{\mathrm{surv}}$ is one by the empty product convention and $F_{\mathrm{selected}}$ is $f_{\mathrm{sel}}$. A product assembled without that test is optimistic, because it includes credit the selected cause does not leave available.

Defining the cause at that point does not assume that every attempt succeeds. It places the preceding controls, the unsuccessful attempts and the incomplete sequences upstream, inside the frequency assigned to the cause. The selected cause is the completed action, and a sequence in which the demand is created while the function remains unaffected is a different branch that is not the subject of this line.

The conclusion does not depend on drawing the boundary there. An analyst who prefers to define the initiating event as the attempt reaches the same place. The contribution is then the attempt frequency multiplied by the conditional probability that the attempt both creates the demand and defeats the function, and then by the surviving layers:

$$ F_{\mathrm{selected}} = f_{\mathrm{attempt}} \, P(D \cap X \mid A) \, P_{\mathrm{surv}} $$

where $A$ is the attempt, $D$ the successful creation of the demand and $X$ the successful defeat of the function. The two forms describe the same line:

$$ f_{\mathrm{sel}} = f_{\mathrm{attempt}} \, P(D \cap X \mid A) $$

That conditional probability is a property of the malicious sequence, the architecture and the controls acting on that sequence, and it is where the effects of security measures are represented under this decomposition. It is not the function’s $\mathrm{PFD}_{\mathrm{avg}}$, which characterises the average probability that the implemented function fails to act on a demand under the assumptions of its integrity assessment, and which conditions on a different event. If it is decomposed further, the dependence has to be carried explicitly:

$$ P(D \cap X \mid A) = P(D \mid A) \, P(X \mid D, A) $$

and not as a product of unconditional terms, since the probability of defeat given that the demand was created is not the unconditional probability of defeat. Wherever the boundary is drawn, the figure that does not appear as an independent multiplier is the same one.

The path is a condition rather than an event. It does not require an engineering activity to be in progress. It is not consumed by use. Its availability is not governed by the frequency of legitimate engineering activity. Under the premise set out above, availability is an enabling condition rather than an independently sampled event, so it does not enter the product as a factor below one. Where a path is available for only part of the time, that fraction enters as a time-at-risk term. It is not a probability of failure on demand and does not belong among the layer terms, and it does not remove the line. Whether it scales the frequency proportionally is a further question: a path open half the time halves the number only where attempts are independent of the window, and where a capable actor can observe or anticipate when the path is open, no proportional reduction is justified at all. During the window the demand-and-defeat action is available in full.

What remains in $f_{\mathrm{sel}}$ is successful reach of the authority, deliberate execution, and the demand-and-defeat effect that follows. A site may hold records of access, attempts and control performance, but those records are not by themselves direct observations of the complete demand-and-defeat chain, and none by itself validates the remaining terms at the order of magnitude required.

The bound on the selected line

Five quantities are in play by this point. $f_T$ is the tolerable frequency for the consequence, $F_{\mathrm{ordinary}}$ the contribution from the causes the study already contains, $f_{\mathrm{sel}}$ the frequency of the malicious chain to the point of demand and defeat, $P_{\mathrm{surv}}$ the product of the $\mathrm{PFD}_{\mathrm{avg}}$ values of the layers whose required response the authority cannot prevent, bypass, reconfigure or disable, and $F_{\mathrm{selected}}$ the product of the last two.

Take a hazardous event whose consequence falls in a severity band for which the company has set a tolerable frequency of $10^{-5}$ per year, so $f_T$ is $10^{-5}$ per year. That figure is illustrative. The standard prescribes no universal tolerable frequency, and criteria differ between companies and between severity bands within a company.

The first plant has a demand rate of 0.1 per year from ordinary causes and three credited layers: a pressure relief device at $10^{-1}$, a process alarm with operator response at $10^{-1}$, and a safety instrumented function at $10^{-2}$:

$$ F_{\mathrm{ordinary}} = 10^{-1} \times 10^{-1} \times 10^{-1} \times 10^{-2} = 10^{-5}\ \text{per year} $$

It meets the criterion exactly and consumes the whole of it.

On the selected line the credited set is smaller. The safety instrumented function is configured through the authority and leaves. So does the alarm, which is generated and presented through systems the same authority reaches, and which can be suppressed or retargeted by the same action that drives the process. Only the relief device survives, and it survives for a reason that has to be tested rather than assumed: it has no configuration, no account and no engineering path, so nothing reaching the administrative infrastructure touches it. That establishes independence from the specified authority, which is necessary for credit on this line and not sufficient for it, since the malicious sequence must also be unable to alter the device’s capacity, routing, isolation or discharge path, and the device must still qualify as a protection layer on the ordinary grounds.

What the selected line has to fall below depends on a convention the paper cannot assume, because practice differs. Where the tolerable frequency is applied per cause-consequence pair, the new line arrives as its own pair and is assessed against the criterion directly:

$$ f_{\mathrm{sel}} \le \frac{f_T}{P_{\mathrm{surv}}} $$

With the relief device alone surviving, $10^{-5}$ divided by $10^{-1}$ gives $f_{\mathrm{sel}}$ at no more than $10^{-4}$ per year. The surviving layer buys exactly its own risk reduction and no more.

This is the treatment used below, because it is the more favourable of the two to retention. The alternative gives no more room, and often less.

Where the criterion is instead applied to the summed frequency of the consequence across its initiating causes, the new line is assessed against what the existing causes have left:

$$ f_{\mathrm{sel}} \le \frac{f_T - F_{\mathrm{ordinary}}}{P_{\mathrm{surv}}} $$

and the numerator is the remaining budget rather than the criterion. The form presumes that budget is positive; where $F_{\mathrm{ordinary}}$ already exceeds the criterion, the ordinary causes fail it before the new line is considered at all. In the scenario above the ordinary line meets the criterion exactly, so the remaining budget is nothing and no positive frequency is available at all. That does not mean the underlying risk is known exactly, and it does not describe every installation. It means the existing allocation supplies no basis for accommodating the new term. Where allocation is less tight, the remaining margin may be little more than that introduced by rounding an allocated risk reduction up to an integrity band.

The guidance does not settle which convention applies, and the argument does not depend on settling it. The admissibility question is prior to that one: whether the coupled function is a protection layer against this cause is settled inside a single cause-consequence pair, and the answer does not change with the convention applied to the results.

Now take a second plant that met the same target differently. Its demand rate is 0.1 per year and the alarm with operator response is credited at $10^{-1}$ as before, but there is no relief path appropriate to the deviation, and the shortfall was closed by allocating more integrity to the instrumented function, credited at $10^{-3}$. The combined product of the credited layers is $10^{-4}$ in both plants, the ordinary line is $10^{-5}$ per year in both, and both meet the criterion exactly.

On the selected line the two plants are not in the same position at all. In the second, every credited layer is configured through the authority, so all of them leave, $P_{\mathrm{surv}}$ is one by the empty product convention, and the bound becomes:

$$ f_{\mathrm{sel}} \le f_T $$

which under the illustrative criterion is $10^{-5}$ per year. Its reciprocal is a hundred thousand years, though that is an average rate expression rather than a prediction that events recur at regular intervals. The bound here is whatever the severity of the consequence and the company’s criteria have already fixed, and under a different applicable criterion it changes accordingly.

An order of magnitude separates the two plants, and the one that is worse placed is the one that relied on greater instrumented integrity rather than on a surviving mechanical layer. Conventionally they are equivalent. On the selected line the first has a layer the authority cannot reach and the second has none. That is not an argument against instrumented protection. It is an observation that only protection independent of the selected cause buys margin on that cause’s line. The threshold itself is arithmetically trivial in both. Establishing that the chain sits below it is the entire difficulty, and it is a difficulty of evidence rather than of calculation.

A mechanical relief device is the strongest case for a surviving layer, not a typical one. Where alarms, interlocks or secondary trips are configured through the same engineering infrastructure, they leave the line on the same reasoning as the function did. Surviving independent protection is what buys margin here, and only protection that genuinely survives counts as surviving.

What does not restore the credit

The surviving-layer test disposes of one proposal directly. Controls over who may reach the authority, whether authentication, role restriction or approval of access, are not surviving layers on this line, because the initiating action as defined is the one that got past them. Crediting such a control downstream credits it against its own failure. Its effect belongs where every other effect on reach and execution belongs, inside $f_{\mathrm{sel}}$.

The discount fails on the same reasoning, and it fails structurally rather than by producing a wrong number. Assigning the coupled function a reduced figure of $10^{-2}$, derived from the strength of the security measures in place, produces a line at $f_{\mathrm{sel}}$ multiplied by $10^{-2}$ and appears to relax the bound by two orders of magnitude. The difficulty is not that $10^{-2}$ is the wrong value. It is that the function is not admissible as a protection layer against this cause, so a $\mathrm{PFD}_{\mathrm{avg}}$ for it has no position in the product at all. Reducing a term that does not belong in the expression does not improve the expression.

A $\mathrm{PFD}_{\mathrm{avg}}$ is conditional on a demand existing. What a security measure changes is the frequency or the conditional probability of one or more stages of the malicious sequence, whether reach, execution, demand creation or defeat. Those two quantities condition on different things, and the fact that both fall between zero and one does not make them interchangeable. The measure may reduce the frequency of the chain, and any defensible numerical effect belongs in the characterisation of $f_{\mathrm{sel}}$, where it is already inside the quantity being bounded. It does not belong downstream as a figure assigned to the function. Entering it in both places counts it twice; entering it only downstream places it against the wrong conditional event. This is not a gap in method. A better estimate of the same quantity is still an estimate of the wrong conditional event, and no improvement in how it is derived moves it to the right one.

Measures that detect an altered configuration after the fact act on how long an alteration persists. That is not a $\mathrm{PFD}_{\mathrm{avg}}$ for the function either. Where the sequence has duration, such a measure may interrupt it, and the effect belongs in the frequency term.

Why the parameter is not on the same footing

The informative test described in the guidance runs on a dangerous failure rate for the shared equipment. What that figure rests on, and what it would rest on for a configuration authority, are not the same kind of thing.

The guidance is explicit that common cause, common mode and dependent failures can be evaluated and shown to be sufficiently low, and the examples named earlier are instructive. They are identifiable failure mechanisms for which failure modes can be defined, observed, tested and informed by accumulated service experience. That is the class of cause for which the guidance supplies an explicit parameterisation.

Conventional equipment failure rate information is not exact, and the contrast is not between certainty and ignorance. Transmitter and valve figures involve generic data transferred to a specific service, sparse populations, proof test coverage assumptions and a good deal of order of magnitude judgement. What they have is a framework: the failure mode is defined, the population can be observed, failures can be classified, proof testing reveals the relevant state in the installation the figure is applied to, and accepted methods exist for qualifying and transferring the data.

What $f_{\mathrm{sel}}$ counts is successful reach of the authority, deliberate execution, creation of the demand and defeat of the function. That is not an equipment failure mode, so the evidential basis used to qualify an equipment failure rate is not directly transferable to it.

The claim is not that no estimate can be constructed. Quantitative security risk methods produce estimates, and they may be the best available characterisation of the chain. Evidence about attempted access, control effectiveness and dwell time accumulates too, and is useful for other purposes. The narrower point is the one that matters. None of it, by itself, validates a complete chain frequency at the order of magnitude the criterion demands, because none of it is an observation of the complete chain. Retention therefore requires an independently reviewed model and evidence package capable of supporting a conservative upper bound at that order, and deferring the question until operating records answer it is not an assurance strategy, since they require a defensible model connecting them to the complete chain before they can.

In the instrument-only worked case, that order of magnitude is where the difficulty actually sits. Under its illustrative criterion, retaining the arrangement requires justifying an upper bound at $10^{-5}$ per year, and available operating experience cannot establish that bound by direct observation alone. A different applicable criterion produces a different numerical bound without changing the evidential question.2 The arithmetic of the threshold is trivial. The demonstration is not.

Changing the method does not dissolve the difficulty. IEC 61511 prescribes no single hazard and risk analysis methodology. Fault tree analysis, Markov modelling and full quantitative risk analysis all represent dependency explicitly and may be used where they produce the required results. Each requires a defensible quantitative characterisation of the shared-authority chain appropriate to that model, whether as a frequency, a transition rate or a conditional probability, and each therefore requires the same underlying demand-and-defeat mechanism to be characterised in the terms that model uses. A more capable model does not supply a missing characterisation. It relocates the places where the required inputs have to be entered.

What closes the test

Two principal dispositions follow. For a shared device the clause distinguishes prohibition from retention on the basis of acceptable overall risk; for a shared authority the same alternatives arise by the extension marked above.

One disposition is to remove the dependency, either by eliminating the shared authority or by providing functional separation in the hardware and software configuration so that the initiating authority can no longer write to, defeat or bypass the safety function through the selected sequence. The guidance identifies functional separation in the hardware and software configuration as the basis for distinguishing a functionally separated shared interface, and identifies management as an SIS component as one available treatment where that separation is absent. It identifies restriction of writes as a consideration for preventing unauthorised or unintended writes. Where separation is provided and demonstrated, the demand-and-defeat coupling is removed. Intentional manipulation of the control function may remain a credible initiating cause, and where it does the line remains, but the separated function can then be assessed for ordinary credit against it because the initiating authority no longer reaches it. What separation restores is the credit, not the absence of the cause. Where separation is not provided, managing the shared interface as an SIS component is a substantial obligation in its own right and may fall on an asset that was not specified, procured or maintained on that basis.

Two measures commonly offered in place of separation are neither disposition. Under the standing-access premise, reducing legitimate engineering activity does not reduce the time for which the path is available. No reduction in this term follows from that measure alone, unless it also removes or independently gates the authority between authorised uses, or its effect is separately represented and justified within the characterisation of $f_{\mathrm{sel}}$. Comparing configuration against a recorded baseline finds an alteration after it has been made. Where the demand arrives with the alteration, that is too late. Where the sequence has duration, detection independent of the compromised authority and fast enough to interrupt before the hazardous event may reduce the chain frequency, and that reduction belongs in $f_{\mathrm{sel}}$ rather than in the function’s figure. Neither measure, by being present, restores the function’s independence or removes the need to characterise the term.

The other disposition is to retain the arrangement and discharge the analysis, which is what the normative shared-device clause permits where acceptable overall risk is demonstrated, and which the informative guidance illustrates using a shared-equipment dangerous failure rate. That route requires a figure for the shared-authority chain that the site is prepared to state, defend and have reviewed, against whatever bound follows from the applicable criterion, which in the worked case with no surviving layer is $10^{-5}$ per year. Adding further protection that survives the selected authority relaxes the bound by exactly the credited risk reduction it supplies, and by nothing else, but it does not avoid the figure. It changes what the figure has to be smaller than. Every quantitative version of that route ends at the same evidential requirement.

One object in this analysis is determinable without a frequency. Whether a standing authority reaches the configuration of both the control function and the credited safety function, and whether a sequence completing the demand and the defeat is available through it, is a property of the installation as built. It is established by inspecting what the authority can write to, and it either holds or it does not. That is the premise this paper takes as given and does not settle for any site. Everything downstream of it is a frequency, and the frequency route has been worked here to show where it stops rather than to supply a method for it. Once the premise holds, no characterisation of $f_{\mathrm{sel}}$ has been validated at the order the criterion demands, and for the reason set out above direct observation cannot by itself supply one at that order. That the modelling route does not close the gap either is argued in The Control Nobody Argues About. The first question is therefore not an estimate. It is whether the reach exists, and that question is answered before any number is written down.

A general objection follows from the argument as stated. If the complete chain frequency cannot be established for the selected cause, the analysis appears to conclude against retention wherever the premise holds, and it may be objected that an analysis tending toward the same disposition whenever the premise holds is not discriminating between installations. Two things bear on that. The premise is site-specific and does not necessarily hold, and whether a standing authority can reach the configuration of both domains and complete the sequence is the question about the installation this paper does not settle. Where it does hold, the disposition is not created by the argument. Where the cause was represented and the coupled function credited against it, that credit was taken without the characterisation needed to support it on that line. Where the cause was absent, the analysis surfaces the missing line. In neither case does the argument create a deficiency; it exposes the status of the existing claim.

The asymmetry in that is deliberate. Nothing here shows that the chain frequency exceeds the budget. Failure to establish the required upper bound does not establish that it is exceeded, but neither does it supply the characterisation on which retention depends. A safety case does not rest on the absence of a demonstration that it is wrong. It rests on a demonstration that it is adequate, and the burden of supplying that demonstration falls on the party claiming the risk reduction.

Where the characterisation cannot be supported to the level of assurance the safety case requires, removing the coupling is the disposition that closes the question without relying on it, whether by eliminating the path or by establishing separation such that the initiating authority can no longer write to the safety function. The standard does not mandate that conclusion. The quantitative route has nowhere further to go without the characterisation.

The claim throughout concerns what a figure is licensed to assert. Where the selected cause uses a shared authority to create the demand and defeat the credited layer in the same causal sequence, the ordinary product is not a slightly optimistic account of that scenario. It is either an account of a different scenario, one in which the function is unaffected by the initiating cause, or an incorrect representation of the shared-authority scenario. Where the site-specific premise is established and the malicious cause is absent, it has to be added; where it is present and the coupled function is credited as an independent layer, that credit has to go, and any residual protective effect has to be represented through a dependency-aware model instead.


The wider claim this paper sits inside, that a compliance programme and a consequence-bounding demonstration are separate obligations and that the second can remain unperformed, is developed in Compliance Has a Working Range.


  1. Clause references are to BS EN 61511-1:2017+A1:2017 and BS EN 61511-2:2017, published by BSI and identical to IEC 61511-1:2016 incorporating Amendment 1:2017 and IEC 61511-2:2016. Part 1 clauses are normative; Annex A of Part 2 is informative. The quoted phrase in the opening section is reproduced for criticism and review. ↩︎

  2. Under a Poisson model, an event-free record of duration T supports a one-sided ninety-five percent upper bound of roughly three divided by T, so supporting a one-sided upper bound of $10^{-5}$ per year from observation alone would require something like three hundred thousand event-free exposure years. A hundred thousand years without an event would not establish it. The illustration is generous, since it assumes a stationary rate and comparable exposure, and neither assumption sits easily with a threat whose capability and connectivity change over time. ↩︎