OT Security

As at: July 2026

Risk Analysis in Production Networks: Deciding with Incomplete Information and a Tight Budget

Hardly any operator starts a risk analysis with a complete asset inventory, documented data flows and a budget that covers every wish. The rule is the opposite: controllers without manufacturer support, network diagrams from the last refit, a maintenance manager who carries the plant in their head, and a budget that stretches to a third of the action list. This article describes how to reach robust decisions under these conditions regardless, and why solution or network architecture is the place where gaps in knowledge are compensated for most cost-effectively.

By Jens Thies · IT by PASSION

Why the analysis almost always begins with gaps

In office IT, an inventory can be generated in days using agents and directory services. Unfortunately, that doesn't work in production. Controllers, frequency converters, scanners and operator panels run without agents. An active scan can crash an older PLC. Many plants arrived as a package from the machine builder, whose internal wiring was never documented by anyone. On top of that come remote maintenance access points set up years ago as a "temporary" measure, and network transitions that no one can identify any more. How often do you come across the typical mobile connections used for remote maintenance purposes, terminating right in the middle of production and connected directly to the internet there. Aside from any PAP infrastructure that may have been installed upstream, there is essentially no real security at this point, and updates are rarely given a thought here.

These are not omissions that need to be fixed before the risk analysis can begin. They are the normal state that the analysis must deal with. Anyone who waits until the inventory is complete will never analyse anything (and the statutory obligation doesn't wait either). The NIS2 Implementation Act has been in force without a transition period since 6 December 2025; Section 30 BSIG requires risk analysis and risk management measures, and Section 38 BSIG obliges management to approve and oversee these measures. Affected entities were required to register with the BSI by 6 March 2026.

From business process to network architecture, not the other way round

The most common methodological error is to start with the vulnerability. A scanner produces a list of outdated firmware, and an action list is derived from that. The result is a prioritisation by CVSS (Common Vulnerability Scoring System) score that has little to do with the actual damage to the business. The sound sequence is: business process, then the plant and systems that support it, then the network segments and transitions through which these systems are reachable.

In practice, this means clarifying three questions with production management, none of which require technical data:

  • Which lines or plants can be down for how long before delivery deadlines, contracts or contractual penalties take effect? A number of hours will do, a euro amount is better.
  • Which plants are safety-relevant, i.e. pose a danger to people, the environment or the equipment itself?
  • Where is there no longer a manual fallback, because recipes, order data or quality releases now only exist digitally?

These answers produce a ranking of processes. Only then are the systems behind them examined from the perspective of reachability: from where could an attacker or a misconfiguration even influence this system? This question can be answered at the network level, even if little is known about the end device itself.

Working with incomplete information

Collect passively rather than scan actively

What is inside the end device is often unknown. What runs over the network, however, can be observed without any risk to the plant. A mirror port or a TAP on the central production switch, a few days of recording, plus the switches' MAC and ARP tables, DHCP leases and (if available) NetFlow or sFlow: this provides a picture of which addresses communicate with whom, over which protocols. This picture is incomplete, but it reliably reveals the transitions that matter for the risk analysis: connections from the office network into the cell level, direct internet access from within plants, remote maintenance tunnels that remain permanently established.

Making assumptions visible

Every statement in the analysis is annotated with what it is based on: measured, taken from documentation, verbally confirmed by the plant supervisor, or assumed. This may sound bureaucratic, but it is the single most important step. A risk assessment based on the assumption that the packaging line only talks to the MES is a different thing from one that has actually seen this relationship in a recording. If the assumption later turns out to be wrong, you know immediately which assessments need to be reviewed.

Treating the unknown as its own zone

Devices that cannot be identified (unknown MAC addresses, address ranges with no contact person, a switch in a cabinet with no management access) are not ignored but managed as an "unclassified" zone. For this zone, the assessment applies the least favourable plausible assumption: no patch level, no authentication, reachable from all adjacent networks. This deliberately generates a high risk, and thus the incentive to either clarify the zone or restrict it at the network transition so that the lack of knowledge can no longer cause harm. This is precisely where network architecture pays off: a segment with a restrictive access rule is more secure even when you don't know what is inside it.

Assessing when the figures are missing

Quantitative models with annual probabilities of occurrence regularly fail in production because no one knows the probabilities. IEC 62443-3-2 therefore relies on zones and conduits. A tolerable risk is defined for each zone, from which a target security level is derived. The assessment is carried out qualitatively, using a few levels. This is entirely sufficient as a starting point, as long as the levels are defined consistently and linked to the business process. The BSI Standard 200-3 describes a comparable approach with four levels for frequency and impact.

Three rules have proven effective:

  1. The impact is derived solely from the business process, never from the technology. An unpatched Windows PC at a test station, whose failure can be manually bridged for two days, has a low impact.
  2. The probability of occurrence is estimated based on reachability: directly from the internet, from the office network, only from the production network, only physically. This classification can be read off the network diagram and is therefore robust even where knowledge of the end device is incomplete.
  3. Worst-case assumptions are kept in check. Anyone who assumes a total plant failure for every gap ends up with a list where everything is red and nothing can be prioritised any more. The realistic damage is that which the specific network path actually allows.

Budget limits as a prioritisation tool

A limited budget is not an obstacle to the risk analysis, it is its actual purpose. Without a budget limit, there would be no need for prioritisation. The question is not "What should we ideally do?" but "Which measure reduces the most risk in the most important processes per euro invested?"

Experience shows that production networks almost always produce the same order:

  1. Establish visibility. A mirror port, flow export and a simple evaluation cost little and improve every subsequent decision.
  2. Control transitions rather than harden end devices. A firewall or an ACL between office and production protects a hundred devices at once; a firmware update protects one and requires a maintenance window. Often the existing layer-3 switch with ACLs and VLANs is sufficient before new hardware is needed.
  3. Consolidate remote access. A single, logged remote maintenance access point, activated by the operator/machine operator, replaces a dozen vendor solutions with permanent tunnels.
  4. Secure the ability to restart. Backed-up control programs, recipes and network component configurations, stored offline and tested once by restoring them. This reduces the impact of almost all scenarios at once.

What is deliberately deprioritised: device hardening on plants with low process relevance, replacing old controllers purely for security reasons, and monitoring platforms with licence costs before the basic structure is in place. These points are not wrong, they simply aren't first in line.

Naming the residual risk and having it owned

What remains open after the budget decision is the residual risk. It does not disappear simply because it isn't funded. It belongs in a document, along with the process, scenario, estimated impact and the reason for deferral, signed off by management. This is not mere formality: Section 38 BSIG makes the approval and oversight of risk management measures a task for the management level, and an auditor will ask not only what was implemented, but why the rest was not.

From a technician's point of view, this document has a second benefit: it shifts the decision to where it belongs. Whether an eight-hour stoppage on line three is acceptable cannot and should not be answered by IT alone. IT provides the technical assessment of which paths lead to this stoppage and what closing them would cost. The trade-off against delivery deadlines and contractual penalties is made by the business.

Updating rather than starting over

The first analysis is rough. That is fine, as long as it follows a clear cycle. Statements marked as assumed are gradually replaced by measurements, the "unclassified" zone shrinks with every plant walkthrough, and the assessment is adjusted with every refit, every new plant and every change to network transitions. Anyone who goes through this quarterly in an hour-long session with production management and maintenance will, after a year, have a picture that a single large one-off project would never have produced – and documentation that any auditor will recognise as lived risk management.

The core message remains: incomplete information and tight budgets are not obstacles to risk analysis in production, they are its boundary conditions. Anyone who anchors the assessment in the business process, uses network reachability as the yardstick, and secures transitions ahead of end devices will, with what is available, arrive at decisions that will hold up before management and before the auditor.

Note

This article reflects the author's personal professional assessment at the time of publication. It does not replace individual consultation. Details regarding standards, deadlines, versions and manufacturer functions should be verified before making any decisions. All content is provided without guarantee.