Platform

What your plant does when the internet goes down

wifi is down

Here is a test worth running on any edge platform, including ours. Unplug the uplink at the demarc and leave it unplugged for an hour. Not a simulated outage in a console. The actual cable.

Then go and look at what the plant is doing. Not whether the dashboard in head office still loads, which it will not and does not matter. Whether the line is running, whether operators can still see their HMI, whether alarms still fire locally, and whether anything that was supposed to be recorded was recorded.

Why this catches so many platforms

Because a great many industrial edge products are a cloud application with an agent at the far end. The agent collects, the cloud decides, and the edge is a delivery mechanism. That architecture is fine until the link is gone, at which point the decisions stop being made.

It rarely fails loudly. More often something narrow stops working and nobody notices for a while. A shift report that never generated. An alarm that was defined centrally and therefore was not evaluated. A historian buffering happily until the buffer is full, and then quietly dropping the oldest data, which is the data you will eventually want.

What local first actually requires

It is more demanding than a cache, and the difference is worth spelling out because "works offline" is claimed by almost everything.

  • Control logic runs on the node, not on a decision service somewhere else. If a rule fires on a condition, the thing evaluating the rule is in the building.
  • Identity and authorisation resolve locally. An operator signing in during an outage is the case that exposes a platform that quietly depends on a remote identity provider.
  • Storage is local and survives a node, so losing one machine does not lose the data that machine was holding.
  • Node to node traffic stays on site. Two machines on the same line talking to each other should never route through anywhere else, because the moment they do, a failure two hops away becomes a production stoppage.

That last one is where most of the real engineering goes, and where the quiet failures live. It is entirely possible to build a system where everything looks local, and two nodes on the same switch are exchanging traffic through a relay in another state. It works perfectly until it does not.

Reconnection is the harder half

Surviving the outage is the part everyone designs for. Coming back is the part that breaks things.

  1. Buffered data arrives all at once, and whatever ingests it now has an hour of backlog delivered in a few seconds.
  2. Two sides may have diverged, and something has to decide which version of a record is correct rather than simply taking the most recent write.
  3. Every node reconnects at the same moment, because the same event released all of them, and the thing they reconnect to has to survive that.

A platform that handles the outage and not the recovery has moved the failure rather than removed it, and moved it somewhere with less attention on it.

Run the test

Whatever you are running, pull the cable and watch. An hour is enough to learn the answer, and it is a much better afternoon than the unplanned version of the same experiment.

If the plant keeps running and the reconnect is boring, the architecture is doing what it claimed. If something surprises you, it is better to be surprised on a Tuesday with the cable in your hand.