Episode 1
What Goes Wrong When Protection Isn't Coordinated
Imagine there's a fault on a panel somewhere in the building, could be anything, a cable issue, a ground fault, nothing dramatic. The breaker closest to that fault is supposed to trip, isolate just that circuit, and the rest of the building carries on as if nothing happened. That's the whole point of a coordinated protection scheme. But what actually happened here was the breaker further upstream tripped first, the one that feeds half the building and the whole lot went dark. Nobody got hurt, nothing was damaged, but the entire site was down for hours because of a fault that should have taken out one circuit on one floor. The reason? Nobody had checked in over a decade whether the breakers were still set up to work together. The settings had drifted apart as the building changed over the years, and nobody noticed because there was never a fault serious enough to test it.
‘The equipment is fine, it's doing exactly what it was told to do, but the problem is it was told the wrong thing, and nobody updated the instructions’
Engineers call this a discrimination failure, or a selectivity failure, same thing. It's one of the most expensive blind spots in electrical design, and it hides in plain sight. Most conversations about uptime obsess over redundancy, generators, UPS, N+1, 2N, whatever resilience tier someone's chasing. All of that is important, but none of it matters if your protection scheme can't figure out which breaker is supposed to trip. You can have the best backup power money can buy, if the wrong breaker goes first, you've still got a blackout on your hands.
Now think about that in a data centre. A fault that should be isolated to one circuit takes out a PDU or a cooling feed. The UPS kicks in, but without cooling it's buying minutes, not hours. Servers start throttling. SLAs breach. What should have been a non-event becomes a conversation with a very unhappy client. Data centres are where this hurts most, because the margins are thinnest and the financial consequences are immediate, but the same thing happens in commercial buildings, factories, hospitals, anywhere the electrical system has grown and changed over time without anyone going back to check whether the protection still works.
Selectivity studies exist to prevent exactly this. So why do they keep failing quietly for years?
To find out, we sat down with Fariha Shahid, Principal Engineer at KIPO. Fariha has been with the company for around three years, working mainly on power studies for clients, with a focus on arc flash and short circuit studies.
Before we get into the technical side, tell us a bit about you. What made you choose engineering in the first place?
Honestly, I've always been curious about how things work. I need an explanation for pretty much everything, and I've never been able to just accept things as they are. That curiosity is what led me to engineering: the chance to see the science behind everything.
When a discrimination study fails, is it usually a design flaw, or something that creeps in over time?
Is that what you see: most failures are drift, not design? How often do you walk onto a site where the study was never wrong, it's just been left behind by years of changes nobody tracked?
Strictly speaking, it's not the study that fails, it's the discrimination between devices. Usually it comes down to the settings of individual breakers changing, or the scheme itself changing over time, without anyone checking that the new settings still coordinate properly with the upstream devices.
Where does miscoordination most often hide, the kind where uptime looks fine right up until it suddenly isn't?
The most common miscoordination I see is between the utility's settings and the downstream MV devices. Very often those utility settings were never provided, were never clearly communicated, or have simply been lost in a pile of documents. That worries me most, because the settings of every downstream device depend on them.
And if a fault does occur, it doesn't only affect the client's site. It can trip the utility's side as well, which means the entire site experiences a blackout.
"The most common miscoordination I see is between the utility's settings and the downstream MV devices."
Everyone talks about selectivity as a safety issue. Why doesn't uptime get the same attention?
What would change if the people holding the budget understood that a coordination study is what stands between a tripped MCB and a site-wide event?
For me, uptime is just as important as safety. The cost of neglecting it can be the same, if not greater. UPS, standby generators and redundancy are there to support the system, but even their selectivity needs to be checked regularly over time, to make sure they'll operate as intended when a fault happens.
"Uptime is just as important as safety. The cost of neglecting it can be the same, if not greater."
Whose job is it to keep discrimination studies current, and why does that ownership keep falling through the cracks?
Is the gap between design, EPC and facilities the real issue, or is it something else?
I believe the responsibility sits with the site engineers, operators and any technical staff on site. They're the ones who should track changes made to the system and, once a change is made, check that it's properly coordinated with the other devices.
But it isn't solely on them. Decision makers also need to understand why regular studies matter, and approve the budget when it's requested.
If you could force one change on how the industry treats protection coordination, what would it be?
Documentation. Documentation. Documentation.
That means proper handover of data if a company is taken over by another, updated schematics whenever the existing system changes, and tracking of any changes to the settings.
Thank you to Fariha for her time and honesty. And if this all sounds a little familiar, here's a question to take back to your own site: when did you last check that your protection still works together?
Get in touch at contact@kipo.co.uk

