We still don't know at a detailed level exactly what the electrical failure affected. I believe with reasonable certainty that it definitely affected the control room workstations, but we don't know whether the actual VCCs shut down.
What we do know:
- There are 4 VCCs for the system, each one a cluster of 3 CPUs. These 3 CPUs all perform the same calculations, and 2 of them must agree before any command goes out to the system. This is sometimes called redundancy, but it's more accurately "error checking". Photos: 1 2 3
- Originally these were powered by OS/2. There's a mention of them being upgraded, but I can't find anything else that confirms that.
- The control operators use workstations running Windows (XP?) and Thales NetTrac SMC and SCADA software. Photos: 1 2
- There are redundant SMC servers, which are normally used for training new control operators, but can be used if the main SMC system goes down.
- If the SMC becomes unavailable, the VCC can still operate via a command-line "degraded" mode, but things like platform destination signs don't function. It's possible that degraded mode would actually involve commanding each train one by one.
We do not know:
- If there are backup VCCs. (Unlikely, since the VCC software is proprietary and it was "prohibitively expensive" to even get more mimic diagrams set up around the control centre).
- If there is a backup power supply for the VCCs.
- If there is a backup power supply for the control center workstations.
Thursday
On Thursday, VCC2 had a network hardware issue that prevented communication with the trains.
This caused all trains on that zone to time out and emergency brake to prevent collisions. Staff attended trains and manually drove them back to stations. Once all passengers were evacuated, staff had to wait for the VCC to be repaired before it could pick up and re-enter the trains.
The total delay was about 5 hours. The majority of that time was spent after evacuating passengers, waiting for the VCC to be repaired.
Monday
On Monday, an electrical failure caused (at the minimum) the control room workstations to lose power. This prevented the use of the on-train and in-station announcement system and all trains were held system-wide. Not sure of the exact time, but it was somewhere around 12:30.
First trains were held (system service hold), but with no ability to make announcements and no working intercom system on the trains, people broke out of trains and started walking, which forced control to de-energize the guideway. I suspect at this point they decided it was already enough of a mess that staff might as well just escort people off of trains via the side walkways.
Note that power had to have been restored to the control room in order to de-energize the guideway
By 13:15, the control room had power, workstations were operational and the SCADA system was being used to de-energize and re-energize the guideway. All field staff were manually driving trains back to stations or escorting passengers off of trains and walking the guideway between stations to ensure it was safe to re-energize. The entire system was re-energized by 14:30 at the latest.
Then all trains needed to be manually driven by staff to be picked up by the VCC for automation. It took several hours to coordinate spacing of the trains, ensuring that they had been re-entered into the system correctly, ensuring there were no on-board computer errors. A handful of trains had door faults due to passengers breaking out, which required vehicle techs to fix before the trains could be automated.
The system was fully operational again by 17:30.
Note, the times and details for this incident are reflective of the Expo line. The Millennium line was running earlier, but that might be just because there were fewer trains to re-enter.
Preventative measures
I'm still not sure we know enough of the details about the electrical failure to know exactly where backup power sources are needed, but for sure there should be a review of the existing systems to ensure that all critical components have more than one source of power with no single point of failure.
As for redundant VCCs, I suspect that all VCCs would need to be upgraded and replaced before that is really feasible. You don't want your backup system to be running on totally different hardware that's untested. I'm not sure that a backup would make problems clear up faster though, since you'd probably still time out all the trains when the main VCC went down.