Every engineering discipline carries the memory of its failures. Bridge designers learned about wind from a bridge that shook itself apart. Offshore oil and gas regulation was rebuilt after a platform fire. A space agency changed how it handles engineering objections after losing two shuttles. These events were tragic, and many people died in them. They were also investigated with unusual thoroughness, and the findings were published, so anyone can learn from them.
The striking pattern in the major investigations is how rarely the cause was a gap in technical knowledge. In most cases, the physical mechanism behind the failure was known, or knowable, beforehand. What failed was how the organisation handled information it already had: an anomaly that had become routine, a test result read in the most convenient way, an objection that never reached the person making the decision, or knowledge that had drifted away from where it was needed.
Those problems are not unique to nuclear plants and space programs. They appear in workshops, factories, construction sites and offices whenever schedule or cost pressure meets uncertain evidence. This article summarises what several well-known investigations found and draws out lessons that apply to a business of any size. It is general information. Where your work involves safety-critical plant, hazardous materials or structural design, follow the relevant work health and safety laws and standards, and involve appropriately qualified engineers.
Seven investigations in brief
| Event | What happened | What changed afterwards |
|---|---|---|
| Three Mile Island, 1979 | A relief valve stuck open while the control room indication suggested it was closed. Operators acted reasonably on wrong information, and the reactor core was partly damaged. | Control room design, indication of actual valve position, alarm handling and simulator training |
| Challenger, 1986 | A booster joint seal failed in cold weather. Seal erosion had been seen on earlier flights, and engineers had objected to the launch. | Launch decision processes; the idea of “normalisation of deviance” became widely known |
| Chernobyl, 1986 | A reactor design that could become unstable at low power was operated in that condition during a test, with safety systems disabled. | Safety culture became a formal concept in nuclear safety |
| Piper Alpha, 1988 | A pump was restarted while a safety valve was removed for maintenance; the permit had not been passed on at shift handover. The fire killed 167 people. | The safety case approach to regulating offshore installations |
| Longford, 1998 | After a process upset at a Victorian gas plant, a very cold heat exchanger fractured when warm oil was reintroduced. Two workers died, and the state’s gas supply was disrupted for weeks. | Scrutiny of on-site hazard knowledge and of the gap between a documented safety system and an effective one |
| Columbia, 2003 | Foam struck the wing during launch. The damage was not treated as critical, and the orbiter broke up on re-entry. | The investigation board found organisational causes closely resembling Challenger’s |
| Deepwater Horizon, 2010 | A pressure test on a well gave anomalous readings that were accepted as satisfactory. The well blew out, and 11 people died. | Barrier management, independent verification of well integrity and handling of ambiguous test results |
Read across the table and the same themes keep returning. The rest of this article takes them one at a time, adds lessons from two older engineering stories, and translates each into practice for an ordinary business.
Lesson 1: a recurring anomaly is a signal, not a reassurance
Before both shuttle losses, the problem that destroyed the vehicle had been seen many times. Seal erosion had occurred on earlier flights; foam had struck orbiters before. Each time nothing catastrophic happened, the anomaly became a little more acceptable, until it was treated as normal. The sociologist Diane Vaughan’s study of the Challenger decision gave this pattern a name: normalisation of deviance, the gradual acceptance of something that departs from the design or the standard because it has not yet caused harm.
The logic is seductive and wrong. Repeated occurrence without consequence does not show that a margin is adequate. It shows only that the margin has not yet been exceeded. The anomaly is still unexplained.
In a business, normalised deviance looks like a machine that “always does that”, a test that usually needs a second attempt, a supplier whose deliveries are “normally” a bit short, or a report that is always late but “fine”. Each one is an unexplained signal. The practical response is to keep a short list of recurring anomalies, give each an owner, and either explain it properly or fix it. The small failures worth explaining article covers deciding which odd events deserve investigation.
Lesson 2: when evidence is ambiguous, lean towards stopping
At Deepwater Horizon, a pressure test meant to confirm that the well was sealed gave readings that did not make sense. An explanation was found that allowed the work to continue. At Three Mile Island and before Challenger, uncertain information was also resolved in the direction that kept things moving.
This is a predictable bias. When people are under schedule pressure, an ambiguous result invites the interpretation that allows progress, and the burden of proof quietly shifts to whoever wants to stop. A sounder rule is that an ambiguous result on a safety-critical or quality-critical check counts as a failed check until it is explained. That rule needs to be written down before the pressure arrives, because it is very hard to invent in the moment.
In a business setting, this applies to inspection results that are borderline, acceptance tests that pass on a retest, unexplained variances in financial reconciliations and readings from instruments that “must be wrong”. The question to ask is what result would have made us stop, and whether we would have accepted that result too.
Lesson 3: show what is actually happening, not what was instructed
The Three Mile Island investigation found that a control room light showed the signal that had been sent to close the relief valve, not whether the valve had actually closed. The operators believed the valve was shut because the indication said so. Everything they did afterwards followed from that one wrong belief.
The distinction between commanded state and actual state turns up everywhere. A machine display shows the set temperature, not the measured one. A project system shows a task as assigned, not as done. A purchase order shows goods as ordered, not as received and checked. An instruction to staff shows a procedure as issued, not as followed. When the commanded state is shown as though it were the actual state, people make confident decisions on false information.
Ask of every important indicator: does this show what we asked for, or what actually happened? Where it matters, measure the actual state directly, and label indicators honestly when they cannot.
Lesson 4: give objections a route that reaches the decision
Before the Challenger launch, engineers who understood the seal problem objected. During the Columbia mission, engineers raised concerns about the foam strike and asked for more information about the damage. In both cases the people with the right concern were present, and their concerns did not carry weight at the point of decision.
The failure was in the route, not the knowledge. Concerns can be filtered as they pass up a hierarchy, softened to avoid conflict, or framed so that the person raising them must prove danger rather than the decision-maker having to show safety. In a small business, the route is shorter but can be just as blocked: a junior employee who notices something wrong may not feel able to tell the owner, or may be told to stop worrying.
Practical measures include stating clearly who can stop work and on what grounds, recording objections and how they were resolved, asking the most junior person present for their view first, and treating a raised concern as a contribution, not a nuisance. When a concern is overruled, record who made the decision and why, so the decision-maker carries the weight of it rather than the person who spoke up.
Lesson 5: keep knowledge where decisions are made
The Royal Commission into the Longford explosion examined, among other things, how engineering staff had been moved from the plant to a central office, and how operators on site had not been equipped to understand the hazard of re-warming very cold equipment. The knowledge existed in the organisation. It was not present where and when the decision was made.
Small businesses face a version of this whenever expertise sits with one person, an external contractor or a manual nobody reads. When the expert is away, decisions are still made, by whoever is there. The protection is to identify the decisions where specialised knowledge matters, make sure the people making them have that knowledge or can reach it immediately, and write the critical points into the work instructions at the point of use rather than in a separate document.
Lesson 6: plan for escalation, not only the first fault
At Piper Alpha, the first leak and fire might have been survivable. What turned it into a catastrophe was escalation: connected platforms kept feeding fuel to the fire, and the accommodation provided no protected escape. Similar patterns appear at Deepwater Horizon and Longford. The initiating event was serious; what followed was worse.
Most risk thinking focuses on preventing the first fault. Equally important is asking: if this goes wrong anyway, what stops it getting worse? In practice, that means isolation points that can be reached and operated in an emergency, clear authority to shut things down, spare capacity, escape routes, backups and communication plans. For a business, it also covers non-physical escalation: a quality problem that reaches many customers because nothing stopped shipments, or a cyber incident that spreads because every system shares one password.
Lesson 7: the temporary condition often governs
The Sydney Harbour Bridge could not be built on supports from below, because the harbour was deep and busy. Each half of the arch was built out from its shore as a cantilever, held back by cables anchored in tunnels, until the two halves met in the middle. During that period, the structure carried loads quite different from those it carries as a finished arch, and it had to be safe at every stage.
Engineers learned long ago that the temporary state, during construction, lifting, transport, installation, commissioning or changeover, often produces conditions more severe than normal service. The same is true outside structures. A business is often most exposed during a system migration, a move to new premises, a changeover between products, a handover between staff, or the first weeks of a new process. Plan and check the temporary state with the same seriousness as the finished one.
Lesson 8: understand the mechanism before choosing the fix
The Tacoma Narrows Bridge in Washington State opened in July 1940 and destroyed itself four months later, in a wind well below the speed it was designed for. Its collapse is often explained as simple resonance, with the wind happening to push at the bridge’s natural frequency. The more accurate explanation is aeroelastic flutter: the bridge’s own twisting changed the way the wind acted on it, which fed energy back into the twisting. The distinction matters because the remedies differ. Changing the bridge’s natural frequency would not have cured flutter; changing the shape of the deck so it does not generate the destabilising force does. Modern long-span bridges use streamlined deck sections and wind-tunnel testing for this reason.
The business lesson is that a fix chosen for the wrong mechanism wastes money and leaves the problem in place. Before choosing a countermeasure for any persistent problem, be sure you understand how it actually happens. The fixing a recurring problem with DMAIC article covers a structured approach to finding causes before choosing solutions.
Lesson 9: design so that failure is safe
The railway air brake offers one of the cleanest examples of fail-safe design. In early “straight air” brakes, air pressure applied the brakes, so a broken air line released them, exactly when a train had split and most needed them. The later automatic brake reversed the logic. Each wagon carries its own reservoir of compressed air, and a valve applies the brakes when the pressure in the train line falls. A broken line therefore applies the brakes rather than releasing them.
The general principle: store the energy for the safe action where it is needed, and use the control signal to release it, so that losing the signal produces the safe outcome. Spring-return valves that close when air pressure is lost and machine controls that stop when the operator lets go follow the same idea. For any protective arrangement, physical or procedural, ask what happens when its own supply fails. If the answer is “nothing”, it is not protecting anything. A backup that depends on the same power supply, person or internet connection as the primary system is a common business example.
Lesson 10: test the hardest case, not the convenient one
When railways tested brakes in the 1880s, the most valuable trials used long, heavy freight trains, where braking was most difficult, rather than short passenger trains where it was easy. After the de Havilland Comet airliner accidents in the 1950s, investigators tested a complete fuselage under repeated pressurisation in a water tank, and metal fatigue was confirmed as the cause.
Tests chosen because they are convenient tend to pass. Tests chosen because they represent the most demanding real condition tell you something. In product development, that means testing at the extremes of temperature, load, usage and user skill. In a service business, it means testing a new process on the hardest customer, the busiest day or the most complex order. The tests that cannot fail article covers checking whether a test could ever detect the problem it is meant to find.
From compliance to a reasoned argument
The most significant institutional change from this period came out of the Piper Alpha inquiry: the safety case. Under a purely prescriptive approach, a regulator lists requirements and an inspector checks them. A facility can comply with every item and still be unsafe, because the list was written for facilities in general. Under a safety case approach, the operator must show, with evidence, that it has identified the hazards of its particular operation, reduced the risks so far as is reasonably practicable, and put in place controls that will actually work. The regulator assesses the quality of that argument.
In Australia, major hazard facilities and offshore petroleum operations are regulated under safety case approaches, and the general duty to eliminate or minimise risks so far as is reasonably practicable runs through work health and safety law. Safe Work Australia and state and territory regulators publish guidance.
The idea is useful well beyond regulated facilities. For any important risk, a business can ask: could we explain, with evidence, why our controls work? A procedure that exists is not the same as a control that works, and only the second protects anyone.
A worked example
This is an illustration. A 12-person business makes hydraulic hose assemblies for mining and agricultural equipment. Each assembly is crimped, then pressure tested with a hold period, during which the pressure must not fall beyond a set limit.
Over about six months, more assemblies need a second test. The usual explanation is that the test oil was cold and a small pressure fall during the hold was expected; on retest, the assemblies pass. At first, retests were about 1 in 200 assemblies. By the end of the period, they are about 1 in 40, and the retest has become an accepted step. One of the testers mentions it at a toolbox meeting and is told it is normal for winter.
Then a customer reports a hose assembly that separated at the fitting in service. Nobody is hurt. The investigation finds that one crimping die has worn, producing crimps at the small end of the acceptable range. The crimping machine’s display showed the diameter it was set to produce, not the diameter actually achieved. Marginal crimps were holding pressure on retest after the fitting settled, which is why they passed the second time.
Mapping the lessons onto the event is uncomfortable but useful. A recurring anomaly had been normalised. Ambiguous test results had been resolved towards continuing. A display showed the commanded state, not the actual state. A concern had been raised and dismissed.
The business makes several changes. Every retest is logged with its reason, and any retest now requires the assembly to be quarantined and inspected before it is tested again. Crimp diameters are measured with a gauge at the start of each shift and at set intervals, and recorded. Die wear checks are added to the maintenance schedule. The weekly production meeting now reviews the retest log, and the owner tells staff that raising a concern about quality is expected, not a complaint. The retest rate falls back below 1 in 200, and the business notifies the affected customer and works through the assemblies supplied from the period of die wear.
How this applies to a small Australian business
- List your recurring anomalies and treat each as unexplained until it is understood.
- Decide in advance that an ambiguous result on a critical check counts as a fail.
- Check whether your indicators show actual or commanded state.
- Give everyone a clear route to raise concerns, and record how they were resolved.
- Keep critical knowledge at the point of decision, not only with one expert.
- Plan for escalation: what stops a problem getting worse?
- Treat changeovers, migrations and moves as periods of higher risk.
- Understand the mechanism before choosing a fix.
- Ask what each protective arrangement does when its own supply fails.
- Test the hardest realistic case.
- Follow work health and safety law, and involve qualified engineers for safety-critical work.
Common mistakes
- Treating “it has always done that” as an explanation.
- Retesting until something passes without asking why it failed.
- Trusting a display or report without checking what it actually measures.
- Dismissing concerns because the person raising them is junior or new.
- Planning for the finished state and not the temporary one.
- Choosing a fix before the cause is understood.
- Building backups that share a single point of failure with the primary system.
- Confusing a documented procedure with an effective control.
Questions to ask
- Which odd things in our business have become normal?
- What result on our key checks would make us stop, and would we really stop?
- Do our displays and reports show what happened, or what was instructed?
- When someone raised a concern recently, what happened to it?
- Where does critical knowledge sit, and is it available when decisions are made?
- If our worst likely problem happened tomorrow, what would stop it spreading?
- Are we testing the hardest case, or the convenient one?
Bringing it together
The major engineering failures of recent decades were rarely caused by physics nobody understood. They were caused by organisations that normalised anomalies, read ambiguous evidence in the convenient direction, trusted indications of what had been instructed rather than what had happened, failed to carry objections to the decision, and let knowledge drift away from where it was needed. Older engineering stories add three more lessons: the temporary condition often governs, a fix must suit the real mechanism, and protective systems should fail safe. None of these lessons needs a large budget. They need a business willing to treat its own odd results as information, and to act on that information before it is forced to.
Source: KEVOS notes, drawing on earlier KEVOS engineering history handbooks covering major accident investigations from 1979 to 2010, long-span bridges and the railway air brake. Accounts of each event are brief summaries of published investigation findings. The worked example is an illustration. This article is general information and does not constitute professional engineering or safety advice.