Topic Hub

Engineering Failures

What an engineering failure means

An engineering failure is commonly imagined as a physical event — a structure falls, an aircraft system misbehaves, a machine jams. Forensically, that event is the end of a chain, not the chain itself. The working definition used across PIA's investigations is broader: an engineering failure occurs when an engineered system behaves outside the envelope its designers, certifiers or operators assumed, and the gap between assumption and reality causes harm, loss or abandonment.

That definition deliberately includes failures in which nothing physically broke. Denver International Airport's automated baggage system, examined in our Denver baggage investigation, was an engineering failure of assumption — the complexity of the real operating environment exceeded what the design and its test regime had accounted for. The system was ultimately abandoned in favour of conventional tugs and carts. No component "failed"; the engineering judgement did.

Key questions the topic raises

  • How do valid design decisions become invalid as a programme evolves, and who is responsible for noticing?
  • What happens to engineering judgement when schedule pressure and certification deadlines coincide?
  • How should safety-critical assumptions be challenged inside an organisation that has already committed to them?
  • Where does responsibility sit when failure emerges at the interface between companies, not within one?
  • What do post-failure investigations consistently find that pre-failure assurance consistently missed?

Assumption risk: the quiet killer

The deepest pattern in engineering failure is not defective calculation but unexamined assumption. Designs rest on stated and unstated premises about loads, environments, human behaviour and maintenance regimes. Programmes change — scope grows, sites shift, schedules compress — and the premises quietly stop being true while the design built on them continues.

The Boeing 737 MAX, examined in our 737 MAX investigation, is the defining modern case. Two accidents — Lion Air Flight 610 in October 2018 and Ethiopian Airlines Flight 302 in March 2019 — killed 346 people and led to a worldwide grounding of the fleet for approximately twenty months. At the centre was MCAS, a flight-control system whose design assumptions about pilot response and single-sensor input were invalidated by the operational reality. The engineering question was inseparable from a governance question: how certification assumptions, delegated authority and competitive pressure on the programme interacted.

Integration and interface failure

Large engineered systems are assembled from systems that were each, in isolation, engineered competently. Failure concentrates at the interfaces — physical, digital and organisational. The Integration Risk Ladder describes how risk escalates as integration is deferred: the later incompatible assumptions are discovered, the fewer and worse the options for resolving them.

Berlin Brandenburg Airport, covered in our Berlin Brandenburg investigation, illustrates the non-structural version of this: a building whose smoke-extraction and fire-safety systems, among others, could not be certified as an integrated whole, contributing to an opening delay of roughly nine years. The individual engineering was rarely the headline problem; the integration, and the governance of it, was.

Production pressure and the erosion of engineering voice

A recurrent finding in post-failure investigations is that engineers identified the problem before the event, and the organisation did not act. The mechanism is rarely a villain overriding a warning; it is the slow re-classification of engineering concerns as schedule risks to be managed rather than safety or performance facts to be respected. The Leadership Blind Spot Matrix maps exactly this condition: leadership teams whose information environment filters out unwelcome technical detail until the detail becomes undeniable.

For practitioners, the operative lesson is structural. Engineering authority must have a route to stop or escalate that does not pass through the schedule owner. Where that route exists and is used — in mature aviation and nuclear regimes — catastrophic failure rates are dramatically lower. Where it does not, the difference between a near-miss and a disaster is largely luck.

From failure to standard

The engineering profession's honest tradition is that failures rewrite codes. Bridge collapses produced modern load and fatigue standards; aviation accidents produced crew resource management and certification reform. The same discipline should apply at programme level: every major engineering failure should change the estimating, assurance and certification practice of the organisations that study it. That is the purpose of PIA's investigations — not to assign theatrical blame, but to extract the transferable correction.

Featured investigations

Related frameworks

Frequently asked questions

What is the most common cause of engineering failure?

Post-failure investigations overwhelmingly point to organisational causes rather than calculation errors: invalid assumptions carried forward, integration left too late, and concerns raised but not acted upon. The technical error is usually the visible tip of a governance failure beneath it.

Are modern engineering failures different from historical ones?

The physics has not changed; the systems have. Contemporary failures increasingly involve software, automation and human–machine interaction — as with MCAS on the 737 MAX — where behaviour is harder to test exhaustively and failure modes are less intuitive than in purely mechanical systems.

What is assumption risk?

The risk that a design premise — about loads, environments, users or maintenance — ceases to be true as the programme or operating context changes, without anyone being tasked with re-validating it. It is the least-tracked and most consequential risk class in most major programmes.

How should organisations protect engineering judgement under schedule pressure?

Structurally, not rhetorically: give the engineering authority an escalation route independent of the delivery chain, record dissent formally, and treat unresolved technical objections as a governance issue for the board, not a scheduling issue for the programme team.

Why do investigations take so long to change practice?

Because the findings are usually institutional and therefore uncomfortable. Adopting them means admitting that assurance structures, incentive systems and certification arrangements — not individuals — produced the failure. Organisations adopt technical corrections quickly and governance corrections slowly, which is why similar failures recur.

Last reviewed: 1 August 2026 · Author: Ramesh Dixit

Related Topics

The Weekly Brief

Get the Next Investigation First

Receive forensic analyses of billion-dollar failures every Monday. No fluff. Just lessons.

Subscribe to The Weekly Brief