Why This Matters: The Methods Are Logical — But Do They Actually Work?
In the previous article, we traced the evolution of quantitative risk analysis from Gantt charts through PERT/CPM to modern Monte Carlo simulation. The methods are intellectually compelling: decompose a project into tasks or cost elements, assign probability distributions to each, simulate thousands of possible project histories, and produce a cumulative probability curve for total cost and schedule.
But here is the uncomfortable truth that the pedagogical literature rarely confronts: there is remarkably little published evidence that these methods actually produce accurate predictions, particularly for the large, technologically complex projects where they are most needed — defence acquisitions, aerospace programs, first-of-class warship construction, and nuclear facility builds .
The independent research organisation critical review of the field found a striking pattern: practitioners universally affirmed that quantitative project risk analysis is "clearly useful, because it is so widely used and so widely recommended." But this endorsement was invariably followed by qualifications — that the methods are not well understood by project management, not well integrated into decision-making, and not easily explainable to senior leaders. When pressed, most experts acknowledged that virtually all evidence for the methods' utility was anecdotal .
This article examines the known problems that plague quantitative risk analysis in practice, explores the two competing philosophies of what constitutes a "good" risk analysis, and considers what a rigorous critical evaluation of these methods would require.
The Known Problems: Five Obstacles to Reliable QRA
Even in the absence of a systematic critical literature, researchers and practitioners have identified a set of recurring challenges that affect the reliability and applicability of quantitative schedule and cost risk analysis. Any honest assessment of QRA must confront these issues head-on.
Problem 1: The Aggregation Dilemma
The first practical question in any QRA is deceptively simple: at what level of detail should you decompose the project?
Consider the extremes. At one end, you could treat the entire project as a single element and ask an expert: "How long will this project take?" This is obviously inadequate — the expert will be under severe pressure to produce a number that fits the client's expectations, and if the estimate doesn't fit, they will struggle to explain what adjustments could bring it into line .
At the other extreme, you could decompose the project into thousands of micro-tasks or tiny cost elements. This creates an enormous elicitation burden, requiring experts to specify probability distributions for activities so small that meaningful uncertainty estimates become difficult to form.
The practical middle ground differs for cost and schedule:
For cost risk analysis, the WBS provides a natural hierarchical structure, and cost estimators have developed rules of thumb for appropriate decomposition levels. One widely cited recommendation suggests using no more than approximately 30 independent WBS elements for the Monte Carlo simulation . For schedule risk analysis, there is no equivalent structural guide. Deciding which tasks to include in the network — and at what granularity — is a matter of experience and judgment, with no published standards for what constitutes an appropriate level.
Problem 2: Probability Elicitation — The Achilles' Heel
Assuming you have settled on an appropriate decomposition, the next challenge is obtaining the probability distributions for each element. This is where the entire analytical framework is most vulnerable, because the quality of the output can never exceed the quality of the input.
There are three possible sources for these distributions:
Historical data from similar completed tasks or projects. This is the gold standard, but such data is rarely available in the open literature, may be proprietary, and for cutting-edge technology programs may simply not exist — there is no historical database for "time to develop a directed-energy weapon system."Parametric models (Cost Estimating Relationships). These can provide baseline estimates grounded in empirical data, but they are typically used at higher levels of aggregation and may not capture the specific uncertainties of the current project. Expert elicitation — the method recommended by virtually all pedagogical sources when data is scarce. This is where the problems intensify.
The pedagogical literature on project risk analysis generally treats elicitation superficially: explain the basic properties of probability distributions, recommend the triangular or Beta distribution, mention some cognitive biases, and move on. A critical quantitative-risk review noted a puzzling disconnect: the project risk literature makes almost no reference to the Bayesian statistical literature, where probability elicitation has been a central research topic for decades .
The Bayesian elicitation literature includes sophisticated techniques for:
- Cross-checking elicited distributions by feeding back their implications to the expert
- Allowing iterative refinement as experts see the consequences of their stated beliefs
- Addressing documented cognitive biases through structured protocols
Yet the project management literature has largely ignored this body of work, relying instead on simple three-point estimation with minimal guidance on how to conduct the conversation with the expert.
Key cognitive biases that distort elicitation include:
| Bias | Effect on Estimates | Defence/Engineering Example |
|---|---|---|
| Anchoring | Estimates cluster around an initial reference point regardless of its relevance | "The last frigate took 6 years, so this one should too" — ignoring that the new design has fundamentally different propulsion |
| Overconfidence | Ranges are too narrow; experts underestimate uncertainty | An engineer specifies 14–18 months for integration testing when the realistic range is 10–30 months |
| Availability | Recent or vivid events are overweighted | A recent high-profile cost blowout makes estimators excessively pessimistic about similar (but actually lower-risk) programs |
| Motivational bias | Estimates are distorted by what the estimator wants to be true or what will be accepted | Contractor engineers shade optimistic to win the bid; government cost analysts pad pessimistic to protect the budget |
Problem 3: Correlations — The Silent Variance Amplifier
Multiple authors have identified the treatment of correlations as one of the most significant — and most commonly neglected — challenges in QRA .
Why correlations matter: If the price of marine-grade steel increases by 20%, it does not affect just one cost element in your WBS — it affects hull fabrication, structural steelwork, deck fittings, and every other element that uses steel. Similarly, if a key subcontractor's workforce is stretched across multiple programs, their delivery delays will affect multiple tasks simultaneously. Ignoring these common-cause dependencies and treating all elements as independent will systematically understate the variance of the total project cost or duration.
The mathematical reality is straightforward: for two positively correlated cost elements X and Y, the variance of their sum is:
When the covariance term is positive (as it almost always is in practice), ignoring it produces an artificially narrow distribution — making the project look less risky than it actually is.
Why correlations are hard to elicit: Specifying a univariate probability distribution (a single uncertain quantity) is already difficult for human experts. Specifying a multivariate distribution — capturing the joint uncertainty between pairs or groups of variables — is exponentially harder. The Bayesian literature has developed conditioning techniques (asking experts to estimate one variable while holding others fixed), but these are almost never used in project risk practice .
Compounding the problem, correlations between positive random variables (which cost and duration always are) face mathematical constraints that are far less intuitive than those for normally distributed variables. A critical review demonstrated that some widely used commercial software tools claimed to model correlations but did not disclose the methods used, and that at least one common approach violated the statistical assumptions that cost risk analysts relied upon.
Problem 4: Feedback Effects and Adaptive Management
A more subtle but equally important challenge is that project managers are not passive observers of unfolding uncertainty — they are adaptive agents who respond to emerging problems with corrective actions. When a schedule begins to slip, managers can reallocate resources, authorise overtime, de-scope requirements, or bring in additional subcontractors. When costs rise, they can value-engineer components, renegotiate supplier contracts, or seek additional funding.
These feedback mechanisms mean that an initial risk analysis — conducted before any adaptive responses have occurred — may systematically overstate the probability of worst-case outcomes. The schedule might slip, but the project manager's response prevents it from slipping as far as the simulation predicted.
Quantitative-risk research proposed modelling these adaptive behaviours using intelligent agents embedded within the Monte Carlo simulation — virtual project managers that would formulate and implement response strategies during simulated project histories. While intellectually appealing, a critical quantitative-risk review noted that implementing this approach faces formidable challenges: it would require encoding not just the project structure but also the subject-matter expertise and management judgment that real project managers bring to bear. For advanced technology programs where uncertainties abound, this seems intractable.
Problem 5: Is There Enough Information?
The final obstacle is the most fundamental: for many complex, technologically novel projects, there simply may not be enough information to support a meaningful quantitative analysis. Not just insufficient historical data — insufficient information even to specify the tasks, the components, or their interrelationships.
When experts are asked to provide probability distributions for activities they cannot fully define, in domains where prior experience is limited, using technologies that have not yet been proven at scale, the resulting distributions may be so wide as to be meaningless — or so narrow as to be overconfident.
One experienced practitioner captured this tension with the observation that the value of an early project risk analysis is "inversely proportional to the number of decimal places" . At the early stages of a complex defence program, pushing for numerical precision can create a false sense of certainty that is worse than acknowledging ignorance.
The Big Question: What Makes a "Good" Risk Analysis?
The five problems above raise a more fundamental question that is rarely asked explicitly in the pedagogical literature: what are we actually trying to achieve with quantitative risk analysis? There are two fundamentally different answers, and the choice between them has profound implications for how we evaluate the methods.
Answer 1: Accuracy — "Tell Me the Right Number"
The most obvious criterion is that the analysis should produce accurate predictions. If the cost risk analysis places the P80 estimate at AUD 280 million, then across a portfolio of similar projects analysed with similar methods, approximately 80% should come in at or below AUD 280 million.
This accuracy criterion is essential when QRA outputs are used for:
- Planning contingency reserves — allocating management reserve in a defence contract
- Comparing competing proposals — selecting a contractor based on their assessed cost risk profile
- Setting milestone commitments — committing to delivery dates with stated confidence levels
- Triggering oversight thresholds — such as the statutory acquisition cost-growth breach criteria for defence cost growth
If the methods cannot deliver accuracy in this sense, they cannot legitimately be used for these high-stakes decisions.
Answer 2: Structured Thinking — "The Process Is the Product"
An alternative view — which has gained currency precisely because of the difficulties described above — holds that the primary value of QRA lies not in the numbers it generates, but in the discipline it imposes on the analytical process.
According to this perspective, the requirement to specify probability distributions forces project teams to:
- Think hard about every aspect of the project
- Put numbers on uncertainties they might otherwise gloss over
- Argue constructively with colleagues who hold different views about risk
- Identify gaps in their understanding that might otherwise remain hidden
- Prioritise by focusing attention on the highest-variance elements
This view finds support in the operations research literature. Modelling research argued that one of the most valuable effects of building a computer simulation model was that it forced analysts to examine carefully the process they were modelling — an exercise that could be worthwhile even if the model was never used or was technically incorrect.
The Tension Between the Two Views
| Criterion | Accuracy View | Structured Thinking View |
|---|---|---|
| Primary output | Reliable probability estimates | Improved project understanding |
| Success metric | Prediction accuracy vs. Actual outcomes | Quality of risk-informed decisions made |
| Evaluation method | Compare predicted vs. Actual cost/schedule | Ethnographic study of management decisions |
| Implication if methods are imprecise | Methods are failing; fix or abandon them | Methods are still valuable; precision is secondary |
| Required investment | High — demands rigorous elicitation, correlation modelling, validation | Moderate — even a rough analysis forces useful thinking |
Neither perspective has been systematically tested in the published literature. The empirical studies of cost and schedule overruns suggest that the methods may not be accurate in the predictive sense — but there has been no systematic investigation of whether the process of conducting QRA improves management outcomes.
What Would a Rigorous Evaluation Look Like?
Testing Accuracy
If accuracy is the goal, a critical evaluation would follow the structure of retrospective empirical studies: collect information from a variety of projects, document the QRA predictions at defined stages, and compare them to final outcomes. The procedure would require:
- Standardised documentation of QRA methods, inputs, and outputs at each project milestone
- Tracking of specification changes that would invalidate earlier estimates
- Comparison of predicted probability distributions against actual cost and schedule outcomes
- Evaluation of accuracy at the element level (individual task durations and cost elements) as well as the aggregate level
The problems with this approach have been well documented: historical project data is often incomplete, milestone stages are inconsistently defined, and project scope changes can make earlier analyses incomparable to later ones .
Testing Structured Thinking Value
If the benefit of QRA is as an aid to structured thinking, evaluation becomes much more complex — and expensive. It would require an ethnographic approach:
- Recording and examining project management meetings
- Documenting information discussed and how decisions were reached
- Tracing specific risks that were identified and acted on back to insights from the QRA process
- Counting risks that were foreseen versus those that were not
- Assessing whether QRA-informed projects made systematically better decisions than those that did not use QRA
Why a Critical Literature Hasn't Emerged
Five barriers have prevented the development of a rigorous critical literature on QRA effectiveness:
Fragmentation of the community. The QRA field spans operations research, business, risk analysis consulting, and software development, with publications scattered across dozens of journals, conference proceedings, and trade publications. Evolving project baselines. As projects change scope, requirements, and technology mid-stream, evaluating the accuracy of earlier analyses becomes ambiguous. Was the original estimate "wrong," or was the project fundamentally different by the time it completed? Proprietary and classified constraints. Detailed project management data is often considered proprietary in industry and classified in defence. Companies view their risk analysis capabilities as competitive advantages; governments classify project details that would enable meaningful comparison. Reputational sensitivity. A critical analysis would reveal who succeeded and who failed in quantitative terms. As long as there is general agreement that QRA "should be done" regardless of evidence about its effectiveness, there is little incentive to assemble evidence that might show poor performance. Resource constraints. Documenting and evaluating the risk analysis process requires investment above and beyond the project itself, for the benefit of the broader professional community — investment that yields no direct return to the organisation making it.
The Chiropractic Analogy: A Provocative Parallel
A critical quantitative-risk review drew a deliberately provocative comparison between the evidence base for quantitative project risk analysis and that for two health practices: classical psychoanalysis and chiropractic treatment.
In all three fields, the proposed methods are not inherently implausible — talking through problems might yield insights, spinal manipulation might help some back conditions, and assessing probabilities and running simulations is a logical approach to quantifying uncertainty. In each case, however, the primary evidence of efficacy is that people seek the treatments, therefore the treatments must be helpful — a circular argument that would not survive peer review in an evidence-based discipline.
The comparison is not meant to dismiss QRA, but to highlight a critical gap: mature professional practices should be supported by empirical evidence of effectiveness, not merely by plausibility arguments and testimonials. The chiropractic field eventually submitted to systematic evaluation; QRA has not yet done so.
What This Means for Defence and Heavy Engineering
The implications of this critical assessment are particularly acute in sectors characterised by high technological complexity, long program timelines, and enormous financial stakes.
For defence acquisition managers: Be a critical consumer of QRA outputs. Understand the assumptions behind the distributions, interrogate the treatment of correlations, and recognise that the S-curve is conditioned on the risks that were identified — it says nothing about unknown unknowns. For contractors: Invest in rigorous elicitation processes that draw on the Bayesian statistical literature, not just three-point estimation conducted in a conference room. The quality of your QRA is only as good as the quality of your input distributions. For project managers: Use QRA outputs as one input to decision-making, not as the decision itself. The structured thinking value of the process may be more reliable than the numerical outputs, particularly in the early phases of a program. For the profession: The risk management community is rich in experience but that experience is primarily anecdotal. Until empirical studies of QRA effectiveness are conducted and published, the methods must be recommended on the basis of plausibility and professional consensus rather than demonstrated efficacy.
Pitfalls: Critical Errors in QRA Practice
Treating the S-curve as a forecast. The CDF shows the distribution of outcomes given the model and its inputs. It does not account for scope changes, unknown unknowns, or adaptive management responses.
Eliciting from the wrong experts. Asking a program manager for task-level duration estimates when the machinist or design engineer has the relevant knowledge produces garbage inputs. Using independence assumptions by default. Software defaults often assume zero correlation between elements. Unless you actively model dependencies, you are systematically understating risk. Ignoring the "inversely proportional to decimal places" rule. Early in a program, presenting QRA outputs to three significant figures communicates false precision and erodes credibility with experienced decision-makers. Confusing the ritual with the substance. Conducting a QRA because the customer requires it — without genuine analytical effort — produces a document that satisfies a compliance checkbox but provides no management value. Failing to update. A QRA conducted at Milestone B and never revisited becomes an artefact, not a management tool. Risk profiles evolve continuously; the analysis must evolve with them.
Key Takeaways
Five known problems undermine the reliability of quantitative risk analysis in practice: the aggregation dilemma, the difficulty of probability elicitation, the challenge of modelling correlations, the inability to capture adaptive management responses, and insufficient information for novel technology programs.
There are two competing views of what constitutes a "good" QRA: accuracy (reliable numerical predictions) and structured thinking (the process forces rigorous examination of uncertainty). Neither has been empirically validated.
The evidence base for QRA effectiveness is almost entirely anecdotal. Despite decades of advocacy and widespread use, there is virtually no published empirical evidence demonstrating that these methods improve project outcomes.
Cognitive biases — particularly anchoring, overconfidence, and motivational bias — systematically distort the expert elicitation process that underpins the entire methodology.
Ignoring correlations is the single most common technical error in practice, and it systematically understates total project risk by ignoring common-cause dependencies between cost and schedule elements.
Critical consumers of QRA should interrogate the input assumptions, understand the limitations, and use the outputs as one input to judgment-based decision-making rather than as deterministic forecasts.
