The Delphi Method in Practice
Interviews that write a survey, a panel that partly overlaps itself, and a label the design does not quite earn. The interesting part is what happens when a work prints every answer it collected: twenty figures become checkable, and three turn out to be a spreadsheet's tie-break.
A hybrid Delphi, and what the hedge is doing
One examined master's work — called Work M here — is the first Delphi-influenced design this library holds. Its whole method chapter is 478 words. It gives two justifications for the design, of two different kinds, and both are worth having.
The methodological justification
- Human research was conducted as "a hybrid Delphi influenced interview and survey"
- Because it is "essential to transform the small number of opinions into a group consensus"
- Cited to a 2000 methods paper on the technique
- The same source is cited again for content validity and for the survey's warrant
The precedent justification
- The structure was selected because a 2014 study succeeded using a Delphi survey in this same industry
- A design argument from prior use in the same setting, not from a textbook
- Precedent justification is the strongest design argument in this batch
- No other work in the library offers one
The word hybrid is doing real work, and the thesis explains it twice — once of the design and once of the interview format, which it calls "a hybrid between both a structured and conversational interview". Both hedges are honest, because this is not a classical Delphi and the thesis never claims it is.
Round one: five interviews that write the survey
Five project management professionals were interviewed for approximately thirty minutes each. The interviews were recorded and transcribed electronically — a detail most works in this library omit — and the twelve-question schedule is printed as an appendix. The stated purpose was to put the literature's major points to people in the industry and have them confirm or deny each one's applicability.
THE INTERVIEW SCHEDULE, AND WHAT EACH BLOCK IS FOR
| Items | What they ask | Design consequence |
|---|---|---|
| 1 | The respondent's experience in the industry | Collected, and never reported per participant — the thesis states only a range from fifteen months to over twenty years |
| 2–3 | Whether the phenomenon is an issue and why; then, open, what the main cause is | Item 3 is the only genuinely open cause question, and it is where the thesis's most-cited interview finding comes from |
| 4–8 | One item per theme from the literature review, in the review's own chapter order, each ending "Why?" | Leading by construction — each names a cause and asks whether the respondent sees it as a cause. Item 3 precedes them, which mitigates but does not remove the problem |
| 9–10 | How the respondent defines success, and whether the phenomenon affects it | Produces the study's most interesting disconfirmation later |
| 11–12 | Experience of preventing it; how the respondent would solve it | Supplies the mitigation items that become survey questions |
Items 4–8 map one-to-one onto the literature review's five themed sections. The mapping is derived here; the thesis does not draw it.
Two things are missing from round one. Recruitment is not described: how the five were identified, how they were approached, how many declined, and on what criteria they were judged experts. And no analysis procedure is stated — no coding, no thematic analysis, no framework, no software. What the 1,503-word results section does is organise each interviewee's position under each of the literature's five themes, which is a framework analysis performed without naming it. That inference is the library's.
Round two: sixteen respondents, twenty items, one scale
The survey collates the points the interviewees raised and puts them to a larger audience. It has twenty items, a single five-point agreement scale throughout, no free text and no demographic questions. Its items are traceable back to round one: one tests a mitigation a single interviewee named, one tests a word an interviewee used, one a phrase two of them used. It is the only instrument in this library whose items can be traced to prior data collected by the same study.
Stating a target n in advance and reporting the achieved number against it is rare: across the seven examined works in this library written to one template, this is the only one that does it. The target is stated without being justified — no power calculation, no saturation argument, no citation for why fifteen to twenty rather than forty. Compare the recruitment accounts at survey design across four examined theses.
Six of the twenty items are double-barrelled or compound: two conditions and two outcomes in one sentence, or a cause coupled to a specified mechanism so that a respondent who accepts the first and not the second has no available answer. The instrument also has no "not applicable" option, and four respondents supplied one anyway — four cells in the printed dataset are marked as no answer provided. Neither instrument was piloted. See questionnaire design for what a compound item costs you at analysis.
The dataset, printed — and what that makes possible
The appendix prints the entire dataset: twenty questions by sixteen respondents, 320 cells, four marked not applicable, giving 316 answers and a completion rate of 98.75%. The thesis states none of those four numbers. It is the only work in this library whose entire quantitative result set can be independently recomputed — and that single fact is why every finding on this page exists.
The means recompute over answered responses, with the not-applicable cells excluded from the denominator. The thesis never says that is what it did, and the four affected items prove it: one item's printed mean of 4.33 is 65 divided by 15, not 65 divided by 16, which would give 4.06. A reader of the results and discussion alone would believe every figure rests on sixteen responses; four of them rest on fifteen.
Three modes that are not modes
TWENTY MEANS, TWENTY MODES, RECOMPUTED FROM THE PRINTED DATA
| Check | Result |
|---|---|
| Means recomputed against the printed value | 20 of 20 exact |
| Modes recomputed against the printed value | 17 of 20 exact |
| Failure 1 — a tie between 3 and 4, six responses each | 4 is printed |
| Failure 2 — responses of 2, 3, 4 and 5 occurring four times each: a perfectly flat distribution with no mode at all | 4 is printed, and the results section calls it "a heightened mode of 4" |
| Failure 3 — a tie between 3 and 4, six responses each, on the item the thesis names as performing worst | 3 is printed |
Recomputed for this library from the thesis's own printed appendix. These are one study's figures.
The consequences are not cosmetic. The item with no mode is the most polarised in the survey — four scale points, four responses each — and is reported as having modal agreement. The item whose mode is a tie is the one the thesis names as performing worst, so a substantive claim rests on a coin-flip between "neither agree nor disagree" and "agree".
Eighteen claims, audited
Every statement the results and discussion make about the survey was checked against the printed dataset. Eighteen are checkable.
THE CLAIM AUDIT
| Verdict | Count | Examples |
|---|---|---|
| Fully correct | 12 | That no item averaged under 3.0; that exactly five items had a mode of 5; the identification and figures for the three highest-ranked items; several individual mean-and-mode pairs |
| Correct on the mean, wrong on the mode | 2 | The worst-performing item, whose mode is a tie; and the flat-distribution item, which has no mode |
| Wrong | 4 | An "equal ninth" that is equal tenth — and a third item with the identical mean is not mentioned; a "three worst" list that is not the three lowest means; four items described as "all three"; and a mode described as "lower" than one it equals |
Derived for this library. Every one of the eighteen is checkable only because the raw data are printed.
The best thing in the work is not in the arithmetic. Three times, round two overturned round one, and three times it was reported: a factor treated as significant in both the literature and the interviews was not confirmed by the survey, and the thesis says so twice. A funnel that lets the second round overturn the first is only worth building if the second round is allowed to. Here it was.
Running a study like this
What to copy, and what to add
- Justify the design twice — once methodologically, once from precedent in your own setting
- State the funnel in one sentence, so a reader knows what each round narrows
- If it is not a classical Delphi, hedge the label and say which features are absent
- State a target n before you field, and report the achieved number against it
- Trace each survey item to the round-one statement it came from
- Print your dataset as an appendix — everything on this page follows from that decision
- Say whether missing responses were excluded from the denominator, and on which items (library addition)
- Report a dispersion statistic per item; in a Delphi, an interquartile range (library addition)
- Declare your ordering rule before you rank anything (library addition)
- Check every tie in your modes by hand — a spreadsheet will silently pick one (library addition)
- Give the panel a "not applicable" option; four of these respondents improvised one
What to carry forward
- Printing the dataset is what converts a claim into a checkable claim. Twenty means here are confirmed; three modes are not; neither result is available without the appendix.
- A tie is a finding. A spreadsheet will resolve it for you silently and by position, and that resolution can reach your conclusions.
- Rank on one statistic and narrate on it too, or declare the rule you are using.
- "Hybrid Delphi influenced" is an honest hedge — but say which Delphi features you do not have, especially the consensus measure.
- Overlapping your two rounds is fine in a Delphi and not fine in a design that claims independent confirmation. Choose one story.
- A second round that is allowed to overturn the first is worth building. This one did, three times, and reported it.
Frequently asked questions
Is this a real Delphi study?
The thesis calls it "hybrid Delphi influenced" and that hedge is accurate. Four features of a classical Delphi are missing: anonymised feedback of round-one results to the panel before round two, a measure of consensus, a third round, and a stopping rule. Described plainly it is an exploratory sequential design — interviews generating survey items — which is a respectable design in its own right.
Why do three modes come out wrong when every mean is right?
Because all three are ties, and a spreadsheet's mode function resolves a tie by returning whichever tied value it meets first in the data. Two items had two values tied at six responses each; a third had four values tied at four responses each, which means it has no mode at all. All three printed values match the first-occurrence rule.
What should you report when an item has no mode?
Report that it has none, and report the distribution instead. The item in question here is the most polarised in the survey — responses spread evenly across four of the five scale points — and it is reported as having modal agreement, which reverses its meaning. A dispersion statistic beside each mean would have made the flatness visible.
Should the same people appear in both rounds?
It depends on the claim you are making. In a classical Delphi retaining the panel is correct, because the point is convergence within one expert group. This study describes round two as putting round one's points "to a larger audience" and as resting on safety in numbers, which is an independence claim — and its respondents included the interviewees. Both rationales appear in one paragraph, and the overlap is never quantified.
How big should a Delphi panel be?
This page cannot give you a number, and the work does not defend one. What it does is state a target of fifteen to twenty in advance and report sixteen achieved — the only work in this library to state a target n at all. Stating your target before you field, and reporting against it, is the transferable move; the figure is one study's.
What is the minimum I should print so my results can be checked?
The instrument, the response scale, the per-item counts and — if you can — the raw response matrix. This study printed all four, which is why twenty of its figures are verified. Without them a reader can only take reported means on trust, which is the position most works in this library are in.
References and source attribution
- An examined master's work supplied as student work: a two-round hybrid Delphi-influenced study combining five semi-structured interviews with a twenty-item survey of sixteen respondents, with both instruments and the complete raw response matrix printed as appendices. Researcher, supervisor, institution, jurisdiction and industry organisation scrubbed. Used as observed practice, not as a model answer.
- The thesis's Appendix C, a 20 × 16 response matrix with four cells marked not applicable, from which every mean and mode on this page was independently recomputed for this library.
- A 2000 methods paper on the Delphi technique, cited twice by the thesis — for the claim that a small number of opinions must be transformed into group consensus, and for the survey's epistemic warrant. Its consensus apparatus is not used by the study.
- The supplied teaching source: weekly study notes, slide decks and assessment activities for a master's-level research methods subject in project management, which teaches neither the Delphi technique nor any consensus measure. Author, institution and year not stated in the supplied files.
Suggested questions for Ask KEVOS
- How do I turn interview findings into survey items I can defend?
- What should I report when a Likert item's mode is a tie?
- Which consensus measure belongs in a Delphi study?
- Should I print my raw response data as an appendix?
- How do I state and defend a target sample size before fielding a survey?
- What is the difference between a Delphi and an exploratory sequential design?
