Sample Size in Quantitative Studies
One published article, from one discipline, supplies every sample-size figure in six weeks of material. This page explains the mechanism behind those figures properly, and then explains exactly why the figures themselves stop at the edge of the discipline they were derived in.
Where these numbers come from, and the rule that governs them
This page draws on a published article set as the week's required reading. It appeared in a sport and exercise science web journal in 2000, states its own length as 4,318 words, and is the only document in six weeks of supplied material that gives concrete sample-size numbers. The subject linked to it; it did not teach it, does not restate it, and offers no sizing rule of its own on any slide.
That is the difficulty. The article is clear, quantified and internally consistent; the week that sets it is none of those things. A reader meeting the two together will read the article's authority backwards onto the teaching material and come away believing the subject taught a set of thresholds. It did not. The deck asks "How large should the sample be?" on one slide and never answers it — see How Large Should a Sample Be?. Sampling is named in none of the week's three learning objectives; across the six weeks on record, fourteen objectives are stated and none is fully delivered (What This Material Does Not Teach).
Three approaches to sizing a sample, and the two that survived
The article opens with a question and a promise: "How many subjects should you study? You can approach this crucial issue via statistical significance, confidence intervals, or 'on the fly'." It then delivers all three. The study notes retain that sentence in full and deliver two.
THE THREE APPROACHES, AS THE ARTICLE SETS THEM OUT
| Approach | What it asks you to size for | In the teaching material? |
|---|---|---|
| Via statistical significance | A sample "big enough for you to be sure you will detect the smallest worthwhile effect". "To be sure" means "detecting the effect 80% of the time"; "detect" means "the p value for the effect has to be less than 0.05" | Yes, in full |
| Via confidence intervals | "enough subjects to give acceptable precision for the effect you are studying", where acceptable means your interpretation would not change whether the true value sat at the upper limit or the lower | No. The whole subsection is deleted |
| On the fly | Start small, then "increase the number of subjects until you get a confidence interval that is appropriate for the magnitude of the effect" | Yes, minus its closing sentence |
The definitions in row one appear nowhere else in the supplied material. 80%, 95% and p < 0.05 are the conventions of significance testing as this article's field applied them in 2000; the teaching material never states any of them in its own voice.
Why a descriptive study needs so many more subjects than an experiment
This is the part worth learning, because the mechanism is general even where the figures attached to it are not. The article puts it in one sentence: experiments need fewer subjects "because it's easier to see changes within subjects than differences between groups of subjects".
Unpack that and it is straightforward. In a descriptive study you look for a relationship across units that differ from one another in every way at once. All of that difference is noise on top of the signal you want, and volume is the only way through it. In a before-and-after experiment each subject is compared against itself, so everything stable about it — size, complexity, team, client — cancels out, the noise shrinks, and far fewer subjects will do. A crossover goes further still: every subject receives both the real and the control treatment, so even the comparison between treatments is made within the subject.
THE ARTICLE'S DESIGN AND SAMPLE-SIZE LADDER, WITH ITS PROVENANCE MARKED
| Design | What the article states | Reached the teaching material? |
|---|---|---|
| Descriptive study | "need hundreds of subjects to give acceptable confidence intervals (or to ensure statistical significance) for small effects". The abstract puts it more strongly: "hundreds or even thousands" | Yes, "hundreds" only |
| Experiment | "generally need a lot less — often one-tenth as many" | Yes |
| Crossover | "even less — one-quarter of the number for an equivalent trial with a control group — because every subject gets the experimental treatment" | No. Deleted |
All three figures are the article's own, for physiological measurement in sport and exercise science, stated in 2000. They are orders of magnitude, not calculations — the article gives no formula and the supplied material contains none. Read them as the shape of the relationship between design and sample size, not as numbers for a proposal.
The practical consequence survives the discipline change even where the numbers do not: if you cannot recruit many units, measuring the same units before and after buys far more than the same effort spent widening a one-off survey. See Descriptive and Experimental Study Designs.
The second lever is measurement quality, and the teaching material cut it out
The article's other determinant of sample size is how well you measure: "the worse your measurements, the more subjects you need to lift the signal (the effect) out of the noise (the errors in measurement)." It splits that across the two design classes. Validity, how well a variable measures what it is supposed to, governs descriptive studies. Reliability, how reproducible a measure is on retest, governs experiments, because an experiment looks for a change in the same measure.
The reproduction, and the five cuts
Two of the five numbered sections of the week's study notes are this article reproduced word for word — about 952 words, roughly a third of the notes. Nothing marks them as quoted, and the attribution for both sits at the foot of the section before them, where a reader takes it as closing what came earlier rather than opening what follows. The first-person voice runs straight through — "I therefore recommend" — belonging to an author the notes never name in the running text.
THE FIVE DELETIONS FROM THE REPRODUCED SAMPLE-SIZE SECTION
| What was cut | Length | What the cut costs the reader |
|---|---|---|
| The whole subsection "Via Confidence Intervals" | 108 words | The concept the retained subsection is built on is never introduced, and the comparison of the two surviving approaches goes with it |
| The closing sentence of "On the Fly", where the author states he had run simulations showing the resulting effect magnitudes are not substantially biased | 17 words | His own evidence for his own recommendation is removed |
| The closing two sentences of "Effect of Research Design", giving the crossover figure | 36 words | The largest of the three reductions goes, leaving the ladder with two of its three rungs |
| The whole subsection "Effect of Validity and Reliability" | 138 words | The only place in six weeks attaching a consequence to measurement quality |
| The whole subsection "Pilot Studies" | 214 words | The only passage addressed to a student without time or resources |
The samples section alongside it carries two further sentence deletions, one a cross-reference to a part of the article the notes do not contain — so the cut conceals the omission rather than marking it. The reproduction is recorded as 952 words in one place and 954 where the sections are counted separately; the difference bears on none of the findings above.
The pilot study passage, and why its removal matters
The deleted subsection opens: "As a student researcher, you might not have enough time or resources to get a sample of optimum size. Your study can nevertheless be a pilot for a larger study." That is the article speaking to the situation most readers of this material are in, and it is the passage that did not survive.
What the deleted passage said a pilot is for
Check the techniques
"to develop, adapt, or check the feasibility of techniques" — whether the instrument works, and whether the data coming back is the data you meant to collect.
Establish reliability
"to determine the reliability of measures". This ties back to the other deleted subsection: knowing a measure's reliability is what lets you size the main study.
Size the main sample
"to calculate how big the final sample needs to be", with one condition: the pilot "should have the same sampling procedure and techniques as in the larger study".
What you would have to establish before reusing any of these figures
The figures are not arbitrary. Each rests on conditions that hold in the article's field and that you would have to show hold in yours. Working through them is more useful than the numbers, because it usually shows quickly that they cannot be borrowed, and tells you what to write instead. The source does not set out this test; it is assembled from conditions the article states.
The test, before any figure from this page enters a proposal
Can you state a smallest worthwhile effect?
The article defines it as "the smallest effect that would make a difference to the lives of your subjects or to your interpretation of whatever you are studying". Every significance-based figure depends on one. If you cannot state yours, say so rather than borrow a number.
Do you know the reliability of your measure?
The 20-per-group and 10-subject figures are conditional on the measure being "highly reliable", where reliability was established by retest on a physical measurement. A schedule variance pulled from a project system, or a five-point satisfaction item, has no retest reliability you can quote.
Is your comparison within units or between them?
The one-tenth and one-quarter ratios exist because the comparison moves inside the subject; if your design compares different projects with each other, neither reduction applies. The article's subjects also come from athlete populations selected on ability, where projects across an organisation vary on far more dimensions — which pushes required numbers up, not down.
Can you get the units at all?
An organisation running forty projects a year cannot produce hundreds, whatever any figure says. The honest move is the deleted passage's: run it as a pilot, report the bounds you can support, say what a larger study would need.
What to carry forward
- The mechanism transfers, the numbers do not. Within-subject comparisons need fewer subjects than between-subject ones, and that is a design decision worth taking early.
- Every figure here — 80%, p < 0.05, hundreds, one-tenth, one-quarter, thousands rather than hundreds, 20 and 10, the first 10 or so — belongs to one article in sport and exercise science, written in 2000. The teaching material states no sizing rule in its own voice anywhere.
- The confidence interval is used nine times in the reproduced section and defined never, because the subsection defining it was cut. Nothing else supplied defines it.
- Measurement quality is the second lever: poor validity costs subjects, high reliability saves them. The only passage saying so was deleted.
- If you are short of units, say so, run it as a pilot and report bounds. That is the article's own advice, and the part the student never sees.
Frequently asked questions
Can I quote the 20-subjects-per-group figure in my proposal?
Not as a threshold. It is stated in a sport and exercise science article, in 2000, and it is explicitly conditional on the measure being highly reliable — a property established there by retesting physical measurements. If you quote it at all, quote it as that article's illustration, name the discipline, and say that the supplied material states no threshold for a project management study.
What is a confidence interval? The material uses the term constantly.
The set reading defines it as the range within which the true population value for the effect is 95% likely to fall, and defines precision in those terms. That definition was deleted from the study notes, and it is the only one anywhere in six weeks of supplied material. The retained material uses the term nine times without ever introducing it.
Why does a descriptive study need so many more subjects than an experiment?
Because it compares across subjects rather than within them. In a descriptive study every difference between your units is noise sitting on the relationship you want to see, and volume is the only way through it. In a before-and-after experiment each unit is its own comparison, so everything stable about it cancels out and far fewer units are needed.
Does the teaching material give any sample-size rule of its own?
No. The deck asks how large a sample should be on one slide and abandons the question; the phrase "sample size" does not appear in the slides at all. The only answer supplied is 387 words reproduced from the set reading, expressed in effect sizes, confidence intervals and p values that the deck never defines.
Is "getting your sample size on the fly" a legitimate approach?
It is one author's explicit personal recommendation — the original reads "I therefore recommend" — and the sentence in which he offered his evidence for it was cut from the study notes. Even in the article the evidence is asserted rather than reported: no simulation design, number of runs or result is given. Attribute it; do not present it as established practice.
What should I do if I can only get a handful of units?
Run it as a pilot and say so. The deleted subsection sets out three purposes a pilot serves and one condition: if the purpose is to size a later study, the pilot must use the same sampling procedure and techniques. A small study still sets bounds on how large and how small an effect can be, which is worth reporting.
References and source attribution
- Hopkins 2000, 'Quantitative Research Design', a sport and exercise science web journal, vol. 4, no. 1. Every sample-size figure on this page is drawn from this article, which states its own length as 4,318 words and carries a notice pointing to an updated version published in 2008; the version set as required reading, and cited here, is the 2000 one.
- The supplied teaching source: the week 8 study notes, in five numbered sections, of which sections 4 and 5 reproduce the article above verbatim with five passages of section 5 deleted.
- The supplied teaching source: the week 8 slide deck on the nature of quantitative research, which poses the sample-size question on one slide and answers it on none.
Suggested questions for Ask KEVOS
- Given twelve available projects and a before-and-after design, what can I honestly claim and what should my limitations say?
- Draft a sample-size paragraph for a method section where no sizing rule from my course material applies.
- Explain the difference between sizing a sample by statistical significance and sizing it by confidence intervals.
- What would I need to know about my measure before any of the set reading's figures could apply to my study?
- Turn my under-sized study into a defensible pilot: what should it establish and what should it report?
