KEVOS
ArticlesServicesCase studiesAboutContact
ArticlesServicesCase studiesAboutContact
← ArticlesSample Size in Quantitative StudiesProject Delivery · Research ProjectsLesson 194/216← PrevNext →
GuidePublished 16 Aug 202616 min readBy KEVOS Editorialsample sizestatistical significanceconfidence intervalstatistical power
On this page

Ask about this page

KEVOS AISample Size in Quantitative Studies

KEVOS knowledge first · trusted web sources when needed

KEVOS/Project Delivery/Research Projects/Week 8 Set Reading
Project DeliveryResearch ProjectsAdvancedQuantitative Research

Sample Size in Quantitative Studies

One published article, from one discipline, supplies every sample-size figure in six weeks of material. This page explains the mechanism behind those figures properly, and then explains exactly why the figures themselves stop at the edge of the discipline they were derived in.

Reading time18 minutes
LevelAdvanced
Topic streamQuantitative Research
Source materialWeek 8 Set Reading
Updated2026-08-16

In brief

  • Every sample-size number in the supplied material comes from one article in a sport and exercise science web journal, published in 2000. The teaching material states no sizing rule in its own voice.
  • The mechanism it explains transfers and is worth learning: a change within a subject is far easier to see than a difference between groups of subjects, which is why descriptive designs need so many more subjects. The figures themselves do not transfer — they were derived for physiological measurement on athletes.
  • About a third of the week's study notes is this article reproduced verbatim, with five passages cut. One of the cuts is the only definition of a confidence interval anywhere in six weeks of material.
  • The passage written for a student researcher without time or resources, the one on pilot studies, is also among the cuts.

Where these numbers come from, and the rule that governs them

This page draws on a published article set as the week's required reading. It appeared in a sport and exercise science web journal in 2000, states its own length as 4,318 words, and is the only document in six weeks of supplied material that gives concrete sample-size numbers. The subject linked to it; it did not teach it, does not restate it, and offers no sizing rule of its own on any slide.

That is the difficulty. The article is clear, quantified and internally consistent; the week that sets it is none of those things. A reader meeting the two together will read the article's authority backwards onto the teaching material and come away believing the subject taught a set of thresholds. It did not. The deck asks "How large should the sample be?" on one slide and never answers it — see How Large Should a Sample Be?. Sampling is named in none of the week's three learning objectives; across the six weeks on record, fourteen objectives are stated and none is fully delivered (What This Material Does Not Teach).

Caution

The rule for every figure on this page

Each number below was derived for studies of athletes and physiological measures — sprint performance, maximum oxygen consumption, muscle mass, blood lipids. They are the positions of one author writing in one discipline in 2000. None is a standard, a benchmark, a minimum or a norm, and none transfers to a project management study without a justification the article does not supply and the teaching material never attempts.

If you carry a figure from this page, the sentence carrying it must do three things: attribute it to the article, name the discipline it was derived in, and say plainly that it is not a threshold for your study — because the supplied material states no such threshold anywhere.

Three approaches to sizing a sample, and the two that survived

The article opens with a question and a promise: "How many subjects should you study? You can approach this crucial issue via statistical significance, confidence intervals, or 'on the fly'." It then delivers all three. The study notes retain that sentence in full and deliver two.

THE THREE APPROACHES, AS THE ARTICLE SETS THEM OUT

ApproachWhat it asks you to size forIn the teaching material?
Via statistical significanceA sample "big enough for you to be sure you will detect the smallest worthwhile effect". "To be sure" means "detecting the effect 80% of the time"; "detect" means "the p value for the effect has to be less than 0.05"Yes, in full
Via confidence intervals"enough subjects to give acceptable precision for the effect you are studying", where acceptable means your interpretation would not change whether the true value sat at the upper limit or the lowerNo. The whole subsection is deleted
On the flyStart small, then "increase the number of subjects until you get a confidence interval that is appropriate for the magnitude of the effect"Yes, minus its closing sentence

The definitions in row one appear nowhere else in the supplied material. 80%, 95% and p < 0.05 are the conventions of significance testing as this article's field applied them in 2000; the teaching material never states any of them in its own voice.

From the source

The deleted paragraph, quoted in full

"Using confidence intervals or confidence limits is a more accessible approach to sample-size estimation and interpretation of outcomes. You simply want enough subjects to give acceptable precision for the effect you are studying. Precision refers usually to a 95% confidence interval for the true value of the effect: the range within which the true (population) value for the effect is 95% likely to fall. Acceptable means it won't matter to your subjects (or to your interpretation of whatever you are studying) if the true value of the effect is as large as the upper limit or as small as the lower limit. A bonus of using confidence intervals to justify your choice of sample size is that the sample size is about half what you need if you use statistical significance."

That is 108 words, and it is the only definition of a confidence interval anywhere in six weeks of supplied material.

Source gap

The term is used nine times in the study notes and defined none

The subsection immediately following the deleted one — "On the Fly", which the notes keep — opens "An acceptable width for the confidence interval depends on the magnitude of the observed effect" and uses the term five times in four sentences. Across the whole of the notes' sample-size section the confidence interval is used nine times and defined never.

The deletion also removes the article's comparison of the two surviving approaches: sizing by confidence intervals costs "about half" the subjects that sizing by significance does. The notes keep both approaches and cut the sentence saying which is cheaper.

Separately: "80% of the time" is statistical power and the word power appears nowhere in the week; "p value" is used and never defined. A reader cannot look either up from what is supplied.

Why a descriptive study needs so many more subjects than an experiment

This is the part worth learning, because the mechanism is general even where the figures attached to it are not. The article puts it in one sentence: experiments need fewer subjects "because it's easier to see changes within subjects than differences between groups of subjects".

Unpack that and it is straightforward. In a descriptive study you look for a relationship across units that differ from one another in every way at once. All of that difference is noise on top of the signal you want, and volume is the only way through it. In a before-and-after experiment each subject is compared against itself, so everything stable about it — size, complexity, team, client — cancels out, the noise shrinks, and far fewer subjects will do. A crossover goes further still: every subject receives both the real and the control treatment, so even the comparison between treatments is made within the subject.

THE ARTICLE'S DESIGN AND SAMPLE-SIZE LADDER, WITH ITS PROVENANCE MARKED

DesignWhat the article statesReached the teaching material?
Descriptive study"need hundreds of subjects to give acceptable confidence intervals (or to ensure statistical significance) for small effects". The abstract puts it more strongly: "hundreds or even thousands"Yes, "hundreds" only
Experiment"generally need a lot less — often one-tenth as many"Yes
Crossover"even less — one-quarter of the number for an equivalent trial with a control group — because every subject gets the experimental treatment"No. Deleted

All three figures are the article's own, for physiological measurement in sport and exercise science, stated in 2000. They are orders of magnitude, not calculations — the article gives no formula and the supplied material contains none. Read them as the shape of the relationship between design and sample size, not as numbers for a proposal.

The practical consequence survives the discipline change even where the numbers do not: if you cannot recruit many units, measuring the same units before and after buys far more than the same effort spent widening a one-off survey. See Descriptive and Experimental Study Designs.

The second lever is measurement quality, and the teaching material cut it out

The article's other determinant of sample size is how well you measure: "the worse your measurements, the more subjects you need to lift the signal (the effect) out of the noise (the errors in measurement)." It splits that across the two design classes. Validity, how well a variable measures what it is supposed to, governs descriptive studies. Reliability, how reproducible a measure is on retest, governs experiments, because an experiment looks for a change in the same measure.

Source gap

The only passage in six weeks that attaches a consequence to validity and reliability was deleted

The article states that if the validity of the main variables is poor "you may need thousands rather than hundreds of subjects", and that with a highly reliable measure "a controlled trial with 20 subjects in each group or a crossover with 10 subjects may be sufficient to characterize even a small effect". Those are the most concrete sample-size figures in the batch and the least transferable of any, both being conditional on measurement properties established by retest on physical measurements.

The whole 138-word subsection was cut, and that matters beyond the numbers. The week defines validity and reliability in one line each and attaches no consequence to either; one slide asserts that error in the data destroys validity. The article says something more useful — poor measurement is compensable, at a price paid in subjects — and the teaching material deleted the only statement of that price. The subject's own treatment is at Validity and Reliability.

The reproduction, and the five cuts

Two of the five numbered sections of the week's study notes are this article reproduced word for word — about 952 words, roughly a third of the notes. Nothing marks them as quoted, and the attribution for both sits at the foot of the section before them, where a reader takes it as closing what came earlier rather than opening what follows. The first-person voice runs straight through — "I therefore recommend" — belonging to an author the notes never name in the running text.

THE FIVE DELETIONS FROM THE REPRODUCED SAMPLE-SIZE SECTION

What was cutLengthWhat the cut costs the reader
The whole subsection "Via Confidence Intervals"108 wordsThe concept the retained subsection is built on is never introduced, and the comparison of the two surviving approaches goes with it
The closing sentence of "On the Fly", where the author states he had run simulations showing the resulting effect magnitudes are not substantially biased17 wordsHis own evidence for his own recommendation is removed
The closing two sentences of "Effect of Research Design", giving the crossover figure36 wordsThe largest of the three reductions goes, leaving the ladder with two of its three rungs
The whole subsection "Effect of Validity and Reliability"138 wordsThe only place in six weeks attaching a consequence to measurement quality
The whole subsection "Pilot Studies"214 wordsThe only passage addressed to a student without time or resources

The samples section alongside it carries two further sentence deletions, one a cross-reference to a part of the article the notes do not contain — so the cut conceals the omission rather than marking it. The reproduction is recorded as 952 words in one place and 954 where the sections are counted separately; the difference bears on none of the findings above.

The pilot study passage, and why its removal matters

The deleted subsection opens: "As a student researcher, you might not have enough time or resources to get a sample of optimum size. Your study can nevertheless be a pilot for a larger study." That is the article speaking to the situation most readers of this material are in, and it is the passage that did not survive.

What the deleted passage said a pilot is for

PURPOSE ONE

Check the techniques

"to develop, adapt, or check the feasibility of techniques" — whether the instrument works, and whether the data coming back is the data you meant to collect.

PURPOSE TWO

Establish reliability

"to determine the reliability of measures". This ties back to the other deleted subsection: knowing a measure's reliability is what lets you size the main study.

PURPOSE THREE

Size the main sample

"to calculate how big the final sample needs to be", with one condition: the pilot "should have the same sampling procedure and techniques as in the larger study".

Source gap

"Pilot study" appears nowhere else in six weeks of material

The questionnaire week names a pilot stage and never describes it. This deleted subsection is the nearest thing in the supplied dataset to a description of one, and it is not in the teaching material — it is in an article the teaching material links to and then edits.

The passage also carries the article's answer to the underlying anxiety: a study too small for a narrow confidence interval is still worth publishing, "because your study will set useful bounds on how big and how small the effect can be", and an unpublished finding cannot contribute to any later synthesis. Its illustration that a pilot for an experimental design "can consist of the first 10 or so observations of a larger study" is hedged in the original with "or so" — an example in its own discipline, not a rule.

What you would have to establish before reusing any of these figures

The figures are not arbitrary. Each rests on conditions that hold in the article's field and that you would have to show hold in yours. Working through them is more useful than the numbers, because it usually shows quickly that they cannot be borrowed, and tells you what to write instead. The source does not set out this test; it is assembled from conditions the article states.

The test, before any figure from this page enters a proposal

  1. Can you state a smallest worthwhile effect?

    The article defines it as "the smallest effect that would make a difference to the lives of your subjects or to your interpretation of whatever you are studying". Every significance-based figure depends on one. If you cannot state yours, say so rather than borrow a number.

  2. Do you know the reliability of your measure?

    The 20-per-group and 10-subject figures are conditional on the measure being "highly reliable", where reliability was established by retest on a physical measurement. A schedule variance pulled from a project system, or a five-point satisfaction item, has no retest reliability you can quote.

  3. Is your comparison within units or between them?

    The one-tenth and one-quarter ratios exist because the comparison moves inside the subject; if your design compares different projects with each other, neither reduction applies. The article's subjects also come from athlete populations selected on ability, where projects across an organisation vary on far more dimensions — which pushes required numbers up, not down.

  4. Can you get the units at all?

    An organisation running forty projects a year cannot produce hundreds, whatever any figure says. The honest move is the deleted passage's: run it as a pilot, report the bounds you can support, say what a larger study would need.

Practice note

What to write when you cannot size a sample

The source does not prescribe this; it follows from the fact that the supplied material contains no sizing rule you can apply.

State the number of units available to you and why that is the number. State the design and what it lets you claim. State that the only sample-size guidance in your course material comes from a 2000 article in sport and exercise science, cite it, and say you have not treated its figures as thresholds because the conditions they rest on are not established for your measures. Then report your finding with its uncertainty attached rather than as a bare point estimate — that paragraph is defensible; a borrowed threshold is not.

What to carry forward

  1. The mechanism transfers, the numbers do not. Within-subject comparisons need fewer subjects than between-subject ones, and that is a design decision worth taking early.
  2. Every figure here — 80%, p < 0.05, hundreds, one-tenth, one-quarter, thousands rather than hundreds, 20 and 10, the first 10 or so — belongs to one article in sport and exercise science, written in 2000. The teaching material states no sizing rule in its own voice anywhere.
  3. The confidence interval is used nine times in the reproduced section and defined never, because the subsection defining it was cut. Nothing else supplied defines it.
  4. Measurement quality is the second lever: poor validity costs subjects, high reliability saves them. The only passage saying so was deleted.
  5. If you are short of units, say so, run it as a pilot and report bounds. That is the article's own advice, and the part the student never sees.

Frequently asked questions

Can I quote the 20-subjects-per-group figure in my proposal?

Not as a threshold. It is stated in a sport and exercise science article, in 2000, and it is explicitly conditional on the measure being highly reliable — a property established there by retesting physical measurements. If you quote it at all, quote it as that article's illustration, name the discipline, and say that the supplied material states no threshold for a project management study.

What is a confidence interval? The material uses the term constantly.

The set reading defines it as the range within which the true population value for the effect is 95% likely to fall, and defines precision in those terms. That definition was deleted from the study notes, and it is the only one anywhere in six weeks of supplied material. The retained material uses the term nine times without ever introducing it.

Why does a descriptive study need so many more subjects than an experiment?

Because it compares across subjects rather than within them. In a descriptive study every difference between your units is noise sitting on the relationship you want to see, and volume is the only way through it. In a before-and-after experiment each unit is its own comparison, so everything stable about it cancels out and far fewer units are needed.

Does the teaching material give any sample-size rule of its own?

No. The deck asks how large a sample should be on one slide and abandons the question; the phrase "sample size" does not appear in the slides at all. The only answer supplied is 387 words reproduced from the set reading, expressed in effect sizes, confidence intervals and p values that the deck never defines.

Is "getting your sample size on the fly" a legitimate approach?

It is one author's explicit personal recommendation — the original reads "I therefore recommend" — and the sentence in which he offered his evidence for it was cut from the study notes. Even in the article the evidence is asserted rather than reported: no simulation design, number of runs or result is given. Attribute it; do not present it as established practice.

What should I do if I can only get a handful of units?

Run it as a pilot and say so. The deleted subsection sets out three purposes a pilot serves and one condition: if the purpose is to size a later study, the pilot must use the same sampling procedure and techniques. A small study still sets bounds on how large and how small an effect can be, which is worth reporting.

References and source attribution

  1. Hopkins 2000, 'Quantitative Research Design', a sport and exercise science web journal, vol. 4, no. 1. Every sample-size figure on this page is drawn from this article, which states its own length as 4,318 words and carries a notice pointing to an updated version published in 2008; the version set as required reading, and cited here, is the 2000 one.
  2. The supplied teaching source: the week 8 study notes, in five numbered sections, of which sections 4 and 5 reproduce the article above verbatim with five passages of section 5 deleted.
  3. The supplied teaching source: the week 8 slide deck on the nature of quantitative research, which poses the sample-size question on one slide and answers it on none.

Suggested questions for Ask KEVOS

  • Given twelve available projects and a before-and-after design, what can I honestly claim and what should my limitations say?
  • Draft a sample-size paragraph for a method section where no sizing rule from my course material applies.
  • Explain the difference between sizing a sample by statistical significance and sizing it by confidence intervals.
  • What would I need to know about my measure before any of the set reading's figures could apply to my study?
  • Turn my under-sized study into a defensible pilot: what should it establish and what should it report?

Related KEVOS knowledge

How Large Should a Sample Be?Core · quantitative researchDescriptive and Experimental Study DesignsAdvanced · quantitative researchSampling and Selecting ParticipantsCore · data collectionControlling Bias: Randomisation and BlindingAdvanced · quantitative researchValidity and ReliabilityCore · research designWhat This Material Does Not TeachCore · research practice
KEVOS® · Project Delivery · Research Projects Page KVS-PM-RES-0194 · v1.0.0 · content 2026.08 Last reviewed 2026-08-16

Continue learning

Descriptive and Experimental Study DesignsGuide · Research ProjectsNEXT LESSON →Choosing What to MeasureGuide · Research ProjectsPlanning Data Management for a Research ProjectGuide · Research ProjectsControlling Bias: Randomisation and BlindingGuide · Research Projects
KEVOS · Engineering, manufacturing and project improvement
ArticlesServicesCase studiesAboutContact
© 2026 KEVOS®