Sample Size and Statistical Power
Study size should reflect the research goal, desired precision and explicit assumptions rather than a universal numerical rule.
#Start with the result the study needs to estimate
Sample-size planning asks how much information is needed for a particular research question. Some studies aim to detect a difference; others aim to estimate accuracy or another quantity with useful precision. These goals require different calculations. There is no sample size that is appropriate for every study.
Statistical power is the probability that a planned test will detect a specified effect under stated assumptions. It is not the probability that a study's conclusion is correct. A study can have high planned power and still be undermined by biased recruitment, poor measurement or an unsuitable analysis.
#Make the assumptions visible
Planning may depend on the smallest difference worth detecting, expected variation, outcome frequency and the chosen error thresholds. Precision-based planning instead focuses on the expected width of an uncertainty interval. Assumptions should be justified using relevant evidence where available, and uncertainty about them should be acknowledged.
The structure of the data also matters. Repeated measurements from one person are not equivalent to measurements from separate people. Participants grouped within clinics may share important characteristics. Missing data, loss to follow-up and planned subgroup analyses can increase information needs. For prediction models, the number of outcome events can be especially important.
#Interpret study size alongside uncertainty
Checking several plausible planning assumptions shows how sensitive the required size is to uncertain inputs. Record these calculations and any allowances before recruitment or analysis where possible. If practical constraints determine the available sample, explain what precision or detectable differences that sample could reasonably support.
A result that is not statistically significant does not establish that there is no meaningful difference. Equally, a very large sample can make a trivial difference statistically significant. Interpret effect size, uncertainty intervals and practical importance together. Calculating power from the observed effect after a study generally adds little useful information.
#Common misunderstandings
There is no single sample size that makes every study reliable. A large study can estimate an effect precisely while still being misleading because of biased recruitment, poor measurement or an unsuitable comparison group. A small study is not automatically worthless: it may support a feasibility assessment, although its estimates of treatment effects may remain uncertain.
Statistical power is also easy to misread. It describes the chance of detecting a specified effect under stated assumptions, not the probability that a study’s conclusion is correct. A result that is not statistically significant does not establish that there is no effect; the range of effects compatible with the data matters.
Conversely, a statistically significant result may represent a difference too small to matter clinically. Recruiting more participants can help detect smaller differences, but cannot establish their importance. Analyses of subgroups or rare harms may require more information than the main study question.
#Questions worth asking a clinician
- How was the sample size chosen to match the research goal and desired precision, rather than a general numerical rule?
- What effect size was the study powered to detect, why was it clinically meaningful, and what assumptions informed that calculation?
- How did the sample-size calculation account for repeated measurements or participants grouped within clinics?
- How did expected missing data and the frequency of the outcome affect the number of participants needed?
- What do the estimated effect size and its confidence interval tell us about clinical benefit or harm beyond whether the result was statistically significant?