There is no universal number. A defensible undergraduate sample size comes from a calculation or an argument, not a convention — a power analysis for quantitative work, an information-power argument for qualitative work. Markers award credit for the justification, which means a small sample you can defend beats a larger one you cannot explain.
That answer frustrates people who wanted “thirty”. So here is what actually determines the number, how to produce a figure you can defend in your methods chapter, and what the frequently quoted rules of thumb genuinely say when you read them rather than repeat them.
Why does everyone say thirty?
Because it is memorable, and because it is loosely tied to the point at which certain sampling distributions become approximately normal. It is not a sample-size rule for your study. It carries no information about the size of the effect you are looking for, how many groups you have, or how many predictors are in your model — which are the things that actually drive the number.
If you write “a sample of 30 was deemed sufficient” with no further reasoning, you have written the weakest sentence in your methodology chapter.
How do you calculate a sample size for quantitative work?
A power analysis. It requires four quantities, any three of which determine the fourth:
- Significance level (α). Conventionally .05. In his 1992 paper “A power primer” in Psychological Bulletin, Jacob Cohen noted drily that this value is rarely stated at all and is “taken to equal .05 (part of the Fisherian legacy)”.
- Power (1 − β). Conventionally .80. Cohen described .80 as “a convention proposed for general use”, noting that it puts the ratio of β to α at 4:1 — that is, it treats a false negative as four times more tolerable than a false positive.
- Expected effect size. The hard one, dealt with below.
- Sample size. What you are solving for.
The standard free tool is G*Power, distributed by the Department of General Psychology and Work Psychology at Heinrich-Heine-Universität Düsseldorf. Be aware that it has not had a release in some years — the current versions are 3.1.9.7 for Windows, released in March 2020, and 3.1.9.6 for macOS. It remains entirely usable and citable; if you use it, cite Faul, Erdfelder, Lang and Buchner (2007) and Faul, Erdfelder, Buchner and Lang (2009), both in Behavior Research Methods.
Where do you get an expected effect size?
Three routes, in descending order of strength:
- From a comparable published study. Best. Find the closest study to yours and use its reported effect size. Cite it.
- From a meta-analysis in your area. Better still where one exists, since it averages across studies.
- From convention. Weakest, but acceptable when nothing closer exists. Cohen’s benchmarks are d = .20 / .50 / .80 for small, medium and large; r = .10 / .30 / .50; and f² = .02 / .15 / .35 for multiple and partial correlation.
If you use the conventions, quote Cohen’s own caveat rather than presenting them as established fact. He was explicit that “the definitions were made subjectively”, glossing a medium effect as one “likely to be visible to the naked eye of a careful observer”. That candour is worth reproducing in your chapter — it shows you read the source.
A sobering number from the same paper: to detect a correlation at α = .05 with power of .80 you need 783 participants for a small effect, 85 for a medium one and 28 for a large one. Most undergraduate projects are powered only for large effects. Saying so honestly is far better than pretending otherwise.

What about the regression rules of thumb?
The most cited come from Samuel Green’s 1991 paper in Multivariate Behavioral Research (26(3), 499–510), which found support for N ≥ 50 + 8m for testing a multiple correlation and N ≥ 104 + m for testing individual predictors, where m is the number of predictors.
Read the rest of the abstract, though, because Green qualified his own rules in the same breath: the first “yields values too large for N when m ≥ 7”, and both “assume all studies have a medium-size relationship”. So with six predictors and an expected medium effect, 50 + 8(6) = 98 is reasonable guidance. With ten predictors, or an expected small effect, it is not — and quoting the rule without the caveat is a misuse of the source that an attentive marker may spot.
A power analysis is always the stronger move. Use the rule of thumb as a sanity check on the answer, not as a replacement for it.
How many interviews does a qualitative dissertation need?
This is where students most often quote a number they cannot support. The evidence is genuinely more conditional than the folklore suggests.
- Guest, Bunce and Johnson (2006), in Field Methods, analysed 60 interviews and found that “saturation occurred within the first twelve interviews”, with basic elements of metathemes present as early as six. This is the source behind the widely repeated “twelve interviews” figure — from one study, in one setting.
- Hennink and Kaiser (2022), in Social Science & Medicine, systematically reviewed empirical tests of saturation and reported it being reached within 9–17 interviews or 4–8 focus groups. Critically, they attached a condition: this applied “particularly those with relatively homogenous study populations and narrowly defined objectives”. It is not a general rule, and citing it as one misrepresents the paper.
- Malterud, Siersma and Guassora (2016), in Qualitative Health Research, proposed information power instead: the more information your sample holds relevant to your aim, the fewer participants you need. They give five determinants — study aim, sample specificity, use of established theory, quality of dialogue, and analysis strategy — and deliberately attach no number. This is the most defensible framing for an undergraduate project.
- Braun and Clarke go further, arguing that saturation is not consistent with the assumptions of reflexive thematic analysis and that sample-size judgements are “inescapably situated and subjective, and cannot be determined (wholly) in advance of analysis”. Note their date if you cite them: the paper appeared online in December 2019 and in Qualitative Research in Sport, Exercise and Health in 2021.
The practical upshot: if you are doing reflexive thematic analysis, do not claim saturation. Argue information power, and justify your number from the specificity of your sample and the narrowness of your aim.
How do you write the justification?
Two or three sentences in your methods chapter, structured as inputs, output, and constraint. For a quantitative study:
An a priori power analysis was conducted in G*Power 3.1.9.7 for an independent-samples t test, with α = .05, power = .80, and an expected medium effect (d = 0.50) based on the effect reported by [source]. This indicated a required sample of 128 participants (64 per group). Recruitment within the available fieldwork window yielded 96, meaning the study was powered to detect effects of approximately d = 0.58 or larger; smaller true effects may therefore have gone undetected, and this is considered in the discussion.
That paragraph does something most undergraduate chapters do not: it states the shortfall and carries the consequence into the discussion. That is what a marker is looking for. The same principle governs the rest of the chapter — describe the decision, then defend it — as set out in our guide to writing a methodology chapter.
What if you simply cannot recruit enough people?
Common, and not fatal. Your options, in order of preference: extend recruitment; simplify the design so it needs fewer participants (fewer groups, fewer predictors, a within-subjects design instead of between-subjects); switch to secondary data; or proceed and report the study as underpowered with an honest discussion.
What you must not do is run additional analyses until something reaches significance, or drop participants after seeing what that does to your p value. Note also that a within-subjects design typically needs substantially fewer participants for the same power, which makes the choice between independent and repeated measures a sample-size decision as much as a design one — a point covered in our guide to choosing a statistical test.
An underpowered study reported honestly is a pass. A study with undisclosed analytic flexibility is an academic misconduct question.
If the number is settled and the writing is what is stalling, you can draft the methods chapter in Tesify from your own inputs — your design, your calculation, 100% written by you.
Frequently asked questions
Is 50 participants enough for a dissertation?
It depends entirely on your design and expected effect size. For a within-subjects comparison expecting a large effect, comfortably. For a multiple regression with six predictors, no. Run the power analysis and let it answer.
Do I need a power analysis if my study is qualitative?
No — power analysis applies to hypothesis testing. Use an information-power argument instead, naming the features of your sample and aim that justify the number of participants you recruited.
Can I do a power analysis after collecting data?
You can compute a sensitivity analysis, which reports the smallest effect your achieved sample could detect — that is legitimate and genuinely useful. What is not legitimate is a post-hoc “observed power” calculation based on your own result, which is circular and adds nothing.
Does my sample need to be representative?
Ideally, but undergraduate projects almost always use convenience samples, and that is accepted. What matters is that you describe the sample accurately, name the technique honestly, and state clearly what it does and does not allow you to generalise to.
How many participants do I need for a chi-square test?
Enough that expected cell frequencies are adequate — the usual guidance is that expected counts should be at least five in the large majority of cells. Small samples spread across many categories fail this quickly, so keep the number of categories down.
Should I report the sample size I aimed for or the one I got?
Both. State the target and its basis, then the achieved sample, then what the gap means for your conclusions. The gap is not an embarrassment; concealing it is.
