Tag: undergraduate dissertation

  • What Sample Size Do You Need for an Undergraduate Dissertation?

    What Sample Size Do You Need for an Undergraduate Dissertation?

    There is no universal number. A defensible undergraduate sample size comes from a calculation or an argument, not a convention — a power analysis for quantitative work, an information-power argument for qualitative work. Markers award credit for the justification, which means a small sample you can defend beats a larger one you cannot explain.

    That answer frustrates people who wanted “thirty”. So here is what actually determines the number, how to produce a figure you can defend in your methods chapter, and what the frequently quoted rules of thumb genuinely say when you read them rather than repeat them.

    Why does everyone say thirty?

    Because it is memorable, and because it is loosely tied to the point at which certain sampling distributions become approximately normal. It is not a sample-size rule for your study. It carries no information about the size of the effect you are looking for, how many groups you have, or how many predictors are in your model — which are the things that actually drive the number.

    If you write “a sample of 30 was deemed sufficient” with no further reasoning, you have written the weakest sentence in your methodology chapter.

    How do you calculate a sample size for quantitative work?

    A power analysis. It requires four quantities, any three of which determine the fourth:

    1. Significance level (α). Conventionally .05. In his 1992 paper “A power primer” in Psychological Bulletin, Jacob Cohen noted drily that this value is rarely stated at all and is “taken to equal .05 (part of the Fisherian legacy)”.
    2. Power (1 − β). Conventionally .80. Cohen described .80 as “a convention proposed for general use”, noting that it puts the ratio of β to α at 4:1 — that is, it treats a false negative as four times more tolerable than a false positive.
    3. Expected effect size. The hard one, dealt with below.
    4. Sample size. What you are solving for.

    The standard free tool is G*Power, distributed by the Department of General Psychology and Work Psychology at Heinrich-Heine-Universität Düsseldorf. Be aware that it has not had a release in some years — the current versions are 3.1.9.7 for Windows, released in March 2020, and 3.1.9.6 for macOS. It remains entirely usable and citable; if you use it, cite Faul, Erdfelder, Lang and Buchner (2007) and Faul, Erdfelder, Buchner and Lang (2009), both in Behavior Research Methods.

    Where do you get an expected effect size?

    Three routes, in descending order of strength:

    • From a comparable published study. Best. Find the closest study to yours and use its reported effect size. Cite it.
    • From a meta-analysis in your area. Better still where one exists, since it averages across studies.
    • From convention. Weakest, but acceptable when nothing closer exists. Cohen’s benchmarks are d = .20 / .50 / .80 for small, medium and large; r = .10 / .30 / .50; and f² = .02 / .15 / .35 for multiple and partial correlation.

    If you use the conventions, quote Cohen’s own caveat rather than presenting them as established fact. He was explicit that “the definitions were made subjectively”, glossing a medium effect as one “likely to be visible to the naked eye of a careful observer”. That candour is worth reproducing in your chapter — it shows you read the source.

    A sobering number from the same paper: to detect a correlation at α = .05 with power of .80 you need 783 participants for a small effect, 85 for a medium one and 28 for a large one. Most undergraduate projects are powered only for large effects. Saying so honestly is far better than pretending otherwise.

    Handwritten power analysis working showing effect size and required participant numbers
    Record the four inputs to your power analysis as you set them — you will need every one of them for the write-up.

    What about the regression rules of thumb?

    The most cited come from Samuel Green’s 1991 paper in Multivariate Behavioral Research (26(3), 499–510), which found support for N ≥ 50 + 8m for testing a multiple correlation and N ≥ 104 + m for testing individual predictors, where m is the number of predictors.

    Read the rest of the abstract, though, because Green qualified his own rules in the same breath: the first “yields values too large for N when m ≥ 7”, and both “assume all studies have a medium-size relationship”. So with six predictors and an expected medium effect, 50 + 8(6) = 98 is reasonable guidance. With ten predictors, or an expected small effect, it is not — and quoting the rule without the caveat is a misuse of the source that an attentive marker may spot.

    A power analysis is always the stronger move. Use the rule of thumb as a sanity check on the answer, not as a replacement for it.

    How many interviews does a qualitative dissertation need?

    This is where students most often quote a number they cannot support. The evidence is genuinely more conditional than the folklore suggests.

    • Guest, Bunce and Johnson (2006), in Field Methods, analysed 60 interviews and found that “saturation occurred within the first twelve interviews”, with basic elements of metathemes present as early as six. This is the source behind the widely repeated “twelve interviews” figure — from one study, in one setting.
    • Hennink and Kaiser (2022), in Social Science & Medicine, systematically reviewed empirical tests of saturation and reported it being reached within 9–17 interviews or 4–8 focus groups. Critically, they attached a condition: this applied “particularly those with relatively homogenous study populations and narrowly defined objectives”. It is not a general rule, and citing it as one misrepresents the paper.
    • Malterud, Siersma and Guassora (2016), in Qualitative Health Research, proposed information power instead: the more information your sample holds relevant to your aim, the fewer participants you need. They give five determinants — study aim, sample specificity, use of established theory, quality of dialogue, and analysis strategy — and deliberately attach no number. This is the most defensible framing for an undergraduate project.
    • Braun and Clarke go further, arguing that saturation is not consistent with the assumptions of reflexive thematic analysis and that sample-size judgements are “inescapably situated and subjective, and cannot be determined (wholly) in advance of analysis”. Note their date if you cite them: the paper appeared online in December 2019 and in Qualitative Research in Sport, Exercise and Health in 2021.

    The practical upshot: if you are doing reflexive thematic analysis, do not claim saturation. Argue information power, and justify your number from the specificity of your sample and the narrowness of your aim.

    How do you write the justification?

    Two or three sentences in your methods chapter, structured as inputs, output, and constraint. For a quantitative study:

    An a priori power analysis was conducted in G*Power 3.1.9.7 for an independent-samples t test, with α = .05, power = .80, and an expected medium effect (d = 0.50) based on the effect reported by [source]. This indicated a required sample of 128 participants (64 per group). Recruitment within the available fieldwork window yielded 96, meaning the study was powered to detect effects of approximately d = 0.58 or larger; smaller true effects may therefore have gone undetected, and this is considered in the discussion.

    That paragraph does something most undergraduate chapters do not: it states the shortfall and carries the consequence into the discussion. That is what a marker is looking for. The same principle governs the rest of the chapter — describe the decision, then defend it — as set out in our guide to writing a methodology chapter.

    What if you simply cannot recruit enough people?

    Common, and not fatal. Your options, in order of preference: extend recruitment; simplify the design so it needs fewer participants (fewer groups, fewer predictors, a within-subjects design instead of between-subjects); switch to secondary data; or proceed and report the study as underpowered with an honest discussion.

    What you must not do is run additional analyses until something reaches significance, or drop participants after seeing what that does to your p value. Note also that a within-subjects design typically needs substantially fewer participants for the same power, which makes the choice between independent and repeated measures a sample-size decision as much as a design one — a point covered in our guide to choosing a statistical test.

    An underpowered study reported honestly is a pass. A study with undisclosed analytic flexibility is an academic misconduct question.

    If the number is settled and the writing is what is stalling, you can draft the methods chapter in Tesify from your own inputs — your design, your calculation, 100% written by you.

    Frequently asked questions

    Is 50 participants enough for a dissertation?

    It depends entirely on your design and expected effect size. For a within-subjects comparison expecting a large effect, comfortably. For a multiple regression with six predictors, no. Run the power analysis and let it answer.

    Do I need a power analysis if my study is qualitative?

    No — power analysis applies to hypothesis testing. Use an information-power argument instead, naming the features of your sample and aim that justify the number of participants you recruited.

    Can I do a power analysis after collecting data?

    You can compute a sensitivity analysis, which reports the smallest effect your achieved sample could detect — that is legitimate and genuinely useful. What is not legitimate is a post-hoc “observed power” calculation based on your own result, which is circular and adds nothing.

    Does my sample need to be representative?

    Ideally, but undergraduate projects almost always use convenience samples, and that is accepted. What matters is that you describe the sample accurately, name the technique honestly, and state clearly what it does and does not allow you to generalise to.

    How many participants do I need for a chi-square test?

    Enough that expected cell frequencies are adequate — the usual guidance is that expected counts should be at least five in the large majority of cells. Small samples spread across many categories fail this quickly, so keep the number of categories down.

    Should I report the sample size I aimed for or the one I got?

    Both. State the target and its basis, then the achieved sample, then what the gap means for your conclusions. The gap is not an embarrassment; concealing it is.

  • Literature Review or Systematic Review: Which Does Your Undergraduate Dissertation Need?

    Literature Review or Systematic Review: Which Does Your Undergraduate Dissertation Need?

    A narrative literature review synthesises what is known about a topic and argues towards your research question; a systematic review answers a single tightly defined question by searching, screening and appraising the evidence against a protocol written in advance. Most UK undergraduate dissertations need the first. The second is a research method in its own right, with a workload to match.

    The confusion is understandable, because departments use “literature review” loosely — sometimes meaning a chapter inside an empirical dissertation, sometimes meaning a whole dissertation with no primary data collection. Those are different things with different rules, and choosing wrongly costs weeks.

    What is a narrative literature review?

    It is the chapter, usually your second, that establishes what is already known and why your question is worth asking. Its job is argumentative. You are not cataloguing studies; you are building a case that ends in a gap only your project can fill.

    A good one is organised by theme or debate, not by study. If your paragraphs begin “Smith (2019) found…”, “Jones (2021) found…”, you have written an annotated bibliography. If they begin “Two competing explanations have been offered for this effect…”, you have written a review.

    There is no requirement to find every relevant paper, and no requirement to document your search. You choose what is relevant and defend the choice implicitly through the quality of the argument.

    What is a systematic review?

    A systematic review is a research method that treats the published literature as its dataset. The defining features are that the question is fixed in advance, the search is exhaustive and reproducible, the inclusion and exclusion criteria are pre-specified, screening is documented at every stage, and the quality of included studies is formally appraised.

    Two reference points govern the field. Reporting is guided by the PRISMA 2020 statement — Page and colleagues, published in the BMJ in 2021 (372:n71) — which sets out what a review must report, including the flow of records from identification through screening to inclusion. Methods, in health and related fields, are guided by the Cochrane Handbook for Systematic Reviews of Interventions, currently version 6.5 (2024), edited by Julian Higgins and James Thomas.

    Note what those two documents are for. PRISMA tells you how to report a review; it is a reporting guideline, not a licence. Producing a PRISMA flow diagram does not make a review systematic, and markers who know the field can tell the difference immediately.

    A screening flow diagram used to record studies included in a systematic review
    A screening flow is documentation, not decoration — every number in it has to be defensible.

    How do the two actually differ?

    Feature Narrative literature review Systematic review
    Question Broad, may evolve as you read Fixed and narrow before you search
    Protocol None Written and ideally registered in advance
    Search Purposive; undocumented Exhaustive, multi-database, fully reproducible
    Inclusion criteria Implicit Pre-specified and applied consistently
    Screening Not recorded Recorded at every stage, ideally by two people
    Quality appraisal Informal, in the prose Formal, using a named tool
    Reproducibility Not expected Essential
    Where it sits A chapter within a dissertation Can be the entire dissertation

    Why is a full systematic review usually the wrong choice at undergraduate level?

    Three practical reasons, none of them about your ability.

    1. Double screening. Methodological standards expect two independent reviewers to screen records, with disagreements resolved by a third. As a lone undergraduate you do not have that, and the honest write-up has to concede it as a limitation that undercuts the method’s central claim.
    2. Database access. An exhaustive search means several databases plus grey literature. Your institutional subscriptions may not cover what the protocol requires, and you cannot report a search you could not run.
    3. Time. Screening several thousand titles and abstracts, then full texts, then appraising each included study, is a serious workload before you write a single word of synthesis. On a one-semester project it routinely overruns.

    None of this makes systematic review impossible as an undergraduate project — some departments run them well, with narrow scope and explicit supervision. But it should be a deliberate choice made with your supervisor, not a default reached because it sounded rigorous.

    What is the middle option?

    A structured or scoping review: you borrow the transparency of systematic method without claiming exhaustiveness. In practice that means you state your search terms, name the databases and the date range, give your inclusion and exclusion criteria, report how many records you screened and how many you kept, and appraise quality informally but consistently.

    This is very often what departments mean when they advertise a “literature-based dissertation”, and it is the sweet spot: markedly more rigorous than a narrative review, honest about what it is, and deliverable in the time you have. Crucially, you describe it accurately — as a structured review — rather than labelling it systematic.

    How do you know which one your department wants?

    Ask, and read three things before you do: your module handbook, the marking rubric, and one or two past dissertations if your department keeps them. The rubric is the most revealing. If it awards marks for “search strategy” and “quality appraisal”, a systematic or structured review is expected. If it awards marks for “critical synthesis” and “identification of a gap”, a narrative review is expected.

    Whichever you are writing, the review has to end by earning your research question — and if you are doing primary research, the methods that follow must be consistent with the argument the review just made. That link is explicit in a well-built methodology chapter: a review that generates testable propositions calls for a deductive design, and one that opens up meanings calls for a qualitative one.

    What does a quality appraisal tool actually do?

    It gives you a consistent set of questions to ask of every included study, so that “this study is weak” becomes a defensible judgement rather than an impression. Different designs need different tools — randomised trials, observational studies and qualitative research are appraised on different criteria — so pick the one that matches the studies you are including, name it, and apply it to every study rather than only the ones you distrust.

    The output is not a score you average. It is a description of where the evidence base is strong and where it is thin, which then shapes how confidently you can state your conclusions.

    Practical next step

    Decide the type first, write the criteria second, search third. Doing it in any other order produces a review that has to be rebuilt. If you know which type you need and want to get the chapter moving, you can draft and structure it in Tesify around your own sources and your own argument — 100% written by you, with the organisation handled.

    Frequently asked questions

    Can an undergraduate dissertation be entirely a literature review?

    Yes, in many UK departments, and it is a legitimate route — especially where ethical approval for primary data collection would be difficult. It is not an easier option, though: without primary data, the analysis of the literature has to carry the whole project.

    Do I need to register a protocol for an undergraduate review?

    Usually not, and most undergraduate projects are not registered. Writing a protocol for yourself is still worthwhile, because it stops you adjusting your inclusion criteria once you can see which results they produce.

    Is PRISMA compulsory if I use a flow diagram?

    If you present a PRISMA flow diagram you are signalling that you followed the reporting guideline, so you should report the other relevant items too. If you only want to show how many records you screened, call it a screening flow diagram and describe it in your own terms.

    How many studies should a structured review include?

    There is no fixed number, and a small, well-appraised set is better than a large, poorly examined one. What matters is that your inclusion criteria were applied consistently and that the number you ended with follows from them rather than from when you got tired.

    What is the difference between a scoping review and a systematic review?

    A scoping review maps what evidence exists on a broad topic and how it has been studied; a systematic review answers a specific question about what the evidence shows. Scoping reviews do not usually appraise study quality, because breadth rather than verdict is the point.

    Can I include grey literature in an undergraduate review?

    Yes, and in policy, education and management topics you often should — reports, theses and official publications may hold most of the relevant evidence. State that you searched it and how, since grey literature is not indexed like journal articles.

    Should the literature review chapter come before or after the methodology?

    Before, in the conventional UK structure, because the methods have to answer the gap the review identifies. Write them in that order too; a methodology drafted before the review is settled almost always needs rewriting.

    My supervisor called my review “descriptive”. What does that mean?

    That you are reporting studies rather than putting them in conversation. The fix is structural: reorganise the chapter around themes or disagreements rather than around individual papers, and make each paragraph state a claim that the cited studies then support or complicate.

  • How to Write the Methodology Chapter of a Business or Management Dissertation

    How to Write the Methodology Chapter of a Business or Management Dissertation

    The methodology chapter is where business and management dissertations are most often marked down, and the reason is almost always the same: the student describes what they did instead of arguing why it was the right thing to do. A methods chapter is not a diary. It is a defence. Every paragraph should answer an unspoken challenge from your marker — why that, and not the obvious alternative?

    What follows is the chapter written as seven decisions, in the order you should make and present them. Each one narrows the next, which is why writing them out of order produces a chapter that contradicts itself.

    Before you start: check your handbook, not this guide

    Business schools vary more than most disciplines in what they expect here. Some require an explicit philosophy section; others treat it as padding and want you in the design straight away. The QAA’s subject benchmark statements for business and management describe the shape of the discipline, but they are explicit that such statements provide general guidance and are not intended to prescribe set approaches — the operative document is your own module handbook. Read it first, and let it override any structural advice below.

    Decision 1: Research philosophy — and how to write it without padding

    If your handbook asks for a philosophy section, keep it short and make it do work. You are stating your assumptions about what counts as knowledge in your study, because those assumptions determine what evidence you are entitled to collect.

    • Positivist-leaning: you believe there is a measurable regularity to find. You will test hypotheses, measure variables and generalise from a sample. Typically quantitative.
    • Interpretivist-leaning: you believe the phenomenon is constructed through how people understand it. You will explore meanings in depth rather than measure frequency. Typically qualitative.
    • Pragmatist: you believe the research question dictates the method. Common in applied management research and the honest position behind most mixed-methods projects.

    The failure mode is a page of definitions copied from a textbook with no consequence. The fix is one sentence of consequence per position: “Because this study treats employee engagement as something that can be measured consistently across respondents, a positivist approach was adopted, which in turn required a standardised instrument rather than open interviews.”

    Decision 2: Approach — deductive, inductive, or honestly abductive

    Are you testing existing theory against new data, or building an explanation up from data? Deductive projects state hypotheses derived from the literature and test them. Inductive projects gather data and develop themes. Say which you are doing, and make sure your literature review agrees — a deductive study whose literature review does not generate testable propositions is incoherent, and markers notice.

    Decision 3: Design — what kind of study is this?

    Name the design explicitly. In undergraduate business research the realistic options are:

    • Cross-sectional survey — one measurement point, many respondents. Most common, most feasible.
    • Case study — one organisation or a small number, examined in depth. Strong for “how” and “why” questions; weak for generalisation, which you must concede.
    • Comparative — two or more organisations, sectors or countries set against each other.
    • Secondary data analysis — company reports, published datasets, archival records. Often the strongest option when access to participants is uncertain.

    Do not leave the design implicit. “A questionnaire was distributed” is not a design; “a cross-sectional survey design was used” is.

    A research design diagram sketched out for a business dissertation methodology chapter
    Drawing the design as a flow — question, data source, instrument, analysis — exposes gaps before you write a word.

    Decision 4: Sampling — the section markers read most closely

    This is where the marks are. You need four things: the population, the sampling frame, the technique, and the achieved sample with a candid word about its limits.

    Be honest about technique. Undergraduate business dissertations overwhelmingly use convenience or purposive sampling, and pretending otherwise is transparent. What distinguishes a strong chapter is not claiming a random sample you did not have, but naming the constraint and reasoning about its consequences.

    A worked example of the paragraph you are aiming for:

    The target population was full-time employees of UK small and medium-sized enterprises in the hospitality sector. No comprehensive sampling frame for this population is publicly available, so a non-probability convenience sample was drawn through the researcher’s professional network and two sector-specific LinkedIn groups, supplemented by snowball referral. Sixty-three usable responses were obtained. This approach limits statistical generalisation to the wider population, and the sample is likely to over-represent employees of firms with an active social media presence. Findings are therefore presented as indicative of patterns within this sample rather than as population estimates.

    That paragraph concedes three weaknesses and is stronger for it. A chapter that concedes nothing invites the marker to find the weaknesses themselves.

    Decision 5: Instruments and data collection

    Say precisely what you used, where it came from, and why it was fit for purpose.

    1. If you adapted an existing scale, cite it and say what you changed. Established instruments carry published reliability evidence, which is free credibility — use it.
    2. If you wrote your own items, explain how they were derived from your literature review and say whether you piloted them. A pilot with even five respondents is worth reporting.
    3. If you interviewed, state the format (structured, semi-structured, unstructured), how long interviews ran, how they were recorded, and how they were transcribed.
    4. If you used secondary data, name the source, the time period, the unit of analysis and any exclusions you applied.

    Include the instrument itself in an appendix and refer to it. Markers check.

    Decision 6: Analysis — state the technique before the results chapter, not in it

    Your methodology chapter should tell the reader exactly what will happen to the data, so that nothing in the results chapter arrives as a surprise.

    For quantitative work, name the tests and the software, and tie each test to a hypothesis. The logic of matching test to design — difference versus association, level of measurement, independent versus repeated — is the same in management research as anywhere else, and if you are unsure which test your design calls for, our step-by-step guide to choosing a statistical test walks through the decision in the same order you should present it here.

    For qualitative work, name the analytic approach — thematic analysis, template analysis, framework analysis — and describe the coding procedure concretely: how codes were generated, whether anyone checked a subset, how themes were arrived at. “The data were analysed thematically” on its own is not a method.

    Decision 7: Ethics, access and data protection

    Every UK business school requires ethical approval before data collection begins, and the sequencing is not negotiable. Institutional guidance is unambiguous on this point: the University of the West of England, for example, states plainly that you cannot start collecting data until you have full ethical approval for that activity, and warns that its expedited route cannot be used to obtain retrospective approval. Collect first and you may find the data unusable.

    Your chapter should cover informed consent, the right to withdraw, anonymity and confidentiality (which are not the same thing), secure storage, and — if you are surveying people — where their data will physically be held. Under the UK GDPR and the Data Protection Act 2018, sending personal data outside the UK is a “restricted transfer” that needs a lawful basis, which is why universities are prescriptive about which survey platform you may use. Note that the Information Commissioner’s Office has flagged that its research-provisions guidance is under review following changes made by the Data (Use and Access) Act, so check the current position rather than relying on an older handbook.

    If you are still choosing a platform, this is a decision with real compliance consequences rather than a matter of preference.

    A structure you can lift

    1. Introduction — restate the research question and preview the chapter
    2. Research philosophy and approach
    3. Research design
    4. Population, sampling frame and sampling technique
    5. Data collection instrument and procedure
    6. Data analysis strategy
    7. Reliability and validity (or trustworthiness, for qualitative work)
    8. Ethical considerations
    9. Limitations of the methodology
    10. Summary

    Put limitations here as well as in your conclusion. Methodological limitations belong beside the methods that caused them.

    If the structure is clear but the drafting has stalled, you can build the chapter section by section in Tesify, working from your own design decisions rather than a template. Your reasoning, your data, 100% written by you.

    Frequently asked questions

    How long should a business dissertation methodology chapter be?

    Typically 15–20% of the total word count, though your handbook governs. On a 10,000-word dissertation that is roughly 1,500–2,000 words. If yours is much shorter, you are probably describing rather than justifying.

    Do I need a philosophy section if my study is purely quantitative?

    Only if your handbook asks for one. If it does, keep it to a few hundred words and make every claim consequential for a later decision. Quantitative work is not exempt from having assumptions; it is just that they are usually left unstated.

    Is a sample of 60 too small for a business dissertation?

    Not inherently. Feasibility is an accepted constraint at undergraduate level, and what matters is that you justify the number, acknowledge what it does and does not support, and avoid claiming population-level generalisation you cannot back. What is not acceptable is failing to mention the issue.

    Can I use a case study of the company I did my placement with?

    Frequently yes, and it often produces the strongest projects because access is real. You must have the organisation’s permission in writing, address the confidentiality of commercially sensitive information, and be explicit in your limitations about your own proximity to the setting.

    Should the methodology chapter be written in the past tense?

    Yes, once the research is done — you are reporting what was carried out. Passive constructions are conventional in business research, but do not let them obscure who made a decision when the reasoning matters.

    What is the difference between reliability and validity in a management dissertation?

    Reliability is consistency: would the same instrument give the same result again? Validity is accuracy: is it measuring the construct you claim? An instrument can be highly reliable and still measure the wrong thing, which is why both need addressing separately.

  • How to Choose the Right Statistical Test for a Psychology Dissertation (2026)

    How to Choose the Right Statistical Test for a Psychology Dissertation (2026)

    Most psychology undergraduates do not have a statistics problem. They have a sequencing problem: they collected the data first and are now trying to reverse-engineer a test that fits it. The test you need is determined by three things you decided long before you opened your dataset — what your hypothesis claims, how you measured your variables, and how your participants were allocated. Work through those in order and the test chooses itself.

    This guide gives you that order as six steps. It is written for a UK undergraduate psychology dissertation, where you are typically working with one or two independent variables, a sample recruited through your department, and a marking rubric that cares more about whether your analysis is justified than whether it is clever.

    Step 1: Decide whether you are testing a difference or an association

    Almost every undergraduate psychology hypothesis is one of two shapes. Either you predict that groups differ on some measure, or you predict that variables move together. This single distinction splits the entire test family tree in half, and it is the first branch used by the decision guides published by university libraries, including the University of Leeds “Find a test” guide, which separates tests of difference from tests of association before considering anything else.

    Write your hypothesis as a sentence and look at the verb:

    • Difference: “Students who revise with self-testing will recall more items than students who reread.” Two groups, one outcome. You need a test of difference.
    • Association: “Higher trait anxiety will be associated with poorer sleep quality.” Two continuous measures on the same people. You need a correlation or regression.
    • Prediction: “Trait anxiety and rumination will predict sleep quality.” More than one predictor, one outcome. You need multiple regression.

    If you cannot write your hypothesis in one of these shapes, that is the finding — your research question is not yet operationalised, and no test will rescue it. Fix the question before you touch the data.

    Step 2: Identify the level of measurement of your outcome variable

    The second branch is what kind of thing your dependent variable actually is. The UCLA Office of Advanced Research Computing statistical methods guide organises its whole selection table around exactly this: the number of dependent variables, the nature and number of independent variables, and whether the outcome is interval and normal, ordinal, or categorical.

    • Continuous — reaction times in milliseconds, scores on a validated scale, hours of sleep. Opens the parametric family.
    • Ordinal — ranked positions, single Likert items, ordered categories such as “never / sometimes / often”. Points you towards non-parametric tests.
    • Categorical — diagnosed or not, chose option A or option B. Points you towards chi-square.

    One judgement call trips up a large share of psychology dissertations: a single Likert item is ordinal, but a summed or averaged scale built from many items is conventionally treated as continuous. If you are analysing a validated multi-item questionnaire, you are almost certainly in the continuous column. Say so explicitly in your methods chapter and give your reason; markers reward the justification far more than the choice itself.

    Step 3: Count your groups and check whether they are independent or repeated

    Now describe your design in three numbers: how many independent variables, how many levels each has, and whether the same people appear in more than one level.

    • Independent (between-subjects): different people in each condition.
    • Repeated (within-subjects): the same people measured more than once.

    Getting this wrong is the single most expensive error available to you, because an independent-samples test run on repeated-measures data throws away the very thing that makes a within-subjects design powerful — the pairing. It will usually make a real effect disappear.

    A hand-drawn decision tree for selecting a statistical test in a psychology dissertation
    Sketching the decision path by hand before opening your data forces you to commit to a design description you can defend.

    Step 4: Check the parametric assumptions before you commit

    Parametric tests buy you statistical power in exchange for assumptions. The British Psychological Society’s supplementary guidance on research methods for accredited undergraduate and conversion programmes expects students to be able to detect differences in sample means using tests such as t tests and ANOVA, and relationships between variables using chi-square, correlation and regression — and, importantly, to show “familiarity with robust alternatives when those assumptions are not met (e.g. non parametric alternatives)”. Knowing the alternative is part of the competence being assessed.

    For most undergraduate designs you need to check:

    1. Distribution of the outcome within each group, judged from a histogram and a normality test together rather than either alone.
    2. Homogeneity of variance across groups, usually via Levene’s test.
    3. Independence of observations — a design question, not a statistical one. If participants worked in pairs or discussed the task, this assumption is already compromised.
    4. Outliers, identified and handled by a rule you state before looking at whether removing them helps your p value.

    If assumptions fail, you do not have a disaster. You have a non-parametric equivalent and a sentence to write explaining why you used it.

    Step 5: Read the test off the table

    With steps 1 to 4 answered, the choice is mechanical.

    What you are testing Design Parametric test Non-parametric alternative
    Difference between 2 groups Independent Independent-samples t test Mann–Whitney U
    Difference between 2 conditions Repeated Paired-samples t test Wilcoxon signed-rank
    Difference between 3+ groups Independent One-way ANOVA Kruskal–Wallis H
    Difference between 3+ conditions Repeated Repeated-measures ANOVA Friedman test
    Two independent variables Either or mixed Factorial / mixed ANOVA No clean equivalent — consider transformation or a robust method
    Association between 2 continuous variables Pearson’s r Spearman’s rho
    Prediction from 2+ predictors Multiple regression
    Association between 2 categorical variables Independent Chi-square test of independence

    The right-hand column is not a consolation prize. Reporting a Mann–Whitney because your data were skewed, and saying so, reads as competence. Reporting a t test on visibly skewed data reads as not having looked.

    Step 6: Report it so a marker can verify it

    A result is only worth the marks if someone can reconstruct it. Report the test, the degrees of freedom, the test statistic, the exact p value and an effect size — the last of these is what turns a significance claim into a meaningful one. Jacob Cohen’s widely used conventions, set out in his 1992 paper “A power primer” in Psychological Bulletin, give benchmarks of .20, .50 and .80 for small, medium and large d, and .10, .30 and .50 for r. Cohen was candid that “the definitions were made subjectively”, describing a medium effect as one “likely to be visible to the naked eye of a careful observer” — so treat them as reference points, not verdicts.

    A worked example of how the sentence should look:

    Participants in the self-testing condition (M = 24.10, SD = 4.32) recalled significantly more items than those in the rereading condition (M = 19.85, SD = 5.01), t(58) = 3.52, p = .001, d = 0.91.

    On referencing style, check your handbook rather than assuming. APA style — currently the seventh edition of the Publication Manual of the American Psychological Association, published in 2020 — is the working convention across UK psychology and is required by BPS journals. But the BPS’s own supplementary guidance for accredited undergraduate programmes states plainly in a footnote that “in respect to referencing of work, alternatives to the APA style are acceptable”. Your department decides. Ask, and follow what it says.

    Which software should you actually use?

    All three of the packages UK psychology departments commonly point at will run everything in the table above.

    • IBM SPSS Statistics is commercial software, currently at version 32, and is the package most UK psychology departments teach; access normally comes through a university site licence rather than a personal purchase.
    • jamovi is free and open source — released under the AGPL3, with its analysis package jmv under GPL2+ — and is “powered by the R statistical language”. It runs the standard undergraduate test set through a point-and-click interface very close to SPSS.
    • R is a free software environment for statistical computing maintained by the R Core Team, with copyright held by the R Foundation for Statistical Computing. It has the steepest learning curve and the longest payoff.

    If your analysis is the standard undergraduate set and you have lost access to a campus licence over the summer, jamovi is the pragmatic answer: it is free, it is legitimate to cite, and its output maps onto what your supervisor expects to see.

    Where students actually lose marks

    Not on the test. On the justification. A methods chapter that says “a t test was conducted” earns less than one that says “an independent-samples t test was selected because the design compared two independent groups on a continuous outcome, and inspection of histograms and Levene’s test indicated the parametric assumptions were tenable.” Same analysis, visibly different competence.

    If you have your design settled and the blank methods chapter is the thing standing between you and a submission, you can start drafting it in Tesify — it works from your own design decisions and your own data, so the argument stays yours. The writing is 100% written by you; what you get is structure and momentum.

    Frequently asked questions

    Do I need to run a normality test if my sample is large?

    Report one, but do not let it decide alone. Normality tests become very sensitive at larger sample sizes and will flag trivial departures as significant. Read the test alongside a histogram and, where relevant, skewness and kurtosis values, and say in your write-up what you looked at.

    Can I use a parametric test on Likert data?

    On a single Likert item, no — it is ordinal. On a summed or averaged multi-item scale, it is conventional in psychology to treat the total as continuous. State which you have and justify it in one sentence.

    What if my assumptions fail and there is no non-parametric equivalent?

    This happens most often with factorial designs. Your options are to transform the outcome variable, to use a robust or bootstrapped procedure, or to simplify the design to one your data can support. Whichever you choose, report what failed and why you responded as you did.

    Is a non-significant result a failed dissertation?

    No. Undergraduate projects are frequently underpowered, and a well-designed study with a null result and an honest discussion of power is a legitimate piece of work. What loses marks is presenting a null result as though the hypothesis had been confirmed, or quietly running additional tests until something reaches significance.

    How many statistical tests should a psychology dissertation contain?

    As many as your hypotheses require, and no more. Each hypothesis should map to one planned analysis. Running a large number of unplanned comparisons inflates your false-positive rate, and markers can see it in your results section.

    Should I report effect sizes even when the result is not significant?

    Yes. An effect size with a confidence interval tells the reader how precise your estimate was, which is exactly the information a null result needs in order to be interpretable.