Tag: coding

  • How to Do a Thematic Analysis for Your Dissertation: The Six Phases, Step by Step (2026)

    How to Do a Thematic Analysis for Your Dissertation: The Six Phases, Step by Step (2026)

    You have eight interview transcripts, a highlighter, and a methods chapter that says “thematic analysis” as though that settled something. It does not. Thematic analysis is the most commonly used qualitative method in UK undergraduate dissertations and the most commonly done badly, because most students learn a six-step recipe without ever learning which version of the method they are cooking.

    This guide walks the six phases as Braun and Clarke set them out, with the expected output of each step stated so you can tell when you have finished it. It also covers the decision that has to come first, and the single error that costs more marks than all the others combined. If your analysis is done and the chapters are the bottleneck, you can structure and draft them in Tesify.

    Step 0: Decide which thematic analysis you are actually doing

    Braun and Clarke are emphatic that thematic analysis is “a family of methods, not a singular method — there is no ‘standardised TA’”. In their 2022 paper on good practice they set out a typology, and knowing where you sit in it is what stops your methods chapter contradicting itself:

    1. Reflexive TA — the version Braun and Clarke themselves developed. Researcher subjectivity is treated as a resource rather than a bias to be eliminated. There is no notion of “accurate” coding, and no second coder checking your work.
    2. Coding reliability TA — emphasises objectivity, structured codebooks and intercoder agreement, usually reported as a statistic such as Cohen’s kappa.
    3. Codebook TA — including framework analysis and template analysis; structured procedures combined with qualitative research values.

    The mistake Braun and Clarke call out by name is mashing these together: running a reflexive analysis and then bolting on a second coder and a kappa value because it sounds rigorous. They describe that as methodological incoherence, and a marker who knows the literature will read it as exactly that. Pick one, name it in your methods, and behave consistently.

    Expected output of this step: one sentence in your methods chapter naming your variant and citing the source you followed. For most undergraduate projects that sentence is “This study used reflexive thematic analysis (Braun and Clarke, 2006; 2022).”

    Phase 1: Familiarise yourself with the data

    Transcribe your own interviews if your timetable allows it. Transcription is slow and it is also the single most efficient familiarisation exercise there is — by the end of it you know your data in a way that no amount of later reading recovers. If you used a transcription service or software, read every transcript against the audio at least once, both to catch errors and to hear the tone that plain text loses.

    Then read everything through without coding anything. Keep a separate document for notes: what surprises you, what recurs, what one participant says that flatly contradicts another. These notes are not codes and should not try to be. They are the first sighting of the arguments your analysis will eventually make.

    Expected output: a set of accurate transcripts you have read at least twice, and one or two pages of unstructured observations.

    Phase 2: Generate initial codes

    Now work systematically through the entire dataset and attach a short label to every segment that is relevant to your research question. A code is a compact description of what is interesting in that extract — “reluctance to ask for help”, “describes deadline as external”, “compares self to coursemates” — not a topic heading.

    Three practical rules. Code the whole dataset, not just the quotations you already like; selective coding produces an analysis that confirms what you thought before you started. Let one extract carry more than one code where it genuinely does. And keep the surrounding context attached to each coded extract, because an isolated line loses the meaning that made it interesting.

    Hand coding with coloured pens and margin notes is entirely acceptable at undergraduate level and many students find it faster than learning software mid-project. If you prefer software, use whatever your department supports and can help you with.

    Expected output: a complete coded dataset and a list of codes — typically several dozen for a small undergraduate study — each with the extracts filed against it.

    Phase 3: Search for themes

    Collate your codes into candidate themes. This is where sticky notes, a large table and an afternoon beat any piece of software: lay the codes out, move them around, and look for the ones that share an underlying idea rather than a subject matter.

    A theme in reflexive TA is a pattern of shared meaning organised around a central concept. It is not a bucket. Some codes will not fit anywhere and should go into a holding pile rather than being forced; some will turn out to be themes in their own right.

    Expected output: three to five candidate themes with their codes and extracts gathered under each, plus a leftover pile you have not deleted.

    Colour-coded codes grouped on a table into candidate themes
    Phase three is physical work: codes that share an underlying idea cluster together, and the ones that refuse to cluster are telling you something too.

    Phase 4: Review the themes

    Test your candidates at two levels. First, read all the collated extracts for each theme and ask whether they genuinely form a coherent pattern; if a theme’s extracts pull in two directions, it is probably two themes, and if a theme has four extracts from one participant it is probably that participant’s view rather than a pattern. Second, read your whole dataset again against your thematic map and ask whether the map represents the data as a whole.

    Expect to lose themes here. Collapsing two into one, splitting one into two, and abandoning a candidate that looked promising in phase three are all signs the review is working, not that the analysis is failing.

    Expected output: a revised thematic map you can defend, with any subthemes identified.

    Phase 5: Define and name the themes

    Write a short paragraph for each theme stating what it is about, what aspect of the data it captures, and what it contributes to your research question. If you cannot describe a theme’s scope in a couple of sentences without listing its contents, it is not yet a theme.

    Names should be informative and, where the data supports it, drawn from participants’ own language. “Barriers” tells a reader nothing. “It felt like admitting I could not cope” tells them what the theme argues before they read a word of it.

    Expected output: final theme names and a written definition of each, ready to become the subheadings of your findings chapter.

    Phase 6: Produce the report

    Your findings chapter is the analysis, not a preamble to it. Each theme gets its section; each section makes an argument and uses extracts as evidence for that argument. The prose does the analytical work and the quotations support it — a chapter that is 60 per cent block quotation with a linking sentence between each is a data display, not a findings chapter, and it is marked as one.

    Choose extracts that are vivid and representative, keep them short, and always follow a quotation with your own interpretation of why it matters. Say who each extract came from using your pseudonyms, and never let one talkative participant carry a whole theme.

    Expected output: a findings chapter organised by theme, with interpretation leading and evidence supporting.

    The mistake that costs the most marks: topic summaries

    Braun and Clarke identify confusing “themes-as-meaning-unified-interpretative-stories with themes-as-topic-summaries” as the most common problem in reflexive thematic analysis. It is worth seeing the difference on the page.

    Topic summary (weak) Meaning-based theme (strong)
    Views on workload Workload is described as weather — something that happens to you
    Support from staff Asking for help is read as an admission of not coping
    Use of technology Tools are trusted for facts but not for judgement
    Advantages and disadvantages The cost is accepted because the alternative is unthinkable

    Everything in the left-hand column is a heading under which you could file quotations. Everything in the right-hand column makes a claim that the extracts can support or undermine. The left column produces a chapter that describes what people talked about; the right column produces a chapter that says what your data means. Markers reward the second, and the gap between them is where most of the available marks in a qualitative dissertation actually sit.

    Saturation and second coders: what not to write

    Two conventions from other traditions get imported into student thematic analyses where they do not belong, and both are worth handling explicitly.

    Data saturation. Braun and Clarke question the concept’s usefulness for reflexive TA altogether, on the reasoning that meaning is generated by the researcher rather than waiting in the data to be exhausted. Writing “saturation was reached after seven interviews” in a reflexive TA is a claim your own cited method does not support. The defensible alternative is to justify your sample by your research question, your design and your practical constraints, decided in advance — the same logic set out in our guide to sample size for an undergraduate dissertation.

    Intercoder reliability. Reporting a Cohen’s kappa belongs to coding reliability TA, where accurate coding is a meaningful idea. In reflexive TA there is no single correct coding to agree with, so a kappa value is not rigour, it is a category error. If your department requires a second coder, that is a legitimate instruction — follow it, and then describe your method as coding reliability or codebook TA rather than citing a reflexive source.

    And never write that themes “emerged”. Braun and Clarke explicitly reject language in which themes are “identified, found or discovered”, because it implies they lay in the data independently of you. Themes are generated, created or constructed. Changing that one verb throughout your chapter signals to a marker that you understand the method you named.

    A handwritten reflexive research journal kept alongside a thematic analysis
    A dated journal kept through coding is what turns a reflexivity paragraph from a statement about who you are into evidence of how it shaped the analysis.

    How to write this up in your methodology chapter

    Your methods chapter needs six things: the variant of TA and its citation, whether your coding was inductive or theory-driven, whether you analysed at a semantic or latent level, who coded and how, the software or lack of it, and a reflexivity paragraph. On that last point, Braun and Clarke are unimpressed by “shopping list” identity statements — the paragraph should link your positioning to your actual analytic practice, saying how being a final-year student on the same course as your participants shaped what you noticed, not merely that you were one.

    The chapter’s overall architecture — philosophy, approach, design, sampling, instruments, analysis, ethics — is the same one we set out for writing a methodology chapter, and your analysis section slots into it. Remember too that interviews mean human participants, so your ethics approval had to be in place before you recruited anyone, and your methods chapter should say so. If your project turned out to be literature-based rather than interview-based, thematic analysis is not the tool you need — see literature review versus systematic review instead. And if you are still deciding between a qualitative and a quantitative design, our guide to choosing a statistical test shows what the other route commits you to.

    A realistic schedule

    For eight to twelve interviews, budget roughly a week for transcription and familiarisation, a week for coding, a few days for phases three to five, and two weeks for writing. The phase students consistently underestimate is coding, and the phase they consistently skip is the review — which is unfortunate, because reviewing is where a set of topic headings turns into an argument.

    When the analysis is finished and the chapters need writing, Tesify can structure and draft your dissertation around your own themes and extracts — the interpretation and every word stay 100% yours, with the structure and the bibliography handled.

    Frequently asked questions

    How many themes should a dissertation have?

    There is no rule, but three to five works for most undergraduate projects. Fewer usually means your themes are too broad to say anything; more usually means you have produced topic summaries and are listing subjects rather than making arguments.

    How many interviews do I need for a thematic analysis?

    Decide it in advance from your research question, your design and what you can realistically recruit, and justify that reasoning in your methods. Six to twelve is common for an undergraduate project. Do not justify it by claiming saturation if you are doing reflexive TA.

    What is the difference between a code and a theme?

    A code labels one interesting feature of one extract. A theme is a pattern of shared meaning across the dataset, organised around a central concept, that codes are gathered into. Codes are the raw material; themes are the argument.

    Can I use thematic analysis on open-ended survey responses?

    Yes — it works on any qualitative text, including open-text survey answers, forum posts and policy documents. Short survey responses give you less context per extract, so expect more semantic and less latent analysis.

    Do I need software like NVivo for an undergraduate thematic analysis?

    No. Hand coding is perfectly acceptable and often faster for a small dataset than learning a new package mid-project. Use software if your department teaches and supports it; do not adopt one three weeks from your deadline.

    Should I report inter-rater reliability?

    Only if you are doing coding reliability or codebook TA, where the concept makes sense. In reflexive TA there is no single correct coding, so a kappa statistic is methodologically incoherent. Follow your department’s instruction, and describe your variant accordingly.

    Is it acceptable to say my themes “emerged” from the data?

    Braun and Clarke explicitly reject that language because it implies themes existed independently of the researcher. Write that themes were generated, constructed or developed. It is a small change that markers who know the method notice immediately.

    What is the difference between semantic and latent coding?

    Semantic coding stays with what participants explicitly said; latent coding interprets the assumptions and ideas underneath it. Most dissertations use both. State which you emphasised, because it tells your marker how to read your claims.

    Can I combine thematic analysis with quantitative data?

    Yes, in a mixed-methods design — but analyse each strand with its own appropriate method and be explicit about how you integrate them. Do not count how many participants mentioned each theme and present that as findings; frequency is not what thematic analysis measures.

    What do I do with the codes that did not fit any theme?

    Keep them. They are useful in your discussion as evidence of variation or contradiction, and a marker who sees a leftover pile acknowledged reads it as honesty rather than untidiness. Deleting inconvenient data is the problem; not every code becoming a theme is not.