Thematic analysis is a way of finding, testing and describing patterns of meaning across qualitative data, and learning how to do thematic analysis step by step means working in a deliberate loop: read the data, code meaningful units, group codes into candidate themes, check those themes back against the raw material, name them, then write the account up. Most students follow Virginia Braun and Victoria Clarke’s six-phase framework, because it leaves a decision trail an examiner can follow. Expect several weeks of part-time work for a standard dissertation dataset, and expect to go back and forth between stages rather than moving in a straight line.
The workflow below runs seven stages. Seven rather than six because separating code refinement from theme generation removes the single most common failure in student analysis, where a pile of codes gets renamed and called themes.
Table of Contents
- 1What You Need
- 2Step-by-Step: How to Do Thematic Analysis
- 3How to Do Thematic Analysis Step by Step: Define the Analysis
- 4Step 2: Familiarize Yourself With the Data
- 5Step 3: Generate Initial Codes
- 6Step 4: Review, Compare, and Refine Codes
- 7Step 5: Build and Name Candidate Themes
- 8Step 6: Define and Check Each Theme
- 9Step 7: Write the Thematic Analysis
- 10Common Mistakes
- 11Frequently Asked Questions
- 12Should thematic analysis coding be inductive, deductive, or both?
- 13How do I make qualitative coding less subjective?
- 14Can I do thematic analysis in Excel, Word, or Google Sheets?
- 15How many themes should a thematic analysis have?
- 16Do I need direct participant quotations in my results?
- 17Conclusion
What You Need
You need five things before you touch a single transcript, and only one of them is software.
- A precise research question. “What do students think about online learning?” is too wide to code against. “How do first-year students describe the moment they decided whether to keep attending online lectures?” gives you a boundary you can actually apply.
- Data that can carry themes. Interviews, focus groups, open-ended survey responses, open-text feedback, diaries, documents, social media posts. A dataset of 15 to 30 interviews is typical for a master’s study; a PhD study of the same method can run to 60 or more.
- Transcripts that are ready to code. Verbatim, labelled with participant ID and line numbers, and formatted so you can point a reader at any extract. A transcript you still have to clean up is not ready.
- A coding workspace with backups. This can be a spreadsheet with one row per quote, a document with a code column, or a CAQDAS package. Version history matters more than the tool: your analysis should never live in one uncopied file.
- Optional software. NVivo, MAXQDA, ATLAS.ti and Delve all handle coding, retrieval and visual maps. Choose one only if its features match what your analysis actually needs.
Two housekeeping points that students skip and regret. Keep your ethics approval and consent forms in the same folder as the data, because a viva question about how consent shaped your sampling is easier when the answer is in front of you. And check whether your institution already licenses a CAQDAS package: students on r/PhD repeatedly assume they need to buy one, when a campus licence often sits unused in the library.
Step-by-Step: How to Do Thematic Analysis
The framework is iterative. Each stage feeds back into the one before it, and Braun and Clarke call these phases precisely because they overlap. Budget roughly 10 percent of your time on Stage 1, 30 percent on coding and refining, 25 percent on themes, and 25 percent on writing up.
| Stage | What you do | Artefact you produce | Rough time for 20 interviews |
|---|---|---|---|
| 1. Define the analysis | Fix the question, unit of analysis, inclusion boundaries, approach | One-page analysis plan | Half a day |
| 2. Familiarize | Read or listen in full, take notes, no coding yet | First-pass memo | 1 to 2 weeks |
| 3. Generate initial codes | Label meaningful units of text | Codebook or coded transcript | 3 to 5 weeks |
| 4. Review and refine codes | Merge, split, rename, discard; compare across cases | Revised codebook with definitions | 1 to 2 weeks |
| 5. Build candidate themes | Cluster codes into broader patterns | Thematic map | 3 to 5 days |
| 6. Define and check themes | Write a definition, test coverage, handle negative cases | Theme definitions and supporting extracts | 1 week |
| 7. Write up | Organise around themes, quote, interpret, report method | Findings and methodology sections | 2 weeks |
How to Do Thematic Analysis Step by Step: Define the Analysis
Define the analysis by writing one paragraph that states your research question, your unit of analysis, what counts as evidence, and what you are deliberately excluding. The unit of analysis is the decision students forget: a sentence, a paragraph, a speaker turn or a whole transcript each produce a completely different codebook.
Write the boundaries down where you will see them. Something like “semantic unit of a complete clause; minimum three words; only data about the transition from in-person to hybrid study; descriptive background about course structure excluded” removes dozens of arguments from later stages.
The stage worked when every coding decision downstream can be judged against this page. If a code cannot be explained as relevant to your stated question, it belongs to a different study. The step below covers generating codes without letting a framework dictate them.
Step 2: Familiarize Yourself With the Data
Familiarization means reading or listening to the entire dataset at least twice before you code anything. The first pass is for immersion, the second for noticing. Organise files so each transcript carries a participant ID, a date and a line-numbered body.
Take notes as you read. A first-pass memo on a study about first-year online learning might read: “Everyone mentions a threshold moment. Most describe it as social rather than academic, and two frame it as a money decision. P07 contradicts the group pattern and deserves a closer look. Nobody raises assessment timing directly, which surprises me. Question wording on Q4 may have led them.”
That memo is the artefact. It should show broad patterns, at least one surprise, and a list of questions you intend to carry into coding. If you reach the end of your first read with nothing written, you read too fast.
Step 3: Generate Initial Codes
Code meaningful units, not categories you already believe will appear. A code is a label attached to a short piece of text so that you can find it later, and in reflexive thematic analysis it is a prompt for thinking rather than a permanent category.
Four coding styles give you the most to work with. In vivo coding uses the participant’s own words. Descriptive coding labels what the content is about. Process coding marks something happening over time. Values coding marks what the participant believes matters.
Take one extract from a student wellbeing study: “I knew I was failing in October, but I did not tell anyone because I could not face my tutor asking why. I just stopped opening the module.” Four passes give you stopped opening the module (in vivo), withdrawal from module content (descriptive), hiding decline until crisis (process) and avoiding judgement from staff (values). That is four routes into the same ten words, and the differences between them are where your themes come from.
How to do thematic analysis step by step means letting the codes stay open. Write a short definition beside each code the first time you use it, because definitional drift, where the same label quietly changes meaning across transcripts, is the hardest fault to spot later. The stage worked when codes are consistent enough to group but loose enough to be revised.
Step 4: Review, Compare, and Refine Codes
This is the stage most people skip and the one that separates a defensible analysis from a pile of notes. Work through your codes in four passes: merge codes that do the same job, split codes covering two ideas, rename codes that describe the code rather than the data, and discard codes applied once with no clear analytical value.
Use three practical checks. Look for similar codes used inconsistently, where one analyst would have labelled a passage “isolation” and another “withdrawal”. Look for codes used exactly once across twenty interviews and ask whether they are a genuine minority case or a slip. Look for quotations that fit no code, because an uncodeable passage usually means your codebook is missing something.
In the wellbeing example, “withdrawal from module content” and “stopped opening the module” collapse into one focused code, disengagement expressed through material avoidance. “Avoiding judgement from staff” and a second code “protecting self-image” merge into concealment of decline. The code count drops from 40 to 31 and the codes get sharper, which is the point.
Step 5: Build and Name Candidate Themes
A theme is a pattern of meaning across extracts that answers part of your research question. A cluster of codes is not a theme, and renaming a code with capitals does not make it one. If your theme name could sit above an individual code with no change in meaning, it is a code.
Build a theme table while you cluster. One row per theme, with the name, a one-sentence definition, the codes inside it and the extracts that support it. Keeping the supporting extracts visible while you name things stops you from creating a theme that no evidence backs.
| Candidate theme | Definition | Included codes | Supporting extracts |
|---|---|---|---|
| Threshold moments of withdrawal | A recognisable point at which participants describe deciding to stop engaging | engagement collapse, silent exit, disengagement through avoidance | 12 extracts across 9 participants |
| Concealment of decline | Participants describe hiding academic trouble rather than seeking help | protecting self-image, avoidance of staff judgement | 7 extracts across 6 participants |
| Relief after disclosure | Participants describe the moment they told someone and the relief that followed | confiding in peers, tutor contact as turning point | 5 extracts across 5 participants |
The quality check at this stage is simple: each candidate theme should answer a recognisable aspect of your research question. “Threshold moments of withdrawal” clearly does. “Student experiences” does not.
Step 6: Define and Check Each Theme
Write each theme definition in one or two sentences stating what it captures, what it does not capture and why it matters to the question. Then run two checks that students rarely do.
The first is internal homogeneity and external heterogeneity: codes inside one theme should hang together, and two themes should not overlap. If your definitions require the word “also” to tell them apart, they are one theme. The second is support: a theme with a single extract across twenty interviews is not a theme, though it may be worth a sentence as a negative case.
A visual thematic map closes the stage. Students on r/DissertationSupport struggle most with showing the hierarchy between research question, categories and themes, and a simple diagram with boxes and connecting lines settles it faster than another page of prose. Braun and Clarke also distinguish semantic themes, which stay close to participants’ explicit meaning, from latent themes, which capture the underlying assumptions and ideas a participant did not state.
Negative cases refine a theme rather than killing it. If one participant left and described the decision as purely financial, you do not discard “threshold moments of withdrawal”, you narrow its definition and say that the financial framing is a boundary condition rather than a contradiction.
Lincoln and Guba’s four criteria from 1985 are a practical check on this stage: credibility, transferability, dependability and confirmability. For reflexive thematic analysis, credibility comes from honest reflexive positioning rather than inter-coder agreement, and confirmability comes from a clear audit trail of your coding decisions.
Step 7: Write the Thematic Analysis
Write the analysis around themes, not around participants and not around codes. One section per theme, each opening with a definition and then an argument that the extracts support, using quotations as evidence rather than as decoration.
Select two to four extracts per theme and make each one earn its place. A useful paragraph pattern is: claim, extract, interpretation, connection back to the question and to other themes. Participants on r/QualitativeResearch ask how much worked example to include, and the answer is enough to let a reader trace your reasoning, not enough to reproduce your dataset.
Report the method transparently. Your methodology section needs your approach and epistemology, sampling and participants, the phases you followed, your coding process and coding style, reflexivity practice, measures of trustworthiness and any software used. A single student wellbeing dataset of 20 semi-structured interviews, analysed reflexively following Braun and Clarke’s six phases, with four coding styles and no software beyond a spreadsheet, is a complete and honest paragraph. Be specific enough that someone could repeat it.
Common Mistakes
These are the errors that come up repeatedly in student analysis, each with the correction that fixes it.
Using your interview questions as themes. If your interview guide has five sections, your findings will not have five themes called the same thing. Questions shape the conversation; themes come from what participants did with them. Rename anything that maps one-to-one onto your guide.
Renaming codes as themes. A theme sits above codes and needs support across participants. If your theme has one code in it, you have a code.
Saying themes emerged. In reflexive thematic analysis you construct themes using your own judgement, so “themes emerged from the data” misrepresents your role. Write “themes were generated” and say how.
Overcoding. Labelling every clause produces 400 codes and no analysis. Keep codes close to meaning units and merge early; undergrads who ask for help after coding are usually drowning in granularity.
Undercoding. Two codes called “stress” and “anxiety” across an entire dataset is a missed analysis, not a clean one. Split them when the extracts differ.
Vague code definitions. “Emotional stuff” cannot be applied consistently and cannot be audited. One sentence per code, written when you create it.
Skipping contradictory evidence. The extract that does not fit is usually the most informative one in the dataset. Chase it, report it, and narrow your theme around it.
Forcing one theme hierarchy onto every dataset. A framework that worked on your pilot study may not fit your main data. Nothing obliges you to reuse last year’s structure, and forcing it is how weak themes get defended as “consistent”.
Claiming saturation without documenting it. “No new codes emerged” is an assertion, not a finding. State what you did: how many late interviews you coded, how many new codes they produced and what you concluded from that. And if your institution uses COREQ-style reporting checklists, community advice on r/academia is to work through the items early rather than at submission.
A short quality check before submission: can a reader trace every claim to a coded extract, can you explain why each theme is distinct from its neighbours, and does your reflexive journal show you engaging with the moments that contradicted you? Those three answers cover most examiner objections.
Frequently Asked Questions
Should thematic analysis coding be inductive, deductive, or both?
Both is the usual answer, and it depends on whether you are generating codes from the data or testing a prior framework. Inductive coding starts from the data with no fixed categories, which suits exploratory questions. Deductive coding applies an existing framework, such as a theory of coping or a policy model. A hybrid approach codes inductively first, then maps the resulting codes onto the framework, and it is the most defensible when your research question names a model but the data came first.
How do I make qualitative coding less subjective?
You cannot remove subjectivity in reflexive thematic analysis, because the researcher is the instrument, and you should not try to. What you can do is make it visible and auditable. Keep a reflexive journal from the start, write a definition beside every code the first time you use it, keep dated versions of your codebook, and record why codes were merged or split. Community practice points consistently toward COREQ-style checklists as a discipline for exactly this. A reader should be able to follow your reasoning even when they disagree with it.
Can I do thematic analysis in Excel, Word, or Google Sheets?
Yes, and plenty of published student work was done that way. The spreadsheet layout that works is one row per quote, with columns for participant ID, transcript and line number, the extract itself, the initial code, the focused code after refining, the theme, and a notes column for your reasoning. Sort and filter on the theme column to pull supporting evidence. Word works if you put a code column beside each extract, and it suits smaller datasets better. A post on r/skimle describes the same one-row-per-quote structure.
How many themes should a thematic analysis have?
There is no fixed number, and any count you choose before analysing will distort the work. A typical defensible range runs from four to eight themes for a 15 to 30 interview dataset, with sub-themes underneath when a pattern is too broad to write about clearly. The real test is coverage and distinctness: every theme should answer part of your research question, sit at the same level of abstraction as the others, and be supported by extracts from more than one participant. Students asking how many is enough usually have a coverage problem, not a counting problem.
Do I need direct participant quotations in my results?
Yes, and they do specific work rather than decorating a point. Each quotation supports a claim you have made about a pattern, and it lets an examiner check your interpretation instead of taking it on trust. Two to four extracts per theme is usually enough, chosen because they are vivid, typical or analytically useful, not because they agree with each other. Where a participant’s words are ambiguous, paraphrase and flag the ambiguity. Include the participant ID and a line reference so any reader can locate the extract.
Conclusion
A dependable thematic analysis rests on four habits: fix the question and your boundaries before coding, define every code the moment you create it, keep moving between codes and themes until each theme is distinct and supported, and check your final interpretation back against the coded material every time. None of that requires expensive software. It requires a clear question, disciplined notes and an honest record of the decisions you made.
Start with the part most people postpone: transcribe your recordings, number the lines, organise the files and write your first-pass familiarization memo. Everything after that stage gets easier once you have seen what is actually in the data.


