A codebook is simply a written list of the labels you will use on your data, with a definition, an inclusion rule, an exclusion rule and a sample quotation attached to each one. Writing one takes a few focused sessions: read a sample of your material, draft provisional codes, define them tightly, test them on real excerpts, then document the final structure. Knowing how to write a codebook for qualitative coding is what turns coding from a private habit into something a supervisor, a second coder or an examiner can follow.
The version below works for interview transcripts, focus group discussions and open-ended survey responses. It also holds up for content analysis of documents and news coverage, where the same logic applies at a faster pace.
Table of Contents
- 1What You Need
- 2Step-by-Step: Build a Qualitative Coding Codebook
- 31. Define the Unit of Analysis
- 42. Review the Research Question and Existing Material
- 53. Start With an Initial Set of Codes
- 64. Define Each Code Precisely
- 75. Separate Codes From Categories
- 86. Add Decision Rules for Overlapping Data
- 97. Test the Codebook on a Sample of Data
- 108. Finalize the Structure and Document It
- 119. Create a Finished Codebook Table
- 12Common Mistakes
- 13Frequently Asked Questions
- 14How many codes should a codebook have?
- 15What should a qualitative codebook include?
- 16Should my codes be inductive or deductive?
- 17What is the difference between a code and a theme?
- 18Can AI or ChatGPT write a codebook for me?
- 19How do I keep my codebook reliable across coders?
What You Need
Most people who stall at the blank-codebook stage are missing inputs rather than skill. Gather these five things before you open a blank document.
- A one-sentence research question. Every code has to earn its place against this sentence. If a code does not connect to it, it is usually a topic you noticed rather than a pattern you can defend.
- A readable sample of your data. Ten to fifteen transcripts, responses or documents is plenty for the first pass. Reading all of your data before coding is not required, and nobody has time for it anyway.
- Your interview guide or topic list. The prompts you actually asked shape the codes you will end up with.
- Analysis software. NVivo, ATLAS.ti, MAXQDA and Dedoose all hold code definitions. A spreadsheet works too, especially for small studies.
- A versioned file and a change log. Name it codebook_v1 and keep every later version. Renaming a code halfway through an analysis without recording it destroys the audit trail, and you will not notice until write-up time.
One honest warning before starting: the codebook you draft today is a draft. It gets revised. Researchers on graduate forums describe revision as the normal condition of coding, not as a sign the method failed.
Step-by-Step: Build a Qualitative Coding Codebook
There is no single right sequence, but these nine steps keep the process transparent and repeatable. Skip a step and the gaps usually show up as disagreements between coders later.
1. Define the Unit of Analysis
Decide first what one click of the code applies to: a single word, a sentence, a paragraph, a speaker turn or a whole response.
This choice sets the size of the smallest unit you will ever store, and it changes what your findings mean. Sentence-level coding suits fine-grained language analysis and supports counts. Whole-response coding suits quick content analysis across hundreds of survey answers, where you record one dominant theme per respondent and can never claim how often a phrase appeared. Speaker-turn coding fits focus groups where the exchange between two people is the thing you care about.
Write the choice into the codebook’s opening paragraph. Two researchers coding the same interviews with different units produce different counts from identical data, which is exactly the kind of problem a supervisor will spot.
2. Review the Research Question and Existing Material
Read your research question, then your interview guide, then whatever theory you are drawing on.
Mark every part of the interview guide that points at a construct in your question. Those marks are your first candidate codes. If your framework already names concepts, such as the categories in a content analysis scheme or the domains in a healthcare interview study, you are starting deductively, and you should write those framework terms in as placeholder codes.
If you have no framework and you want to follow the data instead, start inductively with an empty list. Both routes are legitimate. What matters is that you say which one you used in the codebook’s introduction.
3. Start With an Initial Set of Codes
Read your sample twice without coding anything. The first pass is for comprehension, the second for noticing repetition.
Every time you spot an idea three or more times, write it down in plain words. Aim for something like eight to fifteen provisional codes to begin. Here is a realistic example from a student study on how first-year engineering students use office hours. One interview produces phrases like they just go there when they are stuck, the queue is the problem not the tutor, and honestly I would rather ask ChatGPT first. From those, a first draft of six codes emerges: seeking clarification, avoiding being seen to struggle, peer referral, time and queue pressure, alternative help sources, and tutor approachability. That is a workable starting set, and every code is traceable to something a participant actually said.
Where the codes come from matters, and the choice has consequences for the rest of the analysis.
| Approach | Where codes come from | Strength | Weakness |
|---|---|---|---|
| Deductive | Your theory, framework or interview guide | Comparable across participants and sites from day one | Forces the data into boxes that may not fit |
| Inductive | Recurring ideas in the transcripts themselves | Grounded in what participants actually say | Slow, and can drift away from the research question |
| Mixed | Deductive skeleton, inductive fill | Keeps the framework while letting the data surprise you | Requires you to document where each code came from |
Record the source of every code in the notes column from day one. Six months in, nobody remembers which six codes were framework-derived.
4. Define Each Code Precisely
A code name is a label. A code definition is a rule, and that is where codebooks are won or lost.
Give each code four fields: a name, a definition in one or two sentences, an inclusion rule stating what qualifies, and an exclusion rule stating what almost qualifies but does not. Add one example quotation so another coder can see the unit in practice. The exclusion rule is the part students skip and the part that solves most inter-coder disagreement, because it names the neighbouring code that would otherwise attract the same excerpt.
Compare two weak and strong versions to see the difference in practice.
| Weak code | Strong code |
|---|---|
| “Student attitudes” | “Avoidance of visible struggle: participant describes choosing not to ask for help because it signals poor performance” |
| “Help seeking” | “Help-seeking behaviour: any reported instance of contacting staff, peers or services for academic support, including the decision not to make contact” |
The strong labels are longer, and that is the right trade. You can shorten a label later. You cannot reconstruct a vague definition after three hundred excerpts have already been filed inconsistently.
5. Separate Codes From Categories
A code is the smallest label you attach to a piece of your data. A category groups several codes. A theme is an analytical pattern you argue about in your findings section.
Confusing the three is the most common structural error, and it is what makes codebooks either collapse into four enormous buckets or explode into two hundred micro-labels. Keeping the layers apart solves it. Build one codebook of codes, and keep a separate, much shorter layer of categories or themes above it. Graduate researchers on r/GradSchool describe exactly this two-layer pattern in practice: one codebook for specific codes, one for themes, so the two never get mixed.
In the office-hours example, six codes such as queue pressure and tutor approachability sit under one category named barriers to attendance. The category is what you count and discuss; the codes are what you attached to individual excerpts.
6. Add Decision Rules for Overlapping Data
Some excerpts legitimately carry more than one idea, and some pairs of codes will fight over the same sentence. Write the tie-breakers into the codebook instead of resolving them in your head each time.
Three rules cover nearly everything. First, order matters: the first rule the reader meets is the primary code, and any secondary code goes in a separate field. Second, declare mutual exclusivity where it genuinely exists, such as positive versus negative mention, and say so explicitly. Third, use the hierarchy for genuine containment: if a parent code already covers the excerpt, apply the parent and stop.
A worked decision rule for queue pressure reads: apply when the participant names waiting time, queues, timetables or limited opening hours as the reason for not attending. If the participant describes the tutor as unapproachable instead, use tutor approachability, not this code. If both appear in the same turn, queue pressure is primary and tutor approachability is secondary.
7. Test the Codebook on a Sample of Data
Drafting a codebook proves nothing. Testing it does, and the test is cheap.
Code ten to fifteen responses with the draft scheme. Keep a separate pilot log with four columns: excerpt, code you applied, codes you considered, and why the decision was hard. Do not change definitions while you are pilot coding, or you lose the record of what the original rule did.
Then look at the log for three signals. Codes that were never applied were too narrow or simply redundant. Codes that attracted almost identical excerpts need merging. Codes that kept appearing in the considered column overlap and need clearer boundaries. Coding a calibration sample with a second person and comparing decisions surfaces the same problems faster, and discussion of the disagreements is what actually improves the definitions.
Nowell and colleagues summarise the trustworthiness criteria researchers are asked to document, and this pilot pass is the evidence behind three of them.
8. Finalize the Structure and Document It
The finished document needs a short introduction, the code list with full definitions, the decision rules, and a record of its own history.
The introduction should state the research question, the unit of analysis, the approach you used, the software, the date, and the version number. The history matters as much as the codes: a change log with date, author, what changed and why gives an examiner your full audit trail in one page. So does a short analytic memo explaining the decisions behind the harder calls, which is the part most students omit and later wish they had written.
Give the file a name that sorts usefully, such as study2_codebook_v3_2026, and keep the previous versions. Teams working across sites share one codebook and reconcile it weekly; that reconciliation only works if every change is dated and attributed.
9. Create a Finished Codebook Table

One row per code, one column per field. Anything you add beyond these columns should earn its place.
| Code ID | Code name | Definition | Inclusion criteria | Exclusion criteria | Example quote | Category |
|---|---|---|---|---|---|---|
| C01 | Seeking clarification | Participant describes asking someone for an explanation of course material or an assessment. | Any reported request for explanation from staff or peers. | Requests for deadline extensions or administrative help. | “I just go there when I am stuck on the question.” | Help-seeking behaviour |
| C02 | Avoidance of visible struggle | Participant chooses not to ask for help because asking would signal poor performance to peers or staff. | Any stated decision to withhold a question for reputational reasons. | Not asking because the participant simply had no question. | “I would rather work it out badly than ask.” | Barriers to attendance |
| C03 | Queue and time pressure | Participant names waiting time, timetables or limited opening hours as the reason for not attending. | Explicit reference to queues, hours or scheduling. | Descriptions of the physical space that contain no time element. | “By the time I get there, the queue is already long.” | Barriers to attendance |
| C04 | Tutor approachability | Participant describes a tutor’s manner, tone or attitude as encouraging or discouraging. | Explicit comments about how a tutor interacts. | Comments about teaching quality of content only. | “She is the sort of person you can interrupt.” | Barriers to attendance |
| C05 | Peer referral | Participant was directed to office hours by another student. | Any mention of another student recommending attendance. | Self-initiated attendance with no referral mentioned. | “A friend of mine told me to go, and it worked.” | Help-seeking behaviour |
| C06 | Alternative help sources | Participant uses a non-staff source such as a textbook, a forum, an AI tool or a study group. | Any named non-staff resource used for understanding material. | Staff support of any kind, which belongs under C01. | “Honestly I would ask ChatGPT first.” | Help-seeking behaviour |
Saldaña’s coding manual is the standard reference for this kind of code definition and hierarchy, and Flick covers the logic behind selecting codes. Both are worth a library trip before you finalise anything.
Common Mistakes
Each of these has a quick correction, and each shows up in student codebooks more often than any other problem.
- Labels that describe a topic instead of a pattern. “Student attitudes” tells the next coder nothing. Fix: name the specific stance or behaviour, then write the definition that makes it decidable.
- Overlapping codes with no tie-breaker. Two people file the same quote differently and neither is wrong, because the codebook permitted both. Fix: write a primary and secondary rule for the pair, and delete the weaker code if the overlap is total.
- Too few codes, or four giant ones. A codebook of four buckets cannot support a claim about patterns. Fix: check granularity by asking whether any two codes could apply to the same sentence for the same reason. If yes, split.
- No exclusion criteria anywhere. This is the single largest source of disagreement between coders. Fix: complete the exclusion column before pilot coding, not after.
- Building the codebook before reading enough data. Framework codes that never match the transcripts waste months. Fix: draft deductively, then pilot on ten real responses and retire anything that never fires.
- Changing codes mid-analysis with no record. Fix: version every change, note the date and the reason, and note whether previously coded excerpts were recoded.
- Treating codes and themes as the same thing. Fix: keep the code layer and the category layer in separate documents, as described in step five.
- Never stopping. Codebooks that grow every week are usually a symptom of over-splitting. Fix: apply the stop rule, which is that a new code needs at least three distinct supporting excerpts and cannot be expressed as a subcode of something you already have.
On automation: AI tools can produce a draft list of candidate codes from a transcript, and they are useful for a sanity check or a first pass over a large document set. They cannot judge whether a definition matches your research question, and they will not notice that two of your codes mean the same thing. Treat generated codes as raw material that a person with your data has to accept, merge and discard on the record.
Frequently Asked Questions
How many codes should a codebook have?
There is no fixed number, but dataset size predicts the range. Start with 8 to 15 codes for a small interview study, expect 25 to 60 for a large multi-site project, and use a smaller codebook with a data-driven procedure for document analysis across hundreds of items. Treat any count as a starting estimate, and let the pilot pass set the final number.
What should a qualitative codebook include?
A codebook needs an introduction stating the research question, unit of analysis, approach and version, then one row per code with a name, definition, inclusion criteria, exclusion criteria and an example quotation. Add a category column if you use a hierarchy, plus decision rules for overlapping excerpts, a change log and the date and author of each revision.
Should my codes be inductive or deductive?
It depends on whether your research question tests an existing framework or describes what participants say. Deductive coding starts from theory or an interview guide and makes findings comparable across sites. Inductive coding starts from the transcripts and stays grounded in participants’ language. A mixed approach is common: a deductive skeleton with inductively derived codes filling the gaps.
What is the difference between a code and a theme?
A code is the smallest label you attach to an excerpt of data. A theme is a broader analytical pattern built by grouping codes and explaining what the pattern means for your research question. In practice you count codes and argue about themes. Keeping the two in separate documents stops your findings section collapsing into a list of labels.
Can AI or ChatGPT write a codebook for me?
It can draft candidate codes from a sample of your transcripts, which saves hours of the first reading pass. It cannot judge whether a definition fits your research question, spot two codes that mean the same thing, or write exclusion criteria that survive contact with a second coder. Use generated codes as raw material and document which ones you accepted, merged or rejected.
How do I keep my codebook reliable across coders?
Pilot the scheme on ten to fifteen responses, log every hard decision, then code a shared sample with a second person and compare. Where decisions differ, discuss the excerpt and fix the definition rather than the coder. That agreement, plus a dated change log, is what makes the analysis auditable and reportable in a dissertation or an ethics report.
Start with the smallest possible version. Read ten responses, write six codes with definitions and exclusion rules, code those same ten again, and see where the definitions fail. That single loop teaches you more than a longer draft will, and it is how to write a codebook for qualitative coding that another person can actually follow.


