How to Create a Codebook for Your Dataset: A Practical Guide 2026

A codebook, also called a data dictionary, is a document that lists every variable in your dataset and explains what it means: its name, label, data type, the codes used for every possible value, and how missing data is recorded. To create a codebook for your dataset, you work through one variable at a time in a spreadsheet and fill in those fields.

Most people already have half the work done. The variable names are in the header row, and the codes are hiding in your SPSS syntax file, your Stata do-file or your questionnaire. Assembling them into one table takes an afternoon, and it is the difference between a dataset someone else can use and a dataset that comes back to you with an email asking what homealon means.

This guide is for students, graduate researchers and analysts producing data that someone else will analyse. It works in a spreadsheet first, then covers the faster routes through IBM SPSS Statistics, Stata, R and NVivo. Updated for 2026.

Table of Contents
  1. 1What You Need
  2. 2Step-by-Step: How to Create a Codebook
  3. 31. Inspect the dataset before you create a codebook
  4. 42. Create one row for every variable
  5. 53. Write clear variable labels and descriptions
  6. 64. Document values, codes, and missing values
  7. 75. Record variable types, sources, and measurement details
  8. 86. Review the codebook against the dataset
  9. 97. Export, save, and maintain the codebook
  10. 10A worked example
  11. 11Common Codebook Mistakes and How to Fix Them
  12. 12Confusing variable names with labels
  13. 13Treating missing values as zero
  14. 14Listing a range when the values are categories
  15. 15Letting labels drift across variables
  16. 16Documenting a recode but not the original
  17. 17Updating the data and forgetting the codebook
  18. 18Describing a sensitive variable carelessly
  19. 19Frequently Asked Questions
  20. 20How do I create a codebook?
  21. 21How to create a codebook in Excel?
  22. 22How do I create a codebook in SPSS?
  23. 23How do I create a codebook in NVivo?
  24. 24What should a codebook include?
  25. 25How do you document missing values in a codebook?
  26. 26Conclusion

What You Need

What You Need

Gather four things before you open a blank spreadsheet. Without them you will end up documenting variables twice, once from memory and once from the real file, and the two versions will not match.

  • The raw data file itself. CSV, SAV, DTA or RDS. Work from the file that holds your actual values, not a cleaned copy you plan to make later.
  • A list of the variable names. Usually the header row, exported as a single column so each name can become one codebook row.
  • The questionnaire, instrument or interview guide. This is where the wording behind each variable lives. A code built from memory drifts from the wording you actually asked.
  • Your missing value rules. Which numbers mean don’t know, which mean refused, which mean not applicable. Decide these once, at the top, rather than per variable.
  • A place to write it. Any spreadsheet will do. Stata, SPSS and R users can generate a draft automatically first and then correct it.

Two preparation steps save time. Write your unit of analysis on a sticky note and put it in the codebook header, because a variable called hours_worked means something different at the person level than at the job level. And note the version date of the data file the codebook describes.

Step-by-Step: How to Create a Codebook

1. Inspect the dataset before you create a codebook

Open the data and look at it before writing a single description. In Excel, select the header row and check for duplicates, trailing spaces and inconsistent capitalisation; SPSS, Stata and R all report variable names and types without you having to read the header yourself.

You want three answers: what does one row represent, which variables came straight from the instrument, and which ones you computed. Sorting variables into raw and derived matters because derived variables need a formula, not just a label.

A useful check is a frequency distribution for each column. Numeric variables give you a minimum, maximum and value count in seconds; categorical variables show exactly which codes exist and which are unused. Those codes are your value labels, already sitting in the file.

2. Create one row for every variable

The standard layout is one row per variable and one column per field. That orientation is what repositories expect, what makes filtering possible, and what lets you check the codebook against the data row by row.

These are the columns worth having. Anything beyond this list is optional.

ColumnWhat goes in itExample
Variable nameThe exact name in the file, character for characterhomealon
LabelPlain-language name, usually the exact question wordingIs the respondent living alone?
TypeNumeric, string, date, or coded categoryNumeric
Measurement levelNominal, ordinal, or scaleNominal
CodesEvery valid value and its meaning1 = Yes, 2 = No, 9 = Don’t know
Valid rangeLowest and highest legitimate value1 to 9
Missing valuesWhich codes are not real answers9 = Don’t know
SourceQuestion number or derivationQ4, or mean of Q7 and Q8
NotesSkip logic, reverse coding, redactionsFree-text names removed for confidentiality

Three more columns earn their place on real projects: valid count, percent missing, and a derived flag marking whether the variable is computed. Those last three come free once you have generated a draft with a script.

3. Write clear variable labels and descriptions

A good label lets a stranger understand the variable without opening the questionnaire. Paste the question text where you can, then trim it to a single line that fits in a column header.

Avoid descriptions that repeat the name, like “age: the age of the respondent”. Say what was measured and in what units: “Age at interview in years, calculated from date of birth”. For scales, name the instrument and the range: “Perceived stress scale, 10 items, 0 to 40, higher is more stress”.

Keep the wording consistent across every variable in the same block. If one income variable says “annual household income” and another says “income”, a reader will guess. Deciding on one phrasing early prevents a search through your notes later.

4. Document values, codes, and missing values

List every code the file actually contains, including the ones nobody should answer. Value labels and variable labels do different jobs, and the codebook needs both: the variable label says what is being measured, the value labels say what each number means.

Missing values get their own column because software treats them differently, and the difference decides whether a row lands in your denominator.

SoftwareSystem missingUser-defined missing codes
SPSSA single period (.)Declare up to three discrete values, e.g. 9 and 99
StataEmpty cells, shown as .Extended missing values .a to .z, e.g. 9 = don’t know
RNA of any typeNaN for undefined, or a factor level such as “DK” kept as data
ExcelBlank cellAny agreed sentinel number; blanks and zeroes are ambiguous without a stated rule

The nine-series convention is worth stating explicitly in the codebook: 9 usually means don’t know, 99 means refused, and 999 means not asked because of a skip pattern. Whichever scheme you pick, document it once in a note at the top of the codebook and reference it from every row that uses it.

Never let a real zero hide among your missing codes. If respondents could genuinely answer zero and you have coded 0 as missing, one of those two is wrong and the codebook is where that gets caught.

5. Record variable types, sources, and measurement details

Type and measurement level are separate columns for a reason. Data type is how the value is stored, numeric or string, while measurement level is what you may do with it: nominal for categories with no order, ordinal for ordered categories like agreement scales, scale for continuous quantities where differences and ratios hold.

Getting this right before analysis saves rework. Treating an ordinal agreement scale as scale is the single most common statistical mistake in survey data, and the codebook is where you would have caught it.

Also record units, the direction of any scored index, the collection method, and any transformation applied. For a derived variable, the source column should hold the formula, not the word “calculated”. If the mean of two items becomes a score, write the items, the range check, and what happens when one item is missing.

6. Review the codebook against the dataset

A codebook written from memory drifts from the file. The fix is a pass that compares the two directly, one column at a time.

Work through this list:

  • Every variable name in the file appears exactly once in the codebook, spelled identically.
  • Every code listed for a variable exists in the data, and every code found in the data is listed.
  • The label matches the wording on the instrument, not a paraphrase of it.
  • Minimum and maximum fall inside the stated valid range.
  • Missing codes are declared as missing in the software settings, not only described.
  • No duplicate variable names, and no variable documented twice under different labels.
  • Derived variables carry a formula that another person could type and re-run.

Where your software can automate this, use it. Scripts that compare unique values against your declared codes will catch a mistyped code that your eye will skip a dozen times.

7. Export, save, and maintain the codebook

Save the codebook in two formats. A spreadsheet or CSV is the working copy you can sort and filter, and a PDF or Word file is the read-only version you attach to a deposit or hand to a collaborator.

Keep the file next to the data it describes, with the data file name in the codebook title, so nobody can end up with data from March and a codebook from November. If your institution or repository asks for machine-readable metadata, keep the CSV too; that is what gets harvested.

Then treat it as version-controlled. Every time you recode a variable, rename one or merge two, update the codebook in the same sitting and bump the version and date in the header. A codebook that lags the data is worse than none, because people will trust it.

If you work in a specific package, you can skip much of steps 2 to 4 by generating a draft. In Stata, codebook writes a full variable listing to a text file with codebook, compact giving the short form. In R, the dataMaid package function makeCodebook() produces an HTML or PDF report, and codebookr builds a table from a data frame. In SPSS, Utilities andgt; Codebook Authoring exports variable and value labels for the whole file. In NVivo, the codebook describes categories and node codes rather than numeric variables.

Whichever route you take, treat generated output as a first draft. It captures structure, not meaning, so you still add the labels, the source questions and the notes by hand.

A worked example

This is what a finished row set looks like for a small survey dataset. Homealone comes from a single yes or no question; stress_score is computed.

Variable nameLabelTypeLevelCodes and valid rangeMissingSource
homealonLiving alone in householdNumericNominal1 = Yes, 2 = No; 1 to 29 = Don’t knowQ4
hh_incomeAnnual household income in GBPNumericScale0 to 250000999 = Refused, 9999 = Not askedQ11, banded top code
stress_scorePerceived stress scale total, 0 to 40NumericScale0 to 40Missing if fewer than 8 items answeredMean of PSS10_1 to PSS10_10

Note the two different missing conventions in three rows. That is not sloppiness. It is what happened to the data, and writing it down is the point of the exercise.

Common Codebook Mistakes and How to Fix Them

Confusing variable names with labels

The name income is a technical handle. The label is what the variable measured, in what units, over what period. Filling the label column with the name gives a reader nothing. Write the question, then trim it.

Treating missing values as zero

A blank income is not zero income, and recoding 9 to 0 quietly puts a non-respondent into your lowest income group. Keep 9 as a declared missing value in the data, and keep the rule in the codebook. Recode only at the analysis step, and say so in your methods.

Listing a range when the values are categories

Writing “1 to 5” for an agreement scale hides everything useful. List each code with its label: 1 = Strongly disagree through 5 = Strongly agree, and note that the codes are ordered but the gaps are not equal.

Letting labels drift across variables

Once two similar variables have different phrasings, searches for one term miss the other. Set the wording rules once, at the start, and apply them to every row.

Documenting a recode but not the original

For a derived or recoded variable, keep both: the source variable and its codes, and the new variable with its transformation. Write the rule as a formula someone could type, not as “standardised” or “cleaned”.

Updating the data and forgetting the codebook

This is the most common and the most damaging, because the codebook still looks authoritative. Fold codebook updates into the same cleanup session as the recode, and record a version number and date so drift is visible.

Describing a sensitive variable carelessly

Some variables cannot be documented literally. For free-text fields or small geographic areas, describe the variable and its processing without reproducing identifying content: “Free-text employer name, removed prior to deposit under the disclosure review”. A codebook is a public document once the data is public.

Frequently Asked Questions

How do I create a codebook?

Start with your data file open and build a spreadsheet with one row per variable and one column per field: variable name, label, data type, measurement level, codes with their value labels, valid range, missing values, source and notes. Work through the variables in order, using value labels and skip logic already stored in your statistical software. When you finish, check every documented code against the values actually present in the file.

How to create a codebook in Excel?

Put variable names in column A of a workbook, one per row. Build headers for label, type, measurement level, codes, valid range, missing values, source and notes. Use Data andgt; Text to Columns on the header row if names arrived combined, and a Data Validation list on the type and measurement level columns so entries stay consistent. Freeze the top row, then save as .xlsx for editing and export a PDF copy for sharing.

How do I create a codebook in SPSS?

Use Utilities andgt; Codebook Authoring, which lists every variable with its label, type, measurement level, value labels and missing values, then exports the file. If your labels already live in a syntax file, open it instead: the VARIABLE LABELS and VALUE LABELS blocks contain exactly what the codebook needs. Export from Codebook Authoring and add the descriptions and source questions by hand, since software can read labels but cannot infer meaning.

How do I create a codebook in NVivo?

A NVivo codebook documents categories and node codes rather than numeric columns, and the built-in reports cover part of it. Export your code frame and node attributes to Excel, then add a column defining each category, its parent, and examples of coded excerpts. Keep definitions short and consistent, because qualitative codebooks are read by people applying the codes later, often in a different project.

What should a codebook include?

At minimum: variable name, plain-language label, data type, measurement level, every valid code with its value label, valid range, missing value codes, the source question or derivation, and any notes on skip logic or transformations. Add the dataset-level details separately, including title, creator, date, unit of analysis, sample size and a citation. Those headers are enough to build the codebook from, and they match what repositories ask for.

How do you document missing values in a codebook?

Give every missing code its own entry in a missing values column and state what each one means, such as 9 for don’t know and 99 for refused. Then declare those codes as missing in your software so they are excluded rather than treated as numbers, and note whether blank cells represent not applicable or a skipped question. Record the same rule once at the top of the codebook so every variable follows it.

Conclusion

Open your data file, put one variable per row in a spreadsheet, and fill in the name, label, type, measurement level, codes and missing values for each one. Then read the codebook back against the file before you analyse anything. That check is the whole point, because it is the moment you find the code nobody ever meant to answer and the recode you forgot to document.

Leave a Comment

Practical guides to statistics, surveys and research data

Read the latest guides