How to Write a Data Management Plan for a Thesis (October 2026)

A data management plan (DMP) is a short document that states what research data you will create or use, how you will organise, store, secure and document it, who may access it, and what happens to it once your thesis is finished. Knowing how to write a data management plan for a thesis comes down to answering a standard set of questions about your data, usually in your faculty template or in an online tool, and then updating that plan at each stage of the research.

Most students find the task easier once they stop treating it as a piece of writing and start treating it as a list of decisions. Every dataset you will collect raises the same handful of questions: where does it live while I work on it, who can open it, what document explains it, and what do I do with it in three years. Answer those and the plan writes itself.

Budget three to five hours for a master’s thesis and rather more if you are handling interviews, clinical records or partner data. Do it before you collect anything, not after. Several universities require a plan as part of your research ethics application, and an ethics board that reads the plan alongside your consent form will spot contradictions between the two straight away.

Table of Contents
  1. 1What You Need
  2. 2Step-by-Step: How to Write a Data Management Plan for a Thesis
  3. 3Step 1: Confirm the Requirements and Scope
  4. 4Step 2: Create a Data Inventory
  5. 5Step 3: Choose File Formats and Naming Conventions
  6. 6Step 4: Plan Storage, Organization, and Backups
  7. 7Step 5: Address Security, Privacy, and Ethics
  8. 8Step 6: Document the Data and Research Process
  9. 9Step 7: Plan Sharing, Archiving, and Reuse
  10. 10Step 8: Set Retention and Disposal Rules
  11. 11Step 9: Review and Approve the Final Plan
  12. 12Common Mistakes
  13. 13Frequently Asked Questions
  14. 14Do I need a data management plan for a literature review thesis?
  15. 15How long should a thesis data management plan be?
  16. 16What should be included in a data management plan?
  17. 17Can I share my interview transcripts publicly?
  18. 18What happens to my thesis data after I graduate?
  19. 19What is the difference between a data management plan and a data analysis plan?
  20. 20Conclusion

What You Need

What You Need

Gather these before you open a blank document. Twenty minutes of reading saves an afternoon of guessing what your faculty actually expects.

  • Your faculty’s research data management policy. This usually sets the approved storage tools, the minimum retention period and who owns the data once you finish.
  • Your supervisor’s expectations. Some supervisors want a two-page plan, others want a section in the thesis proposal with the codebook attached.
  • The ethics or research permit requirements. Many committees ask for the DMP as an annex to the application, so read the checklist before writing.
  • Your consent forms and any participant information sheets. Read what you actually promised, because sharing plans cannot go beyond it.
  • A list of your collaborators. Co-authors, supervisors, lab technicians and partner organisations all affect who needs access and where files sit.
  • Access to repository options. Check whether your university runs its own repository and whether you can deposit in a subject repository such as Zenodo, Dryad or ICPSR.
  • Storage and backup tools approved by your institution. A managed cloud drive with versioning and multi-factor authentication usually covers this, and it is not the same as a personal account.
  • Retention rules for consent and ethics documentation. These are frequently kept for longer than the data itself, and often by a different person.

If your faculty has a Word template, use it rather than a generic one. Reviewers score against the questions they expect to see, and a clean answer in their format beats a well-written answer in the wrong structure.

Step-by-Step: How to Write a Data Management Plan for a Thesis

Step 1: Confirm the Requirements and Scope

Open the template your institution uses and read every question before you answer any of them. The questions tell you the scope, and they tell you the order your reviewer expects. Then write down the length limit, the due date and the person who approves it.

You also need to decide what the plan covers. For most theses that means the survey responses, interview recordings and transcripts, observation notes, consent records and analysis scripts you create yourself. If you are reusing an existing dataset, say where it came from, under what licence and whether the original provider permits redistribution. Check that it works: if you can name the exact file location and the exact access conditions for every dataset, your inventory is already accurate.

Step 2: Create a Data Inventory

The inventory is the backbone of the plan, and it is the part most students skip. List every item you will create or hold, and for each one record its source, format, approximate volume, sensitivity, intended use, owner and expected retention period. A spreadsheet with one row per dataset is enough.

DatasetSourceFormatVolumeSensitivityRetention
Survey responsesOnline form, self-completedCSV plus .savRoughly 400 rows, under 2 MBLow once identifiers removedFive years after submission
Interview recordings15 semi-structured interviewsWAV, MP3About 20 GBHigh, contains voice and namesTen years, restricted access
TranscriptsTranscribed from recordingsDOCXAbout 120 pagesHigh until pseudonymisedTen years, restricted access
Analysis scriptsWritten by youR scripts12 filesNonePermanently archived with data

Volume matters more than students expect. Twenty gigabytes of audio behaves differently from a two-megabyte survey file, and it changes which storage and transfer tools are reasonable. Estimate it, round it up generously, and write the number down.

Step 3: Choose File Formats and Naming Conventions

Prefer open, non-proprietary formats for anything you intend to archive or share, because they can still be opened when the software that produced them is long gone. A CSV file stays readable. A proprietary export from a licensed package may not, unless you also export a copy in a neutral format and document that you did.

For statistical work, keep both the native project file and an export. Your SPSS project file, Stata dataset or R workspace has a proprietary extension, and that is fine for daily use as long as the accompanying script reproduces it. A reader in ten years should be able to check your cleaning steps by reading the script rather than by guessing at the menus you clicked.

Then write down a naming convention and stick to it, including the date and a version marker. Something like 20260418_interviews_v02_anonymised.wav tells a future reader the date, the content, the version and the processing state. Vague names like final_final_v3_USE tell them nothing. Check that it works by asking someone else to identify the current version of a file from its name alone.

Step 4: Plan Storage, Organization, and Backups

Name one primary storage location, one backup location and one archive location, and be specific about each. Institution-approved managed storage is usually the right answer for active data. Personal desktops and personal cloud accounts are usually not, because they lack multi-factor authentication, access logs and a recovery path after you graduate.

Organise folders by project and stage rather than by date, so that finding the cleaned dataset is not an archaeological exercise. Keep raw data read-only, and never overwrite it. Analysis outputs go in a separate folder from raw data, which saves you from a very specific and very common kind of panic.

Backups need a rule and a test. A daily automatic sync with version history covers most theses; a weekly separate copy covers the rest. State who checks that it worked, and how often. Then actually restore one file from the backup and open it. A backup nobody has ever tested is a hope, not a backup.

Plan for the awkward cases. What happens when your laptop dies in month seven, when the lab closes, when a co-author leaves the project or when you move institutions? Naming an exit plan for each keeps your supervisor’s most practical question from becoming a research interruption.

Step 5: Address Security, Privacy, and Ethics

Start from your ethics approval rather than from general good intentions. The plan has to match what the committee approved, and any difference between the two invites a question you would rather avoid.

Cover the technical controls in plain language: full-disk encryption on laptops, multi-factor authentication on shared accounts, access restricted to named people, and a defined process for moving files. If you use a third-party tool for survey responses or transcription, name it and confirm it is approved by your institution, because data leaving your institution’s environment changes your obligations.

Cover the data itself. Keep the link between a participant’s identity and their responses separate from the responses, and store that key in a different location with stricter access than the data. Where a participant could be identified from free-text answers, sensitive topics or small subgroups, write down how you will redact or withhold those fields before any release. Secure deletion means documented, verifiable deletion rather than dragging a file to the recycle bin, and your plan should say which of the two you mean.

Step 6: Document the Data and Research Process

Documentation is what separates a reusable dataset from a stack of files nobody can read. At minimum, produce a README file explaining the folder structure, a codebook listing every variable with its definition and units, and a short description of the steps that turned raw data into the analysed dataset.

Write the codebook as you go rather than afterwards. Twenty minutes after each collection wave is far cheaper than reconstructing six months of variable recoding at submission time. Include value labels, missing-value codes, the scale of every numeric variable and the exact question wording behind each column, since a variable called q7_sat tells a future reader nothing about what it measured.

Version differences between two releases of a statistics package cause real reproducibility problems, so record which software and version produced each result and cite them in your thesis. That detail belongs in the data management plan, not only in the methods chapter.

Step 7: Plan Sharing, Archiving, and Reuse

Accessible is not the same as open, and that distinction resolves most student anxiety. Restricted data can be perfectly FAIR if the restrictions are described clearly and access can be requested. What you cannot do is promise open access that your consent form or a partner agreement forbids.

Go through each dataset and classify it: open, restricted with conditions, embargoed until a publication date, or retained and not shared. Write down the reason for the restricted category, since a bare refusal to share is the answer reviewers question most often. An anonymised dataset with a codebook, a licence and a citation instruction is usually publishable even when the raw recordings are not.

For archiving, prefer a subject repository that your field already uses, since that is where a reader of your thesis will look first. A general repository with a persistent identifier is the fallback when no disciplinary option exists. Record which repository, which licence and which identifier type, and state what travels with the deposit: data, codebook, scripts and the README.

Step 8: Set Retention and Disposal Rules

Decide the retention period per dataset rather than per project, because interview material and survey exports rarely deserve the same answer. The usual pattern is a few years for de-identified quantitative data and considerably longer for sensitive qualitative material, with consent and ethics records often retained longest of all. Your faculty policy sets the floor; your own plan states the number.

Name the responsible person for each period, including the period after you graduate, because the common failure is an unowned dataset that nobody is permitted to delete. Say where the data will be held during the retention period, usually institutional storage rather than a personal account, and how deletion will happen at the end. Secure deletion from backups needs a specific note too, since some systems expire backups on a fixed cycle whether or not anyone acts.

Step 9: Review and Approve the Final Plan

Step 9: Review and Approve the Final Plan

Read the finished plan as if you were the reviewer, and then read it as if you were a stranger who has to reuse your data. The first read catches vague promises. The second catches missing documentation.

Then go through it with a checklist. Search for every instance of the words will, proper, appropriate and securely, and replace each with a concrete action. Check that every dataset from your inventory appears, that each file format is named, that each person has a role, and that no statement about sharing contradicts your consent form or your data use agreement. Get the approval in writing, record the date, and store the signed copy with your ethics documentation.

Treat the plan as a living document rather than a finished form. Update it at the stage gates: proposal, ethics approval, the end of data collection, submission, and archive. A change of method mid-thesis usually changes the data inventory, and a plan that silently disagrees with your methods is worse than no plan at all.

Common Mistakes

Vague security statements. Data will be stored securely says nothing a reviewer can act on. Replace it with the named service, the access list and the encryption method. If you cannot name them, that is the work the plan is asking you to do.

A laptop as the only copy. One device is one failure away from losing a year of work. Institution storage plus a separate backup, with a restore you have actually tested, is the minimum.

Ignoring the consent wording. Participants agreed to specific terms, and you cannot retroactively widen them. Where the wording is ambiguous, either restrict access or seek ethics approval for a revised consent process before collecting more data.

Unclear file names and version chaos. Files called final2_FINAL cannot be restored, and nobody can tell which version an analysis used. Write the convention down in the plan and follow it from the first week.

Undocumented variables. A dataset without a codebook is a dataset with a short shelf life. Build the codebook as you collect, not at submission.

Blanket promises to share everything. Promising open data you cannot legally or ethically release damages your credibility. Classify each dataset and give a reason for each restriction.

A plan nobody can act on. The four-to-six page boilerplate version gets skimmed and ignored. Cover your actual datasets, in the actual formats, with the actual tools you will use.

Two habits help more than any of these fixes. Keep the plan next to the data in the same folder structure, so it is updated as work happens rather than reconstructed afterwards. And ask your supervisor to read the draft early, because the comments about scope cost less to act on before you have written four pages.

Frequently Asked Questions

Do I need a data management plan for a literature review thesis?

Usually not. A pure literature review creates no primary data, so there is nothing to store, secure, restrict or archive. You may still need a short statement describing how you will manage the reference sources, any licensed datasets you reuse, and any copyrighted material you quote. If your review involves screening records, coding them into a spreadsheet or extracting data into a table, that counts as data and a short plan is sensible. Check your faculty’s submission checklist rather than assuming.

How long should a thesis data management plan be?

For a master’s thesis, two to four solid pages usually covers it. PhD plans with interviews, partner data or multiple datasets run longer. Length matters far less than coverage: a dense plan that names your files, formats, tools and retention periods is stronger than six pages of generic sentences. If your faculty imposes a limit, treat it as the ceiling and cut background explanation before cutting specifics.

What should be included in a data management plan?

Six things: a description of each dataset and its volume, the formats and naming conventions you will use, where data is stored and backed up, how it is documented, who may access it and under what conditions, and what happens to it at the end. Add roles, responsibilities, consent and ethics constraints, and the repository and licence for anything you share. Answering these in your institution’s template order keeps the plan easy to score.

Can I share my interview transcripts publicly?

Only if your consent form allowed it and the transcripts carry no identifying detail. Most thesis consent forms permit reuse by the researcher but not public release, so transcripts usually stay restricted. A common workable compromise is to archive pseudonymised transcripts with a codebook under restricted access, then publish a smaller open dataset of derived, aggregated or heavily redacted material. If identification remains possible from free-text answers or rare descriptions, withhold those parts and document what you withheld and why.

What happens to my thesis data after I graduate?

Follow the retention period in your plan and your faculty policy, and keep the data in institutional storage during that period rather than a personal account. Typically the dataset and its documentation are archived with a persistent identifier, and identifiable material is destroyed when the retention period ends. Deposit first where you can: archiving in a repository makes deletion of your working copies straightforward, and a published dataset with a DOI is easy to cite. If your institution never claimed the data, ask in writing before you assume either way.

What is the difference between a data management plan and a data analysis plan?

They answer different questions and are usually separate documents. The data management plan covers the whole lifecycle of the data: collection, formats, storage, security, documentation, sharing and retention. The data analysis plan covers the analytic strategy: which variables, which tests, which subgroup comparisons, how missing data will be handled and what will count as a meaningful result. A thesis often contains an analysis chapter, a management plan and sometimes a short management paragraph in a funder application, and none of them substitutes for another.

Conclusion

Start with the data inventory. One row per dataset, with its source, format, size, sensitivity and owner, and you have already answered most of the plan without writing a paragraph of prose.

Then confirm which requirements actually apply to you, and turn each research activity into a specific decision about storage, security, documentation, sharing and retention. A plan built from named files and named tools reads as competence, because it is.

Leave a Comment

Practical guides to statistics, surveys and research data

Read the latest guides