A randomized controlled trial assigns participants at random to an intervention group and a comparison group, then measures the same outcome in both. That is the whole idea: any difference you find between the groups can be attributed to the intervention instead of to who happened to end up where. Learning how to design a randomized controlled trial for a student project takes about a week of planning and one semester of data collection, and the hard part is rarely the statistics. It is the recruitment, the ethics approval, and the discipline to write down your plan before you see any data.
Most published guidance on this assumes thousands of participants, a funded lab, and a statistician on the team. None of that is true in a classroom. What follows is the version that fits a term project: a small trial you can actually finish, built to the same design standards that make an RCT trustworthy, with the compromises named rather than hidden.
Table of Contents
- 1What You Need
- 2Step-by-Step
- 3How to Design a Randomized Controlled Trial: Start With the Question
- 4Select Participants and Set Eligibility Criteria
- 5Choose One Primary Outcome
- 6Plan a Sample Size Without Guessing
- 7Standardize the Intervention and Control
- 8Randomize and Protect Against Bias
- 9Predefine the Statistical Analysis
- 10Pilot, Approve, and Run the Trial
- 11Report Results Transparently
- 12Common Mistakes
- 13Tips for a Feasible Student Project
- 14Frequently Asked Questions
- 15Is a randomized controlled trial realistic for a student project?
- 16How many participants do I need for a student randomized controlled trial?
- 17What should the control group receive in a student RCT?
- 18Do I need IRB approval for a classroom research project?
- 19How do I interpret a null result in my trial?
- 20Conclusion: Start With a Feasible Question
What You Need

You cannot enroll anyone until nine things exist. If any of them is missing, the study is not ready, no matter how enthusiastic the recruitment plan looks.
- A written research question in PICO form. Population, intervention, comparison, outcome. If you cannot fill in those four blanks, the project is a topic, not a study.
- Supervisor approval. Your instructor signs off on the design before you collect anything, and again on the final protocol. Ten minutes now saves a rewrite later.
- An ethics determination. Contact your institution’s IRB or human subjects protection program. Ask specifically whether your study qualifies for exempt review. Start this early; approvals routinely take several weeks.
- A defined eligible population. Who can take part, who cannot, and how many of them exist at your source.
- A recruitment plan with numbers. Channel, expected response rate, and how you will replace people who decline.
- Intervention materials. Written protocol, dose, duration, and a way to check that participants followed it.
- An outcome measure. Validated where you can find one, with units and timing specified.
- Analysis software access. R, Stata, or SPSS, plus the version of the software so your output is reproducible.
- A data-management plan. Where the file lives, who can open it, how identifiers are separated from data, and how backups are made.
Two of those nine items deserve a warning. The ethics determination is not optional paperwork, and the analysis software decision should be made before recruitment, because choosing software after you have data is how people end up running tests that do not match their design.
Step-by-Step
How to Design a Randomized Controlled Trial: Start With the Question
Write the question as PICO before anything else. Among undergraduate students in one introductory statistics course, does a structured retrieval practice session each week compared with rereading the same material for six weeks improve exam performance?
That version is answerable. The vaguer version, studying the effect of study techniques on grades, is not, because study technique has fifty candidate definitions and grades have fifty possible influences. Narrow it until each of the four PICO elements fits in a single sentence.
Feasibility is part of the question, not something you check later. Can you recruit that population? Can you deliver that intervention for that long? Can you measure that outcome with something other than your own opinion? A question that fails any of those tests is a question for a different project.
Select Participants and Set Eligibility Criteria
Write your inclusion and exclusion rules as a numbered list before you approach a single person. The most common student mistake here is being accidentally selective: recruiting only the students who answer your message promptly, or only your friends, which turns a randomized trial into a convenience sample with extra steps.
Count backwards from your target. If your power analysis says you need 120 complete cases and you expect 15% attrition, you need to enroll about 142. If your pool of eligible students is 90, you have the wrong sample size, and the honest fix is to drop to a pilot design rather than to hope attrition stays low.
Decide how people get assigned before you recruit, too. If participants choose their own group, you no longer have randomization, whatever the write-up claims. This is the single design flaw that reviewers and instructors check for first.
Choose One Primary Outcome
Pick exactly one primary outcome and write down its units and its timing. Percentage points scored on the final exam, measured at the end of week six. Not “engagement” and not “academic performance,” which are too vague to analyze.
Everything else you measure is a secondary outcome, and secondary outcomes are allowed to disappoint. Keeping the distinction honest is what stops a project from becoming a fishing expedition where whatever looks significant gets reported.
Prefer measures that are hard to bias. A scored test blind-coded by someone who does not know the group is better than a self-report survey the participant knows is being collected for a study of the technique they were assigned to. If you grade it yourself, blind-code the participant numbers and strip the group labels before you open the file.
Plan a Sample Size Without Guessing
Sample size comes from a power calculation, not from the number of people who happen to be in your class. The inputs are the primary outcome, the test you will run, the effect size you consider meaningful, the alpha level, and the desired power.
Standard conventions for a two-arm parallel trial are alpha at 0.05 with a two-tailed test and power at 0.80. Those conventions give you the numbers below, which you can reproduce in G*Power or in the pwr package in R.
| Expected effect size (Cohen’s d) | Participants per group | Total, before attrition | Total with 15% attrition |
|---|---|---|---|
| 0.30 (small) | 88 | 176 | 208 |
| 0.50 (medium) | 64 | 128 | 152 |
| 0.80 (large) | 26 | 52 | 62 |
Now be honest about the row you land on. A trial with 30 students total has roughly 80% power only to detect effects around d equal to 0.7, which in test-score terms is a gap most teaching interventions do not produce. Your study will probably be underpowered, and that is a design fact, not a failure.
The correct move is to reframe. Run a pilot study whose stated purpose is to estimate feasibility, refine your assumptions, and justify a larger trial later. Report the effect estimate and its confidence interval rather than a significance verdict, and say plainly which effect sizes your sample could and could not detect. A well-framed pilot earns more credit than an over-claimed definitive trial.
Standardize the Intervention and Control
Write the intervention as a protocol someone else could follow without asking you a question. Same materials, same duration, same instructions, same time window. If two participants get different versions because you changed your mind mid-semester, that is a protocol deviation and you have to report it.
The control condition decides what your result means, so choose it deliberately.
| Control type | What the control group receives | Use it when |
|---|---|---|
| Untreated | Nothing beyond usual practice | Usual practice is the thing you are comparing against and withholding it causes no harm |
| Active | A different established approach | You want to know whether your method beats the usual method, not whether it beats nothing |
| Placebo or attention-matched | Something inert or equally demanding | You suspect motivation or time-on-task explains any difference |
| Waitlist, or delayed intervention | The same intervention later | You cannot ethically withhold something that plausibly helps, or your pool is small |
For student projects the waitlist control is often the best option. Every participant eventually receives the intervention, no one is denied a benefit, and your comparison stays clean.
Randomize and Protect Against Bias
Randomization has two separate jobs, and students constantly merge them into the word “blinding.” Sequence generation is who decides the assignments. Allocation concealment is whether the person enrolling a participant can predict the next assignment. Blinding is who does not know the group, after assignment. Only the first two are under your control before enrollment; full blinding is often impossible in a classroom study.
| Method | How it works | Use this when | Student-scale caveat |
|---|---|---|---|
| Simple randomization | Each participant gets an independent coin flip or random number | You have no baseline variables that matter much | Small samples can produce lopsided group sizes by chance alone |
| Blocked randomization | Randomize in blocks of 4 or 6 so each block splits evenly | You have a small sample and want predictable balance | The block size must stay secret until enrollment ends, or the last assignment becomes guessable |
| Stratified randomization | Block separately within strata such as baseline score band | One variable strongly predicts the outcome | Keep the strata simple; too many strata leave tiny cells |
| Matched pairs | Match two similar participants, then randomize within the pair | Very few participants, but you can pair on a strong predictor | Halves your effective sample and makes matching errors hard to undo |
| Cluster or section-level | Randomize whole class sections or table groups, not individuals | Contamination is likely, as when classmates share study materials | You need more clusters than individuals, which is hard at student scale; analysis must account for clustering |
Contamination is the practical problem in a classroom. Classmates talk, forward the study guide, and guess which condition their friend received. When that risk is real, randomize at the section level rather than pretending individual assignment held.
Concealment you can actually run: generate the full allocation list in advance with a free web randomizer such as randomizer.org, print it, and put each assignment in a numbered opaque envelope. The person handing out envelopes cannot see the next one. Sequentially numbered sealed envelopes are the standard student-scale version of allocation concealment, and they are far better than a public coin flip at enrollment time.
Participants should not choose their group, and neither should the person collecting the outcome data if you can help it. Where full blinding is impossible, blind at least the outcome assessor and say exactly who was blinded, in those words. Saying “double blind” without naming the people is a dead giveaway that nobody was blinded.
Predefine the Statistical Analysis
Write the analysis plan before enrollment closes, ideally in a dated document you send your supervisor. For a standard two-arm parallel trial with a continuous outcome, the primary analysis is a two-sample t-test on the post-test scores, or better, a regression or ANCOVA that adjusts for the baseline score. Adjusting for baseline is usually more efficient, and it costs nothing.
State in advance what you will report: group means, the difference in means, the effect size, and a 95% confidence interval. Report the interval whether or not it excludes zero. Also decide your missing-data plan before you see how much data is missing.
The default analysis is intention-to-treat: analyze each participant in the group they were randomly assigned to, not the group they actually ended up in. Dropouts are still dropouts. A per-protocol analysis comparing only the people who finished is fine as a secondary analysis, and it must be labeled as one.
In software terms, the sequence is short. Create two variables, group and score. Check the baseline comparison. Run the test. Then run the adjusted version. Record the package versions and save the session so the numbers can be regenerated.
Pilot, Approve, and Run the Trial
Run a pilot with a handful of people before the real thing. You are testing the forms, the instructions, the timing, the randomization slip procedure, and your data-entry setup. Most of what breaks in student trials breaks here, and a pilot is cheap.
Get supervisor sign-off and your ethics determination before recruitment starts. Ask in writing whether your study is exempt from review, because classroom research often is, and an exemption letter is worth more to your write-up than a paragraph of justification.
Pre-register the protocol on the Open Science Framework if you can. Registration is cheap credibility: a dated record of your hypotheses, exclusions, and analysis plan is very hard to argue with later. If the project is not hypothesis-driven, an analysis plan is still worth posting.
During the trial, keep a fidelity log recording whether each participant actually received the intervention as written. Deviations happen, and a log turns a surprise into a documented fact.
Report Results Transparently

Start with a participant flow diagram showing how many people were assessed, how many were randomized, how many received each condition, how many completed, and how many were included in the analysis. CONSORT 2025, the reporting standard for randomized trials, expects this diagram and it doubles as your attrition accounting.
Then give a baseline table of the two groups before reporting any outcome. A baseline imbalance at random is normal and is not evidence of a flawed randomization; what matters is that you looked.
For the primary outcome, give the effect size with its confidence interval, and put the two groups’ means and standard deviations next to it. Discuss deviations from the protocol and the limitations of the design, and name the ones you know: small sample, single course, contamination risk, self-reported secondary measures.
If the result is null, do not treat it as proof that the intervention does not work. A study with 30 participants and a wide confidence interval is consistent with a small real effect. Say what your data can and cannot distinguish, which is what the interval is for.
Common Mistakes
| Mistake | Why it damages the study | Fix |
|---|---|---|
| Research question with no comparison group | A pre/post change cannot separate the intervention from everything else happening in the term | Add a second group that does not receive the intervention yet |
| Participants choose their group | Self-selection reintroduces every confounding variable you randomized away | Generate the full allocation sequence in advance and conceal it |
| Randomizing by public coin flip at enrollment | The enroller knows the pattern and can steer who gets what | Use a pre-generated list inside numbered sealed envelopes |
| Sample size chosen after collection | Powering on observed data inflates the chance of a false positive | Run the calculation before enrollment and report the assumptions you used |
| Changing the primary outcome after seeing results | It converts a confirmatory test into a search across many outcomes | Fix the primary outcome in the protocol and report every change you make |
| Dropping dropouts with no accounting | Differential dropout biases the estimate in the direction you cannot predict | Analyze by assigned group, show the flow diagram, run a sensitivity check |
| Saying double blind without naming who | Readers cannot judge the bias risk from a label | State who generated the sequence, who enrolled, who assessed, and who analyzed |
| Reporting only the significant analyses | Selective reporting is the easiest way to turn noise into a finding | Report the primary outcome first, whatever it shows |
Tips for a Feasible Student Project
Recruit from more than one source if you can. A single course, a single club, or a single instructor’s section puts your whole result on one person’s mood that week.
Use a validated measure where one exists for your outcome. A published scale usually has scoring instructions, norms, and a psychometrics section you can cite, which makes your methods section stronger than a custom three-item questionnaire.
Keep the intervention to something you can deliver consistently for the whole study. Twenty minutes a week for four weeks is deliverable. Two hours a week for a semester is where student trials quietly fall apart.
Build retention into the schedule rather than hoping. Short sessions, a fixed measurement date, a reminder the day before, and one make-up window beat any amount of goodwill.
Frame a small trial as a pilot and say so in your title and abstract. It removes the expectation that you will produce a definitive answer and lets you report feasibility data, which is genuinely useful to whoever runs the study next.
Keep a dated folder with your protocol, your ethics correspondence, your randomization list, and your analysis plan. When your supervisor asks in week 12 where a decision came from, that folder is the whole answer.
Talk to your supervisor or your ethics board early, before you have built anything. Most of the expensive mistakes are decided in a fifteen-minute conversation that costs nothing.
Frequently Asked Questions
Is a randomized controlled trial realistic for a student project?
Yes, at a small scale, as long as you frame it honestly. Students on research forums often assume trials are only for funded teams, but a pilot trial in one course with 30 to 60 participants is a legitimate design. What you give up is statistical power, so describe it as a pilot and report feasibility, adherence, and effect estimates with confidence intervals rather than claiming a definitive answer.
How many participants do I need for a student randomized controlled trial?
Run a power calculation based on your primary outcome. For a two-arm trial with a two-tailed alpha of 0.05 and 80% power, a medium effect of d equal to 0.5 needs about 64 participants per group, while a small effect of d equal to 0.3 needs about 88 per group. Add roughly 15% for attrition. With a much smaller pool, plan for a pilot and state your minimum detectable effect.
What should the control group receive in a student RCT?
A waitlist control works well in most classroom projects. The control group follows normal practice and receives your intervention at the end of the study, which keeps the comparison fair without withholding something that might help. An active control, where you compare your method against an established alternative, is better when you want to show your method beats the usual approach rather than beating nothing.
Do I need IRB approval for a classroom research project?
Ask your institution’s human subjects protection program in writing before you recruit anyone. Studies that collect identifiable data, manipulate participants, or involve vulnerable groups usually require review, and some classroom projects qualify as exempt. Approval takes weeks, so start early. Even when review is not required, getting written confirmation that your study is exempt gives your write-up an approval letter to cite.
How do I interpret a null result in my trial?
A null result in a small trial usually means the study could not detect the effect, not that no effect exists. Look at the confidence interval around your estimate: if it spans both a small benefit and a small harm, your data simply cannot tell them apart. Report the interval, name the effect sizes your sample could have detected, and describe what a larger trial would need to answer the question.
Conclusion: Start With a Feasible Question
Three things to do before anything else. Write the question in PICO form and check that you can recruit the population, deliver the intervention, and measure the outcome within the term. Ask your human subjects protection program in writing whether your study needs review or qualifies as exempt. Then run the power calculation and accept the design that results, including a pilot, if the numbers are small.
Everything after that is execution, and execution goes faster than most students expect once the protocol is written. A working answer to how to design a randomized controlled trial for a student project is not a big sample or a fancy tool. It is one focused question, one honest comparison group, a concealed allocation, and a written plan you follow even when the result is not what you hoped.


