What Are Control Variables and How to Choose Them (October 2026)

What are control variables and how to choose them? Control variables are the factors a researcher holds constant, or adjusts for statistically, so they cannot explain the relationship between the independent variable and the dependent variable. Picking them is a matter of causal judgement rather than guesswork, and the six-step method below is the version I hand to students who are stuck.

Updated for October 2026. The examples cover SPSS, R and Stata, and the reporting wording is written for a methods section you can paste into.

Table of Contents
  1. 1What Are Control Variables?
  2. 2Control variable or control group: two different things
  3. 3Is a control variable the same as a covariate?
  4. 4Independent, dependent and control variables side by side
  5. 5Why Do Researchers Use Control Variables?
  6. 6Why it is not important to control all variables
  7. 7How to Choose Control Variables Step by Step
  8. 8Step 1: Write your hypothesis as a causal claim
  9. 9Step 2: Draw a simple causal diagram
  10. 10Step 3: Classify each candidate by its causal role
  11. 11Step 4: Check timing and plausibility
  12. 12Step 5: Check feasibility and power
  13. 13Step 6: Justify each one in writing
  14. 14Examples of Control Variables in Student Research
  15. 15Control Variables vs. Confounding and Moderating Variables
  16. 16Must a control variable be correlated with both variables?
  17. 17How Many Control Variables Should You Include?
  18. 18How to Report Control Variables in Your Study
  19. 19A methods sentence that works
  20. 20In the results section
  21. 21Entering covariates in SPSS, R and Stata
  22. 22Frequently Asked Questions
  23. 23What makes a variable a good control variable?
  24. 24How do I know whether a variable is a confounder?
  25. 25Should I control for every variable related to my outcome?
  26. 26Can a control variable be categorical or continuous?
  27. 27When should I include an interaction or moderation term instead of a control variable?
  28. 28How do I explain control variables in my results?
  29. 29What to Do First

What Are Control Variables?

A control variable is any factor deliberately held constant or statistically adjusted for so that it does not distort the relationship between the independent variable and the dependent variable. Holding it constant protects internal validity: the confidence that the effect you observed comes from what you manipulated rather than from something else that moved at the same time.

In plain terms, a control variable is everything in your study that is not your independent variable and not your dependent variable, but could still influence the dependent variable.

Common control variables include:

  • Time of day or season when data were collected
  • Room temperature, lighting or laboratory conditions
  • Age, gender, education level or income of participants
  • Prior attainment or baseline test scores
  • Plant species, soil type and pot size in a growth experiment
  • Dose, duration and placebo allocation in a clinical trial
  • Region, industry or firm size in business data
  • Word count or question order in a survey instrument

Control variable or control group: two different things

A control group is a group of participants who receive no treatment, used mainly in experiments. A control variable is a factor held steady across conditions. They are related but not interchangeable: an experiment can have a control group and no control variables at all, and a survey can have control variables and no control group.

Is a control variable the same as a covariate?

Functionally, yes. A covariate is the statistical term for a control variable you could not hold constant, so you entered it into the analysis instead. Some researchers distinguish them, reserving “control” for experimentally standardised conditions and “covariate” for statistical adjustment. The useful rule is to state which one you mean in your methods section, because a supervisor reading your thesis will not guess.

Independent, dependent and control variables side by side

Variable typeRole in the studyExample (plant growth experiment)Held constant?
Independent variableThe factor you manipulateType and amount of fertiliserNo, it varies by design
Dependent variableThe outcome you measureChange in height after six weeksNo, it is the result
Control variableEverything else kept steadyWater volume, light hours, pot size, plant speciesYes, by design
Confounding variableA third factor linked to bothSoil nutrient content differing by pot groupHeld constant or adjusted

Why Do Researchers Use Control Variables?

The reason is the third-variable problem. Without controls, any factor that influences your outcome and also varies with your independent variable produces a relationship that looks causal and is not. Coffee sales and exam performance move together in plenty of datasets, and the honest answer is usually deadline season rather than caffeine.

Controls do four practical things. They remove alternative explanations, so your claim survives a sceptical reader. They improve precision, because a covariate that explains part of the variance in your outcome leaves less noise for your effect to hide in. They let you compare groups fairly, which is why matching on age or baseline score matters when a randomised allocation is unavailable. And they make results reproducible, since a second analyst who controls the same variables reproduces the same number.

Why it is not important to control all variables

You will see “control for everything relevant” repeated across advice sites, and it is bad guidance. Adding a variable that sits on the causal path between your predictor and your outcome blocks the effect you are trying to measure. Adding a common effect of both creates a spurious association that was not there before. More controls is not safer; the right controls are, and sometimes the correct set is short.

How to Choose Control Variables Step by Step

This is the part most guides skip. Six steps, and the order matters, because the cheap mistakes happen early.

Step 1: Write your hypothesis as a causal claim

Say it as a sentence with a direction: “higher weekly practice hours raise exam scores.” If your claim has no direction, no control strategy will rescue it, and if it is purely associational, say so and stop worrying about causal adjustment.

Step 2: Draw a simple causal diagram

Sketch your independent variable, your dependent variable, and every candidate factor you can think of. Draw an arrow for each plausible cause-and-effect link, including the ones you would rather not think about. A ten-minute sketch prevents most of the damage done later.

Step 3: Classify each candidate by its causal role

This is the decision that decides everything. A confounder has arrows pointing into both your predictor and your outcome — it is a good control and you want it in. A mediator sits on the path from predictor to outcome, so controlling for it estimates a direct effect rather than the total effect you probably wanted. A collider is a common effect of both, so controlling for it opens a path that was closed. A moderator changes the strength of the relationship and belongs in an interaction term, not as a main-effect control.

Step 4: Check timing and plausibility

A control measured after the outcome cannot clean it. Baseline scores, pre-treatment characteristics and design-stage conditions work. Post-treatment measures usually do not, and a variable measured at the same time as the outcome invites reverse causality.

Step 5: Check feasibility and power

Drop any candidate you cannot measure well. A badly measured control adds noise and can bias estimates toward zero just as effectively as a collider biases them away. Then count: each control spends a degree of freedom, so with a small sample your list has to be short.

Step 6: Justify each one in writing

Prepare a one-line reason per control naming its causal role. That single sentence per variable is what separates a defensible analysis from a variable dump, and it is exactly what a reviewer or supervisor asks for.

Examples of Control Variables in Student Research

Here is what the procedure looks like in practice across five designs. Notice that in every case the control is there for a stated reason, not because it sounded relevant.

FieldTypical control variablesHow you would justify one
Laboratory and biologyTemperature, light cycle, water volume, species, pot sizeHeld constant so growth differences cannot be attributed to growing conditions
Clinical researchAge, sex, baseline severity, prior treatment, placebo allocationAdjusted because baseline severity predicts both allocation and outcome
EducationPrior attainment, attendance, class size, parental education, SEN statusAdjusted because prior attainment predicts both course choice and final grade
Business and marketingFirm size, industry, region, prior sales, seasonAdjusted because season predicts both campaign exposure and sales independently
Survey and social scienceAge, gender, income, education, employment status, region, question orderAdjusted because income and age relate to both the predictor and most attitude outcomes

In a plant growth experiment, you hold water volume, light hours, pot size and species constant, and you measure height change as the dependent variable. Hold enough of these steady and the comparison is clean; hold too few and a light-cycle difference explains your whole result.

In a clinical trial, randomisation is doing the heavy lifting, but age and baseline severity still get adjusted for, because with randomisation you average imbalances out in expectation, and adjustment restores precision in any given trial.

In education research, prior attainment is almost always a control, and it is almost always a good one. Students with higher prior scores both select certain courses and score higher regardless of teaching method, so leaving it out makes weak teaching look effective.

In a marketing study, prior sales and season matter more than firm age. A campaign that runs in December will beat the same campaign run in July for reasons that have nothing to do with the campaign.

In a survey, question order and the platform used are controls that researchers routinely forget. Both can shift responses without touching the topic being measured.

Control Variables vs. Confounding and Moderating Variables

Confusing these five roles causes most of the mistakes in this area. The table gives the structural difference, and the examples show what each one looks like in a sentence.

RoleStructureShould you control for it?Example
ConfounderCauses both predictor and outcomeYesIncome affects both educational attainment and health outcomes
Plain controlNo causal link, but noise sourceYes, if measured reliablyTime of day in an attention task
ModeratorChanges the strength of the linkNo, enter an interaction termSleep quality moderating the effect of caffeine on alertness
MediatorSits on the path between themNo, blocks the effectDieting affects exercise, which affects weight
ColliderCommon effect of bothNo, creates biasHospital admission caused by both illness and treatment choice

A moderator is easy to recognise because the effect looks different across groups. If practice hours help novices much more than expert students, the right move is an interaction term, not a separate main-effect control for experience. The same goes for a variable that behaves differently by condition, which often signals that a separate analysis or a matching design would serve you better than one pooled regression.

Must a control variable be correlated with both variables?

No, and this heuristic is widely taught and widely misunderstood. A variable related to both the predictor and the outcome is a classic confounder, but a control is justified by its causal role, not by its correlation pattern. Controls related to neither are still worth standardising if they add noise. And a variable related to both can still be a bad control when it is also a collider.

How Many Control Variables Should You Include?

There is no fixed number, and any answer that gives you one is a rule of thumb rather than a result. What governs the count is your research question, your sample size, the quality of each variable, and the model you are fitting.

Practical guidance for the situations students hit most often:

  • Experiments with a small sample. Keep the list short. Five well-measured controls beat fifteen patchy ones, because measurement error in a covariate attenuates your estimate just as surely as a collider distorts it.
  • Large survey or panel datasets. The limit is collinearity rather than sample size. Check the variance inflation factor for every control, and drop anything above roughly 5 or 10 depending on your field’s convention.
  • Regression models. Rule-of-thumb guidance often quoted is around ten observations per predictor, which is conservative but sensible when your controls are numerous and your sample is modest.
  • Panel data. Unit and time fixed effects already absorb a great deal, so time-invariant controls such as gender or region become redundant once those are in the model.

Every additional control costs precision if it carries no confounding information. That is the trade-off, and it is why the defensible list is usually shorter than the list of everything you could think of.

How to Report Control Variables in Your Study

Reporting is where most marks are lost, because students list their controls without saying what they were for. One sentence per variable fixes it.

A methods sentence that works

Write it as: “Control variables were selected a priori based on their role in the causal model and adjusted for in the regression. Age and prior attainment were included as confounders of both teaching method and final grade. Time of day was standardised across all testing sessions. Question order was not included, as it was randomised at the point of administration and showed no association with the outcome in preliminary checks.”

That sentence does three things a reader needs: it says the selection was theory-led rather than data-driven, it names the role of each variable, and it explains the variable you deliberately left out. Explaining exclusions matters as much as explaining inclusions.

In the results section

Report the adjusted model as your primary specification, give the unadjusted estimate alongside it so readers can see how much the controls moved things, and report the confidence intervals. If a control coefficient is not your interest, you can still note whether it behaved sensibly, since a baseline score with the wrong sign usually signals a coding error.

Entering covariates in SPSS, R and Stata

The mechanics are short in all three. In SPSS, use Analyze, then Regression, then Linear, and move your controls into the covariates box; the same model with categorical predictors goes under Analyze, then Regression, then Binary Logistic when the outcome is binary. In R, fit a formula with the outcome first and controls after, for example lm(score ~ method + prior_score + age, data = d), and check collinearity with the car package. In Stata, regress score method prior_score age does the same thing, and estat vif reports the variance inflation factors for the whole model.

Frequently Asked Questions

What makes a variable a good control variable?

A good control variable has a defensible causal relationship to your outcome: ideally it is a common cause of your predictor and your outcome, and it is measured accurately and before the outcome occurs. It should also be recorded in a way that lets other researchers repeat the same adjustment. Correlation with both variables is a useful sign of a confounder, but it is not the test itself.

How do I know whether a variable is a confounder?

Sketch a quick diagram with your predictor, your outcome and every candidate. If arrows run from the candidate into both the predictor and the outcome, you have a confounder and should control for it. If the candidate sits between them, it is a mediator and controlling for it blocks the effect. If it is caused by both, it is a collider and controlling for it introduces bias.

No. Controlling for every related variable can block mediators, open collider paths and dilute your estimate with noise. Include a variable because you can state its causal role, because it was measured reliably, and because it was recorded before the outcome. A short list with written justification is stronger than a long one assembled from everything that seemed relevant.

Can a control variable be categorical or continuous?

Both. Categorical controls such as gender, region or employment status enter a model as indicator variables, with one category as the reference. Continuous controls such as age or baseline score enter as they are. Two cautions apply: a categorical control with many levels and thin groups eats degrees of freedom, and a continuous control measured on a coarse scale, like a five-point band, is often better treated as categorical.

When should I include an interaction or moderation term instead of a control variable?

When the effect of your predictor changes across levels of another variable. If the relationship between practice hours and exam scores is much stronger for novices than for experienced students, experience moderates the effect, and a main-effect control will hide that. Enter the interaction term instead, then report the simple effects at each level so readers can see where the difference lies.

How do I explain control variables in my results?

State that the model was adjusted a priori for named confounders, list those variables, and give the adjusted estimate with its confidence interval. Add the unadjusted estimate beside it so the effect of adjusting is visible. Finally, say which plausible confounders you could not measure and what direction of bias that might produce. Reviewers reward that limitation statement because it shows you know what your design cannot deliver.

What to Do First

Before you collect anything, write your hypothesis as a causal sentence and draw a diagram with every factor you can think of. Sort those factors into confounders, mediators, colliders and plain controls, keep the confounders you can measure reliably, and set aside everything else. Then write one justification sentence per control you kept. That single hour of work does more for your analysis than any amount of model tweaking afterwards, and it is the answer to how to choose control variables that survives a supervisor’s questions.

Leave a Comment

Practical guides to statistics, surveys and research data

Read the latest guides