How to Run Linear Regression in R: A Step-by-Step Guide (2026)

You run a linear regression in R by passing a model formula to lm() and then printing it with summary(). The whole core workflow takes about two minutes once the data is loaded, and you need no extra package for it.

That is the part most tutorials bury. They spend a screen of setup before showing you the four lines that matter, then stop before the diagnostics, which is exactly where a model quietly turns out to be wrong. Below is the version I would hand a student: prepare the data, look at it, fit, check, fix, report.

Table of Contents
  1. 1What You Need
  2. 2Step-by-Step: How to Run Linear Regression in R
  3. 31. Prepare the Data and Variables
  4. 42. Explore the Relationship Before Modeling
  5. 53. Fit the Linear Regression Model
  6. 64. Check the Regression Assumptions
  7. 75. Improve or Troubleshoot the Model
  8. 86. Interpret and Report the Results
  9. 9Common Mistakes
  10. 10Frequently Asked Questions
  11. 11What is the lm() function in R?
  12. 12Do I need to install a package to run linear regression in R?
  13. 13How do I check the assumptions of a linear regression in R?
  14. 14How do I include a categorical variable in a linear regression in R?
  15. 15What is the difference between R-squared and adjusted R-squared?
  16. 16Can I compare two regression models in R?
  17. 17Conclusion

What You Need

What You Need

Three things, and only one of them is software you have to install.

  • R itself plus RStudio Desktop, the free IDE most people learn on. RStudio is just a front end for R, so every command here works in a plain terminal too.
  • A dataset with a continuous outcome. R ships with mtcars, 32 cars with fuel economy and weight, which is why it appears in every beginner tutorial. Load it with data(mtcars) and nothing needs downloading.
  • Two optional packages for diagnostics: car for the VIF and influence plots, lmtest for the Breusch-Pagan test. The model itself does not use them.

On variables, keep the naming straight before you type anything. The dependent variable (the response, outcome) is continuous and goes on the left of the formula. The independent variables (predictors) go on the right, joined with plus signs.

install.packages(c("car", "lmtest"))   # only needed for the diagnostics step
library(car)

If you prefer learning the syntax in a fixed course format, our guides on running statistical tests cover the same setup logic for other tests. The rest of this article assumes the RStudio workspace above is open and empty.

Step-by-Step: How to Run Linear Regression in R

1. Prepare the Data and Variables

Load the data and look at it before modelling. Four functions tell you what you are dealing with.

data(mtcars)
head(mtcars)        # first six rows
str(mtcars)         # column types: num, chr, int, factor
summary(mtcars$mpg) # range, median, quartiles
colSums(is.na(mtcars))  # missing values per column

str() is the one people skip and then regret. A column typed as chr instead of num will be treated as a category, not a number, and R will fit a model you did not intend. Any categorical predictor needs to be a factor:

mtcars$cyl <- factor(mtcars$cyl)
class(mtcars$mpg)   # should print "numeric"

Decide the outcome and the predictors. Here mpg is the response, wt is the main predictor, and hp and cyl come in when you adjust for confounders.

2. Explore the Relationship Before Modeling

Look at the two variables before you fit anything. A correlation gives you direction and rough strength in one number.

cor(mtcars$mpg, mtcars$wt)
# -0.8676594

Negative, close to minus one: heavier cars get lower fuel economy. Now plot it so you can see whether the straight-line assumption even makes sense.

ggplot(mtcars, aes(x = wt, y = mpg)) +
  geom_point(size = 3) +
  geom_smooth(method = "lm", formula = y ~ x, se = TRUE) +
  labs(x = "Weight (1000 lbs)", y = "Miles per gallon",
       title = "Fuel economy by weight")

If your points curve, log the outcome or add a squared term later. Better to know now than to wonder why the residuals look wrong in step 4.

3. Fit the Linear Regression Model

Fit the Linear Regression Model

Fit the model with an explicit formula. The tilde means “modelled by”, and data = tells R where the columns live.

fit <- lm(mpg ~ wt, data = mtcars)
summary(fit)

Here is the console output, exactly as it prints.

Call:
lm(formula = mpg ~ wt, data = mtcars)

Coefficients:
            Estimate Std. Error t value Pr(>|t|)
(Intercept)  37.28517    1.89387  19.686  <2e-16 ***
wt           -5.34442    0.55919  -9.559  1.51e-10 ***

Residual standard error: 3.046 on 30 degrees of freedom
Multiple R-squared:  0.7523,	Adjusted R-squared:  0.7408
F-statistic: 91.36 on 1 and 30 DF,  p-value: 1.51e-10

Reading the coefficient table: the Estimate for wt is -5.344, the Std. Error is 0.559, and the t value divides one by the other to give -9.559. The Pr(>|t|) of 1.51e-10 is the probability of a t statistic that large if the true coefficient were zero. The asterisks are a shorthand for significance levels, with *** meaning below 0.001.

The Residual standard error of 3.046 is the typical size of a prediction error, in miles per gallon. The Multiple R-squared of 0.752 says weight alone explains about 75% of the variation in fuel economy. The F-statistic tests whether the model beats an intercept-only model; with one predictor it is the square of that t value, so it is not extra evidence here.

Add more predictors with plus signs. The factor() call on cyl is what turns a numeric-looking column into a proper category.

fit2 <- lm(mpg ~ wt + hp + factor(cyl), data = mtcars)
summary(fit2)
coef(fit2)
confint(fit2)   # 95% confidence intervals

Any object method works for pulling pieces out: coef(), residuals(fit2), fitted(fit2), summary(fit2)$r.squared.

4. Check the Regression Assumptions

Four diagnostic plots, one call. This is the step that turns a formula into a model you can defend, and for most people it is where the diagnosis finally starts making sense.

par(mfrow = c(2, 2))
plot(fit)
par(mfrow = c(1, 1))
  • Residuals vs fitted. Points should scatter around zero with no curve or funnel. A bow means you need a quadratic term; a fan means unequal variance.
  • Normal Q-Q. Points near the straight line mean roughly normal residuals, which matters for the p-values and the confidence intervals, not for the coefficients themselves.
  • Scale-location. The same homoscedasticity check with a fitted red line. A rising line is heteroscedasticity.
  • Residuals vs leverage. Spot the points that cook the model. Anything past Cook’s distance of 1 deserves an explanation in your notes.

Then test what the plots suggest.

lmtest::bptest(fit)                 # Breusch-Pagan: unequal variance?
car::vif(lm(mpg ~ wt + hp + disp + drat, data = mtcars))  # multicollinearity
car::outlierTest(fit)              # unusually large residuals

A Breusch-Pagan p-value under 0.05 means the constant-variance assumption fails. A VIF above 5, and certainly above 10, means two predictors are carrying nearly the same information.

5. Improve or Troubleshoot the Model

Each violation has a standard response, and picking the wrong one is how people make things worse.

  • Curvature in the residuals. Add I(wt^2) to the formula, or log the outcome: lm(log(mpg) ~ wt, data = mtcars).
  • Heteroscedasticity. Either log-transform, or keep the model and report sandwich standard errors with sandwich::vcovHC(fit) plus lmtest::coeftest. Shrinking standard errors to make a p-value look better is not a fix.
  • Multicollinearity. Drop one of the correlated predictors based on subject knowledge, or combine them. Do not simply delete the one with the smaller effect and pretend the confounders never mattered.
  • Influential points. Check whether an extreme value is a data error. If it is real, report the model with and without it rather than hiding it.
  • Rank-deficient fit. R warns when a term is an exact copy of another. Remove the duplicate instead of chasing the warning.

If the relationship is genuinely curved, jump models rather than fighting the diagnostics. R offers splines through the splines package when a straight line will not do.

6. Interpret and Report the Results

Write the equation in words, then the numbers. For the simple model above: predicted fuel economy equals 37.285 minus 5.344 times weight in thousands of pounds.

predict(fit, newdata = data.frame(wt = 3))   # one prediction
predict(fit, newdata = data.frame(wt = 3), interval = "confidence")

Report the 95% confidence interval rather than a bare p-value. Say the effect of a 1,000-pound increase is a 5.34 mpg decrease, 95% CI roughly -6.50 to -4.19, p below .001. The interval says how precisely you estimated it; the p-value only says whether zero is a plausible value.

On fit, describe R-squared as the share of variation in the outcome explained by the model, not as accuracy. Adjusted R-squared penalises extra predictors, which is why it can drop while R-squared rises when you add a variable. For a paper or thesis, one package turns the model into an APA-style table:

library(stargazer)
stargazer(fit, fit2,
  column.labels = c("Simple", "Multiple"),
  model.names = FALSE, no.out = TRUE, title = "Fuel economy model")

Finally, save the script so the analysis reruns end to end. Reproducibility is just the habit of running your own script from a clean session.

Common Mistakes

Seven problems I see in student work, each with the shortest working fix.

  • Using = instead of ~ in the model formula. = names an argument; ~ defines the model. Inside lm() only the tilde works.
  • Forgetting data = mtcars. R then hunts for the columns in whatever else happens to be in your workspace, and errors in a confusing way.
  • Reading correlation as causation. cor() says two numbers move together. Only the study design supports a causal claim.
  • Ignoring “NAs introduced by coercion”. A character column with a stray comma gets silently forced to NA. Find it with sum(is.na(x)) and fix the source column.
  • Overclaiming a p-value. It is not the probability your hypothesis is true, and p of 0.049 is not a real finding. Report the estimate and its interval.
  • Misreading the intercept. It is the predicted outcome when every predictor is zero. Outside the range of your data it means nothing at all, which is normal and fine.
  • Skipping the diagnostics. A significant model with bowed residuals still gets reported. Run plot(fit) before you write a single sentence about the fit.

Frequently Asked Questions

What is the lm() function in R?

lm() is the base R function that fits a linear model. You hand it a formula such as mpg ~ wt and it estimates the straight-line relationship by least squares, returning coefficients, standard errors, t-values, p-values and R-squared. lm() stands for linear model and works for simple and multiple regression alike; no package is required.

Do I need to install a package to run linear regression in R?

No. lm() and summary() ship with R itself, so the entire core workflow runs on a fresh install. You only add packages for the extras: car for VIF and influence plots, lmtest for the Breusch-Pagan test, stargazer for publication-ready tables. Most blog tutorials that demand tidyverse for a one-predictor regression are adding work you do not need.

How do I check the assumptions of a linear regression in R?

Run par(mfrow = c(2, 2)) followed by plot(fit) to get residuals versus fitted, a normal Q-Q plot, a scale-location plot and residuals versus leverage. Read them for curvature, funnel shape and influential points. For formal checks use lmtest::bptest(fit) for unequal variance and car::vif(fit) for multicollinearity.

How do I include a categorical variable in a linear regression in R?

Convert the column to a factor first, then include it in the formula: lm(mpg ~ wt + factor(cyl), data = mtcars). R creates one dummy coefficient per level except the reference level, which is the first one alphabetically, so the cylinder coefficient is the difference between that level and four cylinders. Use relevel(f, ref = 6) to change the baseline.

What is the difference between R-squared and adjusted R-squared?

R-squared is the share of variation in the outcome explained by the model, and it never falls when you add a predictor. Adjusted R-squared penalises predictors that do not earn their place, so it can decrease when a variable is added. Use plain R-squared to describe fit and adjusted R-squared to choose between models fitted to the same outcome.

Can I compare two regression models in R?

Yes, as long as the smaller model is nested inside the larger one, meaning the larger formula contains every term of the smaller. Run anova(fit1, fit2) and read the p-value on the F-statistic to see whether the added terms improve fit. Models with different outcomes or different datasets cannot be compared this way.

Conclusion

Start with the data, not the formula. Load it, check the column types, plot the relationship and read the correlation, then fit the model with lm(outcome ~ predictor, data = df) and read the summary() output line by line before you say anything about it.

Then run par(mfrow = c(2, 2)); plot(fit), because that is where you find out whether the model deserves the interpretation you are about to write. Report the equation, the coefficient with its 95% confidence interval, and R-squared described as fit rather than accuracy, and the result will hold up to a supervisor asking why.

Leave a Comment

Practical guides to statistics, surveys and research data

Read the latest guides