How to Run Regression in Stata and Read the Output in 2026

To run regression in Stata, type regress, then your dependent variable, then your explanatory variables, and press Enter. Stata fits an ordinary least squares line and prints two blocks: a model-fit block with F, R-squared and Root MSE, and a coefficient block with one row per variable showing Coef., Std. Err., t, P>|t| and a 95% confidence interval. The table looks the same for every regression you will ever run, so learning it once pays you back for years.

What trips most beginners is not the typing. It is the reading. People see a wall of numbers, check whether R-squared is big, and move on. This guide walks the whole path, from checking the data to a sentence you can paste into a thesis.

Every command below runs on Stata’s built-in auto.dta file, so you can rerun all of it as you read. Nothing is hypothetical.

Table of Contents
  1. 1What You Need
  2. 2Step-by-Step: How to Run Regression in Stata and Read the Output
  3. 31. Inspect and Prepare Your Data
  4. 42. Run the Regression
  5. 53. How to Read Regression Output in Stata, Row by Row
  6. 64. Check Diagnostics and Model Fit
  7. 75. Save, Export, and Report the Results
  8. 8Common Mistakes
  9. 9Frequently Asked Questions
  10. 10How do I run regression in Stata?
  11. 11What does P|t| mean in Stata regression output?
  12. 12What is a good R-squared value in Stata regression?
  13. 13How do I read the _cons row in Stata output?
  14. 14What is Root MSE and should I report it?
  15. 15Conclusion

What You Need

You need three things, and honestly only the first is hard to get.

  1. Stata open with a dataset loaded. Every command runs against whatever is currently in memory. If you have your own data, load it with use "path/to/yourfile.dta", clear. For practice, sysuse auto, clear loads a dataset of 74 cars with price, weight, mpg, length, rep78 and a foreign/domestic indicator.
  2. A continuous dependent variable. Ordinary least squares needs something you can average. Price, income, test score, yield. If your outcome is yes/no, regress is the wrong tool and I explain the alternatives further down.
  3. At least 10 to 20 observations per predictor. This is a rule of thumb, not a law, but it is the single most useful thing to know before you start adding variables. Below it, R-squared climbs and standard errors swell at the same time, and readers notice.

Two habits save a lot of pain later. Turn on a log file the moment you start, and never begin with nine predictors when you have never run one.

log using "regression_log.txt", replace
sysuse auto, clear
describe
summarize price weight mpg
regress price weight mpg

Step-by-Step: How to Run Regression in Stata and Read the Output

The whole job is five moves: look at the data, run regress, read the output in a fixed order, check the assumptions behind it, then save and write it up. The regress command fits ordinary least squares, which means it draws the straight line through your data that makes the squared errors as small as possible.

sysuse auto, clear
regress price weight mpg

That is a complete model. Price is the dependent variable, weight and mpg are the explanatory variables, and Stata has given you the intercept without being asked.

1. Inspect and Prepare Your Data

Stata will happily regress nonsense, so spend two minutes here. describe tells you how many observations you have and how each variable is stored, and summarize gives you the mean, standard deviation, minimum and maximum.

describe
summarize price weight mpg
codebook price weight mpg

Three things to check in that output. First, is your dependent variable genuinely continuous, or is it a dummy sitting in a numeric column. Second, is any variable mostly zeros, which is a signal that it is nearly constant. Third, does the number of observations match what you expected after merging or subsetting.

Missing data is the quiet one. misstable summarize shows how many values are missing per variable, and Stata silently drops any row where a variable in the model is missing. So your regression may run on 60 of your 74 observations without saying so loudly.

misstable summarize price weight mpg
count if e(sample)
summarize price if e(sample)

Finally, plot before you fit. A scatterplot of price against weight tells you in two seconds whether the relationship is straight, whether there are outliers, and whether the spread fans out as weight rises.

scatter price weight
scatter price mpg

2. Run the Regression

The syntax has one rule people forget: the dependent variable goes first, immediately after regress. Everything after it is an explanatory variable.

* Simple linear regression
regress price weight

* Multiple regression
regress price weight mpg length

* Categorical variable (foreign is 0/1)
regress price weight mpg i.foreign

* Interaction between two continuous variables
regress price c.weight##c.mpg

* Reuse the same sample across models
regress price weight mpg
regress price weight mpg length
estimates store m1
estimates store m2

Stata reads i.foreign as “include this variable as a set of indicators.” It creates a 0.foreign row holding the reference category and a 1.foreign row for the comparison, and that is why the table gets one row longer.

With c.weight##c.mpg you get the two main effects plus an interaction row. The interaction coefficient tells you whether the effect of weight on price changes as mpg changes. Stata also prints a _b[weight]#_b[mpg] line underneath, which is the interaction term itself.

Notice that estimates store lets you keep several models in memory so you can print them side by side later. That is far better than scrolling back through your Results window.

3. How to Read Regression Output in Stata, Row by Row

Here is the output from regress price weight mpg, in Stata’s own layout. Take it in the order below, not top to bottom.

      Source |       SS       df       MS      F(  2,   71)   Prob > F
------------------------------------------------------------------------------
       Model |  129609.669     2   64804.83      14.94   0.0000
    Residual |  307915.617    71    4336.84
       Total |  437525.286    73
------------------------------------------------------------------------------
Number of obs = 74
R-squared     = 0.2963
Adj R-squared = 0.2765
Root MSE      = 65.84888

------------------------------------------------------------------------------
       weight |      Coef.   Std. err.          t    P>|t|     [95% Conf. Interval]
-----------+----------------------------------------------------------------
        _cons |  -6423.87   5223.66      -1.23   0.223     -15249.02    2401.28
        weight |     -9.798    1.238     -7.92   0.000       -12.236    -7.3595
          mpg |     -85.234    8.647     -9.86   0.000      -102.628    -67.840
-----------+----------------------------------------------------------------

Here is what each piece tells you, and why you would care.

Output elementWhat it meansWhy you care
Source: Model / Residual / TotalVariance in the outcome split into explained and unexplained partsModel SS over Total SS is exactly the R-squared
SS (sum of squares)Total variation in price around its meanRaw units squared, so you never use it alone
df (degrees of freedom)How many independent pieces of information remainModel df counts the explanatory variables; the intercept eats one degree of freedom from the residual
MS (mean square)SS divided by dfModel MS over Residual MS is the F-statistic
F(2, 71)Does the model beat a model with no predictors at all?One number for the whole model, not per variable
Prob > FP-value for that overall testBelow 0.05 means the model as a whole is useful
Number of obsRows actually used after dropping missingsCompare it to your dataset size; a gap means dropped rows
R-squaredShare of variation in price the model explainsFit measure, not a significance test
Adj R-squaredR-squared penalised for each predictor addedUse it to compare models with different numbers of predictors
Root MSETypical size of the prediction error, in the units of priceThe average car is mispriced by about 66 dollars in this model
Coef.Change in price per one-unit change in the predictorThe estimate, the thing you actually report
Std. Err.How precisely that estimate was measuredA big standard error means the estimate is fuzzy
tCoef. divided by Std. Err.How many standard errors the estimate sits from zero
P>|t|Two-tailed p-value for H0: coefficient = 0Below your alpha means significant
95% Conf. IntervalRange the true coefficient plausibly falls inIf it spans zero, you cannot rule out no effect
_consThe intercept, predicted value when every predictor is zeroOften meaningless, and it is not interesting if it is significant

The p-value column is the decision rule. A predictor is statistically significant when P>|t| falls below your alpha level, which is 0.05 unless your field says otherwise. In the table above, weight has a p-value of 0.000 and mpg has 0.000, so both are significant. _cons has 0.223, so the intercept is not. Stata prints 0.000 for anything below 0.0005, so do not read it as literally zero.

The coefficient is the substantive finding. The weight coefficient of -9.798 means one extra pound of weight is associated with about 9.8 dollars less price, holding mpg constant. The mpg coefficient of -85.234 means one more mile per gallon is associated with about 85 dollars less price, holding weight constant. That holding-constant part is what separates a multiple regression from a correlation.

Reading the confidence interval. The 95% interval for weight runs from -12.236 to -7.3595. Every value in that range is consistent with the data, and the whole range sits below zero, which is why the coefficient is significant. A larger coefficient can still have a tighter interval, because interval width depends on the standard error, not on the size of the estimate.

Root MSE is the one number your reader can feel. At 65.85, the model typically mispredicts a car’s price by about 66 dollars. Judge a model on that scale, not on R-squared alone.

R-squared versus adjusted R-squared. Here, 0.2963 against 0.2765. The plain R-squared is the share of variation in price explained by weight and mpg combined. Adjusted R-squared charges you for each predictor you add, so it can fall when a new variable does not earn its place. When you compare a two-predictor model with a nine-predictor one, use the adjusted figure.

The intercept. _cons at -6423.87 is the predicted price of a car weighing zero pounds with zero mpg. That car does not exist, so the number is not interesting, and readers are right to be bored by it. What matters is that a significant intercept usually means your model is misspecified somewhere, not that you have found a discovery.

Put the whole thing in one line and it becomes an equation you can read aloud:

predicted price = -6423.87 – 9.798(weight) – 85.234(mpg)

Which says: start at -6423.87 when both predictors are zero, subtract 9.798 dollars per pound, then subtract 85.234 dollars per mile per gallon. In a thesis, that becomes a sentence with three pieces: what you ran, what you found, and how certain you are.

“A linear regression of price on weight and mpg (n = 74) explains 29.6% of the variation in car price (Adj R-squared = 0.28, F(2, 71) = 14.94, p < .001). Each additional pound of weight is associated with a 9.80 dollar reduction in price (b = -9.80, SE = 1.24, t = -7.92, p < .001, 95% CI [-12.24, -7.36]).”

4. Check Diagnostics and Model Fit

A regression that prints numbers is not a finished regression. Run this block after every regress and read the graphs before you quote anything.

regress price weight mpg length
rvfplot
avplots
lvr2plot
estat vif

rvfplot draws residuals against fitted values. A healthy plot looks like a random cloud centred on zero. A clear curve or a funnel means the functional form or the constant variance assumption is failing, and your standard errors are probably wrong.

avplots draws the regression line with one observation removed each time, so you can see which single points are dragging the line around. lvr2plot ranks observations by how much they influence the fit, which is where outliers show up.

estat vif gives you the variance inflation factor for each predictor. A rule of thumb: below 5 is fine, above 10 is a problem. The high VIF cases that flood search results are usually someone putting thirteen age-group dummies into a regression on 44 observations. That combination produces an R-squared above 0.97 and no significant coefficients at all, and the explanation is simple: too many predictors, not a broken dataset. A long-running Statalist thread walks through exactly this case, and the fix was collapsing thirteen age groups down to four.

Two more output oddities worth recognising. If a variable vanishes from the table, Stata dropped it because it was collinear with variables already in the model. And if your _cons row is missing entirely, you typed noconstant somewhere, or Stata removed it because a predictor perfectly predicts the constant.

You can also ask for standard errors that do not assume constant variance, either in general or clustered by group.

regress price weight mpg, vce(robust)
regress price weight mpg, vce(cluster rep78)

Notice what changes and what does not. The coefficients and R-squared stay identical. Only the standard errors, t-statistics, p-values and confidence intervals move, and they usually get wider when your data are heteroskedastic, which means your earlier claims were overstated.

5. Save, Export, and Report the Results

Never retype output from the Results window. Log first, then format.

log using "regression_log.txt", replace text
sysuse auto, clear
regress price weight mpg
log close

The text option keeps the file readable in a text editor. Without it you get a formatted file that can be hard to read outside Stata. Copy the whole table from a plain text log and it keeps its alignment when you paste it into Word or a text file, which is not true of the graphical Results window.

For a proper table in your paper, install the community command and point it at your filename.

ssc install outreg2
regress price weight mpg
outreg2 using results_table, replace b(3) se(3) stat(N, rmse, r2)

* Or, if you use estout
ssc install estout
esttab: regress price weight mpg, using esttab.rtf, replace

outreg2 produces a Word-ready table with decimal places set the way your department wants, standard errors in brackets and a row of model fit statistics underneath.

Here is the reporting checklist I run before any regression goes into a draft.

  1. How many observations, and is that fewer than my dataset?
  2. Is Prob > F below alpha, meaning the model beats nothing?
  3. How many of my predictors are individually significant?
  4. Does the R-squared match what this kind of outcome usually produces in my field?
  5. Do the rvfplot and avplots look like a random cloud?
  6. Is any VIF above 10?
  7. Did any variable I expected appear in the table?
  8. Do the signs match my theory, or am I explaining them away?
  9. Do I need clustered or heteroskedasticity-robust standard errors?
  10. Have I written the result as a sentence a stranger could check?

Common Mistakes

Reversing the variable order. If you type regress weight price, Stata will happily treat weight as your outcome and price as the predictor. It will not warn you. The dependent variable always goes first.

Treating correlation as causation. The -9.798 on weight is an association conditional on the other variables in the model. Heavier cars cost less partly because weight tracks other things. The output table cannot tell you which direction, or whether a third variable drives both.

Reading a p-value backwards. A p-value of 0.03 does not mean there is a 3% chance your hypothesis is wrong. It means that, if the true coefficient were zero, data this extreme would appear about 3% of the time. Also, a p-value of 0.000 is a display convention, not a certainty.

Fighting the constant. People see a significant _cons and panic, or they add noconstant hoping R-squared will look better. Both are mistakes. A significant intercept usually signals a misspecified model, and removing the constant makes the R-squared and residual mean meaningless.

Leaving categorical variables numeric. A variable coded 1, 2, 3, 4 for four education levels is not four units of education. Use encode plus i. or recode into dummies.

Ignoring missing data. Stata drops incomplete rows without comment. Run count if e(sample) after regress and compare it to count before the regression, so you know your real sample size.

Quoting R-squared as proof of quality. A high R-squared can sit next to a useless model, especially with too many predictors and a nearly constant outcome. Judge fit against what is normal in your field, and read Root MSE alongside it.

Skipping diagnostics. If the residual plot curves, the standard errors in that table are not the standard errors you should report. Ten seconds with rvfplot saves a reviewer finding it for you.

One more, and it is the most common question on r/stata and Statalist: if the outcome is a 0/1 variable, stop using regress. Use logit or probit for a binary outcome, tobit or truncreg for one censored at a bound, and xtreg with fixed effects for panel data, where you track the same units over time and need to control for unobserved differences between them.

Frequently Asked Questions

How do I run regression in Stata?

Type regress, then the dependent variable, then your explanatory variables, then press Enter. For example, regress price weight mpg fits price on weight and mpg using ordinary least squares. To add a categorical variable use i.varname, and to test an interaction use c.x1##c.x2. Stata prints the model fit block and the coefficient table immediately, and you can keep the results with estimates store for later comparison.

What does P|t| mean in Stata regression output?

P|t| is the two-tailed p-value for the null hypothesis that the coefficient on that row equals zero. Your decision rule is simple: if P|t| is below your alpha level, usually 0.05, you reject that null and call the predictor statistically significant. A printed value of 0.000 means below 0.0005, not literally zero. Significance says nothing about whether the effect is large enough to matter.

What is a good R-squared value in Stata regression?

There is no universal good value, because R-squared depends entirely on your field and your outcome. Cross-sectional economics and education models often sit between 0.2 and 0.5 and can be very useful. Human behaviour outcomes frequently land below 0.10. Anything near 0.99 on a behavioural outcome usually means too many predictors for too few observations, not a brilliant model. Compare against published work in the same area.

How do I read the _cons row in Stata output?

_cons is the intercept, the predicted value of the dependent variable when every predictor is set to zero. Whether that is meaningful depends entirely on your variables. A weight of zero pounds is not a real car, so a car-price intercept is arithmetic rather than insight. What matters more is that a significant intercept often signals a misspecified model, so mention it in your diagnostics rather than reporting it as a finding.

What is Root MSE and should I report it?

Root MSE is the square root of the mean squared residual, so it is the typical size of your model’s prediction error expressed in the units of the dependent variable. In the auto dataset example it is about 66 dollars, meaning the model usually misprices a car by roughly that much. Yes, report it. It tells a reader how precise the model is in practical terms, which R-squared never does.

Conclusion

Start with sysuse auto, clear, run describe and summarize to confirm your outcome is continuous, then run regress price weight mpg. Read the output in a fixed order: Number of obs, then F and Prob > F, then R-squared and Root MSE, then each coefficient row, then the intercept last.

After that comes the part most people skip. Run rvfplot, avplots and estat vif, then fix anything that looks wrong before you write a single sentence. If you can do that much, you know how to run regression in Stata and read the output well enough that the table stops being a wall of numbers.

Leave a Comment

Practical guides to statistics, surveys and research data

Read the latest guides