How to Interpret a P Value in Research Results (2026)

To interpret a p value in research results, read it as one piece of evidence about whether the data fits the null model, then check it against the confidence interval, the effect size and the sample size before drawing any conclusion. A p value below .05 tells you the observed result would be unusual if the null hypothesis were true. It tells you nothing about how big the effect is or whether anyone should act on it.

That distinction catches out a lot of readers, including people with a statistics degree. The number is easy to spot in a results table and easy to misread, because it looks like a verdict on the research question when it is really just a single output from one model.

This guide is written for someone holding a paper, a thesis, a report or a news article that reports a p value, and who needs to decide how much weight it carries. Nothing here needs a stats background beyond the null hypothesis. Where I quote a published result, the numbers are the published numbers.

Table of Contents
  1. 1What Does a P Value Mean?
  2. 2Where the null hypothesis comes from
  3. 3How to Interpret a P Value in Research Results Step by Step
  4. 41. Find the hypothesis that was tested
  5. 52. Identify the test and the tail
  6. 63. Locate the significance level the authors chose
  7. 74. Read the reported p value exactly
  8. 85. Cross-check the confidence interval
  9. 96. Find the effect size and the sample size
  10. 10P Value Thresholds and Common Decisions
  11. 11What a P Value Does Not Tell You
  12. 121. It is not the probability that the null hypothesis is true
  13. 132. It is not the probability that the result happened by chance
  14. 143. It is not a measure of effect size
  15. 154. It is not a measure of practical importance
  16. 165. It does not establish causation
  17. 176. It does not promise replication or data quality
  18. 18How to Read P Values With Confidence Intervals and Effect Sizes
  19. 19The confidence interval cross-check
  20. 20What a large p value is telling you
  21. 21How to Report a P Value in an APA Research Paper
  22. 22How to Interpret a P Value for Different Statistical Tests
  23. 23t-tests and one-way ANOVA
  24. 24Chi-square tests
  25. 25Correlation and regression
  26. 26Non-parametric tests
  27. 27The software output is not the interpretation
  28. 28Common P Value Mistakes and How to Avoid Them
  29. 29Threshold shopping
  30. 30Multiple comparisons
  31. 31P-hacking and unplanned peeking
  32. 32Comparing p values across studies
  33. 33Treating 0.05 as a law
  34. 34A Practical Checklist for Interpreting Research Results
  35. 35Frequently Asked Questions
  36. 36Is a p value of 0.05 statistically significant?
  37. 37What is the difference between a p value and a confidence interval?
  38. 38Does a statistically significant result mean the effect is important?
  39. 39What does p u0026lt; .001 mean in a research paper?
  40. 40Can a p value prove that one variable causes another?
  41. 41Why can the same research finding have a p value of .04 in one study and .06 in another?
  42. 42Conclusion

What Does a P Value Mean?

A p value is the probability, assuming the null hypothesis is true and the model behind the test is correct, of obtaining a test statistic at least as extreme as the one your data produced. Three details in that sentence carry all the weight, and most misreadings come from dropping one of them.

First, the probability is conditional. It describes the data given the null, not the null given the data. Those are different quantities and the order cannot be swapped: P(data | null) is not P(null | data).

Second, what gets the probability is the test statistic, not the hypothesis. A p value describes a number like a t, F, chi-square or correlation coefficient sitting in the tail of its sampling distribution. It is not a verdict on a theory.

Third, “at least as extreme” includes the observed result and everything beyond it. A p value of .03 does not mean your result had a 3% chance of occurring. It means that, under the null, results this large or larger show up about 3% of the time in repeated samples.

Two corollaries follow. A p value is never exactly zero, whatever your software prints, because the tail keeps shrinking as the statistic grows. And a p value below .05 does not mean the research hypothesis is 95% likely to be true. It means the null model fits the observed data poorly.

Where the null hypothesis comes from

The null hypothesis, written H0, is the default position that nothing is happening: no difference between groups, no relationship between variables, no effect of the treatment. The alternative hypothesis, Ha, is the claim you are testing for. The p value measures how far the data stray from the H0 model, so it always argues against the null and never directly for the alternative.

How to Interpret a P Value in Research Results Step by Step

When you open someone else’s results section, work through these six checks in order. Most of the time the p value itself is the least informative number in the section.

1. Find the hypothesis that was tested

Look for what the authors actually tested, not what the abstract implies. “Students who meditate scored higher on reported stress” is not a hypothesis; “meditation-group students reported lower stress than control-group students” is. If the tested claim is different from the headline claim, the p value does not support the headline.

2. Identify the test and the tail

Find the name of the test (independent-samples t-test, one-way ANOVA, chi-square test of independence, Pearson correlation, linear regression) and whether the hypothesis was one-tailed or two-tailed. A one-tailed test only counts results in the direction you predicted, so it produces a smaller p value from the same data. If the paper does not say which it used, treat the p value with suspicion.

3. Locate the significance level the authors chose

Almost every paper sets alpha before looking at the data, usually at .05, sometimes at .01 in medical trials or .10 in exploratory survey work. The decision rule is simple: p below alpha means reject the null, p above alpha means fail to reject it. If alpha appears nowhere in the paper, assume .05 but note the omission.

4. Read the reported p value exactly

Note three things: the value itself, whether it is written as a threshold (p < .001) or an exact number (p = .032), and whether it is one-tailed or two-tailed. An exact value is a more informative report than a threshold, so papers that give .032 are easier to judge than papers that only say p < .05.

5. Cross-check the confidence interval

At the .05 level, a p value just below the threshold corresponds to a 95% confidence interval that barely excludes the null value. This is the fastest sanity check available: if the interval for a mean difference runs from 0.1 to 3.9, it excludes zero; if it runs from -2.4 to 1.8, it does not, and the p value should not have crossed .05.

Cross-check the confidence interval

6. Find the effect size and the sample size

The effect size is the magnitude of the finding (Cohen’s d, partial eta squared, Cohen’s f, r squared, an odds ratio, or a raw difference in units). The sample size tells you how easily the test would detect a small real effect. A study with 40 participants per group can miss a meaningful difference; a study with 4,000 per group will flag a difference so small it changes nothing for a patient or a product. If the paper reports neither, that is a gap in the reporting, and it means the p value is carrying more weight than it should.

P Value Thresholds and Common Decisions

There is no official vocabulary for gradations of a p value. People on forums keep asking for one, and the honest answer is that terms like “highly significant” and “mildly significant” are informal labels. What you can do is map the number to the decision it drives and to a sentence you could defend in a meeting.

Reported p valueEvidence against H0Decision at alpha = .05Plain-English wording
p < .001Very strong against H0 under the stated modelReject H0Results this extreme would be rare if the null model were right; check the effect size before treating it as a big finding
p < .01Strong against H0Reject H0The data sit well out in the tail of the null distribution
p < .05Moderate against H0Reject H0The result is unusual enough under the null to count as evidence, though it is not strong evidence
p between .05 and .10Weak to moderate against H0Fail to reject H0Inconclusive. The study did not gather enough evidence either way, especially if it is underpowered
p > .10Little against H0Fail to reject H0Little evidence of a difference. This is not evidence that there is no difference

A p value of .049 and one of .051 are not meaningfully different in quality. Both sit on the threshold, both produce a confidence interval that barely excludes or barely includes the null value, and both should be described as borderline. Treating .049 as a finding and .051 as nothing is a decision about the rounding rule, not about the data.

The word to use for a result above the threshold is “nonsignificant”, not “insignificant”. The first describes the test outcome. The second implies the effect was trivial, which the test never established.

What a P Value Does Not Tell You

Six things, each one a common misreading that I see repeated in theses and in popular coverage of studies.

1. It is not the probability that the null hypothesis is true

A p value of .03 does not mean there is a 3% chance the null is correct, or a 3% chance the result was a fluke. It means that under the null, results this extreme happen about 3% of the time. The probability that the null is true depends on the prior odds of the hypothesis, which the p value never uses.

2. It is not the probability that the result happened by chance

“Chance” is doing no work in this sentence. The p value is computed from a specific reference distribution built by the test, and it describes position within that distribution. It is not a coin-flip probability and it does not mean the study was a gamble.

3. It is not a measure of effect size

p = .001 and p = .049 can sit on top of identical differences in the data. What separates them is sample size, variability, or both. If you only have the p value, you cannot say whether the effect is large or small, because the p value contains no information about magnitude.

4. It is not a measure of practical importance

Statistical significance answers a question about the model, not about your life, your patients or your product. A change of 0.2% on a conversion rate can clear .05 in a large experiment while being far too small to justify the engineering work. Practitioners need the effect estimate and the interval to judge that, not the p value.

5. It does not establish causation

Significance testing says the data are hard to explain under H0. Whether one variable caused another depends on the design: random assignment, controls, blinding, timing, and whether anything else could plausibly produce the same pattern. A significant observational p value with a well-chosen design is still an association.

6. It does not promise replication or data quality

A study can report p < .001 and still rest on a convenience sample, an uncleared question and outcomes chosen after the data were collected. The 2015 Open Science Collaboration project re-ran 100 published psychology studies and found that about 36% of the original significant results replicated, with effects roughly half the size reported. A small p value is a statement about one dataset under one model, and the replication literature is a standing reminder of that.

How to Read P Values With Confidence Intervals and Effect Sizes

This is the part that turns a number into a finding. A worked example is worth more than any definition, so here is one built on the standard drug-versus-placebo pain-relief trial that most textbooks use.

A trial compares pain scores before and after treatment in two groups. The treatment group drops by 4.2 points on a 0 to 10 scale, the placebo group by 3.4, so the observed difference between groups is 0.8 points. An independent-samples t-test returns t(198) = 2.11, p = .036, and the mean difference with a 95% confidence interval is 0.20 to 1.40.

Here is the same result written two ways, and only one of them survives a methods review.

Wrong: “The treatment was 3.6 times more effective than placebo (p = .036), so there is a 96.4% chance the treatment works.”

Right: “The treatment group improved by 0.8 points more than the placebo group, 95% CI [0.20, 1.40], t(198) = 2.11, p = .036, d = 0.30. Assuming there is no true difference between the treatments, a difference at least this large would arise in about 3.6% of samples of this size.”

Read that sentence in three moves. The observed difference is 0.8 points. If there were no real difference, differences of 0.8 points or bigger would show up in about 3.6% of comparable samples, which is unusual enough to count as evidence against H0. And the interval says the true difference could plausibly be as small as 0.20 points, which is exactly why nobody should promise patients a large benefit from this result.

Cohen’s d of 0.30 here is a small effect by convention, and it survives because the effect size came from the same data as the p value. Had the sample been 400 per group instead of 100, the same 0.8-point difference would have produced a much smaller p value with an identical effect size. Same finding, same magnitude, stronger statistical evidence. That is the sample-size sensitivity of p values in one sentence.

The confidence interval cross-check

At alpha = .05, a two-tailed p value just under .05 corresponds to a 95% interval that just excludes the null value, and a p just over .05 to an interval that just includes it. In the example above, p = .036 and the interval runs 0.20 to 1.40, entirely above zero, and the two agree. If a paper reports p = .03 with a 95% interval of -0.4 to 1.9, something is wrong with the reporting, and that mismatch is worth flagging.

Intervals also answer questions p values cannot. How wide is it? The width is set by the sample size, so a wide interval around a small effect is the usual profile of an underpowered study. Where is it centred? An interval far from zero with a small effect means a large study detected something small. Neither of those questions has any answer in the p value alone.

What a large p value is telling you

When p comes back above .05, most readers treat the study as a failure. There are three legitimate readings, and the confidence interval usually separates them.

  • The effect is near zero. The interval is narrow and sits right around the null value. Nothing to find here.
  • The effect is real but the study was too small to detect it. The interval is wide and still includes both a meaningful benefit and a meaningful harm. This is the case where more participants, not a re-analysis, would change the answer.
  • The measurement is too noisy. The interval is wide and the design allowed for plenty of confounds. The study failed to answer its question rather than answering it negatively.

A p value of .06 is the most common version of this problem, and it is why “my p value is 0.06, do I throw my hypothesis out” is asked so often. No. You have weak evidence, not disproof. Either the study gets bigger, the measure gets better, or you report it as inconclusive and move on.

How to Report a P Value in an APA Research Paper

Reporting is where interpretation goes wrong most often in a methods section, usually because authors default to p < .05 regardless of what the software printed. The APA Publication Manual is specific about this.

  • Italicise the p and use a space after it: p = .032, p < .001.
  • Drop the leading zero for probabilities and correlations: .032, not 0.032. Keep it for values that cannot exceed 1, such as p > .001.
  • Never report p = .000. Software rounds, and the p value is never zero. Write p < .001.
  • Give exact values to two or three decimal places rather than rounding everything to < .05.
  • Put the test statistic, its degrees of freedom, the effect size, the confidence interval and the p value in the same sentence.
  • Disclose a one-tailed test as one-tailed, and say whether the test was specified before data collection.
  • Use “nonsignificant” for a result above the threshold, and “fail to reject the null hypothesis” rather than “rejecting the null” when it was not rejected.

A model sentence for a two-sample test: t(118) = 2.34, p = .021, 95% CI [0.4, 1.6], d = 0.43. For a chi-square test: χ2(3, N = 240) = 11.7, p = .009, Cramér’s V = .22. For a correlation: r(98) = -.31, p = .002, 95% CI [-.48, -.12].

One more reporting habit worth building: say the decision, not just the number. “The difference was statistically significant at alpha = .05” is weaker than “at the pre-specified alpha of .05 the result met the significance threshold, with the effect size reported alongside it”. It costs a sentence and it tells the reader you had a rule before you looked.

How to Interpret a P Value for Different Statistical Tests

The reading sequence does not change across tests. What changes is the shape of the test statistic and the degrees of freedom attached to it, so here is how the same six checks translate.

t-tests and one-way ANOVA

These compare mean differences between groups. The p value comes from a t or F statistic compared to its sampling distribution. Because the F distribution is one-sided, an ANOVA p value is never two-tailed in practice. Report the test statistic with its degrees of freedom, and pair the p with the group means and their confidence intervals, because the p value alone hides which group differed.

Chi-square tests

For count data, the p value comes from a chi-square statistic. Interpret it exactly as before: how far the observed table sits from what the null model of independence would predict. Pair it with an effect size such as Cramer’s V or a difference in percentage points between cells, because “significant” on its own says nothing about which cells matter.

Correlation and regression

A correlation p value answers whether the linear association departs from zero. A regression p value answers whether a given predictor contributes beyond the others in the model. In both cases the coefficient and its interval do the interpreting: r = .18 with p < .001 and a 95% interval of .09 to .27 is a small but well-detected relationship, and reporting it without the coefficient would hide the most important fact in the study.

Non-parametric tests

When assumptions fail, the Mann-Whitney U, Wilcoxon signed-ranks and Kruskal-Wallis tests take over. Their p values carry the same conditional meaning against the null of no difference in distributions. The difference is that they detect differences in shape and rank, so the result should be described in those terms rather than as a difference in means.

The software output is not the interpretation

Wherever the p value came from, the row it sits in tells you what it means. In SPSS, look under Sig. (two-tailed) and check whether the test line is labelled independent samples, paired samples or one-way ANOVA. In R, p_value in a tidy output carries a confidence interval in the neighbouring column by default, which is the most convenient pairing available. In Stata, the significance stars (**, ***) sit next to p values, and a star only ever means below the threshold, never large effect. On dataanalysishelp.net we keep separate walkthroughs for reading each of those output formats, since that is where most people meet a p value for the first time.

Common P Value Mistakes and How to Avoid Them

Each row pairs a mistake with what to write instead. The wrong column is the version that shows up in most drafts.

Common mistakeWhat to write instead
“There is a 5% chance the null hypothesis is true”“If there were no true difference, a difference this large or larger would arise about 5% of the time in samples of this size”
“The result happened to be significant, p < .05”Report the exact value, the test statistic, its degrees of freedom, the effect size and the confidence interval
“p = .001 proves the hypothesis”“The data are difficult to reconcile with the null model at alpha = .05; the estimated effect and its interval are the basis for the substantive claim”
“The result was insignificant, so the effect does not exist”“The result was nonsignificant; the study did not gather enough evidence to rule out an effect of the size its interval allows”
“A lower p value means a stronger effect”“A lower p value means stronger evidence against H0 under this model; effect magnitude comes from the estimate and the interval”
“Because p = .04 the finding is established”“Because p = .04, a re-run study of similar size could land above the threshold; the estimate and interval carry the substantive claim”

Threshold shopping

Choosing alpha after seeing the p value turns a test into a fishing trip. Set the level before data collection, or state clearly that the analysis was exploratory and treat the result as a hypothesis for the next study, not a finding.

Multiple comparisons

Run 20 comparisons at alpha = .05 and you expect one spurious hit even when nothing is happening. Corrections such as Bonferroni, Holm, or a false discovery rate control that family-wise error rate, and any paper comparing many groups, many time points or many outcomes should say which one it used. A single significant subgroup out of twelve is usually noise until shown otherwise.

P-hacking and unplanned peeking

Dropping outliers after seeing the outcome, trying several outcome measures until one crosses .05, or stopping data collection the moment the result clears the threshold all inflate the false positive rate. Simmons, Nelson and Simonsohn coined the term false-positive psychology after demonstrating how many such published findings failed to replicate. Pre-registration, reporting every outcome tested and analysing the data as collected are the practical defences.

Comparing p values across studies

A p value of .01 in one study and .09 in another does not mean the first study found something real and the second found nothing. Different sample sizes, different measures and different hypotheses produce different p values from similar underlying effects. Compare the effect estimates and intervals instead, and ask whether the studies were powered for the effect you care about.

Treating 0.05 as a law

The threshold is a convention that traces back to Fisher’s work in the 1920s, and the American Statistical Association statement on p-values (Wasserstein and Lazar, 2016) is explicit that it is not a natural dividing line. It also says a p value does not measure the probability that a hypothesis is true, the size or importance of an effect, or the quality of a study. If your discipline uses .01, follow that convention; just know what the number does and does not license.

A Practical Checklist for Interpreting Research Results

Before you repeat a finding anywhere else, run it through these eight checks.

  1. What was the research question? Not the title, the question the design can actually answer.
  2. What was the null hypothesis? The no-effect or no-relationship claim the test was run against.
  3. Which test produced the p value? Name it, and state whether it was one-tailed or two-tailed.
  4. What alpha was set in advance? The decision rule is p against that pre-set level, not against the number you hoped for.
  5. What exactly does the p value license? Reject or fail to reject H0 under this model. Nothing about probability, importance or causation.
  6. What does the confidence interval say? Its width gives you the precision, its position gives you the magnitude.
  7. What is the effect size, in units that mean something? Points on a scale, percentage points, a standardised d, a variance share.
  8. What are the design and reporting limits? Sample size, assignment method, attrition, number of comparisons, pre-registration, stated limitations.

If you can answer those eight in two sentences, you have an interpretation. If you cannot get past item five, you have a number and nothing else.

Frequently Asked Questions

Is a p value of 0.05 statistically significant?

A p value of exactly .05 sits right on the boundary, because the conventional rule rejects the null when the p value is below alpha, not at or above it. In practice, .05 and .049 are treated the same way and both count as significant at the standard threshold. What the number licenses is only a rejection of the null model under the stated assumptions. Whether the effect is large, important or causal is a separate question answered by the effect size, the confidence interval and the study design.

What is the difference between a p value and a confidence interval?

A p value answers one yes-or-no question: how far does the data sit from the null model, relative to the noise? A confidence interval answers a range question: what values are compatible with the data at a stated level, and how wide is that range. At alpha = .05 the two are mathematically linked, since a 95% interval excludes the null value exactly when the two-tailed p value falls below .05. The interval carries far more information about magnitude and precision, so read it first.

Does a statistically significant result mean the effect is important?

No. Statistical significance measures how unlikely the data would look under the null model, which depends heavily on sample size. A very large study will produce small p values for effects too tiny to matter in practice, and a modest study can miss an effect that would matter a great deal. Deciding importance is a judgement about the effect estimate, the confidence interval, the measurement scale and the cost of acting or not acting. That is why good papers report the effect size and the interval next to the p value.

What does p u0026lt; .001 mean in a research paper?

It means the observed test statistic and anything more extreme would occur in fewer than one in a thousand samples if the null hypothesis were true. It is strong evidence against the null under that model, and it is reported as a threshold rather than an exact value because the true p is too small to print usefully. It still says nothing about the size of the effect, the importance of the effect, or whether the study design can support a causal claim.

Can a p value prove that one variable causes another?

No. Significance testing only asks whether the data are hard to explain under the null model of no effect. Establishing causation requires a design that rules out alternatives: random assignment, an appropriate control group, blinding, sensible timing, and attention to confounding. A significant p value from an observational study remains an association, however small the number is. Reading a p value as proof of causation is the single most common overreach in popular summaries of studies.

Why can the same research finding have a p value of .04 in one study and .06 in another?

Because a p value reflects the effect, the sample size and the variability in that particular dataset. Two studies of the same underlying relationship can land on opposite sides of .05 purely through sampling noise, especially with modest sample sizes. That is why the confidence interval matters more: it shows how precise each estimate is and whether both studies are compatible with the same underlying effect. Treat a borderline p value as inconclusive rather than as a pass or a fail.

Conclusion

Start with the confidence interval, then read the p value against the alpha the authors set in advance, then look at the effect size in units that mean something to you. That order keeps the number in its proper place: one piece of evidence about how badly the data fit the null model, and nothing more.

If a paper gives you only the p value, ask for the interval and the effect estimate. A study that reports all three has done the work of letting you judge magnitude, and how to interpret a p value in research results stops being a guessing game. Updated for 2026.

Leave a Comment

Practical guides to statistics, surveys and research data

Read the latest guides