10 Common Statistics Mistakes in Student Dissertations 2026

The common statistics mistakes in student dissertations fall into two groups: analytical choices that quietly invalidate your results, and reporting habits that stop you defending them. Most students hit at least one — a t-test on data that needed a non-parametric alternative, a p-value quoted with no effect size, a correlation written up as a cause. The good news is that every one of them is fixable before submission, and most are caught in an afternoon of checking. This guide walks through the ten that come up most often, with the reason each one breaks your analysis and the specific change that repairs it.

I have watched the same handful of errors sink otherwise strong drafts. A candidate with a beautiful literature review loses marks because their analysis chapter says only that “data were analysed using SPSS.” Another reports r = .42, p = .031, and concludes the intervention caused the improvement — in a design that could never have shown cause. None of that is stupidity. It is what happens when statistical training stops at clicking buttons.

So here is the audit list, in the order you would actually work through it: clean the data, check the design, check the assumptions, choose the test, run it, then report it honestly. Each section below names the mistake, shows what it looks like in a real results chapter, and gives you the fix.

Table of Contents
  1. 1Common Statistics Mistakes in Student Dissertations at a Glance
  2. 21. Choosing the Wrong Statistical Test
  3. 32. Ignoring the Research Design When Choosing a Test
  4. 43. Failing to Check Statistical Assumptions
  5. 54. Treating a Non-Significant Result as Proof of No Effect
  6. 65. Reporting Only the P-Value
  7. 76. Interpreting Correlation as Causation
  8. 87. Using an Overly Large Sample Without a Clear Plan
  9. 98. Excluding Data Without Documenting the Decision
  10. 109. Misinterpreting Outliers and Influential Cases
  11. 1110. Overclaiming the Results in the Discussion
  12. 12Frequently Asked Questions
  13. 13Can I reject the null hypothesis based on a p-value?
  14. 14What are five common misuses of statistics?
  15. 15What is the difference between a one-tailed and a two-tailed p-value?
  16. 16How do I know if my data is normal?
  17. 17Should I remove outliers from my dissertation data?
  18. 18Can I use ChatGPT for my thesis?
  19. 19Start with a Statistical Review

Common Statistics Mistakes in Student Dissertations at a Glance

Here is the short version. Read the table, then go to the sections that match the analysis you actually ran.

MistakeWhy it weakens the resultsQuick fix
Choosing the wrong statistical testThe test answers a different question to the one you askedMatch the test to the data shape and the design, not to habit
Ignoring the research design when choosing a testIndependent and paired data are analysed as if they were the same thingDecide from your collection procedure whether cases are matched
Not checking assumptionsNormality, homogeneity and independence violations bias the p-valueRun a Shapiro-Wilk test, Levene’s test and a residual plot
Reading a non-significant result as no effectA study with no power will always look like thisReport confidence intervals and power alongside the p-value
Reporting only the p-valueSignificance says nothing about size or importanceAdd an effect size and a confidence interval to every test
Correlation written up as causationConfounding variables are never addressedUse associative language and adjust for known confounders
Large unplanned sample with many subgroupsInflated Type I error makes findings unreliablePre-specify tests and apply a multiple comparison correction
Undocumented data exclusionsLooks like cherry-picking to an examinerKeep a cleaning log and report every exclusion with a reason
Mishandling outliers and influential casesOne extreme case can drive the whole resultScreen with boxplots and Cook’s distance, then run a sensitivity check
Overclaiming in the discussionConclusions outrun the design that produced themMatch every claim to the evidence and name the limitations

1. Choosing the Wrong Statistical Test

The most common mistake in student dissertations is picking a test before deciding what the research question actually asks. Students default to the independent-samples t-test or one-way ANOVA because those are the two they were shown in a lecture, then force ordinal satisfaction scores, pre/post measurements and skewed income data through them.

The fix is to start from the shape of your data, not from the name of the test. This table covers most undergraduate and master’s projects.

Your data and questionRecommended testKey assumption
One continuous outcome, two independent groupsIndependent-samples t-testNormality within groups, homogeneity of variance
One continuous outcome, two matched measurementsPaired-samples t-testDifferences are approximately normal
One continuous outcome, three or more groupsOne-way ANOVANormality, homogeneity, independence
Two categorical variablesChi-square test of independenceExpected cell counts at least 5
Two continuous variables, linear relationshipPearson correlationLinearity, absence of influential outliers
Two continuous or ordinal variables, monotonic relationshipSpearman rank correlationMonotonic relationship, ordinal data treated as ranks
Two groups, ordinal outcome or non-normal continuous outcomeMann-Whitney U testIndependent observations
Three or more groups, ordinal or non-normal outcomeKruskal-Wallis testIndependent observations
Continuous outcome predicted by several predictorsMultiple or logistic regressionLinearity, no severe multicollinearity

A worked example. You ask whether final-year students who attended a study skills workshop scored higher on a 1 to 5 self-rated confidence scale. The response is ordinal, the workshop group has 34 students and the comparison group has 41, and the confidence distribution is visibly skewed. An independent-samples t-test is the wrong tool twice over, and a Mann-Whitney U test is the defensible choice. Report it as a rank comparison, not as a comparison of means.

One practical safeguard: write your research question as a sentence ending in a word that tells you the test. “Is there a difference between” points to a t-test or ANOVA. “Is there an association between” points to correlation or chi-square. “What predicts” points to regression. Students who do this catch mismatches before they touch the software.

2. Ignoring the Research Design When Choosing a Test

The same test name means different things depending on how the data were collected, and this is where many dissertations go wrong silently. An independent-samples t-test and a paired-samples t-test share a label but answer different questions entirely.

Use the paired test when the same participants are measured twice, such as a pre-test and post-test design, or when each member of one group is matched to a member of the other by age, location or baseline score. Use the independent test when the two groups are made up of different people who never shared a measurement occasion. Running a paired test on independent data inflates the apparent relationship between the two scores, because you are comparing two unrelated sets of numbers as if each row were a pair.

Repeated measures with three or more occasions, such as weekly satisfaction ratings across a twelve-week intervention, need a repeated-measures ANOVA rather than three separate t-tests. And clustered data — students nested within classes, or patients nested within clinics — need a multilevel model or at least a clustered standard error. Treating 400 student answers as 400 independent observations when they came from 20 classes is a unit-of-analysis error, and it makes your confidence intervals far too narrow.

Write one sentence in your methodology chapter naming who was measured, how often, and whether the observations are independent. That sentence is what an examiner will use to judge whether your test was appropriate.

3. Failing to Check Statistical Assumptions

Parametric tests such as the t-test and one-way ANOVA carry assumptions about the shape and spread of the data. Skip the checks and the software will still hand you a confident-looking p-value built on a broken foundation. The four assumptions worth checking are below.

AssumptionHow to check itWhat to do when it fails
Normality of the outcome or of the residualsShapiro-Wilk test, plus a Q-Q plot and a histogramUse Mann-Whitney U or Kruskal-Wallis, or report that the t-test is robust at your sample size
Homogeneity of variance across groupsLevene’s test or a boxplot of the groupsSwitch to Welch’s t-test, which does not assume equal variances
Independence of observationsCheck the collection design and the design effectUse a paired test or a clustered model
Linearity for regression and correlationScatterplot and a plot of standardised residuals against fitted valuesTransform the variable or model the relationship non-linearly

Do not treat the Shapiro-Wilk output as a pass or fail switch on its own. With large samples the test flags trivial departures as significant, and with small samples it can miss real skew. Plot the data, look at the shape, then write one sentence justifying your choice. For example: “The Shapiro-Wilk test indicated that attendance scores deviated from normality (W = 0.91, p = .003), so a Mann-Whitney U test was used.”

For regression and correlation, the residual plot matters more than a formal test. If your residuals are not scattered randomly around zero with no funnel shape, the standard errors and p-values will not be trustworthy.

Also check the measurement scale. Likert items with five or seven points are ordinal. Treating them as interval in a Pearson correlation is a defensible convenience that many examiners tolerate, but a mean computed across five-point items should be labelled as a composite score, and Cronbach’s alpha should be reported for any multi-item scale.

4. Treating a Non-Significant Result as Proof of No Effect

A p-value above .05 does not tell you that two groups are the same. It tells you that this study did not detect a difference large enough to be unlikely under the null. With a small sample, a genuinely important difference can produce p = .18 simply because your design could never have found it.

Say it back to yourself with the numbers substituted: with 12 students per group, a difference that matters in practice will go undetected most of the time. That is low statistical power, and it is a design-stage problem rather than an analysis-stage one. Confidence intervals make the same point better than any sentence, because they show you the range of values still compatible with your data. A wide interval that spans a meaningful effect in both directions is an honest result worth reporting.

Write it as an inconclusive finding, not a null one: “The difference in mean satisfaction between groups was 0.4 points, 95% CI [−0.3, 1.1], p = .26. The interval includes effects of practical interest in both directions, so this study cannot distinguish between no difference and a moderate difference.” That is a valid contribution, and examiners respect it far more than a p-value dressed up as a conclusion.

5. Reporting Only the P-Value

A p-value on its own tells an examiner almost nothing. You can have p = .049 with an effect so small it would never matter in practice, and p = .051 with the largest effect in your study. This is the single most frequently criticised gap in student reporting, and it is also the easiest to fix.

Every inferential statistic you report needs three companions: an effect size, a confidence interval, and the exact p-value with its test statistic and degrees of freedom. APA 7th edition requires this format, and it also protects you at viva.

Compare these two write-ups of the same analysis. The weak version: “A t-test showed that the intervention group performed significantly better than the control group (p = .03).” The strong version: “The intervention group (M = 18.4, SD = 3.1) scored higher than the control group (M = 16.9, SD = 3.4), Welch’s t(64) = 2.21, p = .031, mean difference = 1.5 points, 95% CI [0.2, 2.8], d = 0.45.” The second sentence can be checked, questioned and understood by any examiner in the room.

Choose the effect size that matches your test: Cohen’s d for a t-test, partial eta squared or Cohen’s f for ANOVA, the odds ratio for logistic regression or chi-square, and r for a correlation. Report it even when the result is not significant, since that is exactly when its size matters most.

6. Interpreting Correlation as Causation

A significant correlation between two variables tells you the two move together in your sample. It does not tell you that one produces the other, because a third variable may drive both, and because the direction of influence may run the other way.

Suppose you find that students who use the library more often have higher grades, r = .34, p = .002. The conclusion you cannot draw is that library use improves grades. High-achieving students may simply be more likely to attend the library, and they may also attend more classes, submit earlier and study more consistently. That third variable — prior attainment or study habits — is a confounder, and in a cross-sectional survey design there is no statistical adjustment that can fully resolve it.

Adjusting for known confounders helps. Put prior grades, year of study and a measure of study hours into a regression alongside library use and see whether the coefficient survives. If it drops from .34 to .09, the original association was mostly confounding, and your honest write-up says so.

The language discipline is simple. Write “was associated with”, “higher library use corresponded with”, “explained 12% of the variance in grades after adjusting for prior attainment”. Avoid “led to”, “caused”, “improved” and “impacted on” unless your design is genuinely experimental, with random assignment and a control condition.

7. Using an Overly Large Sample Without a Clear Plan

More respondents does not automatically mean better statistics, and the way extra data goes wrong is predictable. Every extra subgroup you examine is another test, another chance at a false positive, and your threshold of .05 stops meaning what you think it means.

Run twenty comparisons at p < .05 and you can expect roughly one false positive even when nothing is there. Pre-specify which tests you will run, group the exploratory ones under a clearly labelled exploratory analysis, and apply a multiple comparison correction such as Bonferroni or Tukey where you legitimately run several pairwise comparisons after an omnibus ANOVA. Adjusting twenty tests at .05 gives a per-test threshold of .0025, and you report that you did it.

Large samples also make trivial effects reliably significant, which is why section 5 matters more as n grows. And a large convenience sample can still be unrepresentative: 800 responses from one university’s engineering faculty tells you about that faculty. State your recruitment route and any known demographic skew, and treat representativeness as separate from size.

Sample size should be justified by a power analysis based on the effect you expect, the alpha you set and the power you want, usually .80. Running it after data collection tells you what you could have had. It is still worth reporting, because it tells the reader what your null result can and cannot mean.

8. Excluding Data Without Documenting the Decision

Data cleaning is where good intentions produce bad practice. Students remove a few awkward cases, the results tidy up, and nothing about that sequence appears in the methodology chapter. An examiner reading it sees undeclared exclusions, and the credibility of the whole results chapter drops.

Exclusions are entirely legitimate when they follow a rule you decided in advance and applied consistently. Legitimate grounds include duplicate responses from the same participant completing the survey twice, straight-lined or implausibly fast questionnaire responses, logically impossible values such as an age of 250, and cases missing the outcome variable entirely. Inadmissible grounds include anything that happens to weaken a result.

Keep a cleaning log alongside your data: raw count, reason for each exclusion, number removed, and the count remaining. On a dissertation with 200 responses, a line such as “10 duplicate entries were removed, 15 incomplete questionnaires were excluded and 6 cases contained impossible values on the age item, leaving 169 valid responses” tells the whole story in one sentence and can be lifted straight into the methodology chapter.

For missing data, do not silently delete incomplete cases or replace them with the mean. That last habit, mean replacement, artificially shrinks the standard deviation and can manufacture significance. Report the amount of missingness per variable, state your handling strategy — complete cases, multiple imputation, or a justified single imputation — and show whether the choice changed the result.

Software silently corrupts data too. A respondent identifier stored as a factor rather than a number, or survey values imported from a spreadsheet as text, will break every downstream calculation without raising an error. Check your variable types before you run anything.

9. Misinterpreting Outliers and Influential Cases

An outlier is an observation far from the rest of the data. An influential case is one whose removal changes your result substantially. The two are not the same, and conflating them leads to bad deletions.

Use a boxplot or a standardised residual plot to screen for candidates, then decide what each one is. A genuine extreme value — a hospital with ten times the turnover of the sample, a respondent reporting 40 years of experience — is real information and stays. A data entry error, where 250 should read 25, goes, with the rule written down. A legitimate case that happens to be extreme is the case your analysis most needs to explain.

In regression, check leverage and Cook’s distance to find influential cases, not just large residuals. Run the model with and without each candidate and compare the coefficients. If removing one case swings a coefficient from 0.05 to 0.42, that is a fragility you must report, not hide.

APA 7 requires you to say what you did with outliers, so decide deliberately and write it down: “Six cases with standardised residuals above 3 were examined. Three proved to be data entry errors and were corrected at source. The remaining three were retained as legitimate extreme values.” Then show the sensitivity check in a footnote if the conclusion depends on them.

10. Overclaiming the Results in the Discussion

The discussion chapter is where valid analyses get talked into invalid claims. A cross-sectional association becomes a recommendation for policy, a small convenience sample becomes “students today”, and a single institution becomes a whole population.

Every claim in the discussion needs to trace back to something in the results. Before you write a paragraph, check three things: does the design support causal language, does the sample support the population I have named, and does the confidence interval support the size of the effect I have claimed? If any answer is no, rewrite the sentence at the level the evidence supports.

Name your limitations plainly. Single-site sampling, self-reported measures, a cross-sectional design, response rates below your target, low power for subgroup comparisons — these belong in the limitations paragraph, written as boundaries on your claims rather than as excuses. Examiners read a candidate who knows exactly what their study cannot show as someone who understands the work.

Separate findings from speculation clearly. “Students who attended the workshop reported higher confidence” is a finding. “The workshop may have improved engagement across the institution” is speculation, and it should be labelled as an implication for future research rather than presented as a result.

Two peer-reviewed lists make useful reading before you start that chapter: van Smeden and colleagues’ A Very Short List of Common Pitfalls in Research Design, Data Analysis and Reporting, and the Top 10 Statistical Pitfalls reviewer’s guide by Green, Smith and Whittle. Good’s Common Errors in Statistics is the longer reference if you want the full taxonomy.

Frequently Asked Questions

Can I reject the null hypothesis based on a p-value?

No. You either fail to reject the null hypothesis, or you conclude that your data are incompatible with it. A p-value is the probability of data this extreme or more, assuming the null is true. It is not the probability that the null is true. So p less than .05 means p less than .05, not that your effect is large or important.

What are five common misuses of statistics?

The five that come up most often are: treating p less than .05 as proof of a meaningful effect, reporting a correlation as a causal relationship, running a parametric test without checking its assumptions, reporting no effect size or confidence interval alongside the p-value, and drawing a conclusion from a study that was underpowered. Each is a reporting or design fault rather than a calculation error.

What is the difference between a one-tailed and a two-tailed p-value?

A two-tailed test splits your alpha of .05 across both tails, so each critical region is .025, and it detects an effect in either direction. A one-tailed test puts the whole .05 in a single tail, which produces smaller p-values but only detects effects in that one direction. Choosing one-tailed after seeing your results is HARKing and is a statistical mistake in itself.

How do I know if my data is normal?

Run a Shapiro-Wilk test and look at the plot rather than the p-value alone. Generate a Q-Q plot and a histogram: points close to a straight line and a roughly symmetric shape support normality. With large samples the formal test flags trivial departures, so describe what you see in the data and justify your choice in one sentence.

Should I remove outliers from my dissertation data?

Only if you can name a defensible rule you applied to every case, not the cases that trouble you. Data entry errors should be corrected at source. Legitimate extreme values are part of your dataset and belong in the analysis. Report what you screened, what you removed and why, and run a sensitivity check so the reader can see whether the conclusion depended on those cases.

Can I use ChatGPT for my thesis?

Check your institution’s policy, because the rules differ widely and change. The safe rule is to use AI tools for explanation and learning, never to submit analysis you cannot reproduce line by line. The specific statistical risk is fabricated output: invented p-values, plausible but nonexistent references and test results that were never actually run, which you will be asked to defend at viva.

Start with a Statistical Review

If you fix only one thing, align your test with your research question and write that alignment into a sentence in the methodology chapter. It is the decision examiners probe most, and the one that everything downstream depends on.

Then work through these in order before submission. Check the assumptions for the test you ran: Shapiro-Wilk for normality, Levene’s test for homogeneity, a residual plot for linearity and independence. Document every exclusion with its reason, and keep your cleaning log. Report an effect size, a confidence interval and the exact p-value for every inferential test, including the ones that were not significant.

Finally, read your discussion chapter against your results chapter sentence by sentence. If a claim cannot be traced to a specific result, soften it or cut it. Honest framing of a null or inconclusive finding builds more confidence than a strong conclusion built on thin analysis.

If your supervisor or a statistics support service disagrees with one of the choices above, ask them why rather than what to do. Being able to articulate the reasoning behind each decision is what carries you through the viva, and it is the same skill that catches these ten mistakes in your next project.

Leave a Comment

Practical guides to statistics, surveys and research data

Read the latest guides