To avoid p hacking in your analysis, fix the analysis before you look at the outcome. Write the primary outcome, the comparison, the test, the covariates, the exclusion rules and the stopping rule into one dated document, put it somewhere public, then run that analysis once and report it whether or not it comes back significant.
Most p-hacking is not a deliberate act. It is the slow drift that happens when a result comes back at p = 0.07 and you start trying things until the number looks better. Every reasonable-sounding tweak after that point is a new test, and the one you report is the smallest of all of them.
A p-value, in plain words: assuming the null hypothesis is true, a p-value is the probability of a result at least as extreme as the one you observed. A p of 0.05 does not mean there is a 5% chance the finding is real or fake. It means the threshold you chose will let through about one false positive in every twenty clean tests.
Here is the chain that makes p-hacking so easy to fall into. At p < 0.05 with a single pre-planned test, your false positive rate is 5%. Run twenty independent tests on the same data and report only the smallest one, and the chance that at least one crosses 0.05 by pure noise is about 64% — and the expected number of “significant” results across those twenty is one. You have manufactured a finding out of arithmetic alone.
Table of Contents
- 1What You Need
- 2Step-by-Step
- 3Define the analysis before looking at results
- 4How to avoid p hacking in your analysis
- 5Prespecify multiple tests and corrections
- 6Keep a transparent decision and change log
- 7Check stability instead of chasing significance
- 8Save and report the full analysis trail
- 9Common Mistakes
- 10Frequently Asked Questions
- 11What is the difference between HARKing and p-hacking?
- 12What does data dredging mean?
- 13What is a simple explanation of a p-value?
- 14Do I need to correct for multiple comparisons?
- 15Is trying an alternative model always p-hacking?
- 16How do I know if I already p-hacked my analysis?
- 17Conclusion
What You Need

Everything below happens before the first model runs. If you cannot produce these items, you are not ready to analyse, and that is a good thing to know on week one rather than month six.
- A single research question. One question, not a theme. “Does the new onboarding flow reduce 30-day churn?” is answerable. “What affects retention?” is a fishing licence.
- Hypotheses written out. Direction and rough expected effect size, in a sentence, before any data is opened. An estimate is needed anyway for a power calculation.
- A variable dictionary. Name, type, coding, units, and allowed values for every variable. Ambiguous coding is the most common accidental source of forking paths.
- A written analysis plan. Primary outcome, predictor or comparison, statistical test, covariates, exclusion rules, missing-data handling, stopping rule. This is the whole ballgame.
- A frozen dataset. One cleaned copy, saved read-only, that every planned analysis runs against. Keep the raw file untouched in a separate folder.
- A reproducible script or project. An R or Python script, a do-file in SPSS or Stata, or at minimum a dated command log. Analysis that cannot be re-run cannot be audited, by you or by a reviewer.
- A decision log. A running table with a row for every analytical choice, planned or not, dated.
- A registry, if your field allows one. OSF, AsPredicted, or a laboratory protocol page. The timestamp is the point: it proves the plan existed before the results.
That last item is optional in most courses and mandatory in clinical trials. Where you cannot preregister, the next best thing is to email the plan to your supervisor or yourself with a date, and to keep it attached to the project.
Step-by-Step
Define the analysis before looking at results

Turn a broad question into a fixed specification. Every field below gets answered in writing, in this order, before data is touched.
- Primary outcome. One variable, one time window, one definition. “Churn within 30 days of signup”, not “engagement”.
- Comparison. Against what? Control group, baseline period, or the same group under a different condition.
- Test. Named and fixed: independent samples t-test, one-way ANOVA with Dunnett post-hoc, logistic regression, and so on. Include the model specification, not just the test name.
- Covariates. Decided now. Baseline value of the outcome is the usual one, and it is the one people most often add later.
- Exclusion rules. Written as conditions, not judgments: “participants who did not complete the survey past item 14”, “duplicate device IDs”, “implausible completion times under 90 seconds”.
- Missing data. Complete-case, mean imputation, or multiple imputation — chosen in advance, with the variables listed.
- Stopping rule. Fixed sample size N, or a sequential design with a pre-specified boundary. Not “until significance appears”.
- Sample size. Set by power analysis for the smallest effect you would still care about, not by what you can afford or what is already in the file.
A worked example. The question is “Does the guided setup reduce 30-day churn?” The plan becomes: outcome is binary churn within 30 days; predictor is assignment to guided setup versus control, randomised at signup; test is logistic regression with baseline churn as the only covariate; exclusions are duplicate device IDs and internal test accounts, both applied blind to outcome; missing data is multiple imputation on three baseline items; sample size is 1,200 per arm from a power calculation for a small odds ratio; stopping rule is complete N, no interim looks. A preregistered hypothesis paper then runs exactly this and reports whatever comes out. It is unexciting and it is the reason the result is worth anything.
How to avoid p hacking in your analysis
Exploration is not the enemy; hidden exploration is. The fix is to label every analysis as confirmatory or exploratory before you run it, and to let the label decide what you may do with it afterwards.
Confirmatory analyses are the ones registered in your plan. They carry the weight of the paper, and they get reported no matter what the p-value does. Exploratory analyses are hypothesis-generating: they can be run freely, but they are labelled as exploratory in the paper and never presented as tests of the original hypothesis.
Split them in practice, not just in the write-up. Keep the frozen dataset for confirmatory work and a copy for exploration, and log which script used which. In R you can enforce part of this with a script that refuses to run unless a PLAN_ID string matches the registered file; in SPSS a saved do-file with a read-only output folder does something similar.
The subgroup case shows why labelling matters. Suppose you have four customer segments and two outcome measures, and you run all eight combinations. Under a true null, the chance at least one of those eight tests returns p < 0.05 is roughly 34%. Report the one that did and you have a headline finding; report all eight, corrected, and the effect usually looks like what it is — noise. The move that prevents this is running the eight tests with a correction in the plan, or naming the eight as exploratory and holding them for the discussion section.
There is a version of this that catches experienced researchers out. Someone eyeballs an EEG topographic map, notices a few channels that look promising, and then runs statistics on only those channels. Nothing was formally specified, and the selection was driven by the outcome. That is data dredging, and the fix is channel selection from priors or an independent pilot sample.
Prespecify multiple tests and corrections
Correct for multiple comparisons whenever you run more than one test on the same outcome, and decide the method in the plan. The table below covers the methods you will actually meet in SPSS, Stata, R and Python.
| Method | What it controls | When to use it |
|---|---|---|
| Bonferroni | Family-wise error rate: chance of one or more false positives across the family | Few tests, fixed in advance, and you need a simple defensible threshold. Threshold becomes alpha divided by k. |
| Holm-Bonferroni | Same family-wise control, but uniformly more powerful | The default when the number of tests is modest and you can defend strict error control. Usually the right answer for pre-planned post-hoc tests. |
| Benjamini-Hochberg FDR | Expected proportion of false discoveries among the results you declare significant | Genome-scale screens, feature selection across dozens of predictors, anything exploratory where you accept some false positives to catch more true ones. |
| Benjamini-Yekutieli | FDR under arbitrary dependence between tests | As BH, when the tests are correlated in unknown ways and you want the conservative guarantee. |
| Sidak | Family-wise error rate, more powerful than Bonferroni for k above about 10 | Independent tests, moderately large families, when Bonferroni is too blunt. |
| Dunnett | Family-wise error against a single control | Several treatment arms compared with one control group, including planned A/B tests with more than two arms. |
Report the corrected result as the result. Many analyses keep an uncorrected p in the text as a secondary note, which is fine as long as the corrected one is the one the conclusion rests on.
Keep a transparent decision and change log
A decision log is a spreadsheet with one row per analytical choice, dated. It costs ten minutes a week and it is the evidence that separates legitimate iteration from p-hacking.
Five fields are enough: date, item, planned or unplanned, reason, and the outcome in terms of what changed in the script or data. Rows look like this.
- 2026-02-11 | covariate set | unplanned | reviewer asked about baseline churn | added baseline churn to the model; primary result moved from p = 0.06 to p = 0.11 |
- 2026-02-11 | primary outcome | planned | as registered | churn within 30 days, no change |
- 2026-02-14 | exclusion rule | unplanned | three duplicate device IDs found during cleaning | rule added: one record per device ID; applied blind to the churn value |
The log does not have to be tidy, but it has to be contemporaneous. A log written after you know the result is a reconstruction, and reviewers have learned to discount it. Log unplanned changes especially; the ones you would defend as reasonable are the ones that need the reason written down.
Check stability instead of chasing significance
Once you have a result, the question is not whether p is below 0.05. It is whether the conclusion survives reasonable variations in the analysis.
That means sensitivity analyses: refit the model with and without the covariates, with and without the borderline cases, with a different but defensible outcome definition, with the alternative test a reviewer would plausibly ask for. It means confidence intervals and effect sizes alongside every p-value, because a narrow interval around a trivial effect is more informative than a bare p of 0.001. It means checking assumptions — linearity, homoscedasticity, independence, residual normality — and reporting violations rather than quietly transforming until they vanish.
There is a line here worth stating plainly. Data peeking, looking at the data and stopping when the number looks good, is invalid; it inflates the error rate and you should not do it. Repeating an analysis because you gathered genuinely new information is a different act. If the second run has a revised expected effect size drawn from independent evidence, a newly recruited sample, or an added control arm, the error rate can be handled properly. On the site, research and application threads about early stopping keep coming back to exactly this distinction, and the practical test is simple: did the new information exist because of the result you saw, or independently of it?
If you genuinely need to look at accumulating data, plan for it. Sequential designs and alpha-spending boundaries, group-sequential trial designs, and anytime-valid confidence sequences all give you legitimate interim looks. They require setting the boundary in advance, which is the entire point.
Save and report the full analysis trail
Report what you did in enough detail that someone else could repeat it, and separate planned from exploratory in the same document.
Save the cleaning script, the analysis script, software and package versions, the frozen dataset or a hash of it, and the raw output files. In your results section, give a table with one row per primary test: outcome, comparison, test, pre-planned or exploratory, effect size, confidence interval, corrected p-value. Then put the exploratory analyses in a clearly marked subsection with a sentence saying they were not pre-specified and are hypothesis-generating.
Two tools are worth knowing about here. p-curve analysis looks at the distribution of p-values across a literature or a set of studies; a right-skewed distribution is a signature of genuine effects, while a left-leaning mass near small p-values is what p-hacking tends to produce. Specification-curve analysis re-runs the same question under every defensible analytic choice and plots the full distribution of results, which is the most convincing single figure you can produce for a contested finding. Neither replaces an honest plan; both make fishing visible.
Common Mistakes
These are the patterns that show up most often, with the fix that actually prevents them.
- Trying many outcomes until one is significant. Fix: one primary outcome named in the plan; everything else exploratory.
- Changing the test after seeing the result. Switching from an independent t-test to a factorial F-test, or to a matched-pairs analysis, because the first one was p = 0.09. Fix: name the test in the plan, and if a matched-pairs design is the correct one, justify it from the design, not from the data.
- Dropping cases after the fact. Removing outliers that happened to break the result. Fix: pre-specified exclusion rules that do not reference the outcome, applied blind, with the count of excluded cases reported.
- Treating exploratory findings as confirmatory. Fix: label them. A hypothesis accepted after the results were seen is HARKing, and saying so costs you nothing.
- Reporting only the significant results. Selective reporting and preferential rounding — reporting 0.049 but not the 0.051 sitting in the same table. Fix: a results table with every test you ran.
- Optional stopping. Fix: fixed N, or a pre-specified sequential boundary.
- Covariate shopping. Adding age, gender and tenure until the model clears the threshold. Fix: covariates set by design or by a pre-registration, with a stated rationale.
Three terms get mixed up constantly, and keeping them apart makes your own reasoning much clearer.
| Practice | What it is | Tell |
|---|---|---|
| p-hacking | Running many analyses and reporting whichever clears the threshold | The analysis count is higher than the plan, and the reported one is the smallest p |
| HARKing | Hypothesising After Results are Known — presenting a post-hoc explanation as an a priori hypothesis | The wording claims confirmation the design cannot support |
| Data dredging | Searching a dataset for any pattern, with no prior hypothesis | Variables were picked after seeing their relationship to the outcome |
| Cherry-picking | Selective reporting of cases, time points or conditions | The excluded cases and dropped time points are not explained |
A short self-audit. Before you submit, run these questions against your own analysis: did I run more tests than I planned? Did any model change after I saw an output? Did I drop observations without a rule written in advance? Did I choose covariates after looking at the data? Can I point to a timestamped document that predates the first analysis? Can someone with my code and data reproduce every number in the paper? Any unaccounted-for yes is a finding you should fix or disclose, not a detail to hope nobody checks.
If the honest answer is that you have already p-hacked, the recovery path is unglamorous and it works. Re-label the analysis as exploratory, report every test you ran including the null ones, apply a correction for the number of tests actually performed, and treat the result as a hypothesis for a fresh confirmatory study or a direct replication on new data. You can also run the detection tools above on your own literature and report the picture honestly. Analysts who publish itemised decision logs are consistently treated as more credible, not less, when the record is complete.
One last thing, and it is the real root of the problem: incentives. Publication pressure and supervisor pressure are what push people toward re-running analyses until something appears, and no checklist fully solves that. Analysing in public, sharing scripts with collaborators early, and being the person who raises the concern in the group meeting all change the local incentives more than any single statistical safeguard.
Frequently Asked Questions
What is the difference between HARKing and p-hacking?
P-hacking is about the analysis: running many variations and reporting whichever crosses the significance threshold. HARKing is about the story: accepting a hypothesis only after the results were known, then writing it up as though it was the original prediction. The same analysis can be one or the other depending on what you claim it shows. Label the hypothesis and the analysis separately, and the confusion mostly disappears.
What does data dredging mean?
Data dredging is searching a dataset for a pattern that your hypothesis did not anticipate: running every variable against the outcome, splitting groups until something separates them, or keeping the combination that works. It differs from p-hacking in emphasis, since the fishing happens in variable choice rather than in repeated tests. Both are fixed the same way: separate confirmatory analyses, planned in advance, from exploratory ones you label honestly.
What is a simple explanation of a p-value?
A p-value is the probability of getting a result at least as extreme as yours, assuming the null hypothesis is true. A p of 0.05 means that in repeated studies where nothing is really there, about five in a hundred tests would look this good by chance. It is not the probability that your finding is real, and it says nothing about its size. Report the effect size and confidence interval alongside it.
Do I need to correct for multiple comparisons?
Correct whenever you run more than one test on the same outcome and only some results get reported as findings, which in practice is almost always. With 20 independent tests and no correction, the chance of at least one false positive at p below 0.05 is about 64%. If the number of tests was fixed and small, use Holm-Bonferroni for strict family-wise control; if you are screening many predictors, use Benjamini-Hochberg to control the false discovery rate instead.
Is trying an alternative model always p-hacking?
No. Running a genuinely different model because it is theoretically better is legitimate research, and running sensitivity analyses to check stability is good practice. The line is temporal: choices justified by design, theory or literature before results are fine, and choices justified by results after they are seen are not. The test is whether you could have given the same reason before running anything. If not, log the change and label the analysis exploratory.
How do I know if I already p-hacked my analysis?
Count your tests. If you ran more analyses than your plan called for, and the one you plan to report has the smallest p-value, that is the pattern. Other tells: models that changed after an output appeared, observations dropped without a written rule, covariates added late, and no timestamped plan predating the first run. The fix is disclosure: report every test you ran, apply a correction, and reframe the result as exploratory until it is replicated on new data.
Conclusion
Write and lock your primary analysis plan before you review a single outcome result: one outcome, one comparison, one test, written exclusion and missing-data rules, a sample size from power analysis, and a stopping rule. Register it if your field allows it, otherwise date it and keep it with the project files.
Then label every later analysis as planned or exploratory, log each change as you make it, correct for multiplicity whenever you run more than one test, and report the null results alongside the interesting ones. That record is what turns a result into something other researchers can trust, and it is the whole answer to how to avoid p hacking in your analysis.


