Skip to content
PlagiatScanner.de
Research integrity · INT-15

P-hacking explained: when analysis flexibility becomes misleading

P-hacking is outcome-led analytical flexibility: variables, models, groups, exclusions or stopping points are repeatedly altered and selectively reported until a desired p-value appears. Alternative analyses are not inherently improper. They remain defensible when their rationale, complete search space, uncertainty and results are reported regardless of outcome, with planned tests clearly separated from exploration.

Legitimate-versus-problematic analysis pathReviewed 4 September 2026

The hidden selection across many tests is the problem

A single statistical test is interpreted under a stated procedure and assumptions. If a researcher tries many outcomes, subgroups, transformations or stopping points and presents only the most favourable result, the uncertainty reported for that final test does not describe the entire selection process. Simmons, Nelson and Simonsohn demonstrated how undisclosed flexibility in collection and analysis can make false-positive findings easier to produce.

P-hacking need not begin as a deliberate plan to deceive. A genuine attempt to understand an unexpected result can expand into repeated analysis. It becomes misleading when the search disappears and the selected finding is presented as the sole pre-planned test. Safeguards include specific decision rules, a complete analysis history, transparent exploration, and interpretation that considers effect magnitude, uncertainty, multiplicity and replication rather than a threshold alone.

Map the researcher degrees of freedom

Potential choices include several outcomes, alternative scale construction, different covariates, subgroups, exclusions, transformations, model families and times at which data collection might stop. None is automatically wrong. A complex or robust analysis may need several specifications. The decisive question is whether the choices have a scientific justification independent of the desired result, or were selected after inspection and incompletely disclosed.

Repeatedly checking results during collection can also change the operating characteristics of a planned procedure. Whether sequential decisions are valid is a statistical-design issue. Students should not invent universal cut-offs or corrections. Agree an appropriate plan with a methodologically qualified supervisor and document it before the relevant outcomes are examined. This guide is not a substitute for statistical advice.

Authority asset: legitimate and problematic analysis branches

Use the branches before analysis and again during reporting
DecisionLegitimate branchProblematic branchEvidence to preserve
Outlierspre-specified domain rule or documented measurement error; sensitivity comparisonremove cases until a preferred threshold is crossedrule, counts, with/without results and timing
Covariatestheory-led primary model plus declared robustness modelssearch combinations and report only the favourable oneall material specifications and result table
Subgroupsspecified in advance or clearly exploratory with interaction assessmenttest many groups and highlight the lowest p-valueall groups considered, rationale and uncertainty
Stoppingplanned sample or properly designed sequential procedurepeek repeatedly and stop on the desired outcomestopping rule, interim checks and method
Reportseparate planned, sensitivity and exploratory analyses including null resultspresent the successful variant as the original testanalysis log, code and deviation table

If a branch can be justified only after learning the result, do not call it an untouched confirmatory test. Preserve the decision and report it as exploratory or sensitivity analysis as appropriate.

Necessary flexibility should remain visible

Data may violate model assumptions, an instrument may partly fail, or an unforeseen coding problem may emerge. Blind adherence to an unsuitable plan would not improve integrity. Keep the initial specification, record the diagnostic that motivated change and state whether target outcomes were already known. Explain the replacement method and, where useful, show how conclusions vary across reasonable alternatives.

Exploratory data analysis has scientific value. It discovers structure and generates hypotheses. Its findings should not use the same evidential language as an independent confirmation. A newly generated hypothesis is best tested with new data or a future analysis fixed in advance. If a dissertation cannot provide that test, state a precise future design rather than presenting the discovery as settled.

Do not confuse a p-value with effect importance

The American Statistical Association explains that a p-value can indicate incompatibility between data and a specified model, but does not measure effect size or the importance of a result. It cannot replace study design, data quality or complete reporting. Interpret estimates with suitable uncertainty intervals or other discipline-appropriate measures, and describe the sources of variation the analysis addresses.

A conventional threshold should not become the target of an optimisation exercise. Values just above and below it do not create two entirely different scientific realities. Report the value in the format appropriate to the discipline, explain the model and assumptions, and avoid turning it into the probability that a hypothesis is true when the method does not support that reading.

Show the analysis search space in the results

Lead with the pre-specified primary analysis and disclose departures. Report central outcomes whether or not they support the hypothesis. Present sensitivity and exploratory work afterwards. A compact specification table can show the relevant alternatives more clearly than selective prose. Link the output to code, data version and software environment where rights and confidentiality permit.

When many analyses were performed, explain their purpose and address multiplicity using methods appropriate to the field. State how the displayed result was selected. If the statistical treatment is uncertain, seek qualified advice rather than tuning the procedure after the fact.

Preregistration is a record, not a quality certificate

A time-stamped plan can preserve hypotheses, outcomes, exclusions, models and stopping rules before results are known. OSF guidance recommends precise decisions and advance “if–then” contingencies. Departures can still occur and should be explained. A vague registration cannot prevent outcome-led flexibility, while unregistered research can still be transparent if its full process is documented accurately.

Preregistration does not certify a good hypothesis or correct statistical model. It strengthens the evidence about chronology. Choose a format consistent with discipline, ethics, confidentiality and the assignment rather than treating registration as a universal administrative requirement.

Audit the analysis folder for hidden branches

  1. List every outcome, subgroup, model, transformation and exclusion rule examined.
  2. Mark decisions made before and after the target results became known.
  3. Compare the primary analysis, code, output files and prose for omitted variants.
  4. Report magnitude, uncertainty, assumptions and multiplicity, not merely threshold labels.
  5. Classify post-result patterns as exploratory and design an independent follow-up test.

Sources and statistical framework

  1. Simmons, Nelson & Simonsohn (2011): False-Positive Psychology – primary research on undisclosed analytical and data-collection flexibility.
  2. American Statistical Association: Statement on Statistical Significance and P-Values – official principles for p-value interpretation.
  3. OSF Support: Effective practices for rigorous preregistration – official prompts for variables, models, exclusions and contingency rules.