Judge the process, not the label
Researchers must correct some defects before analysis. A duplicated import, impossible date or confirmed instrument fault can misrepresent the observations. Calling an operation “cleaning”, however, does not make it neutral. A defensible process preserves the source, applies the same rule to comparable cases, records the evidence and allows another person to reconstruct the analytical file.
Timing matters. A threshold created after inspecting the preferred result may be tailored consciously or unconsciously to that result. Post hoc decisions are not automatically forbidden, but they should be labelled exploratory, justified independently and tested against reasonable alternatives. Only the competent institutional process can make a formal finding of research misconduct.
Write rules before viewing decisive outcomes
Place the rule sheet beside the analysis plan. Identify variables, thresholds, decision owners and the date on which each rule was fixed. If a procedure comes from a disciplinary standard or equipment manual, cite the exact version. “Obvious errors were removed” cannot be reproduced or reviewed.
Build changes from an immutable raw state
Do not edit the only source file. Preserve a read-only original and create the analytical dataset through scripted transformations where possible. Manual correction requires a row-level record: case ID, field, old value, new value, reason, evidence, operator, date and rule identifier.
For example: “Case P017, interview_date, 2062-04-11 changed to missing; source form unreadable and no corroborating record; rule D-03; reviewed by AB, 4 September 2026.” This record does not pretend to know the correct date. It shows why an impossible entry did not enter the model.
Log rejected corrections too. If a source check confirms an unusual observation, retain it and record the check. Otherwise, the trail may suggest that every suspicious value was automatically removed. The honest analysis audit trail connects transformations to interpretation.
Outliers expose result-led choices
An extreme observation may be a measurement error or genuine variation. Its effect on a p-value, coefficient or theme is not evidence of invalidity. Inspect capture, unit, transfer and substantive plausibility first. A technically valid value belongs to the sample unless a justified exclusion rule applies.
Where several rules remain defensible, show a sensitivity analysis. Present the planned main analysis and a clearly defined alternative using the same frozen source. Explain what changes and what remains stable. Robustness information strengthens interpretation; silently running many variants and selecting one creates false certainty.
Qualitative data require the same care. Correcting transcription, merging codes or removing “irrelevant” passages may alter meaning. Preserve original words, the coding decision and its rationale. If quotations are tidied for readability, state the convention and retain the source transcript under its access controls.
Look for asymmetry and missing provenance
- The rule refers to a desired finding rather than a quality defect.
- Comparable observations receive different treatment.
- Raw or intermediate states cannot be recovered.
- Rationales were written only after outcome comparison.
- Several plausible specifications were run but only one is disclosed.
- Manual replacements have no independent source record.
Ask a second person to select and audit a sample of log entries against source material and rule definitions. Compare row counts, exclusions, missing values and key summaries after every stage. Reconciliation should explain each difference rather than merely confirm that the final file opens.
Report planned cleaning and exploration separately
The methods section should identify the raw-data release, transformation rules, software, exclusion counts and location of the change log. State which operations were planned before the main outcomes were inspected. Explain any post hoc procedure and show the effect of credible alternatives. Every reported number must trace to the named analytical version.
A compact account might say: “The raw release was archived unchanged. Duplicate imports were removed only when origin IDs matched. Values outside the collection range were checked against source records; unsupported replacements were coded missing. Valid outliers remained in the primary analysis, with a predefined sensitivity analysis reported separately. Transformations are recorded in the change log.”
The research integrity hub places data decisions within the wider evidence process. Use the good research practice audit for the complete project. The plagiarism-check foundation guide addresses source similarity and cannot decide whether data cleaning is manipulation.
Final dataset audit
- Can the unchanged raw state be recovered under appropriate access controls?
- Were key rules fixed before decisive results were known?
- Does every manual change identify evidence and an operator?
- Are equivalent quality problems treated equivalently?
- Do row, exclusion and missing-value balances reconcile?
- Are reasonable alternative rules visible in sensitivity analysis?
- Do methods, log and analytical release describe one state?