PlagiatScanner

AI transparency · workflow · KIT-12

Disclose and verify AI-assisted data analysis

In brief: Record the exact data state, service and displayed model, instructions and parameters, generated code or output, human decisions and an independent validation. Keep sensitive raw data out of public evidence. Disclosure makes the analytical route visible; it does not replace methodological justification, data protection review or the rules that govern your assessment.

Model–version–prompt–validation record

Fields for one reproducible AI-assisted analysis step
Field Record
Data state File/version, variables, exclusions and pseudonymisation
System Provider, product, displayed model and access date
Task Prompt, context, objective and prohibited operations
Output Code, categories, calculation or interpretation
Decision Accepted, modified and rejected elements with reasons
Validation Reference calculation, second method, test data or human review
Limits Untested conditions, uncertainty and remaining risk

Separate four analytical roles

A system may write analysis code, suggest variables, categorise qualitative material or phrase an interpretation. These are not equivalent. Code generation requires syntax, logic and test review. A method recommendation needs support in the methodological literature. Automated categories require a defined coding process and checks against the material. Interpretation remains constrained by the question, design and limitations.

Use a precise verb for each step: “suggested”, “generated code”, “classified”, “summarised” or “interpreted”. Then state who made the decision. “AI was used for analysis” conceals whether the method, implementation or only the wording of results was affected.

  • Justify method choice before requesting implementation.
  • Freeze the data version with a stable identifier.
  • Record transformations in their actual order.
  • Review generated code as critically as external code.
  • Do not ask the same system to certify its own unsupported assumptions.

Plan validation that can disagree

A plausible answer is not confirmation. For quantitative work, use small reference data with manually checkable results, established software or a second implementation. Check sample size, missing values, measurement level, grouping, rounding and signs. For qualitative coding, define a codebook, examples and decision rules; review real material and preserve disagreements rather than presenting a bare agreement figure.

Validation must not merely repeat the assumptions behind the proposal. If the system recommends a test and writes the check, compare both with official software documentation and methodological sources. Retain failures in the trail: what was detected, how it was corrected and which tables or statements had to be regenerated?

  1. Write the expected result before running the procedure.
  2. Test a transparent reference and a failure condition.
  3. Save intermediate outputs, not only the final chart.
  4. Classify discrepancies and investigate causes.
  5. Rerun every dependent step after correction.

Case: categorising open responses

A researcher wants to organise 240 free-text responses. She first builds a codebook from published concepts and a human-read subset. The system receives pseudonymised text and fixed categories, with permission to return “unclear”. The record identifies data version, instruction, displayed model, date and output schema. She reviews every unclear case and a defined control set, records overlapping categories and revises two definitions.

The methods chapter reports codebook development, machine assignment, human review and revisions. It does not claim “objective” coding. The appendix contains a schema, redacted illustrations and decision record rather than identifiable responses. Actual access follows consent and the data management plan.

Coordinate methods, declaration and technical evidence

The methods chapter explains data, preprocessing, analytical steps, validation and limitations. The AI declaration summarises tool, role and human contribution. The technical record preserves details. These layers should agree without duplicating confidential material.

Do not upload sensitive or unpublished data to an unapproved service. This is not legal or data-protection advice. Institutional policy, consent and the data management plan govern the project.

The data-cleaning log helps preserve transformations. Use the neutral declaration template for a summary. The AI transparency hub joins model, prompt and analysis records. The AI detection overview explains why a finished text cannot prove this process.

Questions before submission

Is the analysed data state unambiguous? Can every transformation be repeated? Is the system's role explicit? Are method and parameters justified outside its reply? Is there a check capable of finding disagreement? Are failed attempts and limits retained? Does the evidence protect confidential material? Do methods, code, tables and declaration agree? Only the combined trail lets a reader judge the result realistically.

Avoid substituting a generic statement that “all outputs were checked”. Name the check and its reach. If only a subset was reviewed, report the selection logic. If a reference calculation covered one range, do not imply coverage of all possible inputs. Honest boundaries strengthen rather than weaken the scientific account.

Map dependencies between outputs

Draw a chain from raw-data version through cleaning, derived variables and model run to table, figure and prose claim. Give each node a file or record ID. When an early step changes, the map reveals every result that needs regeneration and every sentence requiring review. Without it, an obsolete chart can remain after code changes. Also timestamp exploratory analyses separately from pre-planned tests; a generated suggestion discovered later must not be narrated as an original hypothesis.

Sources

Sources reviewed 4 September 2026.