Skip to content
PlagiatScanner

AI transparency · risk · KIT-18

Confidential data and unpublished writing in AI tools

In brief: Do not submit confidential, personal or contractually restricted material to an AI service until purpose, authority, data minimisation, service terms, technical controls and institutional approval have been established. Use the traffic light below as triage, not permission. Red means stop. Amber means obtain a case-specific review. Green only means that this first pass found no obvious confidential data class.

Data-class traffic lightSources reviewed 4 September 2026

Begin with the stop test

Before choosing a tool, ask whether you are entitled to disclose the material at all. Ownership, research consent, employment duties, examination rules, a non-disclosure agreement or a publisher's embargo may restrict use even when no person's name appears. Convenience is not an adequate reason to expose somebody else's unpublished work.

Inspect the whole object rather than the paragraph you can see. A document may carry author names in comments, respondent codes in hidden columns, locations in filenames or earlier wording in revision history. Several ordinary details can identify a person when combined. Classify the proposed upload according to its most sensitive component.

Pause immediately if you cannot state who controls the material, who authorised this processing and what the service will do with it. This is general research-organisation guidance, not legal advice.

The data-class traffic light

Triage for proposed AI input

Traffic-light classes, examples and required next action
ClassExamplesNext action
Red: do not submitIdentifiable interview transcripts, health details, examination records, confidential peer-review files, credentials, another person's unpublished manuscriptStop. Find a controlled alternative and identify the authority needed for any different process.
Amber: review firstPseudonymised research data, internal drafts, unpublished findings, licensed full text, institutional documentsDocument purpose, lawful basis, contract, storage, access, transfers, minimisation and approval.
Green: lower apparent sensitivityNeutral text you wrote for testing, public facts, genuinely synthetic non-personal data, material expressly cleared for this useStill check necessity, terms, intellectual property and whether combinations reveal something restricted.

The colour describes the proposed input in context; it does not rate a provider. An institutional account may have negotiated controls that a public account lacks, yet those controls do not authorise every data type. A publicly reachable article may be low in confidentiality but still subject to copyright or licence restrictions.

Why amber material needs named responsibility

Pseudonymised data remains linkable with additional information. Replacing names with participant numbers therefore does not make a transcript anonymous. Rare occupations, dates, events and quotations can also identify someone indirectly. Small research communities make combination risks particularly important.

Record who is competent to decide the case. Depending on the project, that might involve a data protection officer, information-security team, research-data service, ethics body, library, project owner or principal investigator. A supervisor's approval may be academically necessary without resolving every privacy, confidentiality or licensing issue.

Minimisation changes the question. Instead of submitting a complete transcript to ask about sentence structure, build a synthetic example that preserves the linguistic problem without the participant's story. Instead of uploading a full unpublished thesis, describe the abstract task. Where approved local processing exists, compare that route with external processing. “AI” is too broad a category for a transfer decision; the exact service, account and configuration matter.

Keep the assessment with the research record: data class, purpose, materials checked, responsible reviewer, permitted scope and review date. Do not paste the sensitive source into the risk log. A useful record proves that a decision occurred without creating a second uncontrolled copy.

Green is not a guarantee

A short, invented sentence with no reference to real people is a plausible green input. Carefully designed synthetic tabular data may also qualify when it cannot be traced to participants or confidential cases. These examples reduce privacy risk because the academic task can be tested without moving protected originals.

Other duties remain. Provider terms may not fit the planned use, a licence may prohibit redistribution, or a collection of public snippets may reveal an unpublished research direction. Submit only the smallest amount needed and avoid attached files when a short abstract instruction works. Reassess the colour when the purpose, account, tool or dataset changes.

Recording a green classification also has value. State what was submitted, why no personal or confidential element was found and what was deliberately excluded. The model-and-access metadata record can then identify the system without reproducing the input itself.

Seven gates before any submission

  1. Define the exact output needed from the service.
  2. Decide whether any original material must be transferred to achieve it.
  3. Check direct identifiers, indirect identifiers, confidential obligations and unpublished third-party content.
  4. Identify the person or office with authority over the proposed processing.
  5. Examine the terms and controls for the specific service and account, not the provider in general.
  6. Replace, reduce, synthesise or locally process the material where feasible.
  7. Record the decision, permitted scope and trigger for reassessment.

A settings screenshot is only one piece of evidence. It cannot establish consent, contractual authority, transfer conditions or the content of a hidden attachment. Review the decision if a service changes, a new collaborator joins, a dataset is enriched or the research purpose expands.

Worked case: shortening an interview quotation

A student wants an AI assistant to shorten a quotation from an interview. Names have been replaced, but the extract gives a small town, a unique job and a recognisable incident. The traffic light is red because that combination may reveal the participant. Removing the name did not resolve the identification risk.

The student writes a fictional paragraph containing the same editing difficulty and uses it to explore general shortening strategies. The real passage stays inside the approved research environment and is edited there. The methods record states that external transfer was rejected, gives the data class and names the safeguard. The protected words never enter the public AI log.

The same separation should be applied to other writing tasks. The translation disclosure guide discusses unpublished source text, while the AI transparency hub connects risk decisions to an accountable use record. The AI scan explainer describes a distinct assessment purpose and does not authorise submission of protected material to an arbitrary service.

Sources and scope

  1. EUR-Lex, General Data Protection Regulation – principles, data protection by design and security of processing.
  2. European Data Protection Board, Opinion 28/2024 on personal data in the context of AI models – case-specific analysis of personal-data processing.
  3. Information Commissioner's Office, Guidance on AI and data protection – risk-based governance and data-protection assessment.
  4. UNESCO, Guidance for generative AI in education and research – privacy, human agency and responsible use.

Sources reviewed 4 September 2026. The governing law, research consent, contracts and institutional policy for your actual project determine what is permitted.