Skip to content
PlagiatScanner.de
Examination procedures · question · PRV-04

Can an AI score prove authorship on its own?

In brief: An AI-detector score is the output of a particular system for a particular text state under particular settings. By itself, it does not reveal who wrote the work or reconstruct how it was produced. A fair review requires the complete report, locatable passages, the tool’s relevant limitations, human assessment and a genuine opportunity to consider contrary process records such as drafts, research notes, versions and documented AI use.

Evidence ladder: detector output and process recordsReviewed 11 September 2026

Establish what the displayed number is meant to represent

Depending on the provider, a percentage may describe estimated text coverage, probability, classification confidence or another proprietary measure. These concepts are not interchangeable. Without the product, version, date, input text, supported language, minimum length and documented settings, a bare score is difficult to reproduce or interpret responsibly.

The analysed state matters as well. A title page, bibliography, quotations, tables or appendices may have been included or excluded. A rerun after minor editing or a service update may produce a different output. Any serious discussion therefore needs the exact file and full report rather than a cropped image of one prominent number.

Authority asset: evidence ladder for an AI-use concern

From isolated signal to inspectable overall assessment
LevelMaterialPotential contributionMain limitation
1 · NumberScore without reportRecords an asserted system valueDefinition, text state and context are missing
2 · ReportTool, version, file, settings and marked segmentsMakes the particular run inspectableStill a probabilistic system assessment
3 · Passage reviewLinguistic observations and alternatives for each locationEnables human plausibility checkingA stylistic feature is not secure provenance
4 · Process recordsDrafts, notes, sources, versions, feedback and AI logsDocuments development and decisionsNo single trace usually proves identity or completeness
5 · Overall viewConsistent and conflicting findings plus responseSupports a bounded, case-specific evaluationOutcome depends on rules and competent judgement

False positives and false negatives are part of classification

A classifier can label human writing as AI-generated and fail to identify AI-generated writing. Performance depends on the evaluation corpus, language, genre, length, model family, degree of editing and threshold. A broad promotional metric cannot simply be transferred to a short German methods paragraph or to writing by a multilingual student.

Published research has reported bias affecting non-native English writing. Other work explains why paraphrasing and editing can alter detection results. Such studies do not prove that a specific flag is wrong. They demonstrate why a score requires relevant validation and individual review before it can support a strong provenance claim.

The reverse inference is equally unsafe: a low score does not certify human authorship. Treating both error directions seriously means using a detector, at most, to frame specific questions rather than to automate an assessment decision.

A reviewable report needs more than a headline score

The record should identify provider, product, displayed version or access date, input file, analysed characters or sections, and exclusions. It should define the score and show passage-level output. Where threshold, supported language and limitations are documented, retain those details too.

Ask who generated the report and who examined the flagged passages. Were quotations, formulaic language, bibliographies and standard phrases considered? If multiple runs exist, preserve text states and results rather than selecting only the most dramatic one. A transparent report also distinguishes provider documentation from the reviewer’s later interpretation.

Process evidence addresses how the work actually developed

Detection infers a class from textual patterns. Drafts, version history, searches, annotations, commits, feedback and activity logs instead document events in the workflow. They may show when an argument emerged, which source was opened, how a paragraph changed or where a model suggestion was rejected.

Process traces also have limits. Version history does not necessarily identify the person at a device; timestamps can move; notes may be retrospective. Their value grows through clear provenance, passage-level linkage and agreement among independent records. The guide to using versions as evidence of your work explains that assessment in detail.

Seven questions that improve an evidence discussion

  1. How does the provider define this particular score?
  2. Which unchanged file and sections were analysed?
  3. Which product state, language and settings applied?
  4. Are the marked passages and complete report available?
  5. What validation actually matches the language, length and genre?
  6. What human passage-level review is documented?
  7. Which process records support, contradict or complicate the output?

The examination procedures hub connects these questions to other procedural tasks. Use the evidence inventory after an AI concern to protect existing records. This page has one foundation connection: academic writing and evidence.

Research and official foundations

  1. Liang et al., Patterns: GPT detectors are biased against non-native English writers – empirical study of potential bias.
  2. Sadasivan et al.: Can AI-Generated Text be Reliably Detected? – research on fundamental detection limitations.
  3. UNESCO: Guidance for generative AI in education and research – human-centred assessment, responsibility and privacy.

Sources reviewed 4 September 2026. Study conditions cannot automatically determine an individual case.

Keep each report paired with the exact analysed file. A later edited version needs a new identifier and its own result. Otherwise a score produced for version A can easily be quoted as though it described version B. A recorded checksum, date and export route can support that narrow file-to-report connection when they were genuinely created and preserved.