Skip to content
PlagiatScanner.de
Examination procedures · problem · PRV-13

What does it mean when AI detectors disagree?

In brief: Disagreement shows that different systems classified the submitted material differently under their respective rules. It does not support averaging percentages, counting a majority or selecting the most convenient result. Preserve the exact input and record language, length, formatting, tool, version, date, output label and missing information. Authorship still requires contextual review and independent evidence of how the work developed.

Detector comparison with uncertainty fieldsSources reviewed 4 September 2026

Treat disagreement as information about the instruments

AI-text detectors differ in model design, training material, thresholds and labels. One may return a document-level category; another highlights sentences. A percentage might describe a proprietary confidence score, a claimed proportion of text or something else entirely. The shared percentage symbol does not make the quantities interchangeable.

Results may vary with language, genre, length, quotations, bullet lists, equations, reference sections and formatting. A provider can update a model without preserving the behaviour seen in an earlier report. Even a repeat run is not a replication unless the same bytes, settings and tool version were used. Conflicting scores therefore constrain what can responsibly be inferred.

Make the input auditable before comparing outputs

Preserve the document that was actually tested and calculate a checksum. Record whether the input included quotations, bibliography, tables, headings and footnotes. If a file upload and pasted text produce separate results, keep them as separate test conditions. Do not trim only inconvenient passages and then present the new score as if it described the original submission.

For each service, note the product name, URL, visible version, account tier, date, time, selected language and full output wording. Capture screenshots for visual context but transcribe essential labels into searchable text. Before uploading an unpublished dissertation, check confidentiality, data processing and institutional permission; additional detector runs are not automatically harmless.

Authority asset: detector-result uncertainty matrix

Fields for comparing outputs without pretending they share a scale
FieldSystem ASystem BUncertainty to retain
Input objectFile, characters, checksum …File, characters, checksum …Were the bytes and preprocessing identical?
Language/genreDetected or selected …Detected or selected …Is this population covered by validation?
Tool stateVersion and test date …Version and test date …Could an undocumented update intervene?
OutputExact label and number …Exact label and number …What does each number purport to measure?
Highlighted textSegments …Segments …Do locations overlap or contradict?
Validation evidencePopulation, metrics, threshold …Population, metrics, threshold …Does it match this language and task?
Permitted inferenceNarrow finding …Narrow finding …Authorship remains a separate question

Use “not disclosed” or “not established” rather than leaving blanks. Missing information is part of the result, not an inconvenience to hide.

Why averaging and majority voting create false confidence

An average is meaningful only when quantities measure the same construct on compatible scales. A score of 70 from one detector and 20 from another need not have common units. Reporting 45 can therefore invent a number that neither provider produced and no validation study supports.

Counting systems is not a remedy. Products may share training corpora, underlying services or stylistic assumptions, so their errors need not be independent. Five websites can amount to several interfaces around related technology. A vote also rewards repeated testing until the desired side wins. Selection rules should be stated before results are known, not improvised afterwards.

Independent research is useful for mapping limitations, but a benchmark result belongs to its test set, language mix, model versions and thresholds. Weber-Wulff and colleagues assessed a defined collection of detection tools and documented reliability and evasion problems. Their findings do not supply a conversion formula for one student’s paper.

Move from scores to an evidence-led account

Contemporaneous work records answer different questions. Topic approval, search logs, reference-manager records, outlines, chapter drafts, supervisor comments, data analysis and a final submission receipt can show a development path. A cloud history may preserve milestone states if ownership, export and integrity are documented. Use the cloud-history preservation check to build such a record.

No single item is absolute. A draft proves that a version existed in the preserved form, not necessarily who wrote every sentence. A large paste may reflect legitimate movement between applications. Consistent prose may follow intensive editing. Examine alternative explanations with the same care whether they favour or challenge the allegation.

Where AI assistance occurred, describe what was actually done under the relevant policy: tool, task, timing, retained material and human verification. Never fabricate prompts or backfill a staged version trail. A known documentation gap should remain visible. Honest uncertainty is more probative than a polished but invented history.

Frame a proportionate response in an examination process

Ask for the precise input, system, test date, complete output and inference being advanced. Complete the matrix, then identify the narrow methodological issue: incompatible scale, changed input, absent version or unsupported language. Do not jump from one weakness to the claim that no academic concern could exist.

Present independent evidence by stable attachment ID and state its limitations. The allegation-response framework keeps passages, records and open questions aligned. The examination-procedures hub covers adjacent preservation tasks. The AI-detector foundation page explains the technology at a general level without resolving an individual case.

If an output highlights passages, a subject specialist can inspect them. Are quotations and sources visible? Is the wording conventional for the discipline? Does it fit the surrounding argument and retained drafts? Can the writer explain research and source decisions? These are contextual questions, not a route for turning stylistic impression into certainty.

A defensible conclusion remains narrow

  • List exactly which system produced which output under which conditions.
  • State whether inputs and output scales are genuinely comparable.
  • Preserve contradictions and undisclosed tool characteristics.
  • Do not average, add or vote on proprietary scores.
  • Place independent writing and research records alongside the outputs.
  • Leave the evaluative decision to a transparent human process.

A suitable summary might read: “The documented systems classified the preserved input differently. Because the output scales have not been shown to be equivalent, the comparison does not establish authorship. The accompanying records are supplied for contextual review.” That sentence communicates the real limitation without claiming that detector disagreement proves human writing.

Scientific and institutional sources

  1. Weber-Wulff et al.: Testing of detection tools for AI-generated text – peer-reviewed comparative study of detector performance and limitations.
  2. UNESCO: Guidance for generative AI in education and research – institutional guidance on privacy, oversight and responsible evaluation.
  3. Jisc: Artificial intelligence in tertiary education – sector guidance on responsible institutional approaches to AI.

Sources reviewed 4 September 2026. Current product behaviour and validation claims must be checked for the exact date and version used.