Treat disagreement as information about the instruments
AI-text detectors differ in model design, training material, thresholds and labels. One may return a document-level category; another highlights sentences. A percentage might describe a proprietary confidence score, a claimed proportion of text or something else entirely. The shared percentage symbol does not make the quantities interchangeable.
Results may vary with language, genre, length, quotations, bullet lists, equations, reference sections and formatting. A provider can update a model without preserving the behaviour seen in an earlier report. Even a repeat run is not a replication unless the same bytes, settings and tool version were used. Conflicting scores therefore constrain what can responsibly be inferred.
Make the input auditable before comparing outputs
Preserve the document that was actually tested and calculate a checksum. Record whether the input included quotations, bibliography, tables, headings and footnotes. If a file upload and pasted text produce separate results, keep them as separate test conditions. Do not trim only inconvenient passages and then present the new score as if it described the original submission.
For each service, note the product name, URL, visible version, account tier, date, time, selected language and full output wording. Capture screenshots for visual context but transcribe essential labels into searchable text. Before uploading an unpublished dissertation, check confidentiality, data processing and institutional permission; additional detector runs are not automatically harmless.
Why averaging and majority voting create false confidence
An average is meaningful only when quantities measure the same construct on compatible scales. A score of 70 from one detector and 20 from another need not have common units. Reporting 45 can therefore invent a number that neither provider produced and no validation study supports.
Counting systems is not a remedy. Products may share training corpora, underlying services or stylistic assumptions, so their errors need not be independent. Five websites can amount to several interfaces around related technology. A vote also rewards repeated testing until the desired side wins. Selection rules should be stated before results are known, not improvised afterwards.
Independent research is useful for mapping limitations, but a benchmark result belongs to its test set, language mix, model versions and thresholds. Weber-Wulff and colleagues assessed a defined collection of detection tools and documented reliability and evasion problems. Their findings do not supply a conversion formula for one student’s paper.
Move from scores to an evidence-led account
Contemporaneous work records answer different questions. Topic approval, search logs, reference-manager records, outlines, chapter drafts, supervisor comments, data analysis and a final submission receipt can show a development path. A cloud history may preserve milestone states if ownership, export and integrity are documented. Use the cloud-history preservation check to build such a record.
No single item is absolute. A draft proves that a version existed in the preserved form, not necessarily who wrote every sentence. A large paste may reflect legitimate movement between applications. Consistent prose may follow intensive editing. Examine alternative explanations with the same care whether they favour or challenge the allegation.
Where AI assistance occurred, describe what was actually done under the relevant policy: tool, task, timing, retained material and human verification. Never fabricate prompts or backfill a staged version trail. A known documentation gap should remain visible. Honest uncertainty is more probative than a polished but invented history.
Frame a proportionate response in an examination process
Ask for the precise input, system, test date, complete output and inference being advanced. Complete the matrix, then identify the narrow methodological issue: incompatible scale, changed input, absent version or unsupported language. Do not jump from one weakness to the claim that no academic concern could exist.
Present independent evidence by stable attachment ID and state its limitations. The allegation-response framework keeps passages, records and open questions aligned. The examination-procedures hub covers adjacent preservation tasks. The AI-detector foundation page explains the technology at a general level without resolving an individual case.
If an output highlights passages, a subject specialist can inspect them. Are quotations and sources visible? Is the wording conventional for the discipline? Does it fit the surrounding argument and retained drafts? Can the writer explain research and source decisions? These are contextual questions, not a route for turning stylistic impression into certainty.
A defensible conclusion remains narrow
- List exactly which system produced which output under which conditions.
- State whether inputs and output scales are genuinely comparable.
- Preserve contradictions and undisclosed tool characteristics.
- Do not average, add or vote on proprietary scores.
- Place independent writing and research records alongside the outputs.
- Leave the evaluative decision to a transparent human process.
A suitable summary might read: “The documented systems classified the preserved input differently. Because the output scales have not been shown to be equivalent, the comparison does not establish authorship. The accompanying records are supplied for contextual review.” That sentence communicates the real limitation without claiming that detector disagreement proves human writing.