A percentage without run context cannot be compared safely
Similarity scores may change while the manuscript remains identical. A comparison collection can grow, websites can be revised, exclusion rules can differ, extraction can improve, or the service can deploy an undisclosed update. Treat the report as the output of a particular run rather than a permanent property of the document.
The objective is accurate capture of what was observable. Do not add an unknown model number, corpus size or accuracy claim. A provider name is too vague; a screenshot of the headline score is also insufficient. The responsible academic institution determines procedural use. This is a technical documentation workflow, not legal advice.
Identify the submitted input byte for byte
Preserve the checked file and calculate a checksum. Record format, size, page count and, where available, extracted character count. A checksum confirms whether later copies contain the same bytes; it does not establish authorship or truth. A fresh PDF export often changes the checksum even when the visible prose appears unchanged.
Inspect processing rather than assuming success. Were footnotes, tables, headers and appendices read? Scanned pages require a separate OCR record including language and settings. Password protection, damaged fonts and unusual encodings can omit text. Preserve every warning shown by the system and identify pages that were not evaluated.
Describe corpus access and filters conservatively
Copy categories such as internet material, publications or institutional submissions from the report or official documentation. A marketing statement about total holdings does not prove which sources were queryable during this run. Give the access date and the provider's scope description, but do not claim a frozen snapshot unless one was supplied.
Record the exact state of quotation, bibliography, short-match, self-submission and source exclusions. Where the interface permits it, preserve both the initial and filtered views. Changing a display later must not overwrite the evidence of the original configuration. For manually excluded material, record who excluded it and why.
Interpret re-runs and cross-tool comparisons cautiously
A repeat can gain matches because the corpus grew or because the earlier submission itself entered a repository. Compare individual passages, targets, exclusions and extracted text rather than headline percentages alone. Label self-matches and document how they were treated.
Scores from different products are not interchangeable. Segmentation, source access and exclusion logic vary. Use a passage matrix showing shared and tool-specific leads. Describe agreement between findings, not “accuracy”, unless there is an independently classified reference set and a suitable evaluation design.
When the service exposes no release identifier, retain screenshots or exports of settings, official documentation with access date and the report's internal ID. That bundle narrows uncertainty without pretending to recreate a hidden system. If a provider later changes its interface, preserve the original labels alongside any modern interpretation.
Archive a small, readable run package
Store the exact input, native report, metadata sheet, settings evidence and a short README together. Apply appropriate access restrictions to confidential work. Where permitted, export proprietary output to a durable readable format as well. Keep annotations in a separate review copy so that the original remains untouched.
The method review sheet classifies individual findings. The development timeline connects runs to manuscript releases. More resources are listed in the examination procedures hub. The one foundation route here is the guide to understanding similarity reports.
Capture the operational route as well: browser interface or API, institutional or personal account, interface region and language, and whether prior submissions were enabled as comparison sources. Describe what was observed without inferring a particular corpus entitlement from the account label. For API runs, preserve consequential request parameters and the documented response identifier in a protected appendix. Credentials, private keys and access tokens never belong in the evidence package.
If the report is regenerated for disclosure, label that export separately from the historic run. Record whether annotations, source availability or display filters changed in between. This distinction prevents a newly rendered file from being presented as the original output while still allowing a readable current copy to accompany it.