Define institutional purpose and non-purpose
Will the system support research, demonstrate limitations in teaching or provide a signal within a governed review? A feedback tool needs different controls from a flag in an integrity procedure. State expressly that output cannot automatically determine misconduct or trigger sanctions. Without a legitimate purpose, there is no meaningful criterion for quality.
Describe input populations: languages, disciplines, genres, typical lengths, translations, code and permitted human or AI editing. Supplier benchmarks may use different distributions. A global result cannot be transferred to the university population where corpus, threshold and error costs do not align.
Plan a local, reproducible validation
Predefine outcomes, thresholds and analysis. Separate training, calibration and independent test sets. Document provenance and rights for every text. Do not use student submissions without an assessed basis. Synthetic or licensed corpora still need relevant conditions and must not be called representative where they are not.
Report absolute case counts, sensitivity, specificity or other purpose-appropriate error measures, plus uncertainty intervals. A percentage without a denominator obscures the sample. Evaluate lengths, disciplines, first and additional languages, translation, professional editing, permitted language correction and hybrid text separately. Draw conclusions only for layers genuinely measured.
Record supplier product, model, date, settings and inputs. An update can invalidate earlier results. Require change notice and revalidation before new output enters a procedure. Preserve test plan, raw results and analysis for inspection.
Examine error consequences and group differences before deployment
False positives may burden students unfairly; false negatives can create unjustified confidence. Both require separate evaluation. Research has reported bias affecting non-native English writing. That does not establish failure in every product, but it supports targeted local testing rather than a blanket fairness assumption.
Stratified reporting must avoid identifying people in small groups. Include relevant specialists and student representatives in evaluating test design, accessibility and burden. Define how uncertainty and conflicting outputs are handled. No score can replace a conversation, process records and human passage review.
Require documented supported languages and text types. “Multilingual” without disaggregated evidence is not a quality guarantee. Treat marketing statements as supplier claims until independent or institutional validation supports them for the intended use.
Specify operation, privacy and exit before award
Map upload routes, recipients, retention, training, subprocessors, transfers and deletion. Consider whether reports create personal profiles and how access and correction requests work. A pilot uses authorised material and tests permissions and deletion; it is not production approval.
Assign roles: who initiates a scan, reviews segments, sees reports, communicates with affected people, decides and audits? Training covers error modes and neutral language. Preserve product version with each report. A challenge route must enable a genuine independent check.
Contracts need change notice, data export, support evidence, end-of-contract deletion and a workable exit. The institutional false-positive process describes the required correction route.
Make an evidence-based procurement decision
- Purpose and prohibited decisions are approved.
- The benchmark demonstrably fits local text populations.
- Absolute counts and both error directions are visible.
- Bias, privacy and accessibility have named owners.
- Passage review and human decision are mandatory.
- Updates trigger revalidation.
- Challenge, audit and exit work without supplier dependence.
The higher education practice hub connects procurement and process. The privacy assessment for screening software develops the data-flow review. The one foundation connection is academic writing and evidence.
Record the decision without false precision
The procurement record should preserve the sample, product version, threshold, observed cases and uncertainty for every criterion. Exclusion conditions are set before testing: for example, an opaque storage route, no export for review, no human correction route or inadequate notice of model changes. This prevents one persuasive vendor headline from outweighing governance and privacy weaknesses.
Small samples should not be converted into impressive-looking universal percentages. The team reports which errors occurred under which local conditions and which questions remain open. A defensible outcome may be a restricted pilot or no purchase. Any approval names an owner, review date and reassessment trigger, such as a model update or a materially different use case.