A PhD student put the number in front of the sector
On 20 July the Higher Education Policy Institute published a paper by Brendal Aformeziem, a PhD candidate at the University of Strathclyde, arguing that UK universities should stop using AI detection scores as the primary evidence in academic misconduct cases. The argument is not that students never use AI. It is that the instrument being used to prove it produces a class of error the process cannot absorb, and that the error is not evenly distributed. International students are 24 percent of all UK higher education students and 51 percent of postgraduates, which is exactly the group the tools misread most often.
There is already a paper trail. The Office of the Independent Adjudicator published four case summaries in July 2025 involving accusations of AI-assisted misconduct. Three of the four concerned international or second-language students, and three were upheld or partly upheld against the universities that had brought them. Aformeziem's recommendations follow from that record rather than from the abstract debate: suspend detection as primary evidence pending independent validation, have the QAA and the OIA issue joint guidance that a detection score cannot be the sole basis for a disciplinary finding, and redesign assessment so that less of it hangs on an unsupervised essay.
Why this belongs outside universities. Strip away the academic setting and what is left is a procurement story that any owner will recognise. An organisation bought a scoring tool, treated the score as a finding, built a process on top of it, and then discovered that the error rate it had accepted at purchase was incompatible with the decisions it was making. The subject happened to be essays. It could as easily have been expense claims, insider-risk alerts, or a supplier fraud score.
Sixty one percent is not a defect, it is a setting
The number underneath all of this comes from Stanford. A team led by James Zou, professor of biomedical data science, ran seven widely used AI detectors over 91 TOEFL essays written by non-native English speakers and a comparison set written by US eighth-graders. The detectors handled the American schoolchildren well. On the TOEFL essays they misclassified 61.22 percent as machine-generated. All seven agreed on 18 of the 91 essays, and 97 percent of the essays were flagged by at least one tool.
The mechanism is not mysterious, and it is the part worth internalising. Detectors treat predictable word choice and simple sentence construction as machine signatures. Writing by a competent speaker of English as a second language has exactly those properties. When the researchers rewrote the same essays with richer vocabulary, the false positive rate collapsed. The tool was never measuring who wrote the text. It was measuring how ornate the English was, and then reporting the answer as authorship.
What a threshold really is. Every classifier trades two errors against each other: text it wrongly accuses, and text it lets through. The vendor picks where on that curve the product sits, and a vendor selling into a setting where a false accusation is catastrophic will tune hard toward letting AI writing pass. That is a defensible engineering choice and it is also the reason the output cannot function as proof. The HEPI paper cites a large evaluation across 805 samples that found average accuracy of 39.5 percent on unmodified AI text, falling to 17.4 percent once simple evasion techniques were applied, with some tools misreading half of all human writing.
Four universities reached the same conclusion separately
The University of Waterloo turned Turnitin's AI detection off in September 2025. Curtin University followed in January 2026. UCLA and UC San Diego had already done it in 2024. None of them coordinated, and none of them concluded that AI use had stopped. They concluded that a number they could not explain in an appeal was a liability rather than an asset. The HEPI paper also cites work by Weber-Wulff finding that none of the tools tested met the standard required for reliable high-stakes decisions.
What replaced the tool is more instructive than the removal. Institutions moved detection to advisory status, where a staff member may look at a score privately but cannot cite it in a formal complaint, and they shifted assessment design so that the evidence of authorship is generated during the work rather than inferred afterwards. That is the same move a fraud team makes when it stops arguing about a model score and starts collecting a transaction trail.
The tell. Watch what an institution does with the tool it says it still trusts. When a university keeps the licence but forbids the score from appearing in proceedings, it has quietly reclassified the product from evidence to triage. Most organisations never make that reclassification explicit, which is how a triage signal ends up in a dismissal letter.
The same arithmetic sits inside your fraud and DLP tools
Almost every scoring product an owner buys has this shape. Data loss prevention flags an outbound file. An insider-risk platform scores an employee's behaviour. A payments provider assigns a fraud probability to a customer. A CV screener ranks an applicant. Each of them sits at an operating point somebody chose, each of them has a false positive rate at that point, and in almost every purchase that number is never asked for, because the sales conversation is conducted in accuracy percentages that describe the best case on the vendor's own test set.
Three questions fix most of it. What is the false positive rate at the threshold we will actually run, on a population that looks like ours. What happens to that rate when the input is produced by someone who is competent but atypical, which for a European employer usually means someone working in their second or third language. And what is the documented step between the score appearing and a person being affected by it. If the answer to the third question is that the tool triggers an action directly, the score has already become a decision and the process is carrying a risk nobody priced.
There is a legal edge here too. Under the GDPR, a decision based solely on automated processing that produces legal effects or similarly significantly affects a person is restricted, and human involvement has to be meaningful rather than a rubber stamp. A reviewer who sees only the score and clicks through is not the safeguard the regulation has in mind. The universities that demoted detection to advisory use got to a defensible position by accident. An employer should get there deliberately.
Read next: Sprout Beat Its Numbers and Cut 20% the Same Day | Nobody Forbade It. The Filter Refused Anyway.



