This is one of the harder things sitting on a lot of desks right now, and nobody in the chain is being unreasonable. Faculty want to hold a standard. Conduct officers want a fair process. Institutions bought a detection tool because it existed and the problem felt urgent. Students want to be believed. The friction comes from a mismatch between what the tool produces and what the conduct process needs it to be. And that mismatch is worth understanding before anyone writes the policy.
Why an AI-writing case is different from a plagiarism case
A plagiarism case has a shape everyone can work with. There is the student’s text, and there is a source, side by side. The match is visible to the student, the instructor, and a committee. It is adversarial, but it is legible.
An AI-writing case usually has no comparable artifact. There is the student’s essay and a number, a percentage, or a highlighted passage, from a model that does not show its reasoning. There is no source to produce. The student cannot easily demonstrate authorship, because “I wrote it” is also exactly what someone who did not would say. The practical effect is that the burden of proof quietly drifts. The tool asserts, and the student is left arguing against a number in a room where someone has already half-decided.
For a student who did use AI against the rules, that lands about where it should, for a student who did not, there is often no clean way out at all. A policy has to be built for that second case, because the second case is the one that does lasting damage.
What the research says about accuracy
The detectors are less reliable than their early marketing suggested, and the errors are not evenly distributed across students.
A study published in Patterns found that GPT detectors were substantially biased against non-native English writers. More than half of a set of TOEFL essays by non-native speakers were misclassified as AI-generated, while the same detectors were near-perfect on essays by US eighth-graders (Liang et al., 2023). A broader evaluation in the International Journal for Educational Integrity tested fourteen detection tools and concluded they were “neither accurate nor reliable” (Weber-Wulff et al., 2023). Over the past few years, several detection products have been pulled or switched off, by their own makers and by institutions, after internal accuracy reviews turned up problems.
Combine the uneven error rate with the drifting burden of proof and the pattern becomes predictable. The students most likely to be wrongly flagged are international students, multilingual students, and students who simply write plainly. And those are also the students least equipped to contest a charge once it lands.
Why the trust is so hard to rebuild afterward
A student who is wrongly accused, taken through a conduct process, and eventually cleared does not come back to the same relationship with the institution. They have learned that a piece of software can point at them, that it was believed over their word, and that they had to prove their own honesty just to stay enrolled. That does not un-learn easily. It travels through their friends, and it shapes how they describe the institution for years afterward.
The students who only watch it happen to someone else adjust too. They start writing defensively, keeping every draft, narrating their process, steering around phrasings they suspect a detector dislikes. That is a real cost the honest students carry, purely because the instrument is not trusted.
You can win the conduct case and still lose the exact thing the conduct process exists to protect.
A policy shape that holds
Define what a score is and is not, in writing. A detector result is a reason to look closer and start a conversation. Not, on its own, grounds for a charge. A charge should rest on additional, independent evidence: a documented discussion of the work, a comparison against the student’s known writing, an inability to talk through their own submission convincingly.
Prepare faculty for what the number means. A 90 percent score is not 90 percent certainty, and most people using these tools have not read the accuracy literature behind them. This is a short piece of faculty development, and it is the difference between the tool being a conversation-starter and the tool being treated as a verdict.
Put the bias in the policy, in writing, explicitly. If the policy does not acknowledge that these tools flag multilingual and non-native writers at higher rates, the conduct data will eventually show a disparate-impact pattern that someone will surface anyway. Name it up front, and build the check for it directly into the process.
Ask the procurement questions at renewal. What false-positive rate does the supplier state, and on what test set? What does the contract say about liability if the tool is simply wrong? Is there published, current accuracy data? Would you be comfortable defending this output in a hearing?
Treat assessment design as the durable answer. The response that lasts is not a better detector. It is assessment that is harder to outsource and easier to verify: process work, in-class components, an oral defense of written work, prompts tied to specific course materials and discussions. That is a course-design investment, and it holds up even when the tools change again next year.
The asymmetry worth keeping in view
A missed case of AI misuse costs a grade that should have been lower. A false accusation can cost a student their standing, their trust, and sometimes their enrollment. Those errors are not the same size, and a policy that weights them as if they were will end up producing the more expensive one far more often than it needs to.
Charlie Wrightmann is a consultant at Wrightmann Education Technologists. Part of the Accessibility and Compliance series.
References
- Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7). https://www.cell.com/patterns/fulltext/S2666-3899(23)00130-7
- Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., et al. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19, 26. https://doi.org/10.1007/s40979-023-00146-z