by Robert Cole, Program Director, Reinert Center
As we continue struggling with how generative AI is situated in higher education, there continue to be calls for a way to detect generative AI use to catch students out in violation of academic integrity policies. The marketing for so-called “AI detectors” is strong and appealing, offering an easy solution to a difficult challenge. Before determining if we should use one, however, it is important to know what these detectors are and how they work.
An important thing to know is, “AI detectors” are not actually detectors. They are probability applications. Text is input, the application performs a statistical analysis of factors such as vocabulary, burstiness (variation of sentence type, structure, and length) and perplexity (how predictable it is for one word to follow another). Once the text is analyzed, it is assigned a numeric percentage or score. It does not indicate how much of the text is produced by genAI. Instead, it is the probability that some portion, any portion, is produced by generative AI. For example, if the score on a piece of text is 30%, it would mean that based on the training corpus of that particular application there is a 30% chance that some portion of it is generated by AI. It will highlight examples, again in line with the training corpus. However, those examples may or may not have been produced by a machine. There is a 30% chance. Different applications use different training materials and have different weights for specific criteria; one “detector” may weigh burstiness two times more than another “detector” for example.
This leads to the first reason the Reinert Center does not recommend using generative AI “detectors”. They produce false accusations against students who wrote their own work. Depending on who is writing, humans may very well produce a sample that falls within a “detector’s” parameters regarding the various criteria previously mentioned. How can we, in good conscience, employ an application that we know produces false positive results? What might this do to our students and the relationship between students and instructors? If we knowingly use an application that falsely identifies genuine student work statistically as AI generated, what message are we sending? If high-stakes use creates due-process and relationship risks is it worth it? There is no independent empirical research that finds any application has zero false positives. Liang et al. (2025) finds that all applications produced unequal error rates across writers (especially when multilingual and additional-language, and neurodiverse writers are included). In addition, Weber-Wulff (2023) found considerable variation across fourteen different applications regarding probability of genAI produced writing.
There are independent evaluations that find that “detectors” are not sufficiently accurate or reliable. Weber-Wulff (2023), Perkins, et al (2024), and Pratama (2025) all conclude that the probability score should not be treated as an objective assessment. Researchers found that the applications were not reliable or accurate and that they could be easily manipulated by human paraphrasing or the use of another genAI application to paraphrase (Weber-Wulff, 2023). Others found that lack of accuracy and the number of false positives and false negatives prevent the “detection” applications from being used in determining academic integrity infraction (Perkins, et al, 2024). In addition, in imbalance between accuracy and fairness indicates that “detector” analysis is context-dependent rather than universally reliable (Pratama, 2025). Regarding the Weber-Wulff (2023) finding related to paraphrasing, many other studies have come to this conclusion as well. Perkins, et al (2024), Huang (2025) and David and Gervais (2025) also found that simple – and very little – paraphrasing, rewriting, translation or other changes like those made by humanizing bots markedly bring down the accuracy of the “detectors”.
These applications also disproportionately disadvantage multilingual, additional-language, and neurodiverse writers. As previously mentioned, most “detectors” use a criterion like perplexity to measure the predictability of the text; one word leading to another. Writers in these groups and perhaps other writers may have a more limited or preferred vocabulary or a more formulaic style or writing structure that may be classified as non-human. Liang et al. (2023) found that “detectors” consistently mislabeled human written text as AI generated. Modifying the vocabulary of the sample sometimes changed the outcome of the analysis indicating that the application may be responding more to linguistic expression than identifying how the sample was written.
Other reasons that contribute to our stance on generative AI “detection” include that “detector” scores do not prove who wrote the work or how it was produced, they provide only a probability score. “Detection” addresses the output rather than the quality of the learning. These applications cannot determine if a student understands an assignment, if students can explain the evidence they provide, or if learning has been transferred to another context or met the actual learning outcomes. The use of “detectors” may contribute to students losing sight of what you want them to learn in place of “how do I get this past an AI ‘detector’?” regardless of whether they use generative AI or not. Imagine having to weigh each word and sentence because you’re afraid your work will be identified as AI-generated. Would that contribute to learning about the content and skills of the course?
If using AI “detectors” is not recommended, what can be done? Our first recommendation is to make your generative AI policy clear. Students take multiple classes and each instructor should provide clear direction regarding generative AI use on the syllabus and we would suggest further, for each assignment. You might consider integrating generative AI into one or more assignments to help students understand that it may not be the panacea some think it is. To get quality results, it takes a fair amount of work, especially when just learning how to use them. Reminding students frequently of the Academic Integrity Policy may also help students make better decisions. Finally, if you feel that you have a student or students that have used a generative AI application against your stated policy, I would urge you to talk with them. An in-person conversation can go a long way. I wrote another blog post about what that may look like here.
For more information or to discuss how you might incorporate these ideas into your courses, contact the Reinert Center by email or submit a consultation request form.
Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100779.
Perkins, M., Roe, J., Vu, B. H., et al. (2024). Simple techniques to bypass GenAI text detectors: Implications for inclusive education. International Journal of Educational Technology in Higher Education, 21, 53.
Perkins, M., Roe, J., Postma, D., McGaughran, J., & Hickerson, D. (2024). Detection of GPT-4 generated text in higher education: Combining academic judgement and software to identify generative AI tool misuse. Journal of Academic Ethics, 22, 89–113.
Pratama, A. R. (2025). The accuracy-bias trade-offs in AI text detection tools and their impact on fairness in scholarly publication. PeerJ Computer Science, 11, e2953.
Sadasivan, V. S., Kumar, A., Balasubramanian, S., Wang, W., & Feizi, S. Can AI-generated text be reliably detected? Stress testing AI text detectors under various attacks. Transactions on Machine Learning Research.
Weber-Wulff, D., et al. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19, Article 26.