The research.
A register of published research on how reliably AI-writing detectors read human-written text. Each entry states what a study tested and what its authors reported, in the study's own figures, dated and linked to the original. The studies measure particular tools on particular datasets and do not yield one aggregate error rate; this page reports them and determines nothing about any case. If an outcome could affect enrollment, immigration status, a degree, professional licensing, or accommodations, consider obtaining qualified advice before responding.
What the published studies report
Newest first. Preprints are labeled as preprints; a vendor's statement about its own product is labeled as a vendor statement. A finding about a measured population is stated at that population and not extended past it.
- Keystroke-timing verification tested against forgery — preprint, January 2026
“On the Insecurity of Keystroke-Based AI Authorship Detection” (arXiv, submitted Jan 24, 2026) examined proposals to verify human authorship from typing rhythm rather than from the text itself. Testing five classifiers on 13,000 typing sessions from the SBU corpus, the authors reported that more than 99.8% of forged or transcribed sessions were classified as human under the attacks they specified — including the case of a person simply copy-typing machine-written text — while fully automated injection remained easy to catch. The authors present this as a security failure of timing signals used alone; the study did not evaluate every process-tracking design. arXiv:2601.17280. Retrieved Sep 20, 2026.
- Sixteen detectors tested for demographic bias — preprint, December 2025, revised April 2026
Stowe et al., “Identifying Bias in Machine-generated Text Detection” (arXiv), assessed sixteen detection systems on a curated dataset of student essays across four attributes: gender, race/ethnicity, English-language-learner status, and economic status. The authors reported that biases were generally inconsistent across systems, but that essays by English-language learners were more likely to be classified as machine-generated, with non-White ELL writers disproportionately flagged relative to White ELL writers, while essays by economically disadvantaged students were less likely to be flagged. Human readers annotating the same essays performed poorly at detection overall but showed no significant biases on the studied attributes. arXiv:2512.09292. Retrieved Sep 20, 2026.
- A national survey of who reports being flagged — Common Sense Media, 2024
A survey, not a detector measurement, entered at its own scope: in “The Dawn of the AI Era” (2024), a national survey of teens and parents, 20% of Black teens reported that a teacher had wrongly flagged their schoolwork as AI-generated, against 10% of Latino teens and 7% of white teens — a K–12 finding about reported experience. Common Sense Media, 2024.
- Fourteen tools tested, false-accusation risk computed per tool — Weber-Wulff et al., 2023
“Testing of detection tools for AI-generated text” (International Journal for Educational Integrity, published Dec 25, 2023) tested twelve publicly available tools and two commercial systems, Turnitin and PlagiarismCheck, and concluded the tools were “neither accurate nor reliable” — all scored below 80% accuracy. The study computed, for each tool, the likelihood that a classification would falsely accuse a student: zero for half of the fourteen tools, while six produced false positives, and for one tool half of its positive classifications would have been false accusations. On human-written English text the tools averaged 96% accuracy; on human-written text machine-translated into English, accuracy dropped by 20 percentage points — human writing in another language, translated, read to the tools as more machine-like. doi:10.1007/s40979-023-00146-z. Retrieved Sep 20, 2026.
- Seven detectors tested on non-native English writing — Liang et al., 2023
“GPT detectors are biased against non-native English writers” (Patterns, 2023) evaluated seven widely used GPT detectors on 91 human-written TOEFL essays from a Chinese forum and 88 US eighth-grade essays from the Hewlett Foundation's ASAP dataset. The detectors classified the US essays accurately but labeled more than half of the TOEFL essays AI-generated — an average false-positive rate of 61.3%, with 19.8% of the essays flagged unanimously by all seven detectors and at least one detector flagging 97.8% of them. When the researchers enriched the TOEFL essays' word choice, the average false-positive rate fell to 11.6%; when they simplified the vocabulary of the US essays, misclassification rose. The authors attribute the pattern to detectors' reliance on text perplexity, which reads a limited range of expression — common in non-native writing — as machine-like. doi:10.1016/j.patter.2023.100779. Retrieved Sep 20, 2026.
- The vendor's own published rates — Turnitin, June 2023
A vendor statement about its own product: Turnitin's Chief Product Officer wrote (blog, Jun 14, 2023) that the document-level false-positive rate is “less than 1% for documents with 20% or more AI writing,” and that the sentence-level false-positive rate — the figure governing any single highlighted passage — is “around 4%,” with such errors more common in documents mixing human and AI writing, at the transitions. Turnitin, 2023. Retrieved Sep 20, 2026.
What the register shows, stated at its scope: false positives are documented in every independent test that measured them; measured rates vary widely with the tool, the dataset, and the writer; and the writers most affected in the research — non-native English writers, English-language learners — are often those least positioned to contest a flag.
What institutions themselves publish about these scores
The institution pages record, campus by campus and in each school's own words, which detection tools a school documents and what weight it says their output carries. A recurring pattern in the collection is an institution stating on its own pages that a detector score alone does not settle a case — among the pages recording such statements are Georgia State, Clemson, Northern Arizona, Temple, and LSU. The quotations, their dates, and their sources live on those pages; whether any statement applies to a particular case is for the institution and its process to determine.
Where the records meet the research
Every instrument in the register above reads the submitted text, or its typing rhythm — none of them reads the process by which a piece of writing came to exist. Drafts, version histories, notes, and the people and systems that observed the work happening are records a classifier never sees and cannot score. Preserving and organizing them is covered step by step in the records checklist; capturing the pages a school publishes, as they read today, is covered at what to do with these sources.
How this register stays current
New research is found by citation, not by keyword: a scheduled scan asks the OpenAlex scholarly index for works citing the register's seed studies, which surfaces the literature on detector reliability with far more precision than a search. The scan produces a candidate list and nothing else. A work enters the register only after a person has read it and it meets the register's standing test: it is a primary source — a study, a preprint, or a vendor's statement about its own product; it bears on detector behavior when reading human-written text; and its figures can be read from the work itself. Works the scan surfaces that fail the test are recorded as seen, so a quiet month means no qualifying research was published, not that nothing was checked. Additions and corrections appear in the change log below, on the record.
Revision history of this page
2026-09-20 — Page created. Register entries verified against their originals this date: the Weber-Wulff article read in full at link.springer.com (browser read; the host answers plain fetches with a challenge page), the Liang article at cell.com (browser read), the Turnitin statement at turnitin.com (browser read), and both arXiv abstracts by direct fetch. The Common Sense Media entry is carried at the figures published on this site's homepage; its host did not answer a plain fetch this date and a first-party re-verification is noted for the next review pass. Deliberate omission: the University of San Diego law library guide cited on the homepage as a secondary pointer is not entered here, because the register holds primary sources only. The citation scan (tools/research-scan.py) was seeded this date: 1,126 citing works recorded as seen; triage of that backlog for qualifying entries is pending and will be folded in on the record. Reviewer: Carrie Schluter (review confirmed 2026-09-20).
Spot an error or an outdated quotation? hello@gatheredwork.com — corrections are made on the record, in this log.