AI Safety & Compliance

Do AI Detectors Work? What Teachers Should Do Instead in 2026

AI detectors flag innocent students and miss real misuse. Here is the evidence against them, and the five tools that make student thinking visible instead.

In brief

AI detectors are not reliable enough to act on: a Stanford study found seven GPT detectors falsely flagged non-native English speakers' essays 61.3% of the time, Vanderbilt University disabled Turnitin's detector over false positives, and OpenAI retired its own classifier for low accuracy. Of the 90 tools in the independent aieducator.tools directory, none is a standalone AI-writing detector; the market has moved to process visibility. Use Brisk Teaching to replay how writing was produced, Collaboration Tracker for group work accountability, Snorkl to hear students explain their reasoning, Modality Learning to track understanding over time, and Dr Connor for sanctioned, visible student AI use.

Worth noting

  • A Stanford study published in Patterns found seven GPT detectors falsely flagged TOEFL essays by non-native English speakers 61.3% of the time on average, while native speakers' essays passed.
  • Vanderbilt University disabled Turnitin's AI detector in 2023, calculating that even the claimed 1% false positive rate would wrongly implicate about 750 of 75,000 papers a year.
  • OpenAI retired its own AI text classifier because of its low accuracy, and no detector's output is reliable enough to serve as sole evidence in a discipline case.
  • Of the 90 tools in the independent aieducator.tools directory, only four mention detection at all and none is a standalone AI-writing detector; the market has moved to process visibility.
  • Brisk Teaching's writing replay and Collaboration Tracker's contribution analytics show how a piece of work was produced instead of guessing what produced it.
  • Snorkl and Modality Learning make student reasoning observable, which is stronger evidence of learning than any probability score.
  • Dan Fitzpatrick's rule for schools: outsource the doing, not the thinking, and redesign assessment so the thinking is visible rather than policing the product.

Somewhere in your district right now, a leadership team is weighing up an AI detector subscription for the fall. I want to save them the money. I run an independent directory of 90 AI tools for education, I take no payment for inclusion, and I do not sell a detector. That puts me in a strange position on this results page: almost every other guide to AI detectors was written by a company that sells one, or by a company that sells software to beat one. This post is for teachers, principals and district leaders deciding what to do about AI and student writing this semester.

The short answer: no, not reliably enough to act on

AI detectors cannot tell you, with the certainty discipline requires, whether a student used AI. The evidence here is unusually one-sided. A Stanford study published in the journal Patterns ran seven GPT detectors against 91 TOEFL essays written by non-native English speakers: the detectors falsely flagged them as AI-generated 61.3% of the time on average, and more than 91% of the essays were flagged by at least one detector. Essays by native speakers sailed through. Vanderbilt University disabled Turnitin's AI detector over exactly this risk, noting that even Turnitin's own claimed 1% false positive rate would have wrongly implicated roughly 750 of the 75,000 papers Vanderbilt submitted in a single year. OpenAI, the company with the best possible view of the problem, retired its own AI text classifier because of its low accuracy.

Three years on, nothing structural has changed. Detectors guess at probability. They flag formulaic writing, which means they flag your English learners (EAL students, for UK readers), your students with IEPs who were taught sentence frames, and your most anxious rule-followers. A student who runs AI output through a paraphrasing tool, meanwhile, usually walks straight past them. A tool that misses cheats and accuses the innocent is not a tool a school can build discipline on.

You cannot outsource trust to a probability score. Make the thinking visible and you will not need to guess.

What the market already knows

Here is the part the detector vendors will not tell you: the education AI market has quietly moved on. Of the 90 tools in our independent directory, only four mention detection anywhere in their feature set, and not one of them is a standalone AI-writing detector. The writing feedback and academic integrity category holds 32 tools, and the growth is all in one direction: tools that show teachers how a piece of work was produced, rather than guessing what produced it. I call this process visibility, and it is the working answer to the question detectors failed.

Two honest caveats from our own data before I recommend anything. Only 4 of those 32 tools carry an independent 9ine security assessment, and 13 of the 32 do not state whether they share data with third parties. This category is newer than most and its paperwork is patchier than it should be. That is part of the story, and you should ask vendors about it directly.

How this shortlist was built

The pool was the 32 tools in the directory's writing feedback and academic integrity category. I filtered for one job: giving a teacher honest evidence about how student work came to exist, without turning the classroom into a courtroom. I weighed classroom fit, safety posture and pricing honesty, and whether the tool keeps the thinking with the student. Nobody paid to be here and vendors cannot buy placement. I left out the big commercial detectors (Turnitin's AI checker, GPTZero, Originality.ai, Copyleaks) for the reason this whole post exists: their false positive problem makes them unsafe as evidence, and they are not listed in our directory. Pricing and features verified 25 August 2026.

Tool Best for Pricing Grades Min age Privacy Data hosted Independent security assessment
Brisk Teaching Seeing how writing unfolded Freemium K-12 Not stated FERPA, COPPA (vendor site) Not stated Yes (9ine)
Collaboration Tracker Group work accountability Freemium 6-8, 9-12, Higher ed 18+ (teacher-facing) Collects no student data Canada No
Snorkl Hearing students explain thinking Freemium K-12 Not stated FERPA, COPPA (vendor site) Not stated Yes (9ine)
Modality Learning Watching understanding develop Paid 6-12, Post-sec 13+ GDPR EU No
Dr Connor Sanctioned student AI use Free 6-12, Post-sec 13+ GDPR UK No

One note on the privacy column: it reflects what vendors declare or document, not an audit. The 9ine column is the independent check.

Best for seeing how writing unfolded: Brisk Teaching

If you want one tool that replaces the detector conversation, this is it. Brisk is a Chrome extension that works inside Google Docs, and its writing inspection feature replays how a document was written: the timeline, the pauses, the pasted blocks. A 500-word essay that arrived in one paste at 11:47pm tells its own story, and a teacher can have that conversation with evidence rather than an accusation. Teachers in our directory rate it 4.77 from 35 reviews, the highest engagement of any tool in this category. What sets it apart from Turnitin's approach is that it shows the process rather than scoring the product. Watch out for the ceiling on the free plan: the Free Forever tier covers 20+ tools, but the faster models and admin dashboards sit in quote-based Premium and Intelligence tiers priced by district enrolment, and small schools can find the per-teacher cost steep. Best for Google Workspace schools; not for pen-and-paper assessment. Pricing: free tier, then quote-based for schools (verified 25 August 2026). Safety: FERPA and COPPA compliance documented on the vendor's privacy pages along with SOC 2, independently 9ine-assessed; data hosting location not publicly stated.

Best for group work: Collaboration Tracker

Collaboration Tracker answers a smaller question than AI detection, but answers it properly: who contributed to this Google Doc? It turns revision history into contribution patterns, role balance and plain-language summaries aligned to your rubric, which quietly surfaces the group member whose 800 words appeared in one go. What sets it apart from eyeballing version history yourself is the time: it does in seconds what takes a teacher twenty minutes per document. Watch out for its narrowness; it only works where the writing happens in Google Docs, and it is a teacher-facing tool (minimum age 18+), so students never touch it. Best for project-based classrooms and department heads who assess group work; not for individual handwritten or offline tasks. Pricing: freemium, with paid pricing not publicly listed (verified 25 August 2026). Safety: collects no student data by its own declaration, hosted in Canada, deletion on request; no independent 9ine assessment yet.

Best for hearing the thinking: Snorkl

Snorkl sidesteps the integrity question entirely by changing what students hand in. Students explain their reasoning out loud over a whiteboard recording, in voice, drawing or writing, and the AI gives instant feedback against criteria the teacher sets. Nobody has yet found a way to outsource their own voice explaining their own working. Teachers in our directory rate it 4.75 from 12 reviews, and its misconception-identification is the feature I hear mentioned most. What sets it apart from Flip-style video tools is the feedback layer: it responds to the reasoning, not just the recording. Watch out for the free cap: 20 activities goes quickly with a full timetable. Best for math and science teachers who want to see working (and grading it, or marking it for UK readers, takes minutes rather than evenings); not for long-form essay assessment. Pricing: free up to 20 activities, school and district plans custom-priced (verified 25 August 2026). Safety: vendor states FERPA and COPPA compliance and that it does not sell student data, independently 9ine-assessed; hosting location not publicly stated.

Best for watching understanding develop: Modality Learning

Modality's tagline is "See How Learning Happens", which is precisely the job detectors pretend to do. It tracks concept mastery over time with confidence signals and struggle detection, so a teacher sees the arc of a student's understanding rather than a single suspicious artefact. If the dashboard shows a student never grasped photosynthesis, the flawless photosynthesis essay becomes a conversation. What sets it apart from Snorkl is longitudinal range: Snorkl captures a moment of reasoning, Modality maps the term. Watch out for the commitment: it is a paid platform with LMS integration, a whole-school decision rather than a Tuesday experiment, and it is new (launched 2026), so the evidence base is thinner than Brisk's. Best for secondary schools wanting assessment infrastructure; not for a single teacher solving next week's essay problem. Pricing: quote-based, no public pricing (verified 25 August 2026). Safety: GDPR-declared, EU-hosted, minimum age 13+, collects learning data, shares none, deletion on request; no FERPA or COPPA declaration, so US districts should ask.

Best for sanctioned AI use: Dr Connor

Dr Connor flips the problem: instead of banning AI and policing the ban, give students an AI built for homework that teachers and parents can see into. Born out of University of Birmingham research on AI and critical thinking, it is a chatbot that pushes students to think rather than handing over answers, with a dashboard showing how they used it. The transparency is the point; the AI use happens where you can see it. What sets it apart from ChatGPT is exactly that dashboard, and a design that refuses to do the assignment for the student. Watch out for scope: it is for ages 13 and up, UK-hosted, and the strongest fit is homework rather than assessed classwork. Best for schools wanting a policy that says "use this, not that"; not for under-13s. Pricing: free, with a free-for-a-term school offer running (verified 25 August 2026). Safety: GDPR-declared, hosted in the UK, collects student data with no third-party sharing and deletion on request; no independent 9ine assessment.

How to choose

Start from a rule I use with every school I advise: outsource the doing, not the thinking. The problem with AI misuse is not that a machine wrote some sentences; it is that a student skipped the thinking. So aim your budget at making thinking visible, and apply some if-then rules. If your writing lives in Google Docs, start with Brisk's writing replay this week; it is free and it ends most "did they or didn't they" conversations. If your real pain is group work, Collaboration Tracker answers it narrowly and well. If you can change the assessment itself, Snorkl's spoken reasoning raises the cognitive stretch: the amount of thinking a task demands of the student after the tool has done its part. And if you are setting whole-school policy, pair a sanctioned student tool like Dr Connor with the redesign conversation, then put it in writing; my guide to writing a school AI policy people will read covers that, and teaching AI literacy is the other half of the same job. Teachers in our community are already sharing how AI has changed their assessment design; the redesigns are more inventive than any detector.

What I would do

AI detectors fail the one test that matters: they cannot be trusted when the stakes are real, and the false positives land on the students least equipped to defend themselves. The working answer is process visibility. Brisk Teaching for seeing how writing happened, Collaboration Tracker for group work, Snorkl and Modality Learning for making thinking observable, Dr Connor for AI use in the open. Before any of them, check the safety paperwork against our independent security assessments and compare the rest of the full directory. Then spend the detector money on something that builds trust instead of testing it.

Dan Fitzpatrick is a Forbes contributor, three-time bestselling author and founder of The AI Educator, and has trained more than 150,000 educators across 30+ countries.

Frequently Asked Questions

How accurate are AI detectors in 2026?

Not accurate enough to act on. Independent studies put real-world accuracy well below vendor claims, and false positive rates on non-native English writing have been measured above 60%. A tool that wrongly accuses innocent students while missing paraphrased AI text cannot support a fair academic integrity process.

Can teachers tell if a student used ChatGPT?

Not from the finished text alone, and neither can software. What teachers can see is process: how a document was written over time, whether a student can explain their reasoning aloud, and whether the work matches their tracked understanding. Process evidence supports a conversation; a probability score supports an accusation.

Why do AI detectors falsely flag English learners?

Detectors flag writing that looks statistically predictable. Non-native English speakers, students taught sentence frames, and students with IEPs often write in exactly that controlled, formulaic style. The Stanford Patterns study found detectors flagged non-native speakers' essays 61.3% of the time, which makes the tools an equity problem, not just an accuracy one.

Is Turnitin's AI detector reliable enough for discipline?

No. Vanderbilt University disabled it, noting that even Turnitin's claimed 1% false positive rate would wrongly implicate hundreds of papers a year at scale, and independent testing suggests real-world rates run higher. Most universities that keep it treat the score as a prompt for conversation, never as proof.

Do AI detectors work on paraphrased or edited text?

Mostly not. Running AI output through a paraphrasing tool, or lightly editing it by hand, defeats most detectors most of the time. That is the core asymmetry: students who misuse AI deliberately usually evade detection, while honest students with formulaic writing styles get flagged. The tools fail in both directions at once.

What should schools use instead of AI detectors?

Process visibility tools. Brisk Teaching replays how a Google Doc was written, Collaboration Tracker shows who contributed what to group work, Snorkl records students explaining their reasoning, Modality Learning tracks understanding over time, and Dr Connor gives students a sanctioned AI with teacher oversight. Evidence of thinking beats suspicion about text.

Does this advice also apply to AI marking in UK schools?

Yes. Marking with AI and detecting AI are separate questions, and UK schools face the same detector false positive problem, with GDPR adding a data protection duty when student work is uploaded to third-party detectors. The process visibility tools here include UK-hosted options, and the same redesign principle applies.

Need to compare tools? The directory can be filtered by subject, task and audience.

Open the tools directory
D
Dan Fitzpatrick

Delivered training to 150K+ educators | Founder of The AI Educator and AI Educator Tools | Forbes Contributor | International Keynote Speaker | 4 x #1 Bestselling Author