The problem is the burden of proof, not just accuracy
Detector accuracy gets most of the attention, and it deserves some. A 2023 study from Stanford found AI detection tools flag 61% of writing by non-native English speakers as AI-generated, against roughly 3% for native speakers. If a chunk of your class writes in a second language, a detector-led policy is not applying the same standard to everyone in the room.
But suppose the tool were excellent. You still receive a number with no reasoning attached. You cannot see which passage triggered it. You cannot rerun it under conditions you choose. The student cannot rebut it except by saying they wrote the thing, which is what anyone would say. You end up weighing a classifier's confidence against a person's word, and there is no version of that conversation you are equipped to win.
The result is a role most teachers did not sign up for. Somebody has to open the meeting, and it is you, and the only evidence in the room is a percentage neither party can interrogate.