I want to make a base-rate argument about text detection, because the accuracy figures being quoted make these products sound far better than they behave in the field. Take a detector that is 95% accurate in both directions and run it over a thousand student essays where fifty were genuinely machine-written. You catch roughly forty-eight of them. You also flag roughly forty-eight human essays. Half of everything the tool accuses is innocent, and that is at an accuracy number most vendors would be delighted to print on a landing page. The harm is not distributed evenly either. Formulaic, low-variance prose gets flagged disproportionately, and that describes an enormous amount of careful writing by people working in a second language. The error lands hardest on the group least equipped to contest it. I do not think the tools are useless. I think any deployment that treats a flag as a verdict rather than as a prompt to look closer is doing real damage to real people.
