Nine signals can shift your confidence that text is machine-written, none of them proves it, detector tools are too unreliable to accuse anyone with, and whether a text is true matters more than who wrote it.
- Signals, not proof. Details that do no work, balance with no position, even rhythm and empty transitions raise suspicion, but each has a human explanation too.
- Detectors misfire. Benchmark research found detectors struggle with models and topics they were not tested on, and simple rewording cuts their accuracy sharply.
- Careful non-native writers get flagged. A 2023 Stanford study found detectors wrongly flagged a large share of essays by non-native English speakers, while doing near-perfectly on US eighth-grade essays.
- Ask about process instead. Drafts, sources and the decisions behind a piece are fairer evidence than a score.
- Check claims, not authors. A machine-written guide with correct facts serves you better than a human-written one that is wrong.
Nothing below proves anything. These are signals, and signals shift your confidence rather than settle a question. The moment a signal becomes an accusation, it is being used wrongly — including by software that claims to do it for you. Machine text that a person edited and careful human text look increasingly alike, and that convergence is permanent.
The nine signals
Details appear and then do no work. A year is named but nothing depends on it; an example is offered but the argument would be identical without it. Models produce the texture of expertise because texture is what they learned. Writing that actually knows a subject tends to let its details change the conclusion.
Weak when: the writer is padding to a word count. Humans do this too.Every consideration weighed, every side acknowledged, nothing risked. "It depends on your needs." "Both approaches have merit." A model trained to be inoffensive drifts to the midpoint of everything it has read. People who have actually done the thing usually have an opinion about it.
Weak when: the genre demands neutrality — encyclopedias, official statements, and this site's own stance boxes.Sentences of similar length, paragraphs of similar shape, each section built like the one before. Human writing lurches: a long thought, then three words. Machine writing breathes evenly, all the way down.
Weak when: the writer is trained to a house style, or the text is short.Three items where two or four would do — clear, concise, and compelling. Three is the pattern language settles into when nothing is pushing back, and models learned it from a mountain of marketing prose.
Weak when: three is genuinely the number. Rhetoric loves threes for good reasons."Moreover." "Furthermore." "It is worth noting that." Connectives that announce a relationship the sentences do not actually have. A person adds a transition when the logic needs one; a model adds one because paragraphs usually start with them.
Weak when: the writer learned formal academic English, where these are conventional.Nothing in the text was learned the hard way. No dead end, no revision, no "I assumed X and was wrong." Real expertise carries scar tissue — the thing that was tried, failed, and changed the writer's mind. Machine text has no history to draw on, so it presents knowledge as if it had always been obvious.
Weak when: the format excludes personal experience — reference pages, product copy.A closing paragraph restating the piece without advancing it: "In conclusion, X remains a complex and evolving space." Models are strongly shaped toward wrapping up. Writers who are finished tend to simply stop.
Weak when: the piece is long enough that a real summary helps.Assertions delivered in the same even tone whether they are well established, contested, or invented. Fabricated citations sit beside real ones in identical typography. A model has no internal sense of which of its claims are shaky, so nothing in the prose signals it.
Weak when: the writer is simply overconfident. Also common in humans.The hardest to describe and the most reliable in practice. No particular reader is imagined, no relationship assumed, no particular person addressed. The text is aimed at everyone, which means it lands on no one. This is what people are reacting to when they say a page feels hollow without being able to say why.
Weak when: you are tired or predisposed. This signal is the easiest to hallucinate.Use the signals to decide what deserves a closer look, not to decide who wrote something: each of them also has an ordinary human explanation.
Why detector software cannot settle it
The tools promising a percentage are weaker than their marketing, and the research is not close. The RAID benchmark (ACL 2024) evaluated twelve detectors against 6.2 million generations spanning eleven models, eight domains and eleven adversarial attacks. Its finding: detectors have substantial difficulty generalising to unseen models and domains, and simple changes — a different sampling strategy, a repetition penalty, swapping words for synonyms — produce steep drops in accuracy. The paper opens by noting that commercial detectors routinely claim 99%+ accuracy. Stanford research published in Patterns in July 2023 tested seven detectors on 91 TOEFL essays by non-native English speakers and found an average false-positive rate of 61.3%, against near-perfect accuracy on US eighth-grade essays.Sources: Dugan et al., “RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors”, read at source 11 Sep 2026 — “over 6 million generations spanning 11 models, 8 domains, 11 adversarial attacks and 4 decoding strategies”. ACL 2024 (6.2m generations, 11 models, 8 domains, 11 adversarial attacks; 12 detectors evaluated) · Liang, Yuksekgonul, Mao, Wu & Zou, “GPT detectors are biased against non-native English writers”, Patterns 4(7):100779, July 2023.
The same study then gave the non-native essays’ vocabulary a lift with ChatGPT and ran them again. The false-positive rate fell from 61.3% to 11.6% — same writers, same ideas, fancier words.Liang et al., Patterns 4(7), July 2023, read at source 17 Sep 2026: “the average false-positive rate dropping by 49.7% (from 61.3% to 11.6%)”, against “near-perfect accuracy for US eighth-grade essays”.
Read that second finding again, because it is the one with victims. A detector penalising plain, regular sentence construction is not detecting machines — it is detecting people who write English carefully because it is their second language. Students have been accused on that basis. That is the practical cost of treating a probability as a verdict.
IF YOU ARE A TEACHER, EDITOR, OR EMPLOYER
Do not accuse anyone based on a detector score. Ask about process instead: earlier drafts, sources consulted, why this example rather than another, what got cut. Someone who wrote the work can talk about the decisions behind it for as long as you like. That conversation is fairer than any tool, and it cannot be gamed by a paraphraser.
Do not treat a detector score as evidence against a person. Research on these tools finds them unreliable, and people who write English carefully as a second language are among those they wrongly flag.
The habit that beats all nine
Stop asking who wrote this and start asking is this true, and does it help me. Authorship is nearly unprovable and getting more so; usefulness is testable today. A machine-written guide with correct, checkable facts serves you better than a human-written one that is wrong. The reason AI text is worth noticing at all is that it is frequently confident and wrong — so verification, not attribution, is the skill that pays.
For images and video rather than text, the companion guide is Spot Fake Images & Video.
When a text makes you suspicious, check its claims before worrying about its author: accuracy can be tested today, and authorship mostly cannot.
AI-written text is not automatically bad and using it is not automatically dishonest — this site's own use of these tools is disclosed on its About page. What matters is disclosure and accuracy. The harm is not that a machine wrote something; it is that unchecked machine text spreads confident errors at volume, and that people are being accused of using it on evidence too weak to support the charge. Both things are worth being angry about, and they point in the same direction: check claims, not authors.