For the person who has to justify a decision about AI to a class, a team, a board or a parent. The four questions you will be asked, with sourced answers and dates you can name out loud.
- Jobs: give the measurement, not the forecast. The best agents still fail most professional tasks at the first attempt, so AI is changing tasks faster than it is changing jobs.
- Detection: the honest answer is no. Detectors wrongly flag writing by non-native English speakers at a high rate, so a detector result is not evidence against a person.
- Made-up answers: it was scored for guessing. Benchmarks reward a lucky guess and give nothing for "I do not know", so systems learn to answer everything.
- Permission: the EU labelling rule targets unreviewed output. It covers AI text published without human review or editorial control, not work a person checked and publishes under their own responsibility.
- Scores are not a simple progress bar. When one agent benchmark saturated, a harder one replaced it and scores fell. The fall was the instrument, not the models.
Everything written about AI is written for the person using it. This is for the person who has to explain a decision about it to someone else — a class, a team, a board, a parent, a client.
The four questions you will actually be asked
"Is it going to take my job?"
Answer with the measurement, not the forecast. The best agents complete fewer than one professional task in four, and only 40% even with eight tries.Mercor, Introducing APEX-Agents, 21 Jan 2026, read at source 17 Sep 2026: “Frontier models successfully complete less than 25% of tasks that would typically take professionals hours.” and “Even with 8 tries, the best agents can only complete 40% of the tasks.” An earlier version quoted two other sentences as coming from this post; neither is in it. Autonomous deployment across business functions remains in single digits, and Gartner expects more than 40% of agentic projects to be cancelled by the end of 2027.Stanford HAI 2026 AI Index · Gartner press release, 25 June 2025 — not 2026, as an earlier version of this page said. Both read at source 11 Sep 2026: “AI agent deployment across business functions remains in single digits”, and Gartner predicts “over 40% of agentic AI projects will be canceled by the end of 2027”
What that supports saying: it changes tasks faster than it changes jobs, and the gap between demonstration and deployment is currently very wide.
"Can we just detect it?"
No, and saying so protects you. Seven detectors tested on 91 essays by non-native English writers averaged a 61.3% false-positive rate, against near-perfect accuracy on essays by US eighth-graders.Liang, Yuksekgonul, Mao, Wu & Zou, Patterns 4(7):100779, July 2023, full text read 17 Sep 2026 via Europe PMC: “High misclassification of TOEFL essays written by non-native English authors as AI generated, with near-perfect accuracy for US eighth-grade essays.”
What that supports saying: a detector output is not evidence, and using one to accuse a person is a decision you would have to defend. The full case is on how to spot AI writing.
"Why did it make that up?"
Because it was scored for guessing. Benchmarks award full credit for a lucky guess and zero for "I do not know," so a system optimised against them learns to answer everything.Kalai, Nachum, Vempala & Zhang, Why Language Models Hallucinate, OpenAI, September 2025 (arXiv 2509.04664), read at source 11 Sep 2026: “language models hallucinate because the training and evaluation procedures reward guessing over acknowledging uncertainty”
What that supports saying: it is not malfunctioning, it is doing what it was rewarded for. Full mechanism on why AI makes things up.
"Are we allowed to use it?"
In the EU, since 2 August 2026, transparency obligations are enforceable. Article 50(4) requires labelling of AI-generated text published to inform the public on matters of public interest unless it has been through human review or editorial control and someone holds editorial responsibility for publishing it — so work a person commissioned, checked and publishes under their own responsibility falls outside it. Deepfake images, audio and video are a separate duty in the same paragraph, with no such exemption.Regulation (EU) 2024/1689, Art. 50(4), read at source 17 Sep 2026: “This obligation shall not apply where the use is authorised by law to detect, prevent, investigate or prosecute criminal offences or where the AI-generated content has undergone a process of human review or editorial control and where a natural or legal person holds editorial responsibility for the publication of the content.” · Commission Guidelines on transparency obligations, adopted 20 Jul 2026, read at source 11 Sep 2026 — the obligations began to apply on 2 August 2026, and outputs generated before that date need not be marked retroactively.
What that supports saying: the obligation attaches to unreviewed output, not to assistance. More on the rules arriving.
Three things that hold a room
- Name the source out loud. "Stanford, 2023, seven detectors, ninety-one essays" ends an argument that "studies show" does not.
- Give the date. Half the disagreements in these conversations are two people holding figures from different years.
- Say what you do not know. The person who admits the uncertain part is trusted on the certain part.
Agent success on a computer-use benchmark went from 12% to 66% in a year — then the benchmark saturated, a harder one replaced it, and the best model scored 20.6%.Stanford HAI, 2026 AI Index Report, read at source 23 Sep 2026: “AI agents made a leap from 12% to ~66% task success on OSWorld”. OSWorld 2.0 leaderboard, last updated 4 Sep 2026, read at source 23 Sep 2026, on the benchmark authors’ own run of Claude Opus 4.8, reported June 2026: “Author-run on the official OSWorld 2.0 harness (108 long-horizon tasks)” … “20.6% binary completion.” The same page now lists higher binary scores and warns that partial scores “pass 75%” on a different metric.
The fall was the instrument, not the models. Scores on the harder benchmark have since climbed again, from runs its own leaderboard says are not comparable. That single fact explains most of the disagreement people have about AI progress, and it is on the agents page with its sources.
What not to say
- Not "90% of the internet will be AI by 2026." Repeated for years in the name of a 2022 Europol report, it did not happen, and Europol has since removed the statement behind it.Europol Innovation Lab, Facing reality?, 2022, publication page read 17 Sep 2026: “In the updated version, a statement from an inaccurate source on the expected future share of synthetically generated content was removed.” The current (January 2024) version therefore cannot be cited for the forecast.
- Not “AI detectors are 99% accurate.” That is vendor marketing; independent testing does not support it.Copyleaks, AI Detector product page, read at source 23 Sep 2026: “find AI-written content with over 99% accuracy”. Against it, Liang, Yuksekgonul, Mao, Wu & Zou, Patterns, July 2023, full text read 23 Sep 2026 via the Europe PMC API: “High misclassification of TOEFL essays written by non-native English authors as AI generated”.
- Not one confident number for AI energy use. Credible estimates differ by an order of magnitude; a confident number is a sign the speaker has not read the range.
Most pages here teach what to do. When not to use AI covers the other side — where these tools are worse than doing it yourself, and how to tell in advance.