- "It's just autocomplete" undersells it. Researchers can find internal features that match real concepts and change behaviour when adjusted.
- "It understands, maybe it's conscious" oversells it. Fluent talk about inner states is not evidence of them, and a 2023 assessment concluded that no current AI systems are conscious.
- Its account of its own reasoning is unreliable. Anthropic found Claude describing school arithmetic while doing something different inside.
- It can misjudge what it knows. A familiar-sounding name can switch off its "don't know" response — one route to a confident wrong answer.
- For practical use, the question barely matters. Accuracy can be checked; understanding cannot. Check the output either way.
"It's just autocomplete"
- Technically accurate about the mechanism: the system predicts the next piece of text, repeatedly.
- Whether it remembers you between conversations now depends on the product — as of August 2026, memory is on by default for Claude's Free, Pro and Max plans — but it has no goals of its own and no stake in being right.Anthropic, Claude's memory works everywhere, and you decide what's in it, 25 Aug 2026, read at source 16 Sep 2026: “Memory is on by default on Free, Pro and Max plans across web, desktop, and mobile.” An earlier version of this line said these systems have no persistent memory by default.First-hand: until 22 Sep 2026 this line said Anthropic “switched memory on by default” in August 2026. The 25 Aug 2026 post says memory is on by default; it does not say that was when the default changed.
- It cannot reliably tell recall from invention, which is why it produces confident falsehoods in the same register as facts. The mechanism is on why AI makes things up.
"It understands, it might be conscious"
- Points at real behaviour: these systems handle novel problems, transfer reasoning between domains, and describe their own steps in ways pure lookup cannot — though those descriptions are not always what happened inside.Anthropic, Tracing the thoughts of a large language model, 27 Mar 2025, read at source 16 Sep 2026: “Strikingly, Claude seems to be unaware of the sophisticated "mental math" strategies that it learned during training.” Asked how it added 36 and 59, “it describes the standard algorithm involving carrying the 1.”
- Interpretability work has identified internal features corresponding to recognisable concepts — structure, not just statistics.
What the evidence actually shows
There is internal structure, and it is findable
Interpretability research can locate features inside these networks that correspond to identifiable concepts, and in some cases steer behaviour by manipulating them. That is a stronger claim than "statistical pattern matching" allows, and it is measured rather than argued.Anthropic, "Mapping the Mind of a Large Language Model", May 2024 — read at source 11 Sep 2026: anthropic.com, "We successfully extracted millions of features from the middle layer of Claude 3.0 Sonnet", and manipulating them "causes corresponding changes to behavior… they aren’t just correlated with the presence of concepts in input text, but also causally shape the model’s behavior." Source added 11 Sep 2026; this claim previously pointed only at another page on this site. Also covered as one of the four evaluation layers on how AI is tested.
Whether new abilities "emerge" suddenly is disputed
A 2022 study argued that some abilities appear only in larger models, so they “cannot be predicted simply by extrapolating the performance of smaller models”. A 2023 study replied that such jumps “appear due to the researcher's choice of metric rather than due to fundamental changes in model behavior with scale”. That disagreement should temper confidence in both directions: you cannot claim to know a system is "merely" predicting when researchers still dispute how its abilities grow.Wei et al., Emergent Abilities of Large Language Models, 2022; Schaeffer, Miranda & Koyejo, Are Emergent Abilities of Large Language Models a Mirage?, 2023; both read at source 16 Sep 2026. An earlier version of this section said "Nobody knows why capabilities appear when they do" and that this "is stated openly by the people building these systems", with no source.
Its sense of what it knows can misfire
Anthropic's interpretability team found something like an internal "don't know" default that recognising a familiar name switches off — and found it switching off wrongly. In their words, “such misfires can occur when Claude recognizes a name but doesn't know anything else about that person.” That is one concrete route to a fluent, confident, false answer, and a genuine limitation whatever one concludes about understanding.Anthropic, Tracing the thoughts of a large language model, 27 Mar 2025, read at source 16 Sep 2026. An earlier version of this section said "There is no internal flag for 'I know this'"; this research describes something like one, and how it fails.
The word "understand" was never precise
Much of this argument is definitional. Does a person who can use a word correctly in every context but cannot define it understand it? Does a chess engine understand chess? These questions were unresolved before AI existed, and pointing a new system at them did not resolve them — it just made the vagueness expensive.
There is real structure inside these systems, but their accounts of their own reasoning, and their sense of what they know, are both unreliable. Neither "just autocomplete" nor "it understands" survives the evidence.
Something is happening inside these systems that is more structured than "autocomplete" suggests and less settled than "understanding" implies. Anyone who tells you confidently which it is — in either direction — is describing their intuitions rather than the evidence.
Why it matters less than it feels like it should
For nearly every practical purpose, the question is irrelevant. Whether a system understands has no bearing on whether its output is accurate, and accuracy is verifiable while understanding is not. A doctor does not need to know whether a test comprehends anything; they need to know its false-positive rate.
Where it does matter is narrower and worth naming: questions of moral status, if they ever become live, and how much trust these systems are given — because "it understands" is used as an argument for reducing oversight, and it is not strong enough to carry that weight.
Do not let "it understands" stand in for checking. Judge the output, which you can verify, not the system's apparent grasp of it, which you cannot.
The dismissive answer protects something: if it is only autocomplete, then human thought remains categorically different and nothing needs re-examining. The credulous answer offers something else: significance, and the sense of standing near something historic. Both are emotionally load-bearing, which is why the argument generates more heat than the evidence supports. Noticing which one you want to be true is most of the work.
This site does not know, and says so. What it holds to is the practical consequence, which is unaffected either way: these systems produce confident output regardless of accuracy, so verification is not optional. That is true if they understand perfectly and true if they understand nothing. Everything else here is philosophy — worth having, and not a reason to skip the checking.