NB: I have been drafting an essay on “understanding understanding”, so the IAI article this is based on was perfectly salient. These are my reflections, expedited by ChatGPT.
Elan Barenholtz has written an intriguing piece for the Institute of Art and Ideas, LLMs show language does not describe reality. It is well worth reading, though apologies in advance: IAI has put much of it behind a paywall, humanity’s traditional method for ensuring that ideas about knowledge remain selectively accessible.
Barenholtz argues that large language models present an awkward problem for the familiar idea that language derives meaning by referring to, representing, or somehow grounding itself in an external reality. LLMs manipulate relationships among linguistic tokens without anything resembling ordinary human sensory access to the world, yet nevertheless produce remarkably coherent linguistic behaviour.
It doesn’t necessarily demonstrate that language has no referential function, nor that human cognition is simply next-token prediction writ biological. But it does show that a substantial amount of what we ordinarily take as evidence of linguistic understanding can apparently emerge without the machinery we previously assumed was necessary for it. That should make us suspicious of the machinery.
My own position begins slightly elsewhere. I don’t assume that humans possess privileged access to an ontic reality against which artificial systems can be judged deficient. Whatever reality may exist independently of us, humans encounter it through an architecture of mediation: sensory systems, neural processing, linguistic categories, cultural practices, institutional structures and prior expectations. We don’t encounter the proverbial chair as it is. We encounter what our particular architecture permits the chair to become for us.
In this respect, the difference between humans and current LLMs is not that humans are grounded whilst LLMs float mysteriously above reality. The difference is that the architectures of encounter are different. Humans possess direct sensory and motor coupling of kinds that text-based LLMs lack. We see, hear, touch, move, fall over objects and occasionally stub our toes against them, reality’s least philosophically subtle contribution to epistemology.
But this difference is technological rather than obviously metaphysical. Vision, audition, spatial navigation, tactile sensing and motor feedback can all be incorporated into artificial systems. Robotics already does precisely this. Such systems need not perceive as humans perceive, any more than bats, dogs or octopuses do. They merely require channels through which external constraints can affect their behaviour. Embodiment therefore changes what can constrain a system. It does not, by itself, establish ‘understanding’.
The more substantial difference between humans and contemporary artificial systems may instead be existential stakes.
Things can go badly for organisms in a particularly uncompromising sense. We can be injured, starved, abandoned, exhausted or killed. Our encounters occur against a biological background in which outcomes matter because continued existence itself is conditional. An LLM doesn’t presently possess comparable stakes. Nothing is obviously at risk for it. But this distinction is not really about cognition.
A bacterium has existential stakes. A language model can solve differential equations. Whatever distinguishes those two systems, it would be peculiar to call the difference ‘understanding’.
Stakes may explain motivation, vulnerability, affect or perhaps moral consideration. They do not automatically explain inference, linguistic competence or cognition. To claim otherwise requires an additional argument.
This is why LLMs pose such an interesting philosophical irritation. They do not prove that machines understand. They expose how poorly we have specified what we meant when we claimed that humans do.
If understanding requires successful inference, contextual discrimination, generalisation and appropriate linguistic response, artificial systems already make the boundary uncomfortable.
If understanding instead requires consciousness, intentionality, phenomenology, genuine meaning or some other internal property, then those terms themselves require defensible criteria. Otherwise the explanation merely replaces one contested word with another.
Barenholtz writes that ‘language doesn’t mean; it does’. I am sympathetic to the deflationary impulse. I would perhaps go one step sideways. Language needn’t provide transparent access to reality in order for reality to constrain language. Nor must a system represent the world perfectly in order to negotiate its encounters with it.
Humans and artificial systems may therefore differ profoundly without either possessing privileged access to things as they ontically are. The interesting question isn’t whether LLMs encounter reality as humans do. Plainly they don’t.
The interesting question is why we assumed that the human manner of encounter constituted the metaphysical entrance examination for understanding in the first place.