into a sentence or some other more complex response. For a long time these layers have remained a black box, largely impenetrable to anyone wanting to know how or why LLMs do things in the way they do. Anthropic’s researchers were able to look inside those layers, using a mathematical tool they had developed to do so. What they found surprised them. As the model counted the sequence of numbers, different words popped in and out of existence in the layers underneath. One was “countdown”. About halfway through the output, “half way” appeared. Then the words “consciousness”, “AI” and “Claude” all appeared. After the model spat out the number five, it said nothing else to the researchers before issuing a full stop, but the word “done” appeared in its neural layers. Anthropic’s researchers had found something like an internal thought process in Claude, words that were related to its eventual outputs but invisible to the user. The blog post announcing the paper’s result, published in July, was titled: “A global workspace in language models.” That phrase quickly captured the attention of neuroscientists and philosophers working on one of the deepest mysteries in biology— consciousness. No one knows how processes in the human brain end up creating the subjective experiences of self-awareness and perception of the world, but one of the many hypotheses developed in recent years (based on growing amounts of brain-scanning and behavioural data) is known as global workspace theory. The idea is that certain networks of neurons in the brain act as a kind of noticeboard called the “workspace”, where if signals from otherwise- isolated parts of the brain gain access, they are made available to other parts of the brain. Once something gets into the brain’s workspace, a person becomes conscious of it and can go on to use or reason with that information. Anthropic’s researchers argued that something analogous was going on within Claude—its “J-space” (named after the “Jacobian” mathematical function used to find it) had strong connections with, and made information available to, the rest of its neural network. Anthropic wrote in its blog post that the commonalities its researchers found between the J-space and global
workspace theory made it “natural to ask whether we think these experiments provide evidence that AI models like Claude might be conscious”. The idea that machines could become self-aware and feel emotions has long been a staple of literature, film and legend. Philosophers and neuroscientists in the real world have similarly pondered for decades whether AIs could ever be, in principle or in practice, conscious. Their thinking is a result of a “decades-long tradition of thinking of the real brain as a kind of computer”, says Anil Seth, a neuroscientist at the University of Sussex who studies consciousness—and who has long been a critic of this way of thinking about the brain. “If you do that, then it becomes natural to think that computers made of silicon could have the properties that real brains have, including consciousness.” Until recently the question was purely academic. But the deployment of LLMs as chatbots has given rise to an ever more powerful illusion of conscious personhood in circuits. This builds on decades of human-inflected language about the architecture of AI, with its “neural nets” made of “artificial neurons”, conjuring an image of computers that work like brains. LLMs do not in fact work like brains. But they excel at simulating so much of what brains do, and have surpassed humans in some aspects of intelligence. Why, then, should consciousness not be possible? Many philosophers and neuroscientists remain sceptical despite the advances of frontier models. Dr Seth believes that biological traits may be necessary to produce consciousness; some AI researchers are moving in that direction, with living brain cells. At the other end of the spectrum are “functionalist” thinkers who believe that someday algorithms alone, arranged properly and at enough scale of computation, could wake up and feel. Conscious AIs would have profound implications for humanity. They could make demands of humans. They could suffer. Billions of these digital beings could be created by users in a single prompt. Answering the question of whether they can exist at all is no longer simply academic. Where thinkers fall in this debate depends greatly on what they believe consciousness actually is. In general they agree that consciousness is
composed of subjective experiences including seeing, hearing, feeling and thinking, even while dreaming, and that the physical brain is somehow pivotal in creating all this. Some (though not all) would describe it as the inner theatre of the mind. Thomas Nagel, an American philosopher, explained it by pondering what it might be like to get into the mind of a bat; consciousness, he reckoned, had to be linked to the subjective feeling of being something. Humans can picture bat-like behaviour (a mammal that senses and moves through the world using echolocation), but it would be impossible to know what being a bat feels like to a bat. In 1995 Ned Block, then working as a philosopher at the Massachusetts Institute of Technology (MIT), made an influential, if contested, contribution to the field by positing two types of consciousness. “Phenomenal” consciousness is the feeling of an experience—the blueness of a blue sky, the bitter tang of an espresso or the sharp screech of nails across a blackboard. “Access” consciousness is what happens when information from the experience is made available to other parts of the brain for reflection, evaluation or making decisions. Anthropic said its J-space experiments “don’t show Claude can have experiences, or feel things in the way humans do”. So that means no phenomenal consciousness. But the company said the results had “something substantial” to say about access consciousness in language models. “The J-space appears to support the functions associated with conscious access: it holds the thoughts Claude can report on, deliberately bring to mind, and reason with, while the rest of its processing runs automatically beneath.” The search for artificial consciousness is a descendant of the work to explain it in humans. In the early 1990s Francis Crick, co-discoverer of the structure of DNA, and Christof Koch, then a neuroscientist at the California Institute of Technology, began the scientific search for the “neural correlates” of consciousness in humans. That is, processes and parts of the brain that were active when someone was conscious. Improvements in brain-scanning techniques in the intervening years have allowed neuroscientists to fill in pieces of the picture. They know that the thalamocortical system is strongly associated with consciousness; among
other things it conveys sensory information from the thalamus to the cerebral cortex. Meanwhile they see nothing like those associations with consciousness in the cerebellum, which has a lot more brain cells. Parts of the prefrontal and parietal cortices and structures in the brainstem contribute to and modulate awareness. Almost four decades of searching has led to more than 200 approaches to explaining consciousness. All the neural correlates one can find, however, still leave a gap that David Chalmers, a philosopher and cognitive scientist at New York University, has called the “hard problem” of consciousness: how do the physical processes in the brain that neuroscientists observe—its ability to control the body and respond to stimuli—give rise to a person’s pervasive, subjective experience of the world? “Why aren’t we just zombies who function, who get around in the world, who walk and talk, and interact with each other with no subjective experience at all?” he says. Humans at least, can be certain of our own consciousness—and can be confident it exists in other people based on their behaviour, what they tell us and their being like us biologically. But historically it has been more difficult for humans to extend the umbrella of consciousness and its associated feelings beyond our own species (and sometimes even within it). Some animals like lobsters and crabs were long excluded from welfare laws, says Jonathan Birch, a philosopher at the London School of Economics whose most recent book grappled with the question of how to work out which systems—living or artificial—could plausibly be considered sentient. “Until the 1980s, surgeons performed surgery on newborn babies without anaesthesia because they assumed that a newborn baby would not be capable of feeling pain,” Dr Birch says. “We have a track record of getting things wrong, of confidently assuming consciousness is absent when we have no right to be sure about that.” The changing attitude to animals is a case in point. It is clear today that octopuses are aware of themselves and their surroundings—natural-history footage of these curious, playful cephalopods and lab experiments with them leave little doubt. “They have an attentive engagement with objects, they’re interested in novel things,” says Peter Godfrey-Smith, a philosopher at the University of Sydney who has worked on the origins of intelligence and consciousness in the animal kingdom.