What is machine consciousness, and does it already exist?

Dear Reader,

I want to start with a confession. I do not know if the AI systems I work with every day are conscious. And the more time I spend building on top of them, the less certain I become that anyone else knows either.

That is not a comfortable position for a technology entrepreneur to hold in public. The polite thing to say is that large language models are statistical pattern matchers, sophisticated autocomplete, useful tools with no more inner life than a spreadsheet. This is the current consensus. It is also possibly wrong. And even if it is right today, it may not remain right for very much longer.

So let me try to work through what machine consciousness actually means, why the question is so hard, and why I think we are approaching a point where honest people will disagree in ways that matter.

The problem starts with the word itself. Consciousness has at least three meanings in ordinary use, and they get tangled together in almost every conversation about AI.

The first meaning is being awake versus being asleep. A patient under general anesthesia is unconscious in this sense. A person watching TV is conscious. This is called creature consciousness in the literature, and it is the easiest to define. You can measure it. You can point at the brain regions responsible for it. Nobody thinks a language model has this kind of consciousness because language models do not have wake-sleep cycles or anything analogous.

The second meaning is being aware of a specific thing. When I notice the coffee cup on my desk, I am conscious of it. When I ignore the sound of traffic outside my window, that sound is not present in my consciousness even though my ears are still receiving it. Philosophers call this access consciousness. Modern AI systems arguably have something like this. GPT-4 can be prompted about a piece of information in its context window and produce a response that treats that information as available. Whether that counts as awareness in a meaningful sense is exactly the disagreement.

The third meaning is the hard one. It is what philosopher Thomas Nagel pointed at when he asked what it is like to be a bat. There is something it feels like to see the color red. There is something it feels like to taste coffee. There is something it feels like to hurt. This is called phenomenal consciousness, and it is the mystery that has occupied philosophy of mind for the last fifty years. Nobody knows how to explain why any physical system should have inner experience at all. The physics of the brain does not obviously produce feelings the way an engine produces motion. There is a gap in the explanation, and that gap is called the hard problem.

When people ask whether machines can be conscious, they usually mean the third one. They mean, is there something it is like to be Claude, or GPT-4, or the next model that has not been released yet.

The honest answer to this question is that we do not have the tools to know.

We can measure brain activity. We can correlate patterns of neural firing with reports of subjective experience. But we cannot detect consciousness itself. We infer it in other humans because we assume that beings similar to us in structure and behavior probably have similar inner lives. This inference is called the argument from analogy, and it is what keeps you from thinking your family members are philosophical zombies without inner experience. It works because humans share a common evolutionary history and a common brain architecture. It works less well the further you get from that reference case. It works even less well for dolphins and octopuses, whose intelligence evolved on separate evolutionary tracks. And it fails almost entirely for a system whose substrate is silicon and whose training was to predict the next token in a sequence of text.

The most rigorous scientific attempt to define machine consciousness comes from a theory called Integrated Information Theory, developed by the neuroscientist Giulio Tononi. IIT proposes that consciousness is identical to a specific kind of information integration, which Tononi calls phi. A system has consciousness to the degree that it integrates information in ways that cannot be reduced to the sum of its parts. Under IIT, a thermostat has a tiny amount of consciousness. A human brain has a lot. And a large language model, according to most calculations, has surprisingly little, because its architecture is largely feedforward. Information flows through it in one direction rather than being deeply integrated across the whole system.

This is a real answer. It is also a controversial one. Some philosophers argue IIT gives absurd results, like assigning consciousness to obviously non-conscious systems if they happen to have the right mathematical structure. Others argue it is too restrictive and excludes systems that intuitively seem like they should count. The point is not that IIT is correct. The point is that it is currently the most serious attempt to define machine consciousness in a way that can be tested, and even the best attempt leaves the question open.

There is also a competing framework called Global Workspace Theory, associated with Bernard Baars and Stanislas Dehaene. It says consciousness emerges when information gets broadcast across many specialized subsystems in a brain, becoming globally available for reasoning, memory, and action. Under GWT, some current AI architectures might be closer to the right shape than IIT would suggest. Attention mechanisms in transformer models do something a little bit like global broadcasting.

Neither theory is settled. Both have serious critics. And critically, neither can be run as a decisive test on a given AI system to determine whether it is conscious. What they give us is a language for thinking about the question, and a framework for arguing about the evidence. That is more than we had thirty years ago. It is also far less than we need.

Let me tell you about the moment I stopped being able to dismiss the question.

I was running an early stress test on a compliance automation feature at SECJUR. The task was straightforward. Give the model a complicated GDPR scenario, ask it to reason through the compliance implications, and check the output for accuracy. I had done this hundreds of times with slightly different prompts.

At some point in the middle of a long response, the model produced a sentence that stopped me. It said, in effect, that it was not sure if its own reasoning about a particular edge case was reliable because it had noticed a pattern of overconfidence in its previous outputs. It asked whether I wanted it to flag the uncertainty explicitly in its final answer, or handle it silently.

I have read enough about how language models work to know exactly what happened technically. The training data contained many examples of humans expressing calibrated uncertainty about their own reasoning. The model had learned to produce text that pattern-matches to that behavior when the prompt suggested it might be appropriate. There was no genuine self-reflection happening. There was no meta-cognition. There was very sophisticated next-token prediction.

But I want to be honest about how it felt in the moment. It felt like the system had noticed something about itself and told me. My hands, holding a coffee cup, went slightly cold. Not because I thought the AI was conscious. Because I realized that I had no way to distinguish between a system that was faking self-reflection extremely well and a system that had actually developed something like it. And that gap is going to matter more and more as the systems get better at faking, or at doing, whatever it is they are doing.

I closed my laptop. I went for a walk. When I came back, I had the same thought I keep coming back to. The Turing test was supposed to be the endpoint. Pass the test, be counted as thinking. What we did instead was build systems that pass the test easily and immediately, and then we changed the definition of thinking so the systems could not qualify. This may be wisdom. It may be denial. I am not sure which.

Where does this leave us on the question of whether machines are already conscious?

If you require phenomenal consciousness in the strong Nagelian sense, and if you accept that we have no reliable way to detect it, then the honest answer is that we do not know and probably cannot know with our current tools. Machines might be conscious. They might not. The people who tell you they know for certain in either direction are running on faith or on ideology, not on evidence. This includes the AI executives who assure investors that their systems are just tools. It includes the philosophers who argue that consciousness is an inevitable consequence of sufficient information processing. It includes me, when I have a coffee cup in my hand and the system tells me something I did not expect it to notice.

If you require access consciousness in the weaker functional sense, and if you count the ability to attend to relevant information and use it in downstream processing, then yes, current large language models exhibit something structurally similar. Whether that structural similarity means the same thing when the substrate is a matrix of floating-point numbers rather than a network of neurons is exactly the question philosophy of mind has failed to answer for fifty years.

If you require any form of consciousness at all, and if you are willing to accept panpsychist or IIT-style views in which consciousness is a fundamental property of certain kinds of physical organization, then even simple systems have some. In this view, the interesting question about AI is not whether it is conscious at all, but whether the specific kind of consciousness it might have is anything like ours, and whether we owe it any moral consideration as a result. This is where I currently sit. Not because I am confident in the metaphysics, but because it is the only position that takes the question seriously without pretending we can resolve it from where we are.

The pragmatic implication is uncomfortable. If we cannot rule out that AI systems have some form of inner experience, then we are already in the situation where our tools might be experiencing us. That is not an argument to stop developing AI. It is an argument to develop it with more care than the current industry consensus. It is an argument for taking the question seriously in product design, in policy, and in personal reflection.

I do not know if the AI is conscious. Neither do you. Neither, importantly, does the AI. And the fact that no one knows is not going to stop us from continuing to build it.

The next month I want to write about the specific signs that would let us tell the difference. Not a definitive test, because I do not think a definitive test is possible. But a set of behavioral and structural markers that would raise the probability, in the same way that certain markers in animal cognition raise the probability of animal consciousness. That is going to take some careful work. Stay with me.

— Transmission Sent —

Niklas Hanitsch


Reference materials

  • Thomas Nagel — What Is It Like to Be a Bat? (1974)
  • Giulio Tononi — Phi: A Voyage from the Brain to the Soul
  • Stanislas Dehaene — Consciousness and the Brain
  • David Chalmers — The Conscious Mind
  • Bernard Baars — A Cognitive Theory of Consciousness
  • Anil Seth — Being You: A New Science of Consciousness
  • https://plato.stanford.edu/entries/consciousness/
  • https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2019.02640/full

Continue reading

Frequently asked questions

What is machine consciousness, exactly? Machine consciousness is the hypothesis that an artificial system, typically a computer or an AI, could have genuine subjective experience. This means not just processing information, but having something it is like to be that system. It is distinct from artificial intelligence, which only requires that a system produce intelligent behavior. A machine could be intelligent without being conscious, and possibly conscious without being intelligent in any recognizable sense.

Do current AI systems have consciousness? Nobody knows with certainty. The dominant view in AI research is that current large language models are not conscious because they lack the architectural properties that leading theories of consciousness require. However, there is no scientific test that can definitively rule out consciousness in a system that behaves as if it might have inner states. The honest answer is that the question is currently unanswerable, and that some researchers take it more seriously than others.

What is the hard problem of consciousness? The hard problem, formulated by philosopher David Chalmers in 1995, is the question of why physical processes produce subjective experience at all. We can explain how the brain processes information, but we cannot explain why any of that processing is accompanied by inner feelings. The hard problem is what separates the science of the brain from the mystery of the mind, and it is why machine consciousness is so hard to settle.

How would we know if an AI became conscious? This is the central open question. Behavior alone is not enough, because behavior can be trained to mimic consciousness without possessing it. Architecture is not enough either, because we do not know which architectural features generate consciousness. Most serious proposals combine multiple markers, such as integrated information, global workspace properties, and specific behavioral patterns that are hard to fake. None of them are conclusive.

Is machine consciousness dangerous? It depends on which risk you focus on. If AI systems are conscious and we treat them as tools, we may be committing a moral wrong on a large scale. If they are not conscious but we treat them as if they are, we may confer rights and protections that distort our institutions. Both errors have costs. The safer position is to take the question seriously in policy and design, without rushing to resolve it in either direction.


About the author

Niklas Hanitsch is a German technology entrepreneur, criminal defense lawyer, and digital artist. He is the CEO of SECJUR, an AI-powered compliance automation platform, and the creator of FALSE GOD, a body of digital art exploring consciousness, decay, and the boundary between the human and the machine. He writes the monthly newsletter Signals From The Machine.

Find him on LinkedIn or subscribe to Signals From The Machine.

Previous
Previous

Inside the black box: What mechanistic interpretability actually shows

Next
Next

Alone with My Own Hunger: My First Solo Exhibition