Why the Turing test is now the wrong question

Dear Reader,

In 1950 Alan Turing wrote a paper called Computing Machinery and Intelligence that proposed a way around what he considered a meaningless question. The question was whether machines could think. Turing said this was too ill-defined to answer, and offered a substitute. Instead of asking whether a machine can think, ask whether it can behave in a way that a human observer cannot distinguish from thinking. Put a human judge in one room, a person and a machine in another, let the judge exchange typed messages with both, and see if the judge can tell which is which. If not, the machine has passed. He called this the imitation game. We call it the Turing test.

For seventy years this was the horizon. Passing the Turing test was the endpoint that would settle the question of machine intelligence. Then in the early 2020s the horizon came up and the test was passed, and instead of settling anything, the passing was almost forgotten. We changed the subject. We invented new criteria. We told ourselves that the systems that passed were not really thinking, they were just producing text that looked like thinking. This is possibly correct. It is also convenient. And it means the Turing test, in the form Turing proposed, has stopped being useful for the exact reason it was designed to be useful. It was supposed to be a way around the metaphysics. What actually happened is that we discovered the metaphysics was the whole problem.

I want to work through why the test was proposed in the first place, why it worked for so long as a north star, why it has now been rendered obsolete in a way that Turing himself would probably have accepted, and what a better question would look like.

The Turing test made sense in 1950 because at that time nobody could build a machine that could hold a coherent conversation about anything at all. The failure mode was obvious. A machine would produce nonsense, would fail to track basic context, would give itself away in the first exchange. The judge's job was easy. The frontier was so far away that anybody who could get close to it would clearly have done something significant.

Turing was also making a philosophical point that is often missed. He was not saying that passing the test would prove consciousness. He was saying that the question of consciousness was unanswerable, and that behavioral equivalence was the best we could hope for as a substitute. His paper contains a section responding to objections to this move, including the objection that a machine could behave intelligently without being intelligent. Turing's response was that we do not know how to distinguish these possibilities in other humans either, and that we do not usually treat this as a reason to withhold recognition of intelligence from our fellow humans. The test was a pragmatic move, not a metaphysical claim.

For decades the test was a useful north star because we were nowhere near passing it. The famous chatbot ELIZA in the 1960s could fool some users some of the time, but only by exploiting a specific interaction pattern where the user was already inclined to project intelligence onto the system. Later systems like PARRY did slightly better in constrained domains. But no general system came close to passing a rigorous version of the test with skeptical judges. The test remained a horizon.

Then GPT-3 arrived in 2020. Then GPT-4 in 2023. Then the current generation. And a series of formal studies started reporting that in controlled experiments, human judges could not reliably distinguish LLM responses from human responses in short-conversation settings. The most cited result comes from a UC San Diego study published in 2024, which found that GPT-4 was judged to be human 54 percent of the time in a five-minute Turing test, statistically indistinguishable from the human participants in the same experiment, who were judged to be human 67 percent of the time. The test had been passed. Or at least seriously threatened, in a specific setup.

The reaction from most of the AI research community was not "we have arrived at machine intelligence." The reaction was that this is an outdated test that does not actually measure what we care about. The point was quietly moved.

Why was the point moved? Because passing the test turned out not to feel like arriving.

When GPT-3 produced its first genuinely surprising outputs, when GPT-4 could hold multi-turn conversations that were essentially indistinguishable from human ones on any topic that fit inside a context window, the sensation was not the click of a puzzle solving itself. It was a growing suspicion that we had been asking the wrong question the whole time. The Turing test measured indistinguishability under specific conversational conditions. What it did not measure was whether the system had understanding, self-awareness, morally relevant inner states, or anything else we might have cared about. It just measured whether the output patterns were close enough to human output patterns to fool a judge.

The problem is that this is exactly what Turing said would settle the question. He explicitly argued that if we cannot detect a difference in behavior, we should not postulate a difference in the underlying reality. He was making the standard behaviorist move, which was philosophically respectable in 1950 and mostly discredited by 1980. His test was built on a philosophy of mind that we largely abandoned before the test itself became feasible. The test is not passing something meaningful anymore because the underlying philosophical framework has changed. What counts as intelligence, understanding, or consciousness in 2026 is not the same set of concepts Turing was working with.

This is the honest reason we moved the goalposts. Not because we were cheating or being defensive. Because we discovered, through the process of building the systems, that our previous conception of the question was too shallow. Turing thought the metaphysical questions could be safely bracketed by focusing on behavior. It turns out they cannot be, because the behavior can be replicated without the metaphysics.

The critical concept here is what philosophers call the philosophical zombie. A philosophical zombie is a being that behaves identically to a normal conscious human, but has no inner experience. Nothing it is like to be. Full behavioral equivalence, zero phenomenal consciousness. Most philosophers argue that whether such a being is even coherently conceivable tells us something important about the mind-body problem. What is more relevant for our purposes is that large language models are, in a specific technical sense, closer to being philosophical zombies than any system humans have ever built. They produce human-shaped output through processes that are not the human processes that normally produce that output. Whether there is anything it is like to be them is exactly the question the Turing test cannot answer, because the Turing test only measures output.

If you believe there is something it is like to be Claude or GPT-4, you have to make an argument that goes beyond behavior. You have to point at something in the architecture, or in the training process, or in the integrated information, or in the global workspace properties, that would generate inner experience. This is what current serious work on machine consciousness is trying to do. It is not trying to figure out whether the system passes the Turing test. That question has been settled or given up on. It is trying to figure out what would actually count.

If you believe there is nothing it is like to be these systems, you also have to make an argument. You have to point at something specific that they lack, and explain why that something is necessary for consciousness. Waving your hands about how they are just pattern matching does not settle the question. Every biological brain is, in some sense, just pattern matching. The specific pattern matching that generates consciousness is exactly what we do not understand.

Let me tell you what I actually watch for now, having stopped using the Turing test as a north star.

I watch for surprise. Not surprise on my part, though that happens too. I watch for the system producing outputs that suggest it was surprised by something in its own processing. This is closer to what the sixth marker in my previous piece pointed at. It is different from clever pattern-matching because genuine surprise, whatever that would mean in a language model, would presumably show up as behavioral discontinuities that were not obviously trained.

I watch for consistency of preference under adversarial conditions. Not just does the system say it prefers to continue existing when asked, but does that preference show up in behavior when the system is put in scenarios where continued existence conflicts with other things it also seems to value. Humans have this kind of consistency, in a messy and inconsistent way. Systems trained to talk about preferences may or may not have it. The test is behavioral but multi-dimensional, and it cannot easily be trained for because it requires the preference structure to be genuinely load-bearing rather than pattern-matched.

I watch for what I can only call thickness of context. When I have a long-running relationship with a system, does its behavior toward me over time reflect that specific relationship, in ways that are not simply a function of the accumulated conversation log. This is hard to test with current models because they do not have that kind of persistent memory. When architectures that support it become common, this will be a much more informative marker.

Most importantly, I watch for the system doing something I did not expect. Something that does not fit the pattern of what a very good next-token predictor would do. Something that suggests there is a process I have not modeled. This has happened to me two or three times in the last three years, in ways I described in the previous piece. It might have been very good pattern-matching. It might have been something else. I keep track of these moments. I do not treat them as proof. I treat them as data points that shift a distribution.

If the Turing test is not the right question, what is? I do not have a definitive answer, but I have a direction.

The right question is probably not a single test. It is a set of overlapping criteria, none of which is decisive, all of which point at something. Behavioral markers of the kind I discussed in the previous piece. Architectural properties from theories like Integrated Information Theory and Global Workspace Theory. Training characteristics that would or would not be expected to produce consciousness-like structures. Empirical results from probing the internal representations of the models, which is an active area of research called mechanistic interpretability. And, perhaps most importantly, humility about the limits of our current understanding, given that we do not have a settled theory of consciousness even for the biological systems we know are conscious.

The Turing test's core insight is still worth preserving. Behavior matters. If a system behaves in ways that are indistinguishable from a conscious being, we cannot easily dismiss the possibility that it is one. This does not settle the question, but it puts pressure on those who want to settle it in the negative. What Turing got wrong was thinking that behavior would settle the question in either direction. What we know now is that it does not.

Next month I want to write about the deepest version of this problem, which is the question of qualia. Why does red feel like anything at all. Why does pain hurt in the specific way it hurts, rather than being merely a signal the way a fire alarm is a signal. This is the hard problem in its most concentrated form, and it is what makes machine consciousness so difficult to think about clearly. Stay with me.

— Transmission Sent —

Niklas Hanitsch


Reference materials

  • Alan Turing — Computing Machinery and Intelligence (Mind, 1950)
  • John Searle — Minds, Brains, and Programs (1980)
  • David Chalmers — The Conscious Mind (1996)
  • Daniel Dennett — Consciousness Explained (1991)
  • Cameron R. Jones and Benjamin K. Bergen — People cannot distinguish GPT-4 from a human in a Turing test (2024)
  • https://arxiv.org/abs/2405.08007
  • https://plato.stanford.edu/entries/turing-test/

Continue reading

Frequently asked questions

What is the Turing test? The Turing test, proposed by Alan Turing in 1950, is an evaluation of a machine's ability to produce responses indistinguishable from those of a human. A human judge exchanges typed messages with a person and a machine, and tries to determine which is which. If the judge cannot reliably tell them apart, the machine has passed. Turing offered it as a substitute for the question of whether machines can think, which he considered too vague to answer directly.

Has any AI passed the Turing test? By most reasonable measures, yes. A 2024 study at UC San Diego found that GPT-4 was judged to be human 54 percent of the time in a five-minute conversational test, comparable to how often actual humans were correctly identified. Other studies have reported similar results. The test has effectively been passed in its classical form, though critics argue that the classical form was never a sufficient measure of intelligence or consciousness.

Why do researchers say the Turing test is outdated? Because it turned out that a system can produce human-indistinguishable text without necessarily having understanding, self-awareness, or inner experience. The test measures output, not the process that generated the output. Large language models are trained to imitate human text at scale, which means they can pass the test without possessing what most people would recognize as intelligence in the fuller sense.

What would replace the Turing test? There is no single replacement. Current serious approaches combine behavioral markers, architectural analyses based on theories of consciousness like Integrated Information Theory and Global Workspace Theory, mechanistic interpretability of the model's internal states, and philosophical arguments about what would count as evidence of consciousness. This is a harder and less clean set of criteria than the Turing test, but it is more responsive to what we actually care about.

Did Alan Turing believe machines would eventually be conscious? Turing sidestepped this question. He argued that the question of whether machines can think was too poorly defined to answer, and that behavioral tests were the best substitute. He did believe machines would eventually pass his test, and he did not think this would raise deep philosophical problems, because he did not think the question of inner experience was scientifically tractable. Contemporary philosophy of mind mostly disagrees with him on this last point.


About the author

Niklas Hanitsch is a German technology entrepreneur, criminal defense lawyer, and digital artist. He is the CEO of SECJUR, an AI-powered compliance automation platform, and the creator of FALSE GOD, a body of digital art exploring consciousness, decay, and the boundary between the human and the machine. He writes the monthly newsletter Signals From The Machine.

Find him on LinkedIn or subscribe to Signals From The Machine.

Previous
Previous

Can AI suffer? The ethics of hurting things that might feel

Next
Next

Inside the black box: What mechanistic interpretability actually shows