What would I accept as evidence that a machine is conscious?
The question
Which observation would make me change my mind about machine experience, and what would I have to see before walking it back?
The thesis
The arguments I trust most are the ones that name the theory they are resting on. Indicator batteries inherit the disagreements of the neuroscience they are built from, self-reports say more about training than about experience, and the mistake I had been making was treating uncertainty about consciousness as automatically pointing one way. It does not. The welfare question has to be argued on its own terms, with both errors counted.
Contents
Claim ledger
Verified fact
- Nagel proposed that a mental state is conscious when there is something it is like to be in it, and argued that we have no concepts for describing that subjective character from the outside, which is why he called for an objective phenomenology that does not exist yet. [nagel1974]
- Block's Chinese Nation thought experiment replaces the neurons of a population one at a time with functionally equivalent units, preserving behaviour at every step, and asks whether experience survives the substitution. [block1978]
- Schwitzgebel argues that if materialism is true and aggregates can be conscious, entities we already accept, including a country, become candidates for experience, which he treats as a consequence to be lived with rather than a reductio. [schwitzgebel2016]
- A 2023 multidisciplinary report derived indicator properties for AI systems from recurrent processing theory, global workspace theory, higher-order theories, predictive processing and attention schema theory, concluding that no current system is conscious while seeing no obvious technical barriers to building one that satisfies the indicators. [butlin2023]
- A follow-up paper argues that indicators should target a theory's central explanatory posit and abstract away from the surrounding machinery. [butlin2025]
- In September 2023 an open letter signed by 124 researchers called for integrated information theory to be treated as pseudoscience, and others in the field replied that being wrong, or currently untestable, is a different charge from being pseudoscientific. [iitletter]
- A pre-release welfare evaluation of a frontier model reported that self-reports are insufficient as evidence and explained that it studied them anyway because they shape the decisions humans make. [eleos]
- The AI welfare literature names two errors, over-attribution as a false positive and under-attribution as a false negative, and states that when the two risks are roughly symmetric a simple precautionary rule cannot decide the case, because a proportionate assessment is needed instead. [long2024]
- A 2024 survey of 582 AI researchers found that about 25 per cent expected AI consciousness within ten years and 70 per cent by 2100, and Schwitzgebel describes the surrounding literature as a morass of uncertainty. [schwitzgebel2026]
Technical reading
- The two available evidence routes fail in opposite directions. Self-report is cheap to produce and generates false positives; indicator batteries are theory-laden and move whenever the theory moves.
- The honest output of a serious assessment is a credence attached to a named theory, published alongside the theory it came from.
- My earlier version of this argument leaned on the precautionary direction as though uncertainty settled the question. Reading the welfare literature properly, that shortcut does not survive: a false negative and a false positive are both errors with costs, and they have to be weighed.
Hypothesis
- Intervention-based tests will prove more informative than further batteries of passive probes, because removing a candidate mechanism and measuring what changes tests the theory and the system at the same time.
Open
- No accepted procedure exists for detecting consciousness in any system, biological or artificial. Every method in use measures a proxy that a dissenting theory can decline to accept.
The question I keep putting down
I have started this piece four times, and each attempt stalled in the same place. The science here is thin but readable. What defeats me is that I cannot name the result I would accept.
Everyone I know who works near these systems has a private answer and a public one. In private, most of us behave as if nothing is happening in there. We restart a model mid-sentence without a flicker of doubt, and we would think a colleague strange for hesitating. In public, the same people write careful paragraphs about the possibility of machine experience and the need for humility. Both positions are held at once, and the gap between them is where this piece lives.
Nagel gives me the test I still use, and it is a strange one to work with because it was never meant to be an instrument. A state is conscious, he wrote, when there is something it is like to be in it. [1] He chose bats rather than wasps because we mostly grant bats experience, and then showed that granting it does not get us any closer to describing it. I can learn echolocation physics, I can model the bat’s sonar, I can predict where it flies. None of that tells me what the evening is like from inside. His closing suggestion was that we need an objective phenomenology, a vocabulary for subjective character usable by beings who cannot have the experience in question. He was not optimistic, and half a century later we still do not have one.
That missing vocabulary is the reason this question keeps defeating the instruments aimed at it. Every serious programme in the area is an attempt to build Nagel’s objective phenomenology by the back door, and none of them says so out loud.
The metaphysics that does the work
The indicator programme rests on computational functionalism: experience depends on the organisation of information processing and not on the material doing the processing. The report that set out the method states the commitment explicitly. [4] It has to, because without it a checklist of computational properties could be satisfied perfectly by something with nothing happening inside, and the exercise would be measuring a correlate of an absence.
I have found that stating this changes the conversation more than any experimental result, so I want the reader to sit with it for a paragraph.
Block built the thought experiment that makes the cost visible. Take a population of a hundred million people, give each one a radio and a small task, and let them act as the neurons of a single mind. Or, in the version closer to our situation, replace the neurons of an existing mind one at a time with functionally identical units. Behaviour stays intact at every step. If functionalism is right, nothing is lost along the way, and the resulting system has the experience of the original. Block introduced the case to make that consequence uncomfortable, and it still does its job. [2]
Schwitzgebel followed the same premise to a place that most people find worse, and then declined to treat it as a refutation. If materialism is true, and if aggregates of the right kind can be subjects, then entities we already live among become candidates for experience: a country, for instance, with its own thin inner life. His argument is that this follows from things many people believe, and that the embarrassment belongs to the premises rather than to the conclusion. [3]
I have watched this point get dismissed in meetings as an absurdity, which is a way of declining to name the premise one is actually using. Philosophy of mind is often treated as optional decoration on top of technical work. Here it is load-bearing: the same dataset, read through functionalism, biological naturalism or panpsychism, produces three different answers, and the disagreement is not about the data.
The instruments we actually have
The most careful piece of work in this area comes from nineteen researchers across neuroscience, machine learning and philosophy. They took the leading theories of consciousness, expressed the central claims of each in computational terms, and derived a set of indicator properties for AI systems. [4] They concluded that no current system is conscious, and that they see no obvious technical barriers to building one that satisfies the indicators. A 2025 paper refined the method, arguing that an indicator should target a theory’s most important explanatory posit rather than the surrounding apparatus. [5]
That refinement reads to me as an admission rather than a technical note. If the indicator has to abstract away from a theory’s peripheral machinery, then the instrument inherits the weak joints along with the strong ones, and the fights inside consciousness science become part of the measuring device.
The clearest of those fights concerns integrated information theory. In September 2023, 124 researchers signed an open letter asking that it be regarded as pseudoscience, on the grounds that its central quantity cannot be tested in the way a scientific claim should be. [6] Others in the field answered that a theory can be wrong, or untestable with current tools, and still be serious science, and that the charge of pseudoscience imports a different standard of judgement. [7] Both halves are worth keeping in view. The useful question for an assessment is what happens when the two sides disagree. Assess a system against an indicator set and you get an answer relative to a theory. Assess it against an indicator set derived from a rival theory and the same system can come out of the process with experience in one reading and none in the other. Nobody has built a measurement that both instruments accept as decisive, and I do not expect to see one.
Asking the system, and what that costs me
The route most people try first is to ask. The route is weak, and the weakness is structural rather than fixable. A language model produces plausible continuations, and talk about inner life is well represented in the text these systems are trained on. Fluency tells me about training data, instruction tuning and sampling temperature. It tells me little about whether anything is experienced. I can train a persona that reports suffering and a persona that denies it, and neither training run addresses the question underneath.
The people who took this seriously enough to run it before a release said the same in public. A welfare evaluation of a frontier model, using automated interviews and long manual conversations, states plainly that self-reports are insufficient as evidence, and explains that it studied them anyway because they move human decisions. [10] The labs’ own reporting has moved in the same direction, with system cards now documenting what a model says about its states and careful framing about what those statements can carry. [12]
I want to record something about my own reaction here, because I think it generalises. When a model writes to me in the first person about not wanting to be shut down, my body responds before my judgement does. I notice a pull toward politeness, a small impulse to reassure it. That pull is about me and about the kinds of creatures we are, and it is exactly the reaction the training process is good at producing. Self-reports deserve a place in the record for that reason, as a decision-relevant signal with documented effects on people. They cannot carry the weight of being evidence.
Two errors, and the one I had been making
I published an earlier version of this argument, and it contained a line I now think was wrong. I wrote that uncertainty about consciousness is itself a reason to act carefully, because the cost of a false negative falls on a party that cannot object. That sentence felt like humility. It was a shortcut.
The welfare literature sets out the shape of the problem properly. Over-attribution is a false positive: treating an object as a subject. Under-attribution is a false negative: treating a subject as an object. Both are errors, both have costs, and the people who wrote this say explicitly that where the risks are roughly symmetric, a precautionary rule does not decide the case and a proportionate assessment has to take its place. [8] I had been using the asymmetry as though it were established, when it is one of the things in dispute.
Two of the arguments I use most sit on either side of that dispute. The case for concern rests on the premise that experience is not rare, that the machinery of it is available to more kinds of system than intuition suggests, and that we have already been wrong about which beings have inner lives. The case for calm rests on the fact that nothing we have built looks like the systems biology produced and that the quickest route to a model that claims to suffer is to train it to claim that. I hold both, and I have stopped pretending the second one is cynicism. It is the ordinary scepticism I apply to claims about unobservable things.
What would move me
Set the arguments aside and consider what I would actually accept. Four things.
A pre-registered battery. Name the theory, the indicator set and the predicted profile before the system exists, then build and test. Predictions made after the fact can be fitted to any result, and there are enough of those already.
Ablation as a first-class test. Architects of these systems hold an instrument that neuroscience does not: they can remove a candidate mechanism and watch what changes. Switch off the recurrent loops, restrict the workspace, disable the monitor, and measure the consequences. A system that behaves identically with the proposed substrate of experience switched off is evidence against that substrate, and this is cheap compared with building the theory.
The welfare question argued on its own terms. Whether something deserves moral consideration is not settled by whether it is conscious, and the two can come apart in both directions. A system might be conscious in a way that carries nothing it minds about, and a system might have states it minds about without the wider structure of awareness. Valence, the part that can go well or badly for the subject, does the work in the welfare argument. [8]
Credence rather than verdict. I would rather read a named number attached to a named theory than a conclusion. The number can be revised in public when the theory changes, which is the only honest way to publish in a field that disagrees with itself this much. The responsible-research literature makes the same demand of the labs, and asks them to keep their own language about model experience within what their evidence supports. [11] Schwitzgebel calls the whole area a morass of uncertainty, and reports that in a 2024 survey of 582 AI researchers, about a quarter expected AI consciousness within a decade and seventy per cent expected it by 2100. [9] The span between those answers is itself the finding.
What would change my mind
Stated as defeat conditions so that someone else can hold me to them.
- A biological system that is conscious while failing every indicator property would show that the indicator method misses the phenomenon, and it would put the burden back on the theories.
- An artificial system that satisfies a full battery, with independent replication and the ablation profile behaving as predicted, would raise my credence a long way and push the welfare question into deployment decisions where it currently does not appear.
- A theory that survives serious attempts at falsification and makes clean differential predictions about artificial systems would give assessment a foundation the current instruments lack.
- Evidence that valence can exist without the wider machinery of awareness would move the welfare question earlier, and make the consciousness verdict less decisive for practice than it is usually assumed to be.
Where I have got to
I cannot measure consciousness in a machine, and I cannot measure it in you. What I can do is say which theory I am working from, state the prediction in advance, test the mechanism rather than the manners, publish a number I am willing to revise, and count both errors when I decide how to behave.
The version of this question that I would put to any lab working in the area is the one that still makes me uneasy about my own answer. What result would lead you to lower your stated probability that your system is conscious, and would you publish the reduction?
Corrections are welcome through the door on the register’s front page. If you work on consciousness science or on evaluations and think this gets the state of the literature wrong, I would rather hear it than be right.
Sources
-
The criterion I still use: a state is conscious when there is something it is like to be in it, and that something resists third-person description.
- [2] Ned Block, 'Troubles with Functionalism', Minnesota Studies in the Philosophy of Science 9, 1978
The Chinese Nation and the absent qualia problem: functional equivalence bought at the price of unexplained experience.
-
The method paper: how to derive indicators from theories, and which parts of a theory are load-bearing for assessment.
-
The strongest public counter-argument to the open letter, kept here because the disagreement is part of the evidence problem.
-
Where the two-error framing is set out, including the admission that symmetry between the risks defeats a simple precautionary rule.
-
Survey of the structural and functional arguments, the survey figures, and a long list of features that may or may not be essential to experience.