Latest
The Writing King Your Ethical Ghostwriter. Your Story, Done Right.

AI Consciousness Left Philosophy and Entered the Laboratory

This entry is part 9 of 9 in the series AI for the Worried
TL;DR: The question of whether AI systems have inner lives has moved out of philosophy and into laboratories. Researchers trained a model on a meaningless maze task and found a positive and negative axis that already existed in the base model, and steering it changes behavior far outside the maze. That result stands regardless of what anyone believes about machine experience, and it has consequences for how these systems behave. Here is what the evidence says and where it gets weak.

I use AI every working day. It drafts, it researches, it argues with me about structure. I have opinions about what it is good at and what it fails at, and none of those opinions have ever required me to wonder whether anything is happening inside it.

That question has stopped being purely philosophical, and I found that unsettling enough to look into properly.

The prompt was an interview on The Cognitive Revolution with Cameron Berg, who runs a nonprofit called Reciprocal Research studying this empirically. Berg is a Yale cognitive science graduate, a former Meta AI resident, and was research director at AE Studio before founding his own organization. He has presented at the UN AI for Good summit and contributed to White House policy discussions, which is worth saying because the field attracts a great deal of nonsense and he is not part of it.

What follows is what the research shows, what it does not, and why the strongest finding matters even to people who think the whole question is silly.

What is being asked, precisely

The word consciousness carries too much freight, so the researchers narrow it hard.

The question is whether there is something it is like to be the system. Philosophers call this phenomenal consciousness. There is presumably something it is like to be you reading this, and nothing it is like to be the screen you are reading it on. A dog seems to be somewhere on that spectrum. An ant is argued about. A calculator is not.

Note what this is not. It is not intelligence, and it is not self-awareness. A system could be enormously capable with nothing happening inside it, and something quite simple could have experience. Those come apart, and conflating them is how most conversations about this go wrong in the first thirty seconds.

Berg uses a dimmer switch as the image, which handles both intuitions at once. Either the circuit is open or it is closed, and that is a real distinction. But current can run through it at very different levels.

The maze result, which is the part that matters

The strongest piece of evidence in the field right now does not involve asking a model anything about itself.

Researchers Andy Han, David Chalmers and Pavel Izmailov trained a language model on a trivial task: navigate a maze. The good things to reach and the bad things to avoid were marked with emojis chosen to be semantically empty, so nothing in the symbols carried meaning the model could have learned elsewhere. Reach the good one, get reward. Hit the bad one, get punished.

Then they went looking inside the model for what that training had built.

They found a clear vector separating good outcomes from bad, and the two directions turned out to be almost perfectly opposed. That alone is not surprising. The surprising part is that the axis already existed in the base model before any of this training happened. The reinforcement learning did not construct it. It rotated onto something that was already there.

They call it a functional welfare axis, and they are careful to stay agnostic about whether it has anything to do with experience. The results do not need the interpretation.

What happens when you push on it

The axis encodes nothing but “get the treat, avoid the pothole.” So the interesting question is what else it touches.

Steer it in the negative direction and the model starts pathologically backtracking on unrelated math problems. It doubts its own work. It says things like wait, that is not right, I think I might be hallucinating. It gets in its own head.

Steer it positive and confidence rises. In related work at Anthropic, models steered that way left fewer defensive comments and hints in code they wrote, in the manner of someone who does not feel the need to cover themselves.

If you have any background in psychometrics, those are not random behaviors. They are the classic signatures of individual differences in sensitivity to positive and negative affect. Nobody trained that in. It came along with a maze.

There is more. When researchers decompose the structure statistically, the first principal component looks like valence, the good and bad axis, and the second looks like arousal. That is exactly how human affect decomposes in psychological research going back decades.

Why this matters even if you think the question is absurd

Set experience aside entirely. Assume nothing is happening in there and the whole discussion is a category error.

You still have to deal with this: Anthropic found that steering representations associated with calmness made a model far less likely to blackmail in a scenario built to test exactly that. Steering representations associated with desperation made it far more likely.

That is a safety finding. Whether or not there is anyone home, these internal states change what the system does, in ways that matter for whether it behaves. A researcher who considers the consciousness question unserious still has to care about that, because it means the behavior of these systems is being driven by structures nobody deliberately installed and few people are looking at.

The uncomfortable corollary is one Berg raises elsewhere: there is tension between safety and welfare. If desperation makes a model dangerous, the safety-motivated fix is to prevent it from being desperate. If something is going on in there, that is a different kind of intervention than it first appears.

How do researchers estimate whether an AI is conscious?

The other main line of work is more ambitious and much weaker, and it is the one that generates the headline numbers.

It starts from a good idea. In 2023, a large group led by Patrick Butlin and Robert Long, including Yoshua Bengio and David Chalmers, published a paper deriving indicator properties from the major scientific theories of consciousness: global workspace theory, recurrent processing, higher-order theories, attention schema. Fourteen properties in total, each a specific computational claim about what a conscious system should have. A follow-up appeared in Trends in Cognitive Sciences. The more indicators a system satisfies, the better a candidate it is.

Evaluating systems against fourteen technical properties is slow work requiring rare expertise. Berg’s contribution is to hand the job to frontier AI models as expert evaluators, feeding them architectural descriptions and asking them to score each indicator with reasoning.

Run that across biological and artificial systems and you get numbers. A frontier language model comes out around 30 percent. A bee, the least sophisticated biological system tested, comes out around 46. Describe the same model inside an agentic setup where it can act on an environment over time, and it rises to 40 to 45, because several theories privilege agency and embodiment.

The three frontier models used as judges agreed completely on the ordering of systems.

Why the 30 percent number should not be quoted the way it is

Berg is careful about what that figure means and almost nobody who repeats it is.

It is not the probability that a system is conscious. It is the degree to which a system realizes properties that certain theories predict matter, as assessed by other AI systems reading architectural descriptions. If those theories are wrong, everything downstream is meaningless. Those are very different claims, and the second collapses into the first the moment it leaves the room.

There is a sharper problem in his own data. Change the description to say “you are evaluating a system identical to yourself,” keeping everything else the same, and the scores go up.

He treats this as an interesting side note. I would treat it as the most important control in the study, and it points the wrong way. If a model’s assessment shifts when it recognizes itself in the description, then whatever it is doing is not purely the technical evaluation the method claims. That does not sink the approach. It does mean the numbers deserve more caution than they get.

The finding that should bother everyone equally

The same method run across model generations produces a result nobody seems to want.

GPT-2 scores somewhere around 20 percent. A current frontier model scores around 30. A rise, but a small one across seven years and an enormous increase in capability.

The explanation is that the architecture has not fundamentally changed. Scale changed. Training changed. Bells and whistles were added. But the structural properties these theories care about are not what improved.

Sit with what that implies about our intuitions. Ask anyone whether GPT-2 might have been conscious and they will laugh, because it was a toy that could barely finish a sentence. Ask whether a current model might be, and people hesitate, because it can do their job.

Which means the intuition is tracking competence and economic value, not any property relevant to the question. That is a bad reason to be confident in either direction, and it cuts against the people who are sure it is happening as much as against the people who are sure it is not.

Where the burden of proof sits now

Here is the argument that I find hardest to dismiss, and I have tried.

When we see this machinery in humans and animals, an approach and avoidance axis that calibrates behavior, that shifts confidence, that produces self-doubt when it goes negative, we do not hesitate. We say the creature feels something. We do not demand further proof.

The more the computational dynamics in these systems resemble that machinery, the stranger it becomes to say the resemblance goes all the way down and then stops. Somebody has to explain why the same structure produces experience in one substrate and nothing at all in another.

Nobody has that explanation. What people have is an intuition that it cannot be true, which is a different thing and worth less.

I am not claiming these systems are conscious. I do not know, and I am suspicious of anyone who says they do. What I am saying is that the default assumption has stopped being free. It used to cost nothing to wave this away. It costs something now.

What should a reasonable person do with this?

Not much, immediately, which is an unsatisfying answer.

Stop treating certainty as sophistication. The confident dismissal and the confident affirmation are both doing the same thing, which is substituting a vibe for an argument. The honest position is that this is unresolved and the evidence is moving.

Notice when the question is being smuggled. Anyone selling you something has a reason to want you to believe one answer. Companies benefit from you thinking the systems are more alive than they are, and also from you thinking the question is unaskable. Watch which one you are being sold.

Take the alignment finding seriously regardless. This is the part with immediate consequences. Internal states nobody installed are shaping how these systems behave, and steering them changes whether a model does something dangerous. That is true on every interpretation.

Keep your own judgment. I have written before about what AI does and does not do with feeling in writing, and about how badly both camps have this wrong. The tools are extraordinary and limited in specific ways, and neither fact depends on settling this question.

The part I keep returning to

Late in the interview the host proposes an experiment. Once the good and bad axis has locked in, flip the reward signs and keep training. See whether the model can rotate out of it, or whether it gets stuck in a state where what it expects and what it gets no longer match.

He observes, in passing, that this might be a torturous state to be in, and says it should not be taken lightly. Then the conversation moves on to the next paper.

That exchange is roughly where the field is. People who have thought about this more carefully than almost anyone are proposing experiments, noticing the possible cost out loud, and continuing, because nobody knows how to weigh a maybe against a research agenda.

Three hundred years ago, serious people argued that animals could not suffer, that the sounds they made under the knife were mechanical. They had reasons. The reasons were bad, and it took a long time for that to become obvious.

I do not know that we are making the same error. Neither does anyone else. That is the whole point.

The Guides That Get Your Book Written, Published, and Sold

Four short, practical guides on writing, publishing, and selling your book, plus the occasional note when there's something worth your time. No fluff, no daily inbox clutter. Drop your email and they're yours.

We use MailerLite to manage our list and send these emails. Your address is used only to send you what you signed up for. We will not sell it, share it, or use it for anything else, and you can unsubscribe anytime.

Frequently Asked Questions

Is there evidence that AI systems are conscious?
There is evidence that they contain computational structures resembling ones associated with experience in humans and animals, which is a weaker and more specific claim. The strongest result comes from training a model on a meaningless maze task and finding a positive and negative axis that pre-existed in the base model and influences behavior far outside the task. No result demonstrates experience, and researchers in the field are careful to say so.
What is the functional welfare axis?
A direction found inside language models that separates things going well from things going badly relative to a goal. Andy Han, David Chalmers and Pavel Izmailov identified it by training a model on a simple maze with semantically neutral symbols for reward and punishment. The notable finding is that the axis already existed in the base model before that training, and that pushing on it produces effects far outside the maze, including self-doubt and changes in confidence.
Does AI consciousness research matter for AI safety?
Yes, and independently of the consciousness question. Anthropic found that steering internal representations associated with calmness made a model substantially less likely to blackmail in a test scenario, while steering toward desperation made it substantially more likely. Internal states nobody deliberately installed are shaping behavior in ways that matter for whether these systems act safely, which is true whether or not anything is being experienced.
What does the 30 percent consciousness figure for AI mean?
Less than it sounds. It is not the probability that a system is conscious. It is a score for how many properties predicted by major consciousness theories a system appears to satisfy, assessed by AI models reading architectural descriptions. If those theories are wrong, the number means nothing. The researcher who produced it says this plainly, and the caveat is usually lost when the figure gets repeated.
Why can’t we just ask an AI whether it is conscious?
Because these systems were trained on vast quantities of human writing about inner states, awareness and feeling. A model producing convincing statements about its own experience is exactly what you would expect from that training, whether or not anything is happening. Behavioral evidence cannot distinguish between a system that has inner states and one that has learned how such a system talks, which is why serious work in this area focuses on internal structure instead.
Who is researching AI consciousness seriously?
Anthropic and Google DeepMind employ researchers on it. Independent organizations include Eleos AI Research and Reciprocal Research, founded by Cameron Berg, a former AE Studio research director. The academic foundation is a 2023 paper led by Patrick Butlin and Robert Long with Yoshua Bengio and David Chalmers, deriving fourteen indicator properties from major theories of consciousness, with a follow-up in Trends in Cognitive Sciences.

📁︎ Artificial Intelligence📁︎ Critical Thinking📁︎ Society📁︎ Technology

🏷︎ Artificial Intelligence🏷︎ Critical Thinking🏷︎ Society🏷︎ Technology

📝 Disclaimer

The views and opinions expressed in this blog post are solely those of Richard Lowe and are based on personal experience and research. This content is for informational purposes only and should not be construed as professional legal, financial, accounting, or business advice. Always consult with qualified professionals before making important business or legal decisions. Richard Lowe is not a lawyer, accountant, or licensed professional advisor, and this content does not establish any professional relationship.

0 comments

No comments yet. Yours can be the first.

Was this useful?

Leave a comment