In Part 1, I argued that thinking and understanding don’t require first-person phenomenal consciousness, and that the information processing they do require doesn’t seem tied to human minds in particular. In principle, a machine that’s complex enough, and that does the kinds of things our brains do when we think and understand, could do it too. This post and the next will argue that current AI models meet the threshold for what we consider real thought and understanding.
Like the last post, this one’s long. This series is me imagining I had a few hours with a smart friend to get across what I actually think about AI and philosophy of mind, and trying to do my very best.
To get to my full argument, I first need to argue that it’s possible in principle for computer programs to think and understand, and only then can I claim that specific AI programs can. This post argues for that first claim, which is pretty contentious on its own, and the next post will argue for the second.
It’d help to have Part 1 fresh in your mind before going on, or at least the section on introspection, the argument that physics is causally closed (and what that means for phenomenal consciousness), and the concluding summary.
I’m not claiming in this post that current AI models think (that’s Part 3). I’m not claiming the brain is a computer, which I’ll explain below. And I’m not claiming AI is conscious. I don’t think there’s anything it’s like to be a current AI model. In Part 1 I argued that the phenomenal consciousness question doesn’t affect whether AI models can think anyway.
Here’s how my argument in this post will go. I argue for each of these premises, and the conclusion follows:
Premise 1: Things that think and understand are reducible to physical processes. There’s no extra nonphysical ingredient.
Premise 2: What makes a physical system think is how it processes information, and what it’s made of only matters insofar as it can do that processing.
Premise 3: A computer can simulate the brain’s information processing to whatever level of detail matters.
Premise 4: A detailed enough simulation of an information process is that process, running in a new substrate.
Conclusion: A computer process can, in principle, think and understand.
Most people who reject the conclusion reject premise 2 or premise 4, so those are what the majority of the post will focus on.
Part 1 took qualities we tend to assume are fundamental to minds and tried to demystify them a bit to show that they might be available to machines. I started high up in our conception of what humans can do and slowly worked us down to a more natural and nuanced picture of the human mind. This and the following post will build up from the ground instead. I’ll start with the basic building blocks of mind in physical reality and work my way up from physical systems (premise 1), to non-carbon-based material (premise 2), to computer programs specifically (premises 3 and 4).
Why computers can in principle produce thinking and understanding
What are minds made of?
Can physical objects think?
The first and simplest step in figuring out whether computers can think is to ask whether merely physical objects can, or whether thinking requires something nonphysical, something more than particles, forces, fields, and the ways they relate to each other. I discussed this a lot in Part 1 but will revisit the core points.
I see 3 possible answers:
Everything about the mind is physical, including thought and understanding.
Some parts of the mind aren’t physical, but thought and understanding aren’t among them, and don’t depend on anything nonphysical. Maybe our phenomenal consciousness (what it’s like to be us) is nonphysical, but our thought and understanding are a kind of physical information processing.
Thought and understanding themselves depend on something nonphysical. They happen in part outside of the physical world.
Number 3 would mean we probably can’t build machines that really think and understand, while 1 and 2 leave the possibility open. I argued in Part 1 that while phenomenal consciousness is still a deep mystery, real thought and understanding don’t seem to depend on it. The only kind of consciousness they need is what philosophers call “access consciousness,” which is a specific, elaborate kind of information processing. So for the purposes of this series, I’m going to go on assuming #1 or #2 is true, and both say that it’s possible in principle to build a purely physical object that thinks and understands. Under either view, human brains are proof that this kind of object can exist.
I also see lots of reasons to think that physics is causally closed. As I explained in Part 1, it seems suspicious to assume that the one place in the world where physics is regularly violated is inside our own heads. This would also seem to break basic important physical rules like the conservation of energy, that would turn the world topsy turvy if they were regularly broken. And it seems weird that a purely physical, unguided process like evolution by natural selection could have stumbled on some way of interacting with nonphysical stuff. This makes 1 and 2 seem especially likely to me. A lot of people I meet who believe 3 believe it for religious reasons, and that’s fine, but if that’s your main reason for believing that AI can’t think, it’s important to say so directly instead of making some unrelated argument about the science of physical brains and computers. Some also believe #3 for scientific and philosophical reasons, and that’s also fine, but I think the burden of proof is on them to demonstrate why they do.
Under the view that the mind is reducible to physics, our thoughts and understanding are high-level patterns that emerge in our physical brains. Here a mind is kind of like the national debt, which isn’t a physical entity itself but is ultimately the result of huge numbers of specific tendencies and patterns in the behavior of purely physical systems.
So I think the answer is either the first or the second, and either way, things that think and understand are reducible to physical processes. There’s no extra nonphysical ingredient. That’s premise 1 of my full argument. The fact that minds happen in physical objects has a lot of additional implications for philosophy of mind.
Implications of minds existing in purely physical objects
In classical Newtonian physics, everything that happens is predetermined by the initial state of the physical system you’re in. If you had absolutely perfect information about the beginning of the universe and infinite time to calculate everything, you could in principle perfectly predict everything that would happen, down to the fact that you’d be reading this blog post right now. If you could rewind a Newtonian universe, the exact same thing would happen when you played it back.
We now think this probably isn’t true. On the standard reading of quantum mechanics, events at the quantum level are fundamentally probabilistic. They aren’t determined in advance, but they can be described and estimated using probabilities. If you rewound the universe and started it over, different things could happen. Some interpretations of quantum mechanics (like my favorite, many-worlds) circle back to a kind of determinism, but even here they involve what looks like randomness from the observer’s point of view. Either way, physical reality reduces to strict laws plus randomness.
Everything we know of in physics either follows strict deterministic laws or is random in the ways quantum mechanics describes. This has an important implication for brains and thought, assuming that they are ultimately physical: any brain activity or mental process is either predetermined by physics, random, or some combination of the two. If your requirement for real thinking and understanding is that it can’t fundamentally rely on either purely mechanistic systems or randomness, then either humans have never thought or understood anything, or human thought transcends physics. Both seem wrong to me, so it looks like real thinking needs to be able to be broken down into purely mechanistic systems or include pure randomness, and there’s no secret third thing involved.
This gives us a quick way to check a lot of arguments that AI can’t think, which I’ll call the brain test. Take the argument and apply its reasoning to the human brain instead. If the reasoning also applies to the human brain, the argument proves too much, because it would also show that you can’t think. I’ll come back to this simple test a lot.
Another way of saying this, which I described in Part 1, is that any system that understands is ultimately made up of individual parts that don’t:
Unless thought and understanding are supernatural and happen outside of physics, any system that understands or thinks is ultimately made up of fundamental particles and forces. These on their own don’t understand or think. Thinking is an emergent property of arranging them in specific ways to perform specific functions. At a higher scale, individual neurons also don’t “think” or “understand.” They are building blocks for a specific system that thinks and understands. All that human neurons do is receive basic chemical signals, change their electrical activity, and release some chemicals. Any system that understands will be made up of parts that “only” do simple things. That’s a basic part of living in physical reality.
…
The role of access consciousness in thought, like anything else in the world, also needs to be completely explained by individual processes that are not conscious. If thinking requires some higher-level awareness, that awareness also has to be explainable in the language of physics (again, unless you believe it happens in a separate nonphysical world that interacts with ours). Dennett once observed (in his book Brainstorms) that explaining an intelligent system means breaking it into smaller, dumber systems, then breaking those into even dumber ones, until you get down to parts so simple that they can be “replaced by a machine.” I think this is the only way minds can ever be explained.
Information theory
My argument later on will be that thought and understanding are fundamentally both specific ways of processing information, so it’s important to explain how information fits into a purely physical view of the mind. There’s a field of math and physics called information theory (started by Claude Shannon, who the AI Claude is widely thought to be named after) that gives us a precise way to measure information. Information is anything that reduces uncertainty, and it’s measured in units called “bits,” where 1 bit is the answer to a yes or no question when both answers are equally likely. If I have a number in my head between 1 and 8, it takes 3 bits of information to pin it down, because you can ask “Is it in the top half of the range?” and then “Is it in the top half of the remaining range?” and then repeat that second question one more time. If I thought of 6, the answers would be yes, no, yes. Each bit doubles the number of possibilities you can tell apart. One bit can tell you if a number is 1 or 2, two bits can tell you which of 1, 2, 3, or 4 it is, three can tell you which number from 1 to 8 it is, and so on. This is why the game 20 Questions can get you to one specific thing, person, or place out of all the options, because 20 bits can narrow things down to 1 option out of about a million, 220.
Shannon’s measure tells you how much information something carries and says nothing about what it means. He set meaning aside on purpose, and I’ll get to it in Part 3.
Each bit needs to be stored in or carried by something physical (and since nothing that carries a signal can move faster than light, information can’t either), but the information itself is independent of the specific thing it’s stored in. It can be transferred from one thing to another. For example, the words on your screen right now started off as information in my brain, which became patterns in how I typed the words into my computer, which became both words on my screen and electric charges in my laptop’s memory. When I added them to my blog, the information traveled as radio waves to my router, then mostly as light in fiber optic cables to a data center somewhere, where it was stored again as electric charges in the chips there. Then it reached your device the same way, came out of your screen as light that you read as words when it hit your eyes, and now it’s being stored in your brain. This was the same information all along. It could be carried by patterns in very different physical stuff, whether brain tissue or light or electricity or the movement of my hands on my keyboard. But this information isn’t some special nonphysical substance. It’s just a very complex, specific pattern in physical material that influences patterns in other physical material, the same way an ocean wave is a pattern in water that shapes the water ahead of it without being anything magical over and above the water. A wave can even cross an ocean while each particle of water just bobs in place.
Does an object need to be made of a specific material to think?
Here I’ll argue for premise 2: what makes a physical system think is how it processes information, and what it’s made of only matters insofar as it can do that processing. This is often the most controversial part of the argument.
A belief that’s very popular among everyday people is that thinking depends on biological material. A lot of people believe that there’s something so fundamentally special about the material the brain is made of that a similar object made of anything else couldn’t “really” think or understand. At best it would just be an imitation. Obviously, to make my case I need to convince you that thinking and understanding can happen in non-biological material like silicon.
The word “substrate” will be useful here. It comes from a Latin word meaning “spread underneath.” In philosophy of mind, it means the physical stuff a process happens in. The substrate my computer programs run on is silicon, and the substrate my mental processes run on is my brain tissue. “Substrate dependence” is the view that mental events like thought or understanding can only happen in very specific substrates, and (on most versions of the view) silicon isn’t one of them.
There are two possible reasons why mental activity might depend on specific substrates:
Only specific substrates can play the causal, functional roles needed for specific mental processes. For the same reason that you can’t make a calculator out of a puddle of water, you can’t make something that thinks and understands out of just helium. A lot of substances don’t allow the kinds of complex organization and structures that minds need. Maybe carbon-based biological substances are the only ones that can form the types of causal structures that thought and understanding require.
Specific substrates have some inherent quality that lets things made of them think or understand, over and above the causal roles they play in the systems they’re in. Maybe only biological brains can really understand anything, because the material they’re made of is somehow tied to what understanding is, beyond any functional role it plays. This view says that even if you could perfectly remake every last process happening in the brain using a different material, the artificial brain would only be imitating understanding and thinking.
To see the difference between these two views, we can imagine (following a famous thought experiment from David Chalmers called “fading qualia) that scientists have made a surprising discovery: they can replicate literally every process that happens in your brain using silicon-based material rather than carbon-based material. There is no causal difference in the brain’s behavior when they replace a carbon-based neuron with a silicon-based one.
Suppose that they replaced a single neuron in your brain with one of these silicon neurons. What would happen to your consciousness and thinking? Would you notice? What would happen if they slowly, over the course of months, replaced more and more of the neurons in your head with these silicon-based neurons?
If you hold belief #1, your mind would remain intact, because these artificial neurons can (for the sake of the thought experiment) in fact play the same functional role that carbon-based neurons play. There would be no change in your perception, thinking, or understanding. The only thing about the carbon-based neurons that mattered for your mind was the specific role they played in your overall brain processes. Replacing them with these new neurons would be like replacing individual transistors in your computer with new ones that work the same way. It could still run all the same programs.
If you hold belief #2, replacing your biological carbon-based neurons with silicon neurons would slowly remove key parts of your mental activity. Maybe your field of vision would fade or shrink. Depending on what you believe is reliant on carbon, by the end you either wouldn’t be conscious, or wouldn’t be able to really think and understand, or both. Your outward behavior would stay exactly the same the whole time, because the cause-and-effect relationships that made up your brain would all still be running, but there would be no real thought, understanding, or consciousness behind it. This is almost exactly one of the scenarios John Searle described in his 1992 book “The Rediscovery of the Mind,” where, as your brain is replaced with silicon, you internally want to cry out that you’re going blind while your voice, completely out of your control and now run by the silicon neurons, keeps describing what’s in front of you.
I’ll call belief #2 “the carbon view,” and I think the thought experiment itself shows that it’s wrong. I’ll also argue later that belief #1 is correct in some sense (you can’t make a mind out of helium because it couldn’t maintain the right cause-and-effect relationships), but that silicon (and probably a lot of other materials) checks all the boxes for the specific causal relationships and functions we’d need for real thought and understanding.
What would actually happen in this thought experiment?
I want to start to poke at the carbon view, starting with an inconsistency in its answer to the thought experiment above.
Let’s imagine that 10% of your neurons have been replaced with these silicon neurons. Which of these would your field of vision look like?
If your consciousness were slowly being taken away, would you notice? Would you react? Remember that these new neurons have the exact same cause-and-effect structure as your old ones. With either kind of neuron, if your brain received information that your field of vision looked strange, it would cause you to say “Hey, there’s a bit of my vision missing.” But because the new neurons play the exact same causal role as the old ones, you would do and say exactly what you would have done and said with your old brain. If your old neurons wouldn’t have led you to say “Hey, my field of vision looks weird,” your new neurons won’t either.
So if your field of vision really were dissolving, you’d never say anything about it, and what you said about your vision would have come completely apart from what you actually see. If the silicon neurons cause you to behave in exactly the same way, this strong observation of your consciousness changing should trigger the same behavior the old neurons did, but for some reason they wouldn’t here. It seems like the only way for this to work is for you to somehow not notice your own phenomenal consciousness changing. You’d be completely and permanently wrong about your own experience while your brain went on working exactly as before. This seems to imply some kind of epiphenomenalism I talked about in Part 1, where your consciousness helplessly observes your life but never has any causal influence on it.
Searle accepts that this could happen, but I find it much harder to believe than the alternative, that your field of vision isn’t dissolving at all. (This is basically the philosopher David Chalmers’ “fading qualia” argument). Philosophers really do disagree about the experience part. For example, if experience turned out to be nonphysical (option 2 from earlier), physics being causally closed would mean it couldn’t affect what you say anyway, which is one situation where Searle’s scenario hangs together. But the case is much stronger for thought and understanding. The new neurons do exactly the same things with information as the old ones, and in Part 1 I argued that’s what thinking and understanding come down to, so a brain made of this kind of silicon could think and understand. What matters for these new neurons is the role they play in your broader mental system, not what they’re made of. If they play the same role, they’re the same for your mental activity. I see this as a strong argument in favor of belief #1.
Another way to put this is as a dilemma. Either the special ingredient in carbon makes a difference to what the brain does, or it doesn’t. If it does, then a silicon brain that does everything the carbon brain does has something playing the same role. If it doesn’t, it can’t be what thinking is, because thinking makes a difference to what you do, and nobody could ever tell who had it, including you.
The carbon view as I’ve described it is mostly a folk view. The philosophers and scientists who take substrate seriously today (like Peter Godfrey-Smith, Anil Seth, and Ned Block) don’t think carbon atoms are special in themselves. They believe that the functional biological processes like metabolism might matter for consciousness specifically. The debate among experts seems to be mostly about phenomenal consciousness, and they are often otherwise open to machines thinking and understanding language.
Even if you find the replacement argument convincing, the carbon view can still feel obviously right, because every mind we’ve ever seen has been made of carbon. The next few sections explain where that feeling comes from, and why it doesn’t tell us anything about whether other materials can think. If you already don’t believe the carbon view, you can skip ahead to “Some takeaways from neuroscience.”
Is carbon required for thought?
Evolution and carbon
A lot of people believe the carbon view because of a simple observation: every mind we’ve ever seen is carbon-based, with no exceptions at all. Anything that we’ve ever called a mind, whether it’s in a human or crow or octopus, was in biological matter. This seems like strong evidence!
However, there seems to be a much more obvious reason why minds have only appeared in carbon-based biological matter.
Minds need an enormous amount of very specific, organized complexity. We only know of two processes in the universe that can build very complex things that work:
Evolution by natural selection
Deliberate design by people
Nothing else we know of can create anything as complex as human and animal brains. And for almost all of Earth’s history, only the first one existed. On Earth, as far as we know, natural evolution has only ever acted on carbon-based biological matter, so only carbon-based material could be shaped into the level of complexity required for minds. Thus, the only things we see that have minds are carbon-based.
Why has natural evolution only acted on carbon-based material?
Well, stepping back, evolution isn’t a force in the world with goals. It’s just an emergent pattern that shows up on its own, because things that survive and replicate more end up producing more things like them. We can expect the pattern of evolution by natural selection to emerge whenever objects in the world have these three properties:
They can replicate themselves.
They carry information that gets passed on to their copies, and that information can be modified in the act of replication, leading to variation among the copies. The copying has to be mostly accurate, or the information would get scrambled over the generations, but it can’t be perfectly accurate, or nothing new would ever appear.
Some of that change affects how good the objects are at replicating in their specific environment, so some versions end up making more copies of themselves than others, spreading their variants more.
Over time, this process can stumble on mutations that add more and more complexity to the species they happen in. With enough time, this creates features that look like they were the product of design, like eyes, wings, immune systems, or minds, even though no designer was involved. For almost all of Earth’s 4.5-billion-year history, this was the only possible source of the level of complex organization minds require.
The basic dynamics of evolution don’t depend on any specific substance. In a famous 2003 study using a digital evolution program called Avida, scientists made self-copying computer programs that evolved to do complex logical operations by building on simpler abilities they had evolved earlier, in the same way complex features emerge in living things over many generations. Evolution can act on anything that checks all three boxes I listed above.
So why has natural evolution on Earth only acted on carbon-based material? It turns out that it’s mostly because carbon is more useful for the kinds of replicators evolution needed to get started on Earth specifically, and is a good, stable material for more complex structures.
The main reason carbon’s so useful and versatile as a building block for replicators is that it has four outer electrons (valence electrons) available to form bonds with other atoms. As it happens, silicon is the very next element with 4 valence electrons on the periodic table.
This means silicon can form some of the same kinds of molecules as carbon. Silane (SiH4), for example, is silicon’s version of methane (CH4), but it has some very different properties (it bursts into flame on its own when it touches air). Silicon is much, much more abundant on Earth than carbon. By the most commonly cited estimate, there are around 600 silicon atoms for every carbon atom in Earth’s crust.
If silicon is so much more abundant than carbon, and can form similar molecules and bonds, why were the first replicators that kicked off the process of evolution carbon-based and not silicon-based?
One reason is that when a carbon atom bonds to another carbon atom, that bond is about 50% stronger than the bond between two silicon atoms. This means silicon is much worse at forming the long chains and rings that are required for larger and more complex structures. Chains of carbon atoms are also very stable in water, whereas silicon chains break apart much more easily in both water and oxygen. That matters a lot because the first replicators (which arose at least 3.5 billion years ago) almost certainly formed in water (the earth was mostly covered in it). The silicon versions of carbon’s building-block molecules (like the silane mentioned above) also tend to be much less stable in general.
Silicon also bonds much more strongly to oxygen than it does to other silicon atoms, so in nature it almost always ends up bonded to oxygen. Almost all the silicon in Earth’s crust is locked up in rock as a result. Silicon dioxide (one silicon atom for every two oxygen atoms) is what we know as quartz, the main ingredient in most sand. If silicon-based animals breathed out silicon dioxide like we breathe out carbon dioxide, they’d breathe out sand.
So it seems like the main reason life is carbon-based is that carbon forms strong bonds with other carbon atoms, those bonds hold up in water, and the first replicators almost certainly arose in water. These facts are interesting, but on their own they don’t tell us anything about whether there’s something unique about carbon-based molecules that makes them important for thinking and understanding. What they tell us is that carbon was the basis of the specific molecules that kicked off the evolution of all the complex life we see today. As far as we know, that happened exactly once. Any new replicators that showed up now would almost certainly be eaten or outcompeted by the life that already exists.
So what evidence do we get from the fact that all minds we’ve ever seen are in carbon-based biological brains? All it tells us is that the universal common ancestor of all life (a single-celled organism that lived more than 3.5 billion years ago, roughly 3 billion years or more before the first simple brains evolved) was made of carbon-based stuff, and it was successful at replicating itself and met the three criteria needed to run the pattern of evolution by natural selection, which was the only process on Earth that could produce the level of complexity minds would need much much later on. Until people started building flying machines, the only objects in the world that could fly were also biological and carbon-based, because again, living things were the only things complex enough to fly. This told us nothing about whether there was something inherent to carbon-based biological material that made it necessary for flight. Like flight, thought and understanding can only happen once some process has organized incredibly complex, specific systems. For billions of years the only process that could do that was evolution, and if I’m right, in the last year or so human design has become a second way to create key high level aspects of minds.
If carbon-based biological matter were required for thought, over and above its ability to play the correct functional roles in systems, this would imply a lot of other strange things about the world. For one, on any planet where evolution acted on another type of replicator (maybe one where the environment was more hospitable to silicon-based life), evolution could never produce thinking animals.
This would be a massive missing tool that would’ve been useful to those silicon-based animals. Thinking has obviously been incredibly useful to animals on Earth, to the point that complex thought seems to have evolved independently multiple times, in lineages that split off from each other before any of them had complex brains. For example, humans and other primates, crows, and octopuses can all solve complex problems that require step-by-step planning, and their evolutionary lines all split from each other long ago.
Octopuses and humans last shared a common ancestor about 600 million years ago, long before either lineage had much of a brain. Crows and humans split off about 320 million years ago, when brains were still small and reptile-like. Our shared ancestor looked something like a small lizard, and birds and mammals each separately went on to evolve brains with about 20 times as many neurons as reptiles of the same body size. This means that on at least three separate occasions, relatively simple brains independently evolved to the point that they could solve complex problems like when an octopus carries coconut shell halves across the sea floor to use as a shelter later, and when a crow pulls up a short stick dangling from a string, uses it to fish a longer stick out of a box, and then uses that to get food out of a hole. The wild differences in the look of the brains of the three species are a result of approaching the same broad ability from wildly different starting points:
It is crazy that intelligence has independently evolved this often.
Notice as an aside that when we say a crow thinks, we’re not saying it has the same structure as a human brain. Crow neurons are arranged very differently from humans’. Most of an octopus’s neurons aren’t even in its brain, they’re in its arms. No one concludes from these deep structural differences that octopuses and crows can’t think or only imitate problem solving. We determine that they think because of what their brains do, not whether they have the exact same structure as human brains. We’ll come back to that later.
But according to the carbon view, this ability, which has been useful enough to animals on Earth that it evolved multiple times, could never evolve on that silicon-based planet. Even though silicon could play all the required physical roles in the animal’s broader system, to the point that animals would behave exactly as if they were thinking, it would somehow “not count” as thinking because it wasn’t carbon.
This view makes the history of life look like a bizarre miracle. The first real brains didn’t evolve until a little over 500 million years ago. For more than 80% of life’s existence, there were no brains. What are the odds that the first replicators happened to be built around the one element that, for some separate reason, is the only one that can support real thought and understanding, or consciousness?
Why would this quality of “necessary for thought” even belong to a specific element? Or one that bonds so easily with other elements? Could there be a universe where only neon could produce real thought, but because neon is a noble gas and doesn’t form stable molecules, no thinking things ever evolved anywhere? The universe would be permanently full of this hidden capacity for thought that nothing could ever actually tap into. Does that concept even make sense? It sounds incoherent to me, which makes it hard to believe that any one element has some inherent property that makes brains made out of it “do real thought” while systems made of other stuff that can perform all the same physical functions don’t. Did we just get incredibly lucky that the thought-element happened to be perfect for replicators?
The carbon view looks pretty strange in the context of evolution. But evolution also gives us clues to why we might hold it so strongly.
To indulge in armchair evolutionary psychology for a moment, it seems like any human in the ancestral environment who would treat rocks or other material as potentially other entities like them, entities with complex mental properties, would probably be at a pretty bad disadvantage. They would waste a lot of resources either trying to cooperate with or fight them. It seems like most animals that can perform complex thought have a pretty strong inbuilt sense of what’s a living thing (to either avoid or fight or eat or cooperate with) vs what’s a nonliving physical object. This sense obviously doesn’t directly track carbon, and can misfire. It looks for things that biological life has, like faces or independent motion, which is also why humans often over-attribute minds to things that do either. This inbuilt sense would’ve been shaped by what helped or hurt animals in survival and reproduction, and in a world where carbon-based life was the only stuff that could develop the kinds of complexity that matter for whether animals react to them as other agents in the world, the strong intuition that only things that look like carbon-based life can have minds would have been very useful, regardless of whether it was true about the universe more broadly. There’s no clear causal path from whether other stuff can perform mental activity in principle to the intuitions that humans in the ancestral environment developed. I suspect a lot of our intuitions about philosophy of mind come from something like this evolutionary pressure to not consider nonliving matter as possibly possessing any complex mental behavior, and the thing that created that evolutionary pressure was ultimately the simple fact that carbon is an especially useful base for replicators, which doesn’t actually relate to whether other matter could in principle perform mental activities if arranged in the correct way at all.
Carbon is itself just a specific arrangement of smaller parts
Another problem for the carbon view is that atoms aren’t the most basic physical things reality is built from. They’re made of protons and neutrons and electrons, and protons and neutrons are themselves made of quarks. A normal carbon atom has 6 protons, neutrons, and electrons, and a normal silicon atom has 14 of each. They’re both made of the exact same stuff. There’s just more of it in the silicon atom.
Everything that makes a carbon atom behave differently from a silicon atom is just an emergent pattern resulting from how many parts there are and how they’re arranged. Both have the same number of outer electrons, which is why they can form similar types of molecules. Silicon forms weaker bonds with other silicon atoms mostly because it has one more layer of electrons below its outer electrons than carbon does, which makes it a bigger atom. Every difference in their chemistry comes from differences in how the same underlying particles behave in each.
This, even more than evolution, is what makes the carbon view look so strange to me. It doesn’t make much sense to say that carbon has some quality over and above the role it plays in larger systems, because carbon itself is just the pattern emerging from the individual roles being played by simpler interacting parts, the same parts that make up every other element. If you arrange the simpler parts differently, they do different things, which we call other elements. Carbon is made up entirely of parts that aren’t carbon. How could a mere role that simpler parts play somehow have a fundamental property beyond the role it plays in a broader system?
About 1% of the carbon atoms in your brain are carbon-13, which has an extra neutron. It chemically behaves almost exactly the same as normal carbon, and your body uses it interchangeably. Does carbon-13 also have this special thinking property as normal carbon? If it does, it’s a little suspicious that the only other type of thing that has the thinking property also plays the exact same functional role in the broader system. Adding a neutron is fine, but a proton is bad? A neutron is itself made of one up-quark and two down-quarks, and a proton is the reverse. Is adding one more up-quark and taking away a down-quark the thing that thinking is so tied to? If the quality of “real thought” disappears with the extra neutron (which barely changes anything the atom does), then about 1% of the carbon atoms in your brain, scattered randomly throughout it, aren’t being used to really think. Again, seems bizarre.
If thinking needed something from carbon atoms over and above what carbon does, it’s difficult to see where that would come from. Everything we know about carbon is completely accounted for by what its smaller parts are doing, and it shares those smaller parts with silicon, just with different amounts and arrangements. Carbon itself is just a specific functional arrangement of smaller parts, so if thinking depends on carbon at all, it depends on a particular arrangement of more fundamental particles, the same particles every other element is made of. So thought still ends up as a specific functional process of matter, and only relies on the type of matter if it can perform that function, partly because every type of matter more complex than quarks or electrons is itself just a function of an arrangement of smaller bits and the roles they play in the system. This circles us back to belief #1: thinking depends on a specific substrate only in the sense that the substrate can or cannot perform the specific causal roles required for thought.
So now we need to figure out if silicon-based structures can play roles that are necessary and sufficient for real thought and understanding. To figure that out we need to first look at neuroscience and then computer science.
Some takeaways from neuroscience
Neuroscience is still deeply uncertain about what happens in our brains when we think and understand. However, we can draw some important lessons from it about what’s happening in brains and whether silicon computers could also do it. These are the key takeaways from the field as I see them relevant for my argument.
Thinking and understanding happen in the brain through neurons passing information to each other.
What matters about those signals is the information they carry, rather than the stuff carrying it.
The brain’s information processing can be described with ordinary classical physics, and doesn’t seem to rely on special quantum effects.
These support my core premises 2 and 3:
What makes a physical system think is how it processes information, and what it’s made of only matters insofar as it can do that processing.
A computer can simulate the brain’s information processing to whatever level of detail matters.
Thinking and understanding each happen in the brain through neurons passing information to each other.
A cornerstone of modern neuroscience is the “neuron doctrine,” the idea that the brain is made up of separate cells (neurons) that each receive signals from other neurons (often thousands of them) at contact points called synapses. Some signals make it more likely that the neuron will fire a brief electrical pulse, and others make it less likely. If the combined signals cross a threshold, the neuron fires, sending a pulse down a long fiber that releases chemicals onto the next neurons, pushing them toward or away from firing, and so on. Your brain has around 86 billion neurons and something like a hundred trillion synapses.
The fact that the brain can be broken up into separate units that each take in inputs and produce outputs means that the entire brain can be described in terms of what its units do and how they’re connected to each other. The true way the brain works is messier, with additional cells and chemicals and neurons wired directly to each other, but ultimately all of it is still physical signaling that can be described the same way.
What matters about those signals is the information they carry, rather than the stuff carrying it.
It appears that whatever makes a pattern of activity in your neurons a thought or perception or memory is the information it carries and what the rest of the brain does with it, rather than the particular physical stuff carrying the information. In the 1920s, the physiologist Edgar Adrian found that individual signals in single nerve fibers are all basically alike. A stronger stimulus (like a harder touch or a brighter light) doesn’t lead to larger pulses (or “spikes”). It just makes them happen more often. So what makes some of these signals sights and others sounds is which neurons they reach and what those neurons do with them.
There are already examples of information being processed by silicon before it reaches the brain, the most famous being cochlear implants, which do the job of the cells in the inner ear that normally turn sound into nerve signals, using a computer chip that sends electrical pulses into the auditory nerve. Many implant users can hold phone conversations without lip reading, so the brain is clearly able to build understanding out of information that silicon has processed.
In 2018, researchers worked with eight epilepsy patients who already had electrodes in their brains for medical reasons. They recorded activity in the hippocampus (a part of the brain that’s key for forming memories) and built a mathematical model of the activity patterns that went along with remembering things correctly. They used a computer to stimulate the same patterns into the patient’s brains during later memory tests, and their scores improved by about 35%. A computer model of part of the brain’s information processing was helping with a cognitive job inside a human brain! This seems a lot like the thing the brain was doing was manipulating information that a computer was able to also manipulate in the same way.
For the most part, neuroscientists connect thought and understanding to how information flows through the brain, rather than to which material is carrying it.
The brain’s information processing can be described with ordinary classical physics. It doesn’t seem to rely on special quantum effects.
Most neuroscientists and physicists think the brain doesn’t make use of quantum effects like superposition, and that the way it processes information can be fully described with classical physics. (Quantum mechanics underlies all chemistry, including the brain’s, but that’s different from the brain using quantum effects to process information.) They think this for two reasons:
Classical models of neurons work really well. Every successful simulation of neurons treats them as classical, mechanistic systems.
Quantum states are incredibly fragile, and the brain is warm, wet, and very crowded with jostling molecules. Max Tegmark estimated that quantum states in the brain would fall apart within 10-20 to 10-13 seconds. Neuron activity itself takes thousandths to tenths of a second, at minimum 10 billion times longer than the quantum effects happen.
Two things about classical physics become really important here:
Classical physics is mechanistic. The brain is way too complicated for anyone to actually calculate what it’ll do, but everything that happens in the brain, as a system subject to classical physics, is determined entirely by its starting state, its inputs, and the laws of physics. Some randomness exists in things like ion channels that randomly snap open and shut, or synapses failing to pass a signal on. If you rewound the universe and replayed it, some of those random events might come out differently, but otherwise the brain would give the exact same response to the exact same input if it started in the exact same initial condition.
Every classical physics situation in the everyday world can be modeled with differential equations (equations that describe how things change from moment to moment). In the brain, these can be used to describe how quantities like voltage and chemical concentrations change smoothly over time.
So the brain’s processes, including whatever thought and understanding are, are reducible to classical, mechanistic physics (plus some randomness).
It seems like thought and understanding have to be specific information flows in the brain
I don’t really see what else thought and understanding could be if not very very complex flows and processing of information in the brain. If you sit down to think about something, or understand something, what are you actually trying to produce? Usually you’re trying to change patterns in things outside your brain, like your mouth when you speak, your body when you move, or your hands when you write. The goal of thinking doesn’t seem to be to change anything about the physical stuff in your brain. The stuff inside your brain DOES change (synapses get stronger or weaker, and neurons change how easily they fire), but only to better shape the flow of information toward what you’re after. So it seems like whether our brains (or anything else) can think and understand comes down to the specific way information flows through them. This is not to say that anything could think if it could produce the same outputs from the same inputs as my brain. If I type “Hello, World!” and a computer runs the line of code print(”Hello, World!”), we produce the same output, but obviously the code isn’t thinking. Thought and understanding depend on the specific ways information is processed inside a system, because that’s what determines whether it can do similar useful things with information in very new contexts, which seems like a clear baseline requirement for real thought. It’s also not to say that all information processing is thinking. A thermostat processes information too. The claim is that thinking is one particular, very rich kind of information processing, and the question for this post is whether things other than brains can do that kind.
If thought and understanding are about the flow of information, then they don’t depend on the specific substance carrying it (as we saw in the information theory section, information is substrate independent). The question is whether a given physical system can perform the causal roles needed for the kind of information flow required for thought and understanding. To figure out whether computers can perform these functions, we need to turn to computer science.
Some takeaways from computer science
In 1936, Alan Turing, often called the father of computer science, wrote a paper on what it means to “compute” something, and what could do it. At the time, “computer” was still a human job title, and he was trying to come up with a definition of a machine that could do all the same things a human computer possibly could. He proposed what came to be known as a “Turing machine,” which has a long strip of tape divided into squares that can each hold a symbol, and a head that reads one square at a time. The machine is always in one of a limited number of “states” which is like one specific line of the instruction sheet. At each moment, based only on the combination of the state it’s in and the specific symbol in front of it, it writes a new symbol, move one square left or right, and switch to a new state. The tape’s assumed to be infinitely long to allow the machine to get as detailed as possible. The reasons take a while to explain, but Turing suggested that any calculation a human could do by mechanically following step-by-step rules, with a pencil and unlimited paper, could also be done by this machine. This claim is now called the Church-Turing thesis.
Different Turing machines could perform different tasks. Solving a calculus problem is a step-by-step process, so is finding the optimal route to travel to a location. A system is “Turing complete” if it can simulate any specific Turing machine, which means it can carry out any calculation that can be written as step-by-step rules with enough memory and time.
Your laptop is Turing complete, because it can follow basically any set of instructions you give it, whether it’s running a spreadsheet or a video game with complex physics. The only two things that limit it are what you do with it and how much memory it has.
A lot of other things are Turing complete. For example, the trading card game Magic: The Gathering. In 2019, researchers found that you can set up Magic games where the creature tokens stand in for symbols on the tape of Turing machines and can be made to run any computational process. People have also built working computers out of the circuits you can make inside Minecraft.
A simulated computer could also be Turing complete. Emulators are simulated computers running on other computers. Many people use emulators to play old Nintendo games on their laptops that can’t be played on modern computers. The simulated parts of the computer move information around the same way a physical computer’s parts do, for the same reason people can build working circuits inside Minecraft.
Importantly, a Turing complete computer can simulate any process that follows the laws of classical physics to any level of detail, because classical physics can be written entirely as equations that a computer can work through step-by-step. Since the brain’s processes follow classical physics, a machine that’s Turing complete could in principle simulate a brain as precisely as we like. That gets us to premise 3 of my argument: a computer can simulate the brain’s information processing to whatever level of detail matters.
Even if the brain did depend on quantum physics, this could also be simulated, just much, much more slowly. It’s been shown that quantum computers can’t compute anything a Turing machine can’t, they can just do some kinds of computing much faster. The brain being classical mostly matters for how practical it would be to recreate the processes inside.
Would a simulated mind be able to really think and understand?
Many philosophers have debated whether a simulation of a brain would have phenomenal consciousness, whether there would be something it’s like to be the simulated brain. John Searle seemed pretty confident that a simulated brain would not have phenomenal consciousness, while David Chalmers believes it would.
I’m going to sidestep the question of phenomenal consciousness in simulated brains. As I argued in Part 1, I don’t believe that phenomenal consciousness is required for thought or understanding. What’s required is for a system to handle information in particular ways, using very high-level functions we don’t really understand in human brains yet. The question is whether a simulated brain could think and understand. A common line (going back to Searle) is that a simulated brain wouldn’t really think, for the same reason a simulated rainstorm doesn’t get anyone wet.
Importantly, I’m not claiming that AI models are simulated human brains. When you make an LLM, you’re not mapping the human brain and mimicking all the physical stuff happening inside. I’m instead trying to show that it’s possible in principle for a computer process to think, and will get to the details of how AI systems actually work in Part 3. If you’d agree that a simulation of a human brain with enough accurate detail could really think, you agree that computer programs can in principle think, and the question from there is just what computer systems qualify as really thinking. Right now I’m trying to get you to that first step, using an imagined perfectly simulated human brain to do it.
So far I’ve argued for premises 1 through 3. Thinking is something physical systems do, what matters is how a system processes information, and a computer can simulate the brain’s information processing to any level of detail. Together, I think these get us almost all the way to the conclusion. What matters for thinking and understanding is how information is altered by the system and affects other information in the system, and what the system can do with that information once it’s done processing. Whatever processes in the brain manage this information flow can all in principle be replicated by a Turing complete computer. If a simulation at one level of detail leaves out key subtleties in how those processes interact, it can just add more detail. And because the brain is so noisy (ion channels snap open at random, synapses fail), details far below the level of that noise can’t make any systematic difference to how it processes information (they just add to the noise that’s already there), so there has to be some finite level of detail that captures everything relevant to thought. A Turing complete computer can simulate the brain at that level of detail. The simulated brain would do the exact same type of information processing as the real brain, in the same way that a video game run on a good emulator (a simulated computer) plays exactly like it did on the original console. The simulation is simply another substrate that is handling the information. Because thinking and understanding are a specific type of information processing in the brain, and a simulated brain (like a simulated computer) could perform all the same types of information processing, computers can in principle really think and understand.
The last step is probably the most controversial. Is a computer simulation merely another substrate? That’s premise 4 from the start: a detailed enough simulation of an information process is that process, running in a new substrate.
To simulate a process on a computer is to set up the computer’s physical parts so that they carry the same information as the parts of the original process, and change in response to each other in the same pattern (at whatever level of detail you’re simulating). That’s also just a description of a process running in a new substrate. Earlier I defined a substrate as the physical stuff a process happens in. In a computer simulation of a brain, the process is the brain’s information processing, and the physical stuff it’s happening in is the computer’s silicon.
Let’s circle back to the imagined silicon neurons from earlier. Start with your brain after every neuron has been swapped for a silicon one. Now replace a small cluster of those silicon neurons with a single chip that takes the same inputs and produces the same outputs the cluster would have. Then replace bigger and bigger clusters, until one computer is calculating what every part of your brain would do. Every step changes how the information processing is physically carried out, but none of them changes the processing itself. At the end you have a brain simulation. If the gradual neuron replacement kept your thinking intact, a skeptic has to say at which of these steps it disappeared, and why.
In any other case that involves something that’s fundamentally information rather than matter, it’s clear that if you simulate something that handles the information in the exact same way, the exact same pattern of information emerges, and thus “the same thing is happening.” If you run a calculator program on an emulator, or in Minecraft, it really also calculates. Right now, you’re looking at words on a simulated piece of paper (your screen). Are these “real” words? Did I “really” write them? Yes, because words are fundamentally information, and specific symbols that communicate the information, that don’t depend on the specific substrate that hosts them. If the mind is also about information processing, all this same processing could happen in a simulated brain, and thus it would have the same qualities of mind we call real thought and understanding, for the same reason that you’re reading real words right now.
I think a lot of people’s immediate reaction to this is that the simulated brain wouldn’t have phenomenal consciousness in the way we have. Again, I’m leaving that debate aside, and would just ask you to recall the arguments I made in Part 1 that phenomenal consciousness doesn’t seem to be required for real thought and understanding. What matters is the right ways of handling information, which the simulated brain would have.
“Simulation” is itself just a name for another physical process. If we go back to the example of an emulator, a simulated computer, what’s actually happening is that we’ve basically just added an additional layer of programs and code to another program and code, that are all interacting with each other. There’s no deep difference between a simulation of a computer running a game, and running a game that just happens to have a lot of specific extra layers of code to it that it comes with. This is all ultimately just different specific causal relationships between computer code, which all come down to transistors in computer chips reacting to each other in specific ways, and it’s convenient for us to call a part of that code “a simulated computer” because it happens to be able to also perform the functions of a computer with other programs. These are just two separate methods of doing the exact same information processing, in the same way that the same type of boat can be made from different types of wood, or with a second layer of wood. If what matters for thought and understanding is how the brain works with information, not the specific atoms it’s made of, and a simulated brain would do the exact same things with the information, then what it’s doing seems to also be fully and completely “thought” in the exact same way the video game run on an emulator is fully and completely the same game.
Maybe the most famous objection to this is John Searle’s line about how simulated rainstorms don’t get anything wet, and simulated fire doesn’t burn down real neighborhoods, so why would a computer simulation of understanding actually understand anything?
This looks to me a lot like saying that the words you’re reading are just simulated words, so they can’t really convey information in the way a physical book can. If I’m right that thought is a kind of information processing, it’s mixing up what understanding and thought are. The nature of water and fire is fundamentally bound up with specific physical properties resulting from what they’re made of. But information is independent of any one substrate. These words can exist in my head and as electrical signals in my computer and radio waves to my router etc. Many simulated games have both things that we’d consider “real” and “fake” that are obviously differentiated by whether the thing depends on specific physical substance. Fire in Minecraft can’t actually burn anything, it’s fake, but if I build a calculator in Minecraft it can really do arithmetic, so it’s a real calculator. The realness of information processing clearly does not depend on what substrate is hosting it.
So everything comes down to which kind of thing thought is. Searle compared understanding to biological processes like lactation and photosynthesis. Photosynthesis produces sugar, a physical substance, but as I argued earlier, what thinking produces is patterns in your speech, your movements, and your writing, which is a lot more like what the Minecraft calculator produces.
Maybe real understanding requires a specific embodied interaction with the outside world, so it does rely on our physical bodies in a way computers can’t access. I’m very skeptical of this, because our interactions with the world are entirely mediated through systems in our body that convert them to the type of information our brains can handle: electrical and chemical signals in our nervous system. When I have an embodied experience of walking through a forest, the physical wood the trees are made of isn’t directly involved in my mental processing. Instead, what’s happening is that my body has receptors all over it (touch receptors in my skin, cells in my eyes that detect light, cells in my inner ear that respond to sound, etc.). This all seems obviously there to pick up on informational patterns light or sound or the texture of objects is carrying, and then convert it to the type of information patterns that the brain can work with, in the same way that while I can experience a lot of different patterns in the world that I can communicate in my writing, my wifi router can only work with pulses in radio waves, so the information needs to be translated into those before it can deliver them to you. The question then becomes whether a simulated brain could receive the kind of constant informational input that human bodies receive, and that seems like a difference of degree in access points to information rather than a difference in the kind of thing the simulated brain is doing.
Programs aren’t just in the eye of the beholder
A more sophisticated objection to the idea that a computer simulation of a brain could really think is that whether a computer is running any specific program at all is in the eye of the beholder.
In Borges’s short story The Library of Babel, the universe is a library containing every possible book of 410 pages composed from a fixed set of symbols. Its inhabitants search through overwhelming nonsense for books that explain their lives and their world. Any answer that can be written within the library’s format must be there somewhere, alongside countless false answers. But the narrator also suggests that apparent gibberish may carry meaning in an unfamiliar language. Taking this idea further, any book could be made to say anything, provided we were free to invent the language or code in which it was read. Any book could explain everything about your life if read with the correct language.
Similarly, any object, including rocks or walls, can be said to contain any computer program, if someone just used the right tools to read them in the right way. If I had the exact right scanner for a specific rock, where the scanner read the details of the rock’s surface as inputs, and I taught it to translate those signals in specific ways, I could determine that the rock contains full instructions for running the computer game Roller Coaster Tycoon, and play it off the rock like I would a CD. Searle argued that minds seem much more fundamental to the world than this, so computer simulations of minds can’t be real minds, because whether a simulation is doing anything in particular is just a story we tell about it.
The standard response to this, which David Chalmers worked out in detail in a 1996 paper called “Does a Rock Implement Every Finite-State Automaton?”, is that running a program can’t just rely on physical states happening to line up with a program’s steps. It requires the right cause-and-effect relationships, including what would happen if things went differently. A rock only “runs” Roller Coaster Tycoon in the specific sense that someone could run a very specific linear play through of the game from beginning to end, the rock lacks the responsiveness to enable the player to make any decisions or include any cause and effect relationships.
A simulation of a brain could easily respond to new causes with new coherent effects that a rock or wall couldn’t. If you showed its simulated eyes a different picture, its neurons would fire differently in the way yours would. Any ways it’s still “fake” fails my brain test from earlier, where whether you yourself have a mind is an equally arbitrary matter of perspective.
There are a lot more philosophical arguments about the meaning of understanding and thinking and how it relates to simulated brains, most famously Searle’s Chinese Room. Before getting to those, I need to address a more basic common scientific complaint, that “the brain is not a computer,” and then build an argument against those other philosophical objections based on the fact that, the way they’re usually used, most of them would also rule out any physical system thinking, including the brain.
What do people mean when they say “the brain is not a computer”?
One of the most common ways people push back against what I’m arguing here is to say “but science says the brain isn’t a computer.” There are many popular articles making this claim:
Robert Epstein’s 2016 Aeon essay “The empty brain,” with a subtitle that says “Your brain does not process information, retrieve knowledge or store memories.” That seems bad for my argument!
Matthew Cobb’s Guardian article “Why your brain is not a computer,” which was adapted from his book “The Idea of the Brain”
The neuroscientist Anil Seth has argued that the brain is not a computer, and that the idea that computation is enough for consciousness comes from taking that metaphor too literally.
These articles often get shared to say that AI cannot in principle think or understand. I’ve seen Epstein’s article shared a lot on Twitter recently. However, I think that even if every claim in these articles were true, it wouldn’t change my argument at all, except for two claims. One is that brains don’t process information, which seems obviously false. The other is that the brain is analog in a way computers can’t capture, which is the strongest version of this objection, and I’ll answer it below.
Look back at the five steps of my argument at the top. None of them says the brain is a computer. I’m saying something more like “The game Monopoly can be fully run and played on a computer. All relevant aspects of Monopoly can happen in a computer” rather than “The game Monopoly IS a computer.” Even if the brain is nothing like a computer, computers successfully simulate things that are not at all like computers all the time, like weather, and as I argued above, thought is the type of thing that is really truly thought if it is successfully simulated. Whether a computer can do what the brain does is a completely separate question from whether the brain is a computer, for the same reason that whether my computer can play Roller Coaster Tycoon is separate from whether Roller Coaster Tycoon is itself a computer that can run programs. Large language models aren’t computers either, in the same way Roller Coaster Tycoon isn’t. They’re things computers run.
Part of why this distinction gets lost is that the word “computer” actually has two different meanings that regularly get confused. In everyday speech, it means something like your laptop, that has a very specific architecture involving memory and hand-written code. In computer science, it has a more technical meaning that goes back to Turing, where it means any system that processes information by following a well-defined set of rules or instructions, regardless of its physical makeup. Our brains might be computers in the second sense, and obviously work very differently from the first sense. This is a good paper from 2022 arguing that the whole debate is so semantically confused that it doesn’t seem worth having. This is one reason why I’m interested in constructing an argument that doesn’t hinge on it at all.
Take Epstein’s essay, which has been going around a lot recently. His main point is that the brain doesn’t work like a laptop. It doesn’t store clear, distinct memories that can be easily retrieved, and it doesn’t follow clear internal written rules for how to use words. On most of this, I agree! But his picture of the brain is a much better description of current AI models than of the computers he’s contrasting the brain with.
Take this passage:
As we navigate through the world, we are changed by a variety of experiences. Of special note are experiences of three types: (1) we observe what is happening around us (other people behaving, sounds of music, instructions directed at us, words on pages, images on screens); (2) we are exposed to the pairing of unimportant stimuli (such as sirens) with important stimuli (such as the appearance of police cars); (3) we are punished or rewarded for behaving in certain ways.
We become more effective in our lives if we change in ways that are consistent with these experiences – if we can now recite a poem or sing a song, if we are able to follow the instructions we are given, if we respond to the unimportant stimuli more like we do to the important stimuli, if we refrain from behaving in ways that were punished, if we behave more frequently in ways that were rewarded.
As I’ll discuss in Part 3, AI models don’t store what they learn as separate files in memory that they can easily pull up the way your laptop stores information. The information they “learn” is spread across their weights connecting their “neurons” and it can’t be read off as clearly written rules. Like people, they can still memorize poems word for word with enough training, but this is very different from a computer storing the poem’s text in easily accessible memory. This passage describes current AI models so well that it’s strange to see it shared as evidence that AI can’t think.
The one claim in the essay that would actually hurt my argument is that brains don’t process information at all. Epstein seems to be using “information processing” in a much narrower sense than the formal one. He’s attacking what psychologists call the “information processing” model of the mind, where the phrase means storing, retrieving, and manipulating symbols, a lot like a laptop does. In Shannon’s sense, and information theory in general, information is anything that reduces uncertainty, and Epstein’s own description of memory, where learning a song changes your brain so you can sing it later, is a clear example of processing information. Learning the song reduces uncertainty about what notes to sing when. He also describes a student who can’t draw a dollar bill accurately from memory, as evidence that nothing like a picture of the bill is stored in our brains:
I agree with Epstein that this obviously shows that the mind doesn’t process information in the specific way a laptop does, but it also shows by standard of information theory more broadly that the student clearly had a lot of information about dollar bills, just not perfect information. They could answer lots of yes or no questions that reduce uncertainty, like “Are dollar bills triangular?” or “Do dollar bills have ‘In God We Trust’ written on them?” His fly ball example is a good one too (outfielders really do follow a simple visual rule instead of calculating where the ball will land), but following a rule like that is still processing information about the ball. None of this affects my argument, which never assumed the brain stores perfect copies of anything.
A lot of people seem to be getting tripped up on this semantic confusion between the strict information processing model of the mind, where the phrase means storing, retrieving, and manipulating symbols like a computer, and the more broad definition of information from information theory. It’s undeniable that the brain processes information in the second way, and I agree with the critics that it doesn’t “process information” in the first much more narrow way. But the confusion between these claims is causing some people to make the bizarre assertion that what happens in brains is somehow not really information and not available to Turing complete systems, neither of which actually follows from the correct use of the term.
Anil Seth’s commentary is more interesting and relevant to modern AI, but he’s talking about consciousness, and seems to separate it from questions of machine intelligence like I do. He has asked whether it’s possible to understand something without being conscious of it, and Part 1 was my argument that it is, so if Part 1 is right, his argument is also irrelevant to my claim. Maybe he’d agree? I’m not sure.
One last note on this is that a lot of people will often complain that the mind is always compared to the latest technology at the time. When we were in the Newtonian era, the mind was compared to clockwork, then when electricity and information technology took off it started to get compared to wires, then computers, then AI. I personally think each of these was actually a pretty useful metaphor compared to the common wisdom at the time, maybe with the exception of computers. It was a radical proposition when Hobbes suggested that humans are basically very complex machines in Leviathan. The mechanistic explanation of minds to this day draws a lot of ire and scorn despite being most compatible with our best science. Connections of wires transmitting information again looks like a much better explanation of minds to me than a lot of what was on offer in the 19th Century. AI has renewed debates in how much of our own thinking is about prediction and pattern matching and carrying those forward, and again this is leaps and bounds better than the folk theory people carry around and get mad when you poke at it. I’m not really troubled by these metaphors, and insofar as they’re bad (like the computer metaphor reinforcing the idea of clear easy to access concepts and memories stored in the brain), they are bad insofar as they reinforce our folk pre-scientific concepts of minds that they are otherwise helping us to move away from.
But isn’t the brain analog?
One other point that gets brought up a lot is that the brain is analog, in that it operates with a lot of continuous processes, whereas a computer being digital has to work in discrete steps. Anil Seth makes a point like this about consciousness. Maybe if thinking depends on infinitely fine-grained details of continuous physical processes, a digital simulation of it will never be able to model it exactly. It’s like trying to redraw a circle with lines that can only go up down left or right. You could approach the smooth curve of the circle infinitely but never fully get there.
But all real analog systems can only ever make use of a limited amount of precision, there’s always some random noise too. Your neurons can’t respond differently to two voltages that differ from each other less than the random jiggling of the molecules around them, that gets lost to noise. Anything below that is invisible to the brain’s information processing, so a computer program only needs to get down to this level of detail to perfectly mimic the information processing in the brain. There must be some finite level of digital detail that can capture everything relevant to thought in the analog systems of our brains.
A lot of arguments that computer processes can’t think are arguments that any physical system can’t think
Some of the most common arguments that AI can’t think seem to unknowingly reduce to “it’s just a physical system,” which would also apply to brains. If an argument applies just as well to your brain, it implies nothing can think or understand. This is the brain test I mentioned earlier.
The classic version of this argument came from the philosopher Leibniz in 1714. He imagined a thinking machine blown up to the size of a mill so you could walk around inside it. All you’d find, he said, are parts pushing on each other, and he took this as proof that perception couldn’t be explained mechanically. (The header art in this series is a jokey reference to Leibniz’s mill.) A lot of modern arguments against AI thinking look like versions of Leibniz’s mill. You can look inside an AI model, find only arithmetic and cause-and-effect relationships, and conclude that it’s not actually thinking. But a brain the size of a mill would also look to us like it wasn’t thinking, since it’d just be ions and proteins bumping into each other. The fact that we can’t see where the thinking happens when we look at the individual parts of a mechanistic system tells us nothing about whether the system thinks. Leibniz’s mill fails the brain test.
Two other common arguments that fail the brain test (assuming what happens in the brain is physical) are the claims that AI models are too mechanistic to think, and that they’re just a version of Searle’s Chinese Room. There’s also my argument in Part 1 that we often require AI to have a kind of introspection that seems physically impossible for humans to have.
When you give an AI model a prompt, every calculation it does is fixed by the prompt and the model’s weights. AI models have a randomness setting called temperature, and if you turn it all the way down to zero, the model always picks the single most likely next word, so the same prompt will give you the same answer almost every time. Because the chips don’t always add numbers up in the same order, tiny rounding differences can sometimes change a word, which acts like a little extra randomness. But that doesn’t seem to add anything to the model’s ability to think.
If the temperature is above zero, the only reason you get different answers is that a random number generator is choosing among different words, weighted by how likely the model thinks each one is. The AI model is really just a domino chain plus dice: a completely straight line from cause to effect, with some randomness thrown in if you’d like.
It seems obvious that dominoes and dice can’t think. A straight single line of dominoes from premises to conclusions can’t think and understand. A straight causal deterministic line from one given input to one specific output isn’t thinking along the way, and adding randomness to this process doesn’t make it more like thinking at all. In fact, it makes it even less like thinking because pure randomness doesn’t tell the system anything about the question it’s answering. That means that the AI model must not really think.
The domino-and-dice picture is in fact a pretty accurate description of how AI models work. Any AI model can be understood as an incredibly complex web of domino chains, plus some randomness, and nothing else. The problem is that this is also a complete description of every physical system, including your brain. As I argued earlier, every physical system is made up of unimaginably vast numbers of processes that are either mechanistic or random, and there’s no third option. Your brain is a physical system, and your thinking can be described by classical physics (plus some randomness). Suppose we asked you a question and recorded your answer, then rewound the universe to exactly the same state, atom for atom, and asked you again. You’d follow the exact same chain of cause and effect and give the same “output,” for the same reason a billiard ball hit at the same speed and direction always ends up in the same spot. The only way your answer could come out differently is if some random events (like an ion channel happening to open) went differently the second time. Your brain is dominoes plus dice too, so this argument fails the brain test.
The argument that AI models are too mechanistic to think therefore implies that no purely physical system can think or understand. I’ve given a lot of reasons in this post and the last one to believe that thinking and understanding rely on purely physical processes in the brain. Either our sense that thinking needs something beyond mechanistic cause and effect plus randomness comes from a vague folk theory about what happens when we think, or AI models are missing some third way of processing information that’s neither mechanistic nor random. The first seems much, much more likely.
Many arguments that AI models can never really think are based on the idea that their mechanistic cause-and-effect processing is basically a form of unknowing information retrieval. People usually back this up with two famous thought experiments:
Ned Block’s “Blockhead”
John Searle’s Chinese Room
Blockhead is a computer that simply has a canned response stored for every possible conversation. You say “How are you?” and it just matches this to a pre-written response it’s supposed to give: “I’m well!” This obviously wouldn’t really understand anything. Block meant this thought experiment to show that whether something is intelligent depends on how it produces its answers. So passing a conversation test like Turing’s doesn’t, on its own, show that a machine is intelligent. I agree with that.
The Chinese Room is a bit different. In a 1980 paper, Searle imagined a man who doesn’t speak Chinese locked in a room with a rulebook (a program) for shuffling Chinese symbols. People input questions in Chinese into the room. The man follows the rules, and he passes back answers that are indistinguishable from a native speaker’s, without actually understanding a word. Searle argued therefore that following any step-by-step program, however complicated, can’t by itself be what we mean by understanding what the symbols mean, because programs only work with the shapes of symbols and can’t access their meaning directly. Some people in the AI debate now use the Chinese Room as if the rulebook were just a list of canned answers, which turns it into Blockhead.
I agree that Blockhead doesn’t really understand, and neither would a Chinese Room whose rulebook was just a list of canned answers. But that’s only because they lack the internal causal processes that would let them respond to fundamentally new situations using the meaning of the words. If the rulebook could mimic literally all the causal processes required for thinking and understanding, the man and the rulebook system together would “understand the words” the way we do, even though neither does on its own. It’s the same reason my brain matter and the laws of physics together produce a system called Andy that understands words, even though my individual neurons and the laws of physics don’t understand. Here the room and book are my brain matter, and the man is the laws of physics. I’ll address the comparison to Blockhead first, and then the point the full Chinese Room is making.
My problem is with how the two thought experiments get used in the current AI debate. People correctly point out that AI models can be described as “mere information retrieval,” because some specific answer is waiting at the end of any possible prompt. With the temperature at zero, a given prompt can only go one way, and any deviation from that is due to randomness. So maybe models are like the canned-answer version of the Chinese Room, with the guy inside sometimes spinning a wheel to pick at random between a few similar responses he doesn’t understand at all. But any purely mechanistic system can be described this way. If you let a block slide from rest down a frictionless ramp of height h, it will always reach the bottom at one very specific speed (√(19.6h) meters per second, with h in meters). In this sense, its final speed was “canned” and merely “retrieved” by the input. If you ask someone a question, their brain follows classical physics to arrive at one very specific answer. The only things that can change it are randomness and new input from you or the world while they’re answering. The brain can only respond to any input with systems that are mechanistic or random, and our sense that it’s some secret third thing comes from just how much information it’s handling. It’s hard to even imagine what that secret third thing could be.
Every physical system’s outputs are fixed by its inputs (plus randomness), but that’s different from the outputs being written down somewhere in advance and looked up. Blockhead looks its answers up. Brains and AI models have to work theirs out. There are more possible 30-word prompts than atoms in the observable universe (even if you only used the 10,000 most common English words), so no model could have an answer stored for each one. Block himself noted that a real Blockhead could never be built for this reason. So if “retrieval” means the answer is fixed by the input, brains are retrieval machines too. If it means the answer was stored ahead of time, AI models aren’t retrieval machines at all.
I separately worry that part of the reason the Blockhead and simplified Chinese Room examples are so compelling to people is that most people believe that AI models literally just have a large collection of canned responses. In a Searchlight Institute poll from August of last year, 45% of Americans said ChatGPT looks up an exact answer in a database, and another 21% said it follows a script of prewritten responses:
Only 28% got this right, which might be partly because the options seem kind of poorly written to me. Still…
If your takeaway from the Chinese Room is that anything whose outputs are fixed by its inputs (plus randomness) is by definition not really thinking or understanding, then you don’t believe any physical system can think or understand. That means the mind must hinge on extra nonphysical stuff. My claim in these last two posts has been that thinking and understanding very likely depend on purely physical processes. This makes it look more likely that the intuition that “mere mechanism + randomness” isn’t enough for thought comes from a folk theory that dissolves when you poke at it. And in fact when people try to pin down what these purely rote cause-and-effect processes are lacking, they get very hand-wavy and defer to a strong felt sense that understanding requires something over and above roteness + randomness. As I argued above, that strong felt sense might come from evolutionary pressure that rewarded early humans for only treating carbon-based living things as having minds. That pressure itself comes from carbon forming stable bonds that hold up in water, so carbon-based molecules could work as replicators billions of years before brains existed. This felt sense seems like a bad guide to machine thinking.
Turning now to the actual Chinese Room thought experiment that isn’t just a restatement of Blockhead. Searle meant it as evidence that manipulating symbols can’t on its own produce understanding, so running a computer program can’t either. But “blindly following step-by-step rules that can be written in a book” also describes any physical system that processes information, including the human brain. Searle actually agreed that brains can be described this way. He said we’re machines ourselves, running any number of programs, but that the programs aren’t what make us understand. If the rulebook could somehow contain the rules for everything every physical process in the brain would do given a specific input, I’d argue the system running it could think the way a brain does. Searle considered this case in the same paper. He was responding to the idea that a computer simulating every neuron firing in a Chinese speaker’s brain would understand Chinese, and he imagined the simulation built out of water pipes instead.
He imagined the simulation being built out of water pipes instead. A man who doesn’t speak Chinese operates an elaborate system of pipes and valves, with each pipe connection playing the role of a synapse in the Chinese speaker’s brain. He receives inputs and follows instructions telling him which valves to turn on and off. Eventually, the entire system produces answers in Chinese that come out the other end. He claims that the man operating the system obviously doesn’t understand Chinese. The pipes, “obviously,” don’t understand, and so Searle thought it was absurd to say the man and the pipes together did. He then concluded that a brain simulation only copies the formal structure of the brain’s neurons and leaves out the brain’s ability to understand. Thus, biological brains can understand, but simulations of brains cannot.
I think the man is a distraction here. In the pipe system, he’s playing the role that the laws of physics play in a brain, moving things along according to fixed rules. Asking whether he understands Chinese is like asking whether the laws of chemistry understand English. Searle’s reply to this kind of point was that the man could memorize the whole system and run it in his head, so there’d be nothing in the room but him. But as I’ll get to in a second, that would mean memorizing something like 100 trillion connections and running them at the speed of a brain, which no person could do, and at that point he’d be simulating a second brain in his head, and as I’ve argued before, that simulated brain’s thinking would be identical to a real brain’s thinking, so it could also really think and understand.
In this thought experiment, Searle seems to be leaning on the idea that to really understand a word is to be conscious of it, to have a subjective phenomenal experience in your inner mental theater to use it correctly. He defended this directly elsewhere, with what he called the “Connection Principle,” which says a mental state can only really be about something if it could at least in principle become conscious. I don’t see another reason to say that the pipe system doesn’t understand the words it’s using. I’ve argued in Part 1 that this view of language and understanding as requiring this inner phenomenal theater seems wrong. It doesn’t cohere with physics, or with how words actually behave, or even with our own subjective experience when we use language. It also doesn’t appear “obvious” to me at all that this system wouldn’t understand the words it’s using. To argue for my own view that the pipe system doesn’t need extra-physical phenomenal consciousness to understand the words it’s using, I’ll paint a picture of what it’d actually look like.
There are roughly 100 trillion synapses in the human brain, and since the average neuron fires about once a second or less, a simulation like this would need up to around 100 trillion valve operations per second. This guy managing the pipes would either have to keep up that pace, following unimaginably detailed instructions about which valves to turn depending on what the water is doing, or (operating at a more human speed) each response would take millions of years. Even if each valve and its pipes fit in a box about 4 inches on a side, the whole structure would be a cube almost 3 miles on each side. Suppose he could keep up with the brain’s real speed. You could theoretically remove my brain, hook up my body to signals generated by the outputs of this water system, and I would walk around and interact with the world, and because Searle is assuming that this water system can actually perfectly mimic all the same information processing happening in the human brain, any informational input to my body, whether it be a sight or sound or touch or impulse from my body, would be taken in by this massive system, and it would give out the same information back that my brain would have. I would say and do the exact same things real Andy would, including sitting down to write this long blog post. It’s easy to lose sight of the fact that “having the exact same ability to merely process information” as the human brain would lead to systems that can also perfectly run human bodies to the point that they could type blogs and talk and interact with the world without anyone noticing. This version of me run by the water system would write the same blog posts about data centers and water and go on podcasts to talk about them, no one noticing he was using way more water than all data centers… This watery version of me controlled by the pipes I’d say would “understand” the world just as well as I do. Literally all the ways I can use language, it can as well. This seems like a point in favor of my belief that phenomenal consciousness isn’t necessary for understanding.
Similarly, if someone were in the Chinese Room with a sufficiently advanced rulebook, that could simulate all the processes in my brain with trillions of little individual decisions like this, they could perfectly accurately control my brainless body and I’d live out my life exactly as I’m doing now. I’d use words in all the same ways I do now. The fact that the man in the room is merely following instructions without understanding them tells us as much as the fact that the physical laws governing my brain don’t themselves understand what I’m doing, which isn’t much. Sufficiently-advanced mere rule following can simulate any form of information processing down to an arbitrary level of detail, so I argue that the full Chinese Room rulebook can think and understand the words it’s using with some sufficient level of detail in the rules the book uses (which would add trillions of steps per second per input).
Searle himself anticipated this reaction in the same paper, and called the entity like the version of me controlled by the pipe system or room “an ingenious mechanical dummy.” This seems ridiculous to me. At this point, Searle’s “ingenious mechanical dummy” can do anything and everything humans can do with words. It could live a life identical to mine. It could write all these blog posts and go about its day. It seems at this point that Searle’s relying so heavily on the idea that “real understanding” depends on phenomenal consciousness that he’s willing to say a system that could do literally everything humans can do with words in any and all contexts doesn’t understand the words. At this point the disagreement itself seems pretty semantic to me. If AI is a Chinese Room, it can do literally everything human minds can do, and just happens to lack some additional thing we call “understanding” that doesn’t have any effect on how we behave.
Searle makes a lot more arguments about why computer programs can’t think, including that they merely follow step-by-step rules and lack the “intentionality” of what the words they use are about. I claim that his argument here cannot be made to work with any conception of the brain as a physical system, because all physical systems are reducible to mechanistic step-by-step cause and effect processes. If no combination of these can think, the brain can’t either. Searle’s Chinese Room often looks to me like an argument for a nonphysical Cartesian theater where the true meanings of words can be grasped by something not merely mechanistic (and thus not merely physical). Searle would deny this. He would often say that humans are biological machines. But his demands for what minds can do seem to require something over and above the physical to make them work, and if you agree with my argument so far, it seems much more likely that this itself is a folk theory that masks the possibility that the processes that happen in our brain can be reduced to more simple physical processes.
If thinking is a kind of information processing, and nothing about carbon makes it special, then Searle’s pipe intuition is relying on exactly the two things this series has argued against. One is a sense that only biological stuff can think, and the other is the idea that understanding a word requires a conscious inner observer experiencing its meaning. This whole series so far has been an implicit argument against the view of language and understanding that Searle puts forward as “obvious,” and I hope that after making it this far you can see why I don’t think it’s obvious at all.
Conclusion
To return to my main argument:
Premise 1: Things that think and understand are reducible to physical processes. There’s no extra nonphysical ingredient.
Premise 2: What makes a physical system think is how it processes information, and what it’s made of only matters insofar as it can do that processing.
Premise 3: A computer can simulate the brain’s information processing to whatever level of detail matters.
Premise 4: A detailed enough simulation of an information process is that process, running in a new substrate.
Conclusion: a computer process can, in principle, think and understand.
If you accept all 4 premises, you believe that thought and understanding are specific types of information processing that happens in physical systems. Information is independent of specific substrates as long as that substrate can carry the information and perform the cause and effect behavior required to process it. It would be very weird if there were something fundamental to carbon over and above the causal role it plays in larger systems that made it necessary to thought, and the only reason all minds seem to have run on carbon-based biological matter so far is that it was the only stuff evolution could act on to create complexity in the first place, which minds require. The question is next whether silicon can run the same causal processes required for the type of information processing that makes up thought. Silicon computer chips are Turing complete, which implies they can simulate any classical physical process to an arbitrary level of detail. The type of information processing in a human brain would require a lot of detail, but it relies on classical physics and doesn’t rely on infinitely small details, so it’s a finite amount, so all information processing that happens in a physical brain could in principle also happen in a sufficiently detailed and accurate simulation of a brain, and since thought is just the right type of information processing, this simulated brain could also truly think. Thus, computer processes can in principle think and understand.
So if you’re with me, you agree that computer programs can in principle think and understand. The next question that Part 3 will try to answer is whether AI models specifically are specific instances of computer programs thinking and understanding. This is a long stretch from simulated human brains. AI models are very different from human brains. But remember that octopus and crow brains are also radically different from human brains, and yet they also seem to be able to do what we mean by thinking. It’s on me now to argue that AI models have the same internal qualities that together qualify for this, which will require a deep dive into how current AI models work and how they might successfully mimic all the processes that make up thought and understanding, to the point that they are real in the same way calculators built in Minecraft are real calculators.














Congratulations on another great piece. A few comments.
"Thinking and understanding are things physical systems do." This, in the premises, seems to potentially confuse readers to think you're arguing that all physical systems think and understand, perhaps even some kind of physicalist panpsychism. I don't think this is your intent?
"...literally any physical system is just a Chinese Room" as the logical conclusion of the CR argument against AI thinking: would this not resolve by admitting this conclusion, including for the human mind, naturally sans free will? I await your CR deep dive.
"the mind is always compared to the latest technology..." Tangentially, so have humanity's explanations of their gods: a potter, a weaver, a watchmaker, an engineer; recently, a simulation.
For anyone curious about the 2018 epilepsy-patient study, Opus thinks it's Hampson et al (2018) (https://iopscience.iop.org/article/10.1088/1741-2552/aaaed7/pdf). See also Roeder et al (2024) (https://www.frontiersin.org/journals/computational-neuroscience/articles/10.3389/fncom.2024.1263311/full), which is a follow-up from the same group.