A privilege of being a physics major is learning just how strange fundamental reality is compared to the concepts we developed to make sense of it. As far as we know, all the stuff we’ve ever run into or interacted with can ultimately be written out like this:
In a more legible form (already an imperfect metaphor), this says that all ordinary matter is some combination of these particles (which are really ripples in underlying fields), acting on each other through fundamental forces:
If you put them together in juuuust the right way, you eventually get Antarctica, evolution by natural selection, or tables. These things are our imperfect maps and metaphors for the physical reality underneath, but we often come to believe they’re foundational entities in the world, as if the lamp I see in my room shares some fundamental lampness with all other lamps, and at the heart of reality there are neatly sorted categories like “lamp” or “continent” or “human.” The category of “lamp” is useful, but it would be a mistake to think that lamps are a fundamental part of reality, and that there will be 100% clean and predictable ways all lamps behave that I can infer from any one of them.
A frustration I regularly have in debates about AI is when people act as if the words we use about it are getting at fundamental categories of reality itself, even when they’re clearly especially imperfect, pragmatic metaphors we use to predict our human-scale world. In these moments I kind of want to grab the other person and announce to them that the concepts we use to think about the world in general are just very imperfect social conventions we use to try our best to find these basic but fallible patterns in the incomprehensible complexity lurking beyond our perception. This is obvious in most other circumstances, but it somehow often gets lost in debates about the future of AI.
One example that drives me crazy is the statement that AI is “just a tool.” People seem to infer from this that AI is and always will be completely inert, like a hammer, and that it’s the person using it or designing it who’s responsible for anything bad that happens. This just assumes away the possibility that AI could do things the user doesn’t want, or that the designer didn’t expect and had no way of foreseeing. How do people arrive at this conclusion? By studying how AI actually works? It looks to me that instead they mostly get there because AI, like tools, is something created by humans to be used by humans to solve problems, or that it’s not a literal biological animal. I don’t see anything in how AI models are trained that implies they’ll always behave the way we want, like a hammer does, but this metaphor seems to convince a lot of people that they will. Why?
When people say this, I worry that they believe they’ve accessed some real part of fundamental reality called “tools” and discerned that AI has this same tool-ness as hammers. It seems obvious to me that the word “tool” is a useful but wildly imperfect social game we’re playing to try to make sense of reality. There’s no fundamental tool-ness at the heart of reality. Anything from a rock to a nuclear weapon can, in the right circumstances, be called a “tool.” In some sense a vaccine is a “tool,” but some vaccines contain live pathogens that, in rare cases, behave in ways nobody wants. Does the fact that the vaccine is a “tool” tell us much about how the pathogen inside it will behave? Or that no one could ever have a bad reaction to it unless the person who made or administered it intended that to happen? It seems like the fact that the vaccine shares some properties with other tools doesn’t automatically tell us anything about other unrelated properties it might have. I need some clear additional reason for the fact that a vaccine is a “tool” to imply that it shares any one property with other tools. I can’t just say “Well it’s a tool, so therefore like a hammer it will only cause problems if the person using it intends it to.” That’s obviously silly! Yet in conversations about AI, moves like this regularly fly.
I would think that “AI is just a tool” is meant as a stepping stone to a real debate about whether it’s controllable the way other tools are, but instead most people who say it seem to want to frame the entire conversation around any and all evidence that AI has some inherent property of tool-ness that other things in the universe share. If it’s created by people, that gets counted as evidence that it has this tool-ness, and therefore evidence that it will be inert and controllable. I don’t believe this kind of property actually exists, so this step can’t be taken. It’s all just atoms and the void! So I don’t believe that a fact like “it was created by people” can tell us much about whether AI has some other, unrelated property, like “it will always be easy for the user to control,” just because the completely different objects we call “tools” happen to have both. Drawing this inference seems about as reasonable as saying that because candlesticks aren’t made the way weapons are, a candlestick can never be a weapon. It assumes there’s some part of fundamental reality that literally always connects how a thing was made to what it can do, and I just don’t see the connection.
This has also come up for me in a lot of conversations about when it’s right or wrong to “anthropomorphize” AI. I agree that because AI writes as if it’s human, it’s often tempting, and wrong, to ascribe qualities of the human mind to it that it doesn’t currently have. But critics of anthropomorphizing AI often go way farther and imply that “human” is itself a fundamental category of reality that a system either belongs to or doesn’t, and that because AI isn’t human, it will never be useful to describe it in anthropomorphic terms. This seems completely mistaken to me. I don’t think “human” is a fundamental category of reality either. What we experience as human is the very specific result of incredibly complex systems of biology, neural wiring, and social and economic interaction, each of which is itself made up of simpler nonhuman systems working together in nonhuman ways to give rise to entities like us. To say that we ourselves are human is itself a useful simplification of the complex nonhuman processes that work together to make us up. When I attribute intent to another person, I’m not actually saying “There is some fundamental property of the universe called ‘intent’ that this person is experiencing.” What I’m actually saying is “The complex patterns in this person’s brain and behavior line up in such a way that saying they ‘intend’ to do something helps me predict what they’ll do and experience.” Everything we consider human is a pattern that emerges from simpler nonhuman processes, and the words we use to describe ourselves are imperfect metaphors that help us imperfectly predict and react to those more complex processes in ourselves and others.
Our minds don’t actually work like our folk theories of psychology suggest, and yet our folk language about them still often yields good predictions. Because anthropomorphic language is ultimately a useful fiction imposed on a much more complex process, there are times when it might also be a useful fiction to impose on AI. I think saying an AI “filled out a form” is a useful description of what it did. It often, but not always, yields the same true predictions about the AI’s behavior as it would about a person’s. Similarly, the most recent models check a lot of the boxes in my own definition of what it means to “think.” I don’t think AIs have first-person experience, but they seem to form world models, move step by step through chains of reasoning, and manipulate words in ways that seem impossible without granting that they basically “understand” them. Is it wrong to say that AIs “think”? I think this debate should focus entirely on the pragmatic question of how saying so would affect our predictions about AI. It’s bad to say AI can “think” if it communicates to the other person that the AI is very humanlike, has first-person experience like we do, or has complex emotions and intentions like ours. But it seems useful to say AI can “think” if it communicates that AIs can work with language the way we can and draw reasoned inferences from that language based on world models. In this case, it seems good to anthropomorphize AI for the same reason it’s good to anthropomorphize people: both give us good predictions about the world. In figuring out when it’s right or wrong to use “anthropomorphic” language, it seems useless to frame the question around some abstract fundamental category called “humanness,” existing above and beyond the physical world, that AIs either are completely or not at all tapping into. I don’t think this is something that actually exists, and so trying to figure out whether AI can access it or not won’t tell us anything about what they’re currently like or will be like in the future.
We aren’t able to think without simplified concepts and metaphors. But we have a responsibility to recognize that basically all our concepts are simplified and imperfect. Figuring out the near-term future of AI requires us to work out the specifics of when our imperfect metaphors are useful or not. This debate is not helped at all by conversational games that imply that the speaker has accessed some fundamental category in the universe that AI belongs to along with other entities, so it must always be exactly like those other entities, and any property they have, it must have too. The same game gets played in reverse, where AI is declared to be outside some category like “human,” so it can’t share any property with the things inside it. Neither move tells us anything, because the categories aren’t real. There’s no tool-ness at the bottom of reality that guarantees AI will stay as controllable as a hammer, and there’s no humanness that AI is locked out of that settles whether it can think. Words like “tool,” “human,” and “think” are shorthand we invented to get by in a world far more complicated than we can perceive, and the only real question about any of them is whether applying the word to AI helps us predict what AI will actually do. If it does, we should use it, and if it doesn’t, we should drop it. Whether AI carries within it some fundamental nature of tool-ness or human-ness is a terrible way of approaching a question about it, because like everything else, AI is ultimately just atoms and the void. Or more precisely, the universal wave function.




Great post, but also, I’m guessing a lot of the people you’re responding to would be skeptical that humans can be reduced to a wave function. And while I think they’re directionally wrong, there is something to be said for the fact that the universe and humans remain mysterious in some fundamental ways. Eg, quantum mechanics doesn’t mesh with general relativity, and neither seem to explain how subjective experiences arise. Of course, none of this justifies thought-terminating metaphors about AI.
As usual, Yudkowsky scooped us all by a couple of decades.
https://www.lesswrong.com/s/SGB7Y5WERh4skwtnb