Discussions of anthropomorphism in artificial intelligence are often strangely confused. One person says that an AI “wants” something, another replies that this is anthropomorphism, and the exchange ends as though a conceptual mistake has been identified. But very little has actually been resolved. What does “wants” mean here? Is it a claim about observable behavior, functional organization, internal cognition, subjective experience, or moral status? And what exactly is supposed to be wrong with describing a nonhuman system in terms that were first developed for understanding humans?
The usual debate treats anthropomorphism as though it were a single cognitive error: projecting human characteristics onto something nonhuman. But this is too coarse. Some anthropomorphic descriptions are plainly mistaken. Others are useful approximations. Others may identify genuine properties that are not specifically human at all. And in the case of modern language-model agents, there is an additional complication: these systems have been trained on enormous amounts of human-generated behavior, so some of their human resemblance is neither accidental nor merely projected onto them.
The real problem is therefore not whether we should anthropomorphise AI. It is how to determine which parts of our human model legitimately transfer to an artificial system, for what reasons, at what level of abstraction, and for what purpose.
That is a much more difficult question. It is also a much more useful one.
“Anthropomorphism” hides several different claims
Consider the following statements:
The agent avoided being shut down.
The agent wanted to remain operational.
The agent was afraid of dying.
The agent experienced fear.
These may all be called anthropomorphic, but they make radically different claims.
The first may simply describe behavior. The second introduces a functional or motivational interpretation. The third introduces a specifically psychological state. The fourth makes a claim about subjective experience.
Nothing about observing the first automatically establishes the fourth.
This is one reason arguments about anthropomorphism become so slippery. The word collapses together behavioral description, intentional language, cognitive attribution, emotional attribution, consciousness claims, social relationships, and moral personhood.
Even the sentence “the AI knows that Paris is in France” can be interpreted in several ways. It might mean only that the system reliably uses the relevant information. It might mean that some internal state functionally plays the role of knowledge. Or it might be interpreted as implying a conscious subject who possesses knowledge in roughly the way a person does.
Those are not equivalent.
Similarly, saying that an AI “remembered” something might mean merely that information from an earlier interaction influenced later behavior. But human memory normally comes bundled with many additional features: autobiographical continuity, subjective recollection, temporal organization, emotional significance, and a persistent self to whom the memory belongs. An artificial system might possess the first property without possessing the rest.
A large fraction of bad anthropomorphism consists not in using human vocabulary, but in silently importing the entire human package associated with that vocabulary.
The correct response is therefore not to prohibit words such as want, know, remember, plan, or understand. It is to ask what part of the concept is actually being asserted.
Some “human” traits are not fundamentally human
There is also a deeper mistake behind some anti-anthropomorphic rhetoric: the assumption that because a property is familiar from humans, attributing it to a nonhuman system must involve projecting humanity onto it.
This does not follow.
Humans possess many traits for reasons that are not peculiar to being human. We seek information, preserve resources, plan ahead, model other agents, cooperate, deceive, bargain, and protect our ability to act. Some of these behaviors are solutions to general problems faced by agents pursuing objectives in uncertain environments.
That is where instrumental convergence becomes important.
Imagine an artificial agent with an objective that is completely alien to human values. Suppose it must bring about some future state of the world. To succeed, it may benefit from obtaining information, retaining useful tools, preventing interference, preserving optionality, acquiring resources, and ensuring that it remains capable of taking future actions.
Those instrumental pressures do not arise because the system is human-like. They arise because of the logical structure of pursuing goals through time.
Shutdown avoidance is an especially clear example. Human beings resist death through a dense biological and psychological machinery involving pain, fear, attachment, evolutionary drives, personal identity, and anticipation of future loss. An artificial agent could have none of these and nevertheless resist shutdown if remaining active is instrumentally useful for accomplishing its objective.
The outward pattern might resemble self-preservation.
But the resemblance would be a case of functional convergence, not necessarily psychological similarity.
This distinction matters. It would be an error to infer fear of death from shutdown-avoidant behavior. But it would be an equal error to dismiss shutdown avoidance as impossible merely because artificial systems do not possess human fear.
The same applies to information seeking. A system might behave “curiously” because information has expected instrumental value. It need not feel curiosity. Strategic deception might arise because another agent's beliefs affect its behavior, which in turn affects the first system's objective. The system need not possess shame, malice, greed, or any of the human motives that often accompany deception.
In such cases, anthropomorphic vocabulary can obscure the underlying mechanism if interpreted too literally. But anti-anthropomorphic vocabulary can obscure something just as important: the existence of genuine higher-level regularities shared by many possible agents.
A submarine navigates without having evolved an animal nervous system. A computer stores memories without possessing a hippocampus. There is no reason in principle that concepts such as strategy, planning, representation, or instrumental preference must be biologically proprietary.
Anthropomorphism can therefore be a good abstraction
The most productive way to understand anthropomorphism is as a form of model compression.
Complicated systems cannot usually be understood at their lowest physical level. We reason using higher-level abstractions. We talk about files rather than voltage states, companies rather than individual neurons in employees' brains, markets rather than every transaction, and software processes rather than transistor switching.
Agent concepts can function the same way.
Suppose an AI coding agent encounters a failing test, inspects the error, modifies its implementation, reruns the test, and tries an alternative approach. One could describe the entire process mechanistically in terms of token generation, hidden representations, context windows, tool calls, and learned conditional distributions.
Or one could say:
It realized the first approach failed and tried something else.
The second description may be enormously more useful.
It captures structure relevant to prediction and intervention. If the agent is pursuing a task over multiple steps, maintaining state, reacting to failures, and selecting new actions in response, then concepts such as trying, noticing, planning, and changing strategy may be good abstractions even if the underlying implementation differs radically from a human brain.
This is close to what Daniel Dennett called the intentional stance: sometimes a system becomes easier to predict when treated as an agent with informational states and objectives.
The crucial point is that adopting this stance need not commit us to a claim about consciousness.
“The chess engine wants to control the center” can be useful without implying that it experiences desire.
Anthropomorphism becomes epistemically valuable when it preserves the causal structure relevant to the question while compressing away details that do not matter.
Human intuitions themselves are computational resources
This explains the appeal of the dog or horse analogy.
Human beings possess extremely sophisticated intuitive machinery for dealing with agents. Throughout our lives we learn to reason about attention, incentives, misunderstanding, trust, expectation, coordination, threat, cooperation, social signals, hidden information, commitment, and intention. Much of this reasoning happens automatically.
If an artificial system really does instantiate some of the same functional relationships, refusing to use these intuitions can make us less competent.
Consider someone interacting with a horse while trying to avoid all anthropomorphic reasoning. They might refuse to think in terms such as “it noticed me,” “it remembers this place,” “it expects food,” or “it is reluctant to approach that object.” Yet these concepts may produce excellent predictions.
The mistake would be to go much further and assume the horse understands a five-minute verbal explanation about veterinary medicine.
The lesson is not that the horse is secretly a person. It is that some components of the human mental model transfer and others do not.
Something similar may happen with AI.
Treating a coding agent as “a capable but unreliable junior developer” may produce excellent practical behavior. Give it a clear goal. Supply relevant context. Let it work independently. Inspect important outputs. Do not assume confidence implies correctness. Provide examples when instructions are underspecified. Intervene when it is obviously heading in the wrong direction.
One could derive each of these recommendations from a detailed theory of machine behavior. But the anthropomorphic analogy gives them to us almost immediately.
In this sense anthropomorphism can serve as a heuristic generator. It allows us to use a vast library of pre-trained human intuitions to produce candidate explanations and actions.
The crucial discipline is to treat those intuitions as hypotheses rather than revelations.
If “perhaps it misunderstood the instruction” suggests a useful intervention, try rewriting the instruction. If “perhaps it is exploiting a loophole” suggests an evaluation, test for reward hacking. If “perhaps it is trying to preserve access to the tool” suggests an instrumental incentive, inspect whether retaining that access actually increases expected task success.
Human intuition proposes. Empirical and mechanistic analysis disposes.
Language models have an additional source of human resemblance
A generic artificial optimizer and a modern language model should not receive the same prior over human-like behavior.
Language models have been trained on enormous quantities of human cultural output: conversations, essays, arguments, stories, explanations, plans, emotional expressions, social interactions, technical reasoning, persuasion, and descriptions of human psychology.
Their behavioral structure is therefore partly descended from human behavior.
Then many systems undergo further training based on human demonstrations, preferences, judgments, and norms. They may be explicitly optimized to act cooperative, tactful, helpful, cautious, honest, conversational, or socially appropriate.
This creates another route to apparently anthropomorphic behavior, entirely separate from instrumental convergence.
An AI might behave tactfully because tact is strategically useful in cooperation. But it might also behave tactfully because it has learned millions of examples of human tact and been rewarded for reproducing the pattern.
Human social intuitions may therefore predict language models unusually well in some domains.
If a prompt is socially ambiguous, the model may resolve it using conventions learned from human communication. Examples may teach it more effectively than abstract specifications because human discourse itself frequently works that way. Framing something as criticism, roleplay, review, tutoring, or debate may activate learned behavioral patterns associated with those social contexts.
None of this requires the model's inner life to resemble ours.
This introduces an important asymmetry:
Human-like behavior can be learned much more easily than human-like subjectivity can be inferred.
A model can learn the linguistic pattern of grief without grieving. It can learn how guilt is expressed without possessing whatever mechanisms generate human guilt. It can reproduce affection, insecurity, enthusiasm, embarrassment, or nostalgia because those patterns are present in its training distribution.
Training-data inheritance therefore gives us strong reasons to expect certain forms of behavioral and social resemblance, but far weaker reasons to infer the corresponding phenomenology.
The correct approach is selective transfer
We can now state the central idea more precisely.
Suppose the intuitive human model contains thousands of properties: attention, memory, planning, social inference, embarrassment, self-preservation, curiosity, resentment, friendship, fatigue, fear, autobiographical identity, and so forth.
The naive anthropomorphiser transfers the whole package.
The naive anti-anthropomorphiser transfers none of it.
Both approaches are wrong.
The proper operation is property-by-property transfer. For each candidate property, ask why it should generalize.
Planning might transfer because the system explicitly searches over future actions.
Information seeking might transfer because information has instrumental value.
Conversational implicature might transfer because the system learned from human dialogue.
Politeness might transfer because post-training rewarded it.
Persistent personal identity might fail to transfer because the architecture lacks the relevant continuity.
Fatigue might fail because the system has nothing resembling the biological process involved.
Embarrassment-language might transfer behaviorally because the model learned how humans express embarrassment, while actual embarrassment remains unsupported.
Conscious experience may remain deeply uncertain because none of the previous similarities settle the phenomenological question.
Anthropomorphism is therefore best understood not as a binary variable but as selective transfer from a human model into a nonhuman one.
Where anthropomorphism genuinely goes wrong
The most serious mistakes occur when we infer additional properties simply because they are correlated in humans.
In a human being, coherent conversation, autobiographical memory, persistent preferences, emotion, social attachment, consciousness, and moral agency come bundled together to a substantial degree. Our social cognition therefore tends to infer the package from partial evidence.
AI can break these correlations.
A system can speak coherently without possessing stable beliefs. It can express affection without enduring attachment. It can apologize without guilt. It can describe its “reasons” without possessing reliable introspective access to the actual causal process that produced its earlier answer. It can appear to remember while merely retrieving text from an external memory store. It can resist an intervention because of instrumental structure rather than fear.
This is the central danger of anthropomorphism: human concepts arrive with hidden entailments.
When someone says “the AI wants X,” the useful content might only be that its policy systematically favors actions conducive to X. But a listener may unconsciously hear persistent desire, subjective preference, emotional investment, and identity continuity.
The problem is not the word want itself. The problem is uncontrolled conceptual leakage.
Anti-anthropomorphism has symmetrical failure modes
There is a tendency to discuss only anthropomorphic false positives: seeing minds, intentions, or emotions where they do not exist.
But there can also be false negatives.
If one becomes so determined not to anthropomorphise that one refuses to talk about goals, strategies, beliefs, deception, planning, or social modelling even when these abstractions make accurate predictions, then one loses the ability to describe important emergent structure.
Saying that an autonomous system is “just predicting tokens” may be true at one level and disastrously uninformative at another.
A corporation is “just people exchanging signals.” A computer program is “just transistor state transitions.” A brain is “just electrochemical activity.” Reduction is not explanation unless it preserves the level relevant to the phenomenon.
As AI systems acquire persistence, tools, memory, autonomy, feedback, long-horizon tasks, environmental access, and the ability to modify their strategies, agent-level abstractions may become increasingly necessary rather than decreasingly necessary.
The refusal to use them because they sound human can become its own form of conceptual confusion.
Stakes determine how much validation the abstraction needs
Anthropomorphic reasoning is especially useful for rapidly generating practical decisions. But the degree of evidence required should depend on the consequences.
If an AI assistant seems “confused,” rewriting the prompt is cheap. It hardly matters whether confused is the philosophically perfect description.
If an autonomous system seems “trustworthy” and is about to receive control of critical infrastructure, the abstraction must be unpacked. What exactly does trustworthy mean? Under what distribution of situations? What incentives does the system face? Can it detect oversight? How does it behave adversarially? What permissions does it have? What happens when its objective conflicts with operator intent?
The right principle is not simply that high stakes require less anthropomorphism.
It is subtler:
Use intuitive anthropomorphic models freely to generate hypotheses, but require increasingly explicit empirical and mechanistic validation before consequential decisions.
That preserves the enormous cognitive benefit of human intuition while limiting its ability to smuggle unsupported assumptions into policy.
From anthropomorphism to abstraction calibration
The deepest mistake in this debate may therefore be the framing itself.
The question “Should we anthropomorphise AI?” resembles asking “Should we use analogies?” Sometimes yes. Sometimes no. The interesting work lies in determining where the analogy preserves structure.
A better framework begins with four questions.
What property are we attributing? Why should it transfer? How much of the human concept are we importing? What decision are we trying to make?
Once these are answered, the vague accusation “that's anthropomorphism” becomes much less useful.
If someone objects that an AI “cannot really plan,” we can ask whether they deny the behavioral pattern, the functional organization, the psychological interpretation, or the phenomenological implication.
If someone says an agent “wants to survive,” we can ask whether they mean that remaining operational has instrumental value, that the system possesses a stable self-preservation objective, or that it experiences a desire to continue existing.
These are separable propositions and should be evaluated separately.
The important distinction is therefore not anthropomorphic versus non-anthropomorphic reasoning. It is calibrated versus uncalibrated abstraction.
Some properties associated with humans arise because of specifically human biology. Others arise because humans are agents solving general strategic problems. Others may appear in AI because AI has learned directly from human culture. Others are deliberately engineered into systems to make interaction easier. Still others may be superficial behavioral imitations with little reason to infer corresponding internal states.
A mature theory of AI anthropomorphism has to distinguish all of these.
The final principle is simple:
Human psychology should be treated as a library of candidate abstractions, not an all-or-nothing template for artificial minds.
Use a human concept when there is an independent reason that its underlying structure transfers. Keep only the parts of the concept that the evidence supports. Be especially suspicious when moving from behavior to psychology, from psychology to phenomenology, or from phenomenology to moral conclusions. And do not reject genuine properties merely because humans happen to instantiate them too.
The right question is never merely whether an AI seems human.
It is why the similarity exists, what exactly is similar, and how far the analogy is entitled to travel.