50  The Nearest Agent

Modeling Other Minds, and One’s Own

50.1 The shuffle on the path

Two people approach each other on a narrow path. Each sees that the other must pass. Each begins to step aside—and they choose the same side, correct in unison, choose the same side again, and stop in the small foolish shuffle that nearly everyone has performed. Finally one person holds still long enough for the other to commit.

Nothing failed in either visual system. Both people saw the other clearly. The difficulty arose because each movement changed the evidence on which the other movement was being selected. Two controllers were coupled through a shared path, each continually revising in response to the other.

That description may be sufficient. Perhaps neither person represented anything about the other’s mind; perhaps each merely reacted to the latest visible adjustment. But another possibility becomes important when the stakes rise. Each person may have anticipated which side the other was likely to choose, including how the other would respond to being anticipated. The shuffle itself cannot decide between those accounts. It opens the question that organizes this chapter: when is sensitivity to behavior enough, and when does successful control require an estimate of the goals, knowledge, or beliefs that will generate behavior not yet visible?

The distinction matters in ordinary life. A stranger’s hand moves toward a pocket, and your response depends partly on what you think the movement is for. You remain silent in a meeting because you anticipate how one person will interpret what you are about to say. A lie succeeds only if its teller can influence what a listener comes to believe. In each case, another agent is not merely a moving object. Their next action depends on how they construe the situation, and that construal may include an estimate of you.

This is the destination toward which the book has been moving. Chapter 2 began with movement and control. Later chapters widened the loop to allostasis, learning, and action in a constructed environment. Chapter 48 argued that a human life depends on a social niche made of provisioning, teaching, exchange, and inherited knowledge. Chapter 49 treated language as coupled action between agents. Flexible communication requires some estimate of what a listener can perceive, already knows, expects, or misunderstands.

The word model will recur, but it should not be imagined as a miniature person or an explicit theory stored in a cortical box. A neural system models a variable whenever its physically instantiated state carries enough structure about that variable to guide prediction and action beyond the immediately visible cue. Such a state may be distributed, partial, task-dependent, and unavailable to verbal report. Habit and direct sensorimotor coupling can often do the job more cheaply. The interesting cases are those in which they cannot.

The hardest features of a social niche are therefore not simply other bodies. They are other adaptive control systems—systems whose actions depend on their own histories, goals, information, and predictions, and whose behavior changes when they predict you in return.

Theory of mind, perspective-taking, empathy, and the sense of self should not begin as a list of late cognitive faculties with presumed cortical homes. The adaptive problem comes first. What architecture lets one organism coordinate with, compete against, teach, deceive, trust, or care for another organism whose internal state is not directly observable? And what happens when some of the same operations are directed toward the organism doing the modeling?

The nearest such agent is the self.

50.2 A mind is a hidden cause

Another person’s goals, knowledge, and beliefs are not directly available to your senses. You see a reach, a glance, a hesitation, or an utterance. You also know something about the setting and the person’s history. From those traces, the nervous system settles on an interpretation that makes some future actions more likely than others.

This problem resembles perception in one important respect. Sensory input underdetermines its causes. The retinal image does not uniquely specify the three-dimensional scene that produced it; an acoustic stream does not contain labeled word boundaries; a facial movement does not announce whether it is fear, concentration, or play. Neural systems exploit regularities learned across development and evolution to constrain those possibilities.

A mental state is an especially difficult cause because it belongs to another adaptive system. The cause is not immaterial. A goal or belief is realized in a living brain and body. It is hidden only in the operational sense that the observer cannot sample it directly. The observer must use behavior, context, prior interaction, and knowledge of what the other person could have perceived.

Two directions of analysis are useful. One begins with behavior and asks what state of the agent would make the behavior intelligible. The other begins with an estimated state and asks what action is likely to follow. These are often called inverse and forward problems. The terms identify relations the system must solve; they do not establish that the brain runs an explicit equation, constructs a symbolic sentence about the other person, or implements one uniquely Bayesian algorithm. The solution may emerge from distributed dynamics shaped by learning and recurrent interaction.

This caution matters because a successful prediction does not prove that a mind was represented. A familiar person may always sit in the same chair. A dominant animal may win access to food. A driver approaching a red light will probably stop. Habits, social roles, learned contingencies, and visible affordances can support excellent forecasts without any detached representation of belief.

Mental-state modeling becomes most useful where those regularities underdetermine action. The other person has not moved yet; several actions are physically possible; the relevant information is available to them but not to you, or to you but not to them; deception is possible; and the encounter will recur. In such cases, behavior depends on how the agent construes the situation, including what the agent thinks you know.

This is the property that makes social prediction different from recovering a surface from an image. A surface does not alter its reflectance because it has inferred your plan. Another agent can alter their behavior because they noticed that you noticed them. The causal loop becomes recursive: your estimate changes your action; your action changes their evidence; their revised estimate changes what they do next.

50.3 When predicting behavior is not enough

Most successful social prediction probably does not require explicit reasoning about beliefs.

Visible movement carries information. A running predator has a trajectory. A reaching hand is constrained by the objects within reach. Posture, gaze, and facial movement can reveal where attention is directed. Repeated interaction supplies habits: this individual usually shares, that one retaliates, another avoids conflict. Social roles and conventions add further shortcuts. A person in a queue will probably move forward when the line advances. A familiar greeting calls for a familiar response.

These regularities are valuable because they are fast and cheap. An animal that predicts only from current behavior, learned contingencies, and stable social structure can coordinate remarkably well. The existence of complex social behavior therefore does not by itself prove a theory of mind.

Several conditions increase the value of going beyond those predictors.

The first is action before movement. When no relevant movement has begun, trajectory cannot supply the forecast. Context, habit, and affordance may still suffice, but several incompatible actions can remain possible. What distinguishes them may be the agent’s goal or interpretation of the situation.

The second is informational asymmetry. Two agents may occupy the same physical scene while possessing different evidence. One saw the food moved; the other did not. One overheard the warning; the other was absent. Predicting from the actual world will then fail, because behavior follows the world as represented by the agent rather than the world as known to the observer.

The third is strategic recurrence. In a repeated interaction, today’s action changes tomorrow’s expectations. Cooperation, retaliation, reputation, and deception depend on what each participant expects the other to remember and infer. The recursive depth can increase: I may act on what I think you believe about what I intend. Human institutions often stabilize this recursion with roles and rules, but they do not remove it.

These cases connect social cognition to allostasis. A nervous system gains an advantage by preparing before a consequential event arrives. Estimating another agent’s goal or belief can move preparation earlier, before the body has committed to an action. The benefit is greatest when the other agent is also anticipating and when a late response is costly.

None of this implies that a fully explicit belief model appears whenever social life becomes complex. Mentalizing can be partial and selective. An observer may track what another can see without representing a proposition about belief, or may infer a goal while ignoring knowledge. The system may recruit richer representations only when habit and visible behavior fail. That is one reason social understanding is better treated as a graded collection of capacities than as a single faculty that a species either possesses or lacks.

The evolutionary proposal is correspondingly modest. Recurrent, individualized, cooperative, or competitive social niches should increase the payoff for estimating states that decouple from visible behavior. Whether that prediction explains differences across species remains open. It tells us what evidence would matter: not merely whether an animal lives socially, but whether it can use another agent’s informational perspective when surface regularities are controlled.

Mind-modeling earns its cost where bodies, habits, and conventions leave the next action unresolved.

Figure 50.1: From visible behavior to agent-specific prediction. A, When an action is already underway, its visible trajectory constrains what is likely to happen next; faded positions indicate observed movement, and the dashed path indicates the predicted continuation. B, Before movement begins, learned regularities—including an individual’s habits, social role, or shared conventions—may support an efficient forecast without requiring a detailed estimate of hidden mental states. C, When two agents have different perceptual histories, prediction must be indexed to the information available to each agent. Agent A observed the object’s change of location and is predicted to approach its current location, whereas Agent B missed the change and is predicted to act on outdated information. The occluder represents blocked perceptual access during the earlier event, not a barrier to later action. D, In recurrent interaction, each agent’s action becomes new evidence for the other: an initial prediction produces an adjustment, the adjustment is observed, and both agents may update their next actions. Solid or faded paths represent observed or completed movement, dashed arrows represent predicted movement, and eye symbols indicate perceptual access. The sequence describes a continuum of problem conditions, from cases in which visible behavior and stable regularities may suffice to cases in which agent-specific state estimation becomes especially valuable; it is not a ladder of species, developmental stages, or dedicated cognitive modules.

50.4 Overlapping machinery, different targets

A system that can construct a possible state of another agent may also contribute when the possible agent is oneself.

Before acting, people can rehearse a route, imagine an outcome, compare alternatives, or revisit what might have happened under another choice. Memory supplies fragments, valuation supplies preferences, and motor and perceptual systems help construct a scene that is not currently present. These operations were introduced earlier as prefactual and counterfactual control: using possible consequences to guide a current action.

Social perspective-taking has a related structure. The current sensory scene must be reorganized around another vantage, another history of access, or another set of goals. In both cases the nervous system departs from the immediate present and constructs a state that is not directly given.

This resemblance motivates a working hypothesis: remembering one’s past, imagining one’s future, and estimating another person’s perspective draw partly on shared constructive machinery. The hypothesis predicts overlap among the neural systems recruited by episodic memory, prospection, navigation, and mental-state reasoning. Such overlap exists, especially in medial prefrontal, posterior medial, lateral parietal, and medial temporal regions [@bucknercarroll2007].

Overlap does not establish identity. The same broad territory can contain neighboring or interdigitated networks with different response preferences, and a task can recruit shared operations while also requiring specialized components. Severe episodic-memory impairment need not abolish explicit false-belief reasoning. Belief attribution may depend on representations not required for imagining one’s own future. The safer formulation is therefore overlapping machinery, different targets, not one universal simulator.

That distinction will matter throughout the chapter. At a coarse scale, self-projection and mentalizing repeatedly converge. At a finer scale, they can dissociate. The goal is not to choose in advance between a dedicated “theory of mind module” and a completely domain-general engine. It is to ask which components are shared, which are specialized, and how the architecture changes with the problem being solved.

We will return to the self at the end. First, we need a behavioral test that separates tracking another person’s representation from tracking the world as we ourselves know it.

50.5 False belief is a stringent test

To distinguish representing another person’s perspective from merely tracking the actual world, experimenters need the two to disagree.

Suppose a person leaves their keys in a cupboard and departs. While they are gone, someone moves the keys to a drawer. The keys are really in the drawer, but the returning person will initially search the cupboard because their action follows the information available to them. An observer who predicts the search must preserve two states at once: the true location and the location represented by the other person.

This is why false belief became the central test of mental-state understanding. It creates a controlled decoupling between reality and an agent’s representation of reality. In the classic Sally–Anne form, one character places an object in one location and leaves; the object is moved in the character’s absence; the child is asked where the character will look on returning. Standard verbal versions are usually passed during the preschool years, often around age four, although performance depends on language, memory, inhibitory control, and details of the task [@baroncohen1985].

False belief should not be treated as the moment at which a child or species suddenly acquires a mind. Goals, intentions, attention, knowledge, and ignorance are also mental states. An observer can usefully track what another agent wants or sees even when that state agrees with reality. False-belief tasks test something more specific: whether behavior can be predicted from a representation the observer knows is mistaken.

The capacities can therefore be arranged as a conceptual gradient rather than a rigid developmental staircase: sensitivity to animacy; prediction of goals; tracking of gaze and perceptual access; tracking of knowledge and ignorance; representation of belief; representation of false belief; and recursively embedded belief. The categories overlap, and success at one level may be achieved by more than one mechanism. The gradient marks increasing separation from the immediately observable, not a set of guaranteed stages.

The task is stringent, but it is not pure. A correct verbal answer requires the child to remember the story, understand the question, suppress the true location, and select the relevant agent. A failure can therefore reflect executive or linguistic demands as well as a limitation in belief representation. Conversely, success on one familiar paradigm does not establish a general, explicit theory of all minds.

The controversy about infants and nonhuman animals follows from this measurement problem. Nonverbal looking, anticipation, and helping tasks reduce some demands but introduce ambiguity about what the measure means. The responsible question is not simply “does the subject have theory of mind?” It is “what information did this task require the subject to maintain, and what alternative strategy could produce the same behavior?”

Figure 50.2: False belief as representational separation and compound task performance. A: In the Sally–Anne sequence, Sally places a marble in the basket and sees it there, then leaves before Anne transfers it to the box. When Sally returns, the participant knows that the marble is now in the box but predicts that Sally will first search the basket because she did not observe the transfer. Successful prediction therefore requires maintaining two distinct representations: the current state of the world and the outdated location attributed to Sally. The solid arrow marks the observed transfer, whereas the dashed arrow marks Sally’s predicted first search. B: Related forms of social prediction range from detecting an agent and its goal, through tracking perceptual access, knowledge, and belief, to representing a false belief that conflicts with the known world and an embedded belief about what one agent thinks another believes. This is a conceptual gradient of increasing separation from immediately observable behavior, not a set of fixed developmental stages, an evolutionary ladder, or a sequence supported by only one mechanism. C: A correct verbal false-belief answer is a compound performance. In addition to attributing Sally’s outdated belief, the participant must remember the event sequence, comprehend the story and question, inhibit the prepotent true-location response, and keep Sally’s information distinct from the participant’s own knowledge. An incorrect answer therefore does not, by itself, reveal which component of the task failed.

The standard false-belief task is a wonderful instrument and a treacherous one. Passing it requires not only representing another’s false belief but also holding a goal in mind, inhibiting the prepotent urge to answer with what one knows to be true, parsing the language of the question, and tracking a short narrative. A child could fail for reasons that have nothing to do with mental-state representation, and an experimenter could mistake a memory or inhibition limit for a conceptual one.

This matters because of a genuine empirical puzzle. A body of work using looking-time and anticipatory-looking measures reported that infants well under four — far too young to pass the verbal task — already behave as though they track others’ false beliefs, looking longer when an agent acts inconsistently with what the agent should believe, or looking in anticipation toward the location an agent falsely takes to be correct [@onishibaillargeon2005]. If real, these results imply that something belief-like is in place long before children can answer Sally-and-Anne aloud, and that the verbal task measures the ability to deploy mental-state knowledge under task demands rather than the presence of the knowledge itself.

The trouble is that several of the infant findings have proven difficult to reproduce. Large, preregistered replication efforts have failed to recover some of the key anticipatory-looking effects, and the literature is now openly unsettled about which infant results are robust [@kulke2018]. The summary is that there is probably more than one thing here. One proposal, worth taking seriously, is that humans operate two distinct systems: an early-developing, fast, automatic, but inflexible capacity for tracking what others register about the world, and a later-developing, slow, effortful, but flexible capacity for reasoning explicitly about beliefs [@apperlybutterfill2009]. On that view the contradiction partly dissolves, because the verbal task and the infant looking measures are probing different machinery. The field has not converged.

50.6 Does the chimpanzee have a theory of mind?

The phrase theory of mind entered the literature in 1978, in a paper by David Premack and Guy Woodruff that asked the question above of a chimpanzee [@premack1978]. Their evidence was that a chimpanzee shown a film of a human struggling with a problem — reaching for out-of-grasp bananas, or trapped behind a door — would, in a forced choice, select the photograph depicting the solution, the stick or the key. The chimpanzee, they argued, grasped what the human was trying to do.

Set against the framework of this chapter, what Premack and Woodruff demonstrated sits low on the ladder just described — which is not to say outside it. Reading the target of an action is genuine mental-state attribution: the chimpanzee that selects the key infers a goal, and a goal is a mental state. But goal inference is the rung closest to behavior, because a goal is largely readable from the structure of the behavior aimed at it, and it does not require holding a model of another’s belief, still less a belief the observer knows to be false. Goal-reading buys a great deal of useful prediction without ever decoupling the agent’s representation from reality.

The demanding question — the one higher on the ladder, where representing a mind comes apart from sophisticated behavior-prediction — is whether any nonhuman animal can represent a false belief in another. Here the evidence is genuinely contested, and the shape of the contest is exactly the seam this chapter has been tracing. Studies using anticipatory looking have reported that great apes look toward the location where an agent falsely believes an object to be, as though predicting the agent’s mistaken search — a nonverbal analogue of the false-belief task [@kano2019]. If that interpretation holds, it would place some belief-tracking outside our species. But there is a deflationary reading available, and it is the same deflationary reading that haunts the infant literature: the apes may be projecting their own past experience of being misled onto the agent, or responding to learned regularities in where agents tend to look, rather than representing the agent’s belief as a belief. The data underdetermine the two stories.

This is not a failure of the field so much as the place where the field’s central question actually lives, and the framework lets us say something more useful than “it is unclear.” This is the prediction the earlier section set up. If decoupled mentalizing is favored by niche structure rather than switched on by sociality as such, then it should be graded — richest where social life is most densely recurrent, most strategic, and most dependent on pre-empting others’ actions: in species with stable, individualized, long-term relationships and intense cooperative and competitive interdependence. That is a prediction, and it is the right kind, because it is in principle wrong. It says the decoupled capacity is not a binary human possession to be confirmed or denied but a graded adaptation that should track the structure of a species’ social niche. The comparative data are not yet good enough to test it cleanly. The chapter’s claim is not that the question is answered. It is that this is the question, and that the niche-first frame tells us where to look.

50.7 The mentalizing network, and the overlap

Functional imaging studies of mental-state reasoning repeatedly recruit a recognizable set of regions. Stories or cartoons whose interpretation depends on a character’s belief, intention, or misunderstanding engage the temporoparietal junction, often more strongly on the right; medial prefrontal cortex; posterior cingulate and adjacent precuneus; and other lateral and medial parietal territories [@gallagher2000; @saxe2003]. Lesion, stimulation, and connectivity evidence support the conclusion that these regions make important contributions to social inference.

The finding is reproducible enough to deserve a name: the mentalizing network. The name should not be turned into a claim that one network reads minds in isolation.

Several neighboring systems occupy the same broad geography. Posterior superior temporal cortex responds to biological motion and to socially informative changes in gaze and body movement [@pelphrey2004]. More anterior temporoparietal tissue participates in the ventral attention network, which interrupts ongoing processing when an unexpected event requires reorienting [@corbettashulman2002]. Medial prefrontal, posterior medial, and lateral parietal regions also appear during autobiographical memory, imagination of the future, and unconstrained rest [@raichle2001; @bucknercarroll2007].

At the resolution of a group-average activation map, these tasks can look as though they converge on the same blobs. That convergence is informative but ambiguous. It may reflect a shared operation, such as constructing a displaced perspective or updating a model when socially relevant evidence arrives. It may also reflect distinct populations lying close enough together that spatial smoothing and anatomical variability merge them.

Individualized and connectivity-based mapping increasingly supports a mixed picture. Broad territories contain adjacent or interdigitated subnetworks with different preferences: some are more strongly engaged by belief reasoning, others by episodic construction, attention, or biological motion [@mars2012; @bragabuckner2017]. The boundaries vary across people, which makes a group coordinate an unreliable substitute for a functional map in one brain.

The architecture is therefore distributed without being undifferentiated. Mental-state reasoning depends on interactions among systems that represent agents and actions, track information and perspective, construct alternatives to the present, select relevant evidence, and connect those representations to language and decision. Some components are shared with nonsocial tasks; some show greater social or belief selectivity.

This is exactly the kind of result the book’s architecture-first approach predicts. Evolution did not need to create a cortical island ex nihilo. It could modify the weighting and connectivity of older systems for attention, memory, perception, and control. The resulting specialization can be real even when its boundaries are graded and its components are reused.

Figure 50.3: Distributed mentalizing territories at group and individual spatial scales. A, Group studies repeatedly implicate a distributed, bilateral set of cortical territories during reasoning about other agents, although their precise boundaries and response profiles vary with task and individual anatomy. The right lateral view shows approximate territories along the posterior superior temporal sulcus (pSTS), associated especially with biological motion and socially informative action; the anterior temporoparietal and STS–TPJ junction, associated with attentional reorienting and overlapping social demands; and posterior temporoparietal and inferior parietal cortex, often recruited during belief, knowledge, and perspective reasoning. The medial view shows approximate territories in medial prefrontal cortex (mPFC) and posterior medial cortex, including posterior cingulate cortex and precuneus. Only one lateral hemisphere is displayed for clarity; the broader system is bilateral. B, Spatial scale changes the organization revealed by functional mapping. In group averages, anatomical and functional variability across participants can cause nearby responses to merge into broad territories. Repeated measurements within one person can instead reveal stable neighboring or interdigitated topographies. In temporoparietal cortex, regions showing relative preferences for biological motion and social perception, attentional reorienting, and belief or perspective reasoning may be distinguishable within an individual. In posterior medial cortex, precision mapping can separate adjacent distributed networks with relative preferences for mentalizing and for episodic construction or projection. The same large-scale networks recur across multiple cortical zones rather than forming isolated local modules. Colors indicate approximate response preferences or network affiliations, not exclusive functions. Broad group-level overlap therefore does not demonstrate one undifferentiated mentalizing system, while fine within-person differentiation does not establish sharply bounded modules dedicated to single psychological operations.

Two opposite errors stalk this literature.

The first is reverse inference: observing that a region active during theory-of-mind tasks is now active in some new task, and concluding that the new task therefore involves theory of mind. This does not follow. A region that participates in a function is not thereby exclusive to it. The temporoparietal junction’s appearance in an attention task does not show that the participant was mentalizing, any more than the heart’s involvement in running shows that everyone who runs is in love. Activation localizes where a manipulation changes blood flow. It does not, by itself, reveal what computation the tissue performs, and the same blob can be produced by different processes.

The second error is the mirror image: treating the broad overlap as proof that there is no specialization at all, that it is “all one network” and the regional labels are meaningless. That overshoots too. The “temporoparietal junction” is not one thing — it is an anatomically imprecise label spanning tissue with different connectivity and, on closer parcellation, different functional profiles. Connectivity-based parcellation places an anterior sector with the ventral-attention and salience circuitry — its connections run to ventral frontal cortex, anterior insula, and midcingulate cortex — and a posterior sector with the default and mentalizing network, its connections running to angular gyrus, precuneus, and posterior cingulate [@mars2012; @bzdok2013]. The borders shift with atlas and task, and the two functions overlap rather than separating cleanly: a meta-analysis found the region recruited by both attentional reorienting and false belief sitting anteriorly, while the posterior sector converged more specifically on false belief [@krall2015]. Some evidence further suggests that a portion of the right posterior temporoparietal junction is unusually selective for representing beliefs specifically, dissociable from general social or attentional demands.

The defensible position sits between the two errors. There is real regional specialization, and there is real sharing of machinery across social, attentional, and self-projective tasks. Neither the “dedicated mind-reading module” picture nor the “it’s all the same soup” picture survives contact with the parcellation data. Holding both facts at once is uncomfortable and correct.

No discussion of the neuroscience of social cognition can avoid the mirror neuron, and few topics have been so oversold. In the macaque, neurons in premotor area F5 discharge both when the animal performs a particular goal-directed action — grasping a raisin — and when it observes another individual performing a similar action [@rizzolatti2004]. The finding is real and important: here are cells whose activity is shared between doing and seeing.

The leap that followed was to declare that such cells constitute the understanding of others’ actions — that we grasp what another is doing by covertly running the same motor program, and, by extension, that mirror neurons are the basis of empathy, language, imitation, and the social deficits of autism. That extrapolation outran the evidence in several directions [@hickok2014]. A shared motor response is not the same as a representation of another’s goal or belief; it may be a consequence of understanding the action rather than its cause. People can understand actions they cannot perform, and damage to motor regions does not reliably abolish the comprehension of others’ movements. And nothing in a motor-matching mechanism explains the capacity that this chapter has identified as the stringent test of decoupled mentalizing — the representation of a belief the observer knows to be false — because mirroring an action gives you the action, not the possibly-mistaken model of the world behind it.

The sober residue is worth keeping. Sensorimotor systems are clearly recruited when we observe others act, the coupling between perceiving and producing action is genuine, and it plausibly contributes to imitation and to the front end of reading goals from movement. That is a real piece of the machinery. It is not the whole of social cognition, and it is not the part that does the hard work.

50.8 Self-projection

The broad overlap among remembering, imagining, navigating, and mentalizing motivated the proposal of self-projection: the capacity to shift away from the immediate present and construct a situation from another time, place, or perspective [@bucknercarroll2007].

The proposal captures a real structural similarity. Episodic remembering reconstructs a past scene. Prospection constructs a possible future. Navigation can require imagining a route or viewpoint not currently occupied. Perspective-taking reorganizes a scene around another agent’s access, goals, or beliefs. All require information not supplied directly by the present sensory input.

Many of these tasks recruit medial prefrontal cortex, posterior cingulate and precuneus, lateral parietal cortex, and medial temporal structures. These regions overlap broadly with the default mode network, identified through correlated activity during rest and through relative decreases during many externally demanding tasks [@raichle2001]. The overlap has encouraged the idea that the default network supports internally constructed models that can be used for memory, planning, and social inference.

Two cautions are essential.

First, rest is not a task with one known computation. During an unconstrained scan, people may remember, plan, monitor the room, mind-wander, regulate emotion, or think about other people. Default-network activity should not be equated automatically with self-projection, introspection, or any one mental content.

Second, broad network overlap does not mean that the same neurons perform the same operation in every task. Precision mapping can resolve neighboring subnetworks with different preferences for episodic construction and mentalizing [@bragabuckner2017]. Behavioral and neuropsychological dissociations likewise show that one capacity can be impaired more than another.

Self-projection is therefore best treated as a useful organizing hypothesis at an intermediate level. It identifies a family resemblance among tasks that construct situations beyond the immediate present. It does not settle whether belief reasoning depends on a specialized component, whether episodic memory is necessary for social inference, or whether one domain-general simulator explains the entire family.

The hypothesis nevertheless fits the control-system arc of the book. A body benefits from representing consequences before acting. A social body benefits from representing another agent’s perspective before the agent acts. Neural architecture can support both through partly shared systems for constructing alternatives, while retaining specialized pathways for the different information each problem requires.

The self-projection account is elegant, but elegance is not evidence. The central dispute is whether belief reasoning depends on special-purpose machinery or is one use of a more general capacity to construct displaced situations.

The domain-specific case begins with the computational demand. Beliefs are representations that can disagree with reality, and predicting from them requires keeping track of who represents what. Portions of the right posterior temporoparietal junction often respond more strongly to belief information than to other socially relevant facts about a person. Such selectivity is difficult to explain if the tissue performs only a generic shift of attention or scene construction [@saxe2003].

The domain-general case begins with convergence. Remembering, imagining, navigating, and taking another perspective recruit overlapping medial and lateral networks. Each task constructs a situation that is not the present one, and each could draw on shared operations for scene construction, perspective transformation, and evaluation [@bucknercarroll2007].

The strongest evidence now argues against the simplest form of either story. Group-average maps make the overlap look unitary, whereas precision mapping resolves adjacent or interdigitated subnetworks, one more strongly weighted toward episodic construction and another toward mentalizing [@bragabuckner2017]. Neuropsychological dissociations show that severe episodic-memory impairment need not eliminate standard theory-of-mind performance. Conversely, belief reasoning clearly depends on more than one tiny belief-selective patch.

A plausible architecture therefore contains both shared and specialized components. General constructive systems may supply scenes, alternatives, and perspective shifts; more selective circuits may track agents, informational access, and belief. The unresolved question is not whether social cognition is wholly modular or wholly general. It is how these components are divided, coupled, and recruited across tasks.

50.9 More than one route into another mind

The word empathy is often used for several processes that should be separated.

One route is cognitive perspective-taking: estimating what another person perceives, knows, wants, or believes. This is continuous with the mentalizing described so far. It can be accurate without producing a matching emotion. A negotiator, clinician, teacher, or deceiver may understand another person’s state precisely while remaining calm or even using that understanding against them.

A second route is affective sharing: another person’s expression, voice, posture, or situation recruits some of the observer’s own affective and bodily systems. Distress can become contagious before its cause is clearly understood. A person may feel another’s fear or pain while misidentifying what produced it. Affective sharing is therefore not simply perspective-taking with an emotion added. The routes interact, but they can come apart.

Neither route guarantees compassion. Accurate mentalizing can support cooperation, teaching, and care, but also manipulation. Shared distress can motivate help, avoidance, or self-protection. Moral action depends on valuation, norms, relationships, and control in addition to representing or sharing another person’s state.

The separation also reveals a computational problem. If another person’s state is partly represented through systems used for one’s own actions and feelings, the state must remain tagged as theirs. Otherwise perspective-taking collapses into contagion, and simulation is mistaken for direct self-experience. Networks recruited when people inhibit automatic imitation overlap with systems involved in mental-state reasoning and self–other distinction [@spengler2009]. The overlap is consistent with a need to use shared representational resources without losing track of whose state is being represented.

This chapter therefore treats access to another person as a coordinated set of routes: perception of body and voice, goal and belief estimation, affective resonance, memory for the relationship, and control of self–other boundaries. Different tasks and clinical conditions can alter these routes unequally. Calling all of them “empathy” can conceal more than it explains.

Figure 50.4: More than one route into another mind. Social understanding draws on partly separable but interacting processes rather than a single faculty. 1, Social evidence. Gaze and orientation, action, speech content, vocal prosody, facial expression, and posture provide evidence about another agent. The ongoing situation, social roles, norms, relationship history, and prior interactions supply additional context. These information sources are not assigned exclusively to one route: the same cue can contribute both to an estimate of the other agent and to an affective or bodily response in the observer. 2, Interacting routes within the observer. The upper route, agent-state estimation, uses available evidence to estimate what the other agent could perceive, what outcome the agent is pursuing, what information reached the agent, and what the agent knows or believes. The lower route, affective and bodily resonance, represents responses evoked in the observer, including autonomic arousal, affective experience, facial or motor resonance, and tendencies to approach, withdraw, or otherwise act. Appraisal and perspective can alter the observer’s affective response, while bodily and affective responses can bias attention, inference, and action. The routes can therefore dissociate without being independent. 3, Source attribution and self–other control. The observer must keep an estimate of their state—the other agent’s goal, perceptual access, knowledge, belief, or emotion—distinct from my response to it—the observer’s arousal, feeling, bodily response, and action tendency. Source and ownership may nevertheless be confused or imperfectly regulated. 4, Action selection. Estimates and evoked responses do not determine behavior directly. They are integrated with the observer’s goals, values, norms, relationship to the other agent, expected costs and benefits, empathic concern or personal distress, executive control, and available actions. Possible outcomes include helping or comforting, teaching or coordinating, withdrawing or avoiding, and deceiving or manipulating. The dashed return path indicates that the observer’s action changes the interaction and generates the next round of social evidence. Blue and warm-colored pathways distinguish the two principal routes conceptually; they do not identify isolated neural modules. Neither route, alone or in combination, guarantees compassion or prosocial action, and the distributed architecture shown here is schematic rather than exhaustive.

An influential account proposed that autism involves a specific impairment in representing other minds. Early studies found that many autistic children failed standard false-belief tasks even when comparison groups with other developmental conditions passed them [@baroncohen1985]. The result helped establish theory of mind as an experimentally tractable construct.

The broad deficit claim did not survive intact. Autistic performance is heterogeneous and varies with language, executive demands, familiarity, anxiety, and the form of the task. Many autistic people pass explicit false-belief tests, and some describe rich, effortful strategies for understanding others. Claims that autistic people categorically lack theory of mind misrepresent both the evidence and autistic experience [@gernsbacheryergeau2019].

The double empathy problem shifts part of the explanation from an impaired individual to a mismatched interaction. Autistic and non-autistic people may differ in expressive timing, sensory priorities, expectations, and conversational conventions. Prediction can therefore fail in both directions: each person is estimating an agent whose signals and habits are less familiar. Some studies find smoother understanding or rapport within neurotype than across neurotypes, although the magnitude and generality of that pattern vary [@milton2012].

The relational account should supplement rather than erase individual differences. Autism can involve real difficulties with language, flexible inference, sensory regulation, or social prediction, and those difficulties vary greatly. Non-autistic partners and institutions can also create avoidable failures by treating one communicative style as the only normal one.

The lesson is methodological as well as humane. A task result does not reveal the presence or absence of one faculty in a whole class of people. Social understanding is produced by two interacting systems, a history of experience, and a particular setting. Some breakdown belongs to the person; some belongs to the pairing; some belongs to the task used to measure it.

50.10 The nearest agent is the self

The word self covers several phenomena that should not be collapsed.

The embodied self is the organization of experience around this living body. Interoceptive, proprioceptive, vestibular, tactile, and visual signals continuously constrain estimates of bodily state and location. The access is privileged but not transparent. Signals are noisy, integrated, and sometimes misleading, as illusions and disorders of bodily ownership make clear. Still, the nervous system receives information from its own viscera and motor apparatus that it cannot receive from another person’s body.

The agentive self is the sense that an action is being initiated or controlled by oneself. Motor commands are accompanied by predictions of their likely sensory consequences, allowing self-produced changes to be distinguished from external events and permitting online correction. Agency can also be disrupted or misattributed. It is a construction grounded in special access to action-related signals, not an infallible inner witness.

The narrative self is the temporally extended account of who this person is, what they value, and why they acted. It draws on autobiographical memory, language, social feedback, and projections into the future. This is the layer most directly connected to mentalizing. To explain one’s own completed action, the system uses many of the same kinds of evidence used to explain someone else: the situation, remembered goals, bodily state, prior habits, and the action’s consequences.

The similarity should not be overstated. Self-knowledge has access to bodily and motor signals that an observer lacks. It also has blind spots. The physiological signal that the heart is racing does not specify whether the cause is fear, anger, exertion, or attraction. The feeling of having chosen does not expose every neural and contextual influence that shaped the choice. A verbal reason can therefore be an interpretation rather than a direct readout.

Split-brain confabulation makes that limitation unusually visible. When the speaking hemisphere is asked to explain an action whose determining information was delivered to the disconnected hemisphere, it can produce a plausible reason unsupported by the actual cause. The result shows that a coherent self-explanation can be assembled from incomplete evidence; it does not show that every ordinary explanation is fictional [@dehaan2020]. Anosognosia provides another extreme case in which people may deny an obvious deficit and construct explanations around the denial. Again, a selected clinical dissociation demonstrates possibility, not frequency in healthy life.

The chapter’s working synthesis can now be stated with appropriate boundaries. The embodied and agentive selves are anchored in privileged bodily and motor signals. The narrative self is more inferential and socially constructed. It is shaped by the same cultural language used to describe other people and may recruit overlapping systems for memory, prospection, and mental-state explanation. Whether it is literally a model of the same kind as a model of another person remains an open question.

The title The Nearest Agent names that asymmetry. The self is nearest because its body supplies continuous internal evidence. It is also difficult to explain because the system cannot step outside its own causal history to inspect every process that produced an action. We know ourselves differently from the way we know others, but not completely.

Figure 50.5: The nearest agent: interacting forms of self-representation and incomplete causal access. A, The figure distinguishes three conceptually separable but continuously interacting forms of self-representation. The embodied self is an estimate of bodily state, ownership, location, and orientation constructed from interoception, proprioception, vestibular signals, touch, first-person vision, and other bodily evidence. These signals are unusually available from the first-person perspective, but they must still be integrated and may sometimes be misleading. The agentive self concerns initiating, controlling, correcting, and attributing actions. Goals and intentions guide motor commands and action; motor-related signals support predictions of expected consequences, while actual sensory consequences provide feedback for correction and updating. Predictive, sensory, and contextual evidence all contribute to experienced agency. The narrative self extends identity across remembered past, interpreted present, and imagined future. It draws on autobiographical memory, language, social feedback, roles, values, cultural concepts, bodily states, intentions, actions, and their consequences. Narrative identity in turn influences future goals and action selection. The overlapping bands and bidirectional arrows indicate interaction rather than anatomical compartments, fixed processing stages, or separate self centers. B, The split-brain example illustrates how a coherent explanation can be constructed despite incomplete access to the cause of an action. A determining cue guides the action, but that cue is unavailable to the speaking explanatory system in the task. The person can nevertheless use the observed action and surrounding context to generate a plausible verbal reason. The explanation may be sincere and coherent while remaining unsupported by the cue that actually determined the response. This clinical dissociation demonstrates that reconstruction from partial evidence is possible; it does not imply that ordinary self-explanation is generally false. Solid arrows indicate available information flow, bidirectional arrows indicate mutual influence, the short dashed pathway represents predicted consequences, orange arrows indicate influences on goals and identity, and the interrupted pathway marks unavailable causal information. The self is the nearest agent because bodily and action-related evidence is asymmetrically available, but the causes of one’s own actions remain only partly accessible and may have to be inferred.

50.11 Coda: the system that models systems

We began this book by asking why an animal should have a brain at all. The first answer was movement and control: a nervous system links sensing to action so that a body can remain within the conditions required for life and reproduction.

Each unit widened that loop. Regulation became anticipatory as the organism prepared for disturbances before they arrived. Learning allowed past encounters to alter future control. Organisms modified their environments, and persistent modifications altered later development and selection. Human dependence distributed calories, labor, care, and knowledge across many people. Language allowed one agent’s observation or plan to reorganize the behavior of another.

This final chapter turned the control system toward the agents that make up that niche. Other people are difficult to predict because their actions depend on goals, information, memories, and expectations that are not directly visible. Much of social life can still be governed by behavior, habit, convention, and role. Where those cues underdetermine action—especially under informational asymmetry and recurrent strategy—estimating another person’s perspective or belief becomes valuable.

The neural evidence does not reveal one mind-reading organ. Mental-state reasoning recruits a distributed architecture involving temporoparietal, medial prefrontal, posterior medial, temporal, memory, attention, and language systems. Some components overlap with biological-motion perception, episodic construction, and prospection. Finer mapping also reveals neighboring and interdigitated subnetworks with different preferences. Specialization and reuse coexist.

The same balance applies to the self. A body supplies privileged interoceptive and motor evidence, so self-knowledge is not simply mind-reading turned inward. Yet the narrative explanation of one’s own actions is partial, reconstructive, and culturally expressed. It can draw on machinery used to remember, imagine, and explain other agents, and it can be wrong with confidence.

The book therefore ends without locating the distinctively human in one new cortical box. The more continuous conclusion is also the more demanding one. Older systems for regulation, perception, movement, memory, and social interaction were reorganized within an unusually constructed niche. That niche rewarded brains able to learn from many people, communicate across perspectives, and prepare for agents who were preparing in return.

A honey bee is built to live in a hive that bees built. A human brain is built to develop among other brains and to estimate enough of their hidden states to coordinate, compete, teach, trust, and care. The nearest agent—the self—is known through a different mixture of privileged signal and interpretation, not through perfect access.

Animals create niches that modify animals. For the human animal, the most consequential features of the niche are other agents, including the one doing the modeling.


This chapter has moved from social prediction to false belief, neural networks, empathy, and the self. The evidential foundations are not equally firm.

We are confident that: humans routinely use behavior and context to estimate unobservable goals, knowledge, and beliefs; false-belief tasks create a useful decoupling between reality and another agent’s representation of reality; standard verbal tasks depend on language, memory, inhibition, and narrative tracking as well as mental-state reasoning; temporoparietal, medial prefrontal, posterior medial, and connected regions contribute reliably to mentalizing; mentalizing shares broad anatomical territory with biological-motion perception, attention, episodic construction, and prospection while finer mapping reveals functional differentiation; and coherent self-explanations can be constructed from incomplete evidence in split-brain and anosognosia cases.

We have good reason to think that: visible behavior, learned habit, and social convention solve much social prediction without explicit belief reasoning; informational asymmetry and recurrent strategic interaction increase the value of tracking an agent’s perspective; cognitive perspective-taking and affective sharing are interacting but partly separable routes; self–other attribution is part of using shared sensorimotor or affective resources; and the narrative self draws substantially on memory, language, social feedback, and inferential reconstruction.

We remain genuinely unsure about: which infant false-belief findings are robust and what nonverbal looking measures represent; whether any nonhuman animal represents false beliefs in the same flexible sense tested by explicit human tasks; whether belief reasoning depends on a dedicated component, a domain-general constructive system, or an architecture containing both; how to draw stable functional boundaries within temporoparietal and default-network territories; how broadly the double-empathy account generalizes across autistic and non-autistic interactions; and how literally the narrative self should be treated as a model built by the same machinery used for other agents.

The secure claims support a distributed, architecture-first account of social understanding. The open questions concern the representational depth, comparative reach, and degree of specialization of that architecture.