41  Wandering When Full

Exploration, Information, and Future Controllability

41.1 A controller cannot use an option it has never discovered

Chapter 40 ended with a controller whose actions can be governed by another person’s future. That achievement depends on more than valuing another person’s welfare or representing what another person knows. The partner, promise, reputation, local norm, or institution must first have become part of the animal’s world. A controller cannot plan to recruit a particular partner it does not know exists, follow a route it has neither sampled nor inferred, or choose a refuge whose existence it never learned.

The same problem has been present throughout this unit. Chapter 37 asked how food, water, shelter, social approval, and uncertain alternatives become comparable enough to compete for action. Chapter 38 asked how one selected future remains effective after its cue disappears. Chapter 39 asked how a represented consequence changes the body before it occurs. Every one of those problems presupposes a repertoire of possibilities. Before an outcome can be priced, protected, or felt, the organism must have acquired enough information to represent it.

The original temptation is to imagine a library of futures that exploration fills and valuation later searches. The metaphor captures one truth and hides another. Experience does enlarge what can be represented, but there is no passive index standing apart from the rest of control. Sampling a new route changes spatial and relational memory. Encountering a new partner changes social expectations. Trying an unfamiliar action changes estimates of effort, risk, and consequence. The current body helps determine what is sampled, and what is learned changes later value. Exploration, memory, valuation, and action selection form a recurrent loop.

The central requirement can be stated without installing an inner explorer:

A controller cannot use an option it has never discovered. Exploration changes what future control will be able to do.

The animal may acquire that information while pursuing an immediate need, while testing an uncertain alternative, while playing, or while roaming under conditions of satiety and safety. Those behaviors do not all arise from one drive. What joins them is that action changes the organism’s information about the world and therefore changes the set of policies available later.

41.2 Exploration is not one behavior

The most familiar formulation is the explore–exploit dilemma. To exploit is to select an option largely because experience already supports its value. To explore is to sample in a way that can change what is known about the available options. The terms describe the informational relation between action and later choice. Exploitation is not selfishness, and exploration is not aimless wandering.

Several different policies can produce exploratory behavior. Directed exploration favors an uncertain option because information from that option could improve later decisions. The information has present value because it can alter a future choice. Random exploration increases behavioral variability, allowing less-favored alternatives to be sampled without explicitly adding an information bonus to their estimated value. Both can be adaptive. One biases choice toward information; the other loosens choice enough for alternatives to be sampled, preventing a controller from becoming permanently trapped by its current estimates.

Other behaviors overlap with these strategies but are not identical to them. Novelty seeking favors something because it has not been encountered before. Novelty often implies uncertainty, but the two can be separated. A familiar option may remain poorly understood, and a novel stimulus may reveal little that matters. Investigatory behavior includes orienting toward, approaching, sniffing, touching, or otherwise examining an object or agent. Patch leaving abandons a declining local source and accepts the cost of searching elsewhere. Need-directed search occurs when the current deficit cannot be corrected through a reliable known policy. Play can reveal the mechanical and social consequences of actions without an immediate practical payoff. Social sampling can identify partners, rivals, teachers, conventions, and sources of information.

These distinctions prevent curiosity from becoming another invisible faculty with a single neural address. A mouse inspecting an unfamiliar object, a monkey preferring advance information about a reward, and a person selecting the less-known option in a bandit task all acquire information. They do not necessarily respond to the same variable, compute the same future benefit, or use the same circuitry.

Exploration can also be organized at several timescales. A momentary deviation may look like noise when viewed one choice at a time but become part of an effective search policy across many choices. Conversely, a systematic attraction to novelty can be locally orderly and globally wasteful. The category cannot be defined by whether behavior looks purposeful to an observer. It must be defined by what the sampling changes and how that information can influence later action.

Exploration is a family of information-acquiring control policies, not a unitary drive released whenever other drives fall silent.

41.3 Information can be worth acquiring now

The Horizon task makes the future value of information experimentally visible. Participants chose between two slot machines with uncertain payoff distributions. Before free choice began, the task forced several samples from the machines. In some games, the two options had been sampled equally. In others, one had been sampled more often, leaving the other less well known. Participants then made either one free choice or six free choices.

When only one choice remained, information acquired on that choice could not improve another decision within the game. When six choices remained, an early sample could change what the participant did on the next five. The longer horizon therefore increased the prospective usefulness of information without changing the immediate rewards offered by the machines.

Participants responded in two separable ways. When one option was less well known, they became more likely to choose it under the longer horizon. That change was directed exploration: uncertainty received an information bonus because resolving it could improve later choices. Participants also became more variable when the options were equally well known. That change was random exploration: the longer horizon increased sampling by reducing the determinism of reward-guided choice [@wilsonetal2014exploration].

The result rules out a simple contrast between intelligent exploitation and mindless exploration. Some exploratory variability is controlled by the number of future decisions in which acquired information can be used. The participant need not articulate a theory of information value, but behavior is nevertheless sensitive to it.

Information can also acquire value even when it cannot change the outcome. In an experiment with macaques, one cue revealed the size of an upcoming reward in advance, whereas another left the animal ignorant until the reward arrived. The informative cue did not improve the reward and did not permit a different action, yet the monkeys preferred it. Midbrain dopamine neurons signaled both expected reward and expected information about reward [@brombergmartinhikosaka2009information]. Information seeking therefore cannot be reduced to the instrumental value of improving a later choice. Organisms sometimes value reducing uncertainty about an outcome they cannot control.

Novelty adds another distinction. Cockburn and colleagues designed a human task that separated how recently an option had been encountered from how uncertain its reward probability remained. Uncertainty-directed exploration was sensitive to whether the information could still be useful. Novelty produced a more persistent attraction, modeled as an optimistic initial value for less-familiar options rather than as a calculation of prospective information gain [@cockburnetal2022novelty]. Novelty and uncertainty can cooperate, but they do not enter choice through one common mechanism.

These findings replace the claim that exploration is simply a dumb pull toward the new. Some novelty-directed behavior may indeed be immediate and myopic. Random exploration can arise through increased variability rather than a represented information bonus. Directed exploration, however, depends on the relation between uncertainty and the choices still to come. The mechanism need not understand its evolutionary history. It can still represent that learning now will improve action later.

41.4 Wandering when full—and searching when hungry

The title identifies a real ecological opportunity. A well-fed, well-watered, safe animal can afford actions that do not produce an immediate regulatory return. Failure is less costly. Time can be spent inspecting a new path, manipulating an object, playing with a conspecific, or learning the distribution of resources before any one resource becomes urgent. Wandering when full is exploration conducted with regulatory margin.

That margin matters because information has a price. Sampling consumes energy and time. It exposes the animal to hazards and may require abandoning a reliable source. When dehydration is severe and a known water source is available, unrelated exploration has a high opportunity cost. The familiar route deserves control because delay itself is dangerous.

But need does not switch exploration off. It changes what information is worth acquiring. When the familiar patch is exhausted, hunger makes search necessary. When a route to water is blocked, thirst raises the value of discovering another route. Threat can suppress broad sampling while making information about an escape path urgent. An animal in need may explore less widely and more purposefully, or it may accept risks that a sated animal would reject.

A recent mouse experiment provides a direct corrective to the claim that hungry animals do not explore. Food-restricted mice showed more interaction with a novel, noncaloric object and less risk assessment than sated mice. Hunger suppressed dopamine release in the tail of the striatum during encounters with the object. Manipulating AgRP neurons and tail-striatal dopamine altered both investigation and risk-related behavior, linking caloric state to a specific form of exploratory control [@kamathetal2025hunger].

The result is important precisely because the object did not contain food. Hunger changed how the animals sampled a potentially threatening novelty even in the absence of a caloric reward. It did not merely point locomotion toward a known meal. The experiment does not establish one universal effect of food deprivation across species and settings. It establishes the more important principle: bodily need can increase exploration as well as narrow it.

Three control situations should therefore be distinguished. When an urgent need is paired with a reliable known policy, exploitation is favored. When an urgent need is paired with failure or uncertainty, search can become necessary. When urgent needs are quiet and the animal has regulatory margin, broad exploration becomes more affordable. Satiety is not an on-switch for curiosity. Hunger is not an off-switch. Bodily state changes the target, breadth, risk tolerance, and useful time horizon of information acquisition.

The same logic applies beyond hunger. Fatigue changes the cost of a distant route. Thermal stress changes the value of shade and the danger of delay. Injury changes which actions are feasible. Social isolation can raise the value of information about potential partners while social threat raises the cost of approaching them. The body does not supply exploration with a generic amount of motivation. It changes the consequences of sampling this option, in this place, now.

Bodily state changes the price and the likely use of information.

41.5 Exploration changes what the controller knows

The immediate product of exploration is not always reward. Sometimes it is a change in what the animal can later infer.

The classic latent-learning experiments made this distinction visible. Tolman and Honzik allowed rats to traverse a maze without food reward and later introduced food at the goal. Performance then improved rapidly, indicating that experience acquired before reward had altered what the animals could do once an incentive made that knowledge useful [@tolmanhonzik1930latent]. The experiment did not reveal a complete modern theory of cognitive maps. It established a durable distinction between acquisition and performance. What an animal has learned cannot be read directly from what it is currently motivated to express.

Exploration can change several kinds of knowledge. It can reveal where resources and hazards are located, which landmarks predict a turn, and which routes remain available when the shortest path is blocked. It can reveal transition structure: what state is likely to follow an action, where a corridor leads, or how a tool changes an object. It can reveal volatility: whether the environment is stable enough for an old estimate to remain useful. It can also reveal affordances that are specific to the body. A gap that can be crossed when rested may not be crossable when injured or carrying a load.

The social world enlarges the same problem. Exploration identifies more than mates or competitors. It reveals who provides reliable information, who reciprocates, who dominates access to a resource, which signals are trusted, and which local rules govern exchange. The future of another agent can influence action only after recurrent interaction has made that agent’s likely responses partly predictable.

Exploration also changes estimates of ignorance. Sampling may reveal that a familiar option is more variable than it appeared, that a supposedly safe route has become unreliable, or that the animal’s model omitted an entire class of possibilities. The most important discovery is sometimes not a new resource but the inadequacy of the current policy.

These changes do not remain in an isolated exploration system. Hippocampal and cortical memory systems preserve relational, spatial, episodic, semantic, and social structure. Value-learning systems update expected consequences. Sensory and motor systems become more efficient at recognizing and acting on familiar features. The newly acquired information changes later attention, valuation, route selection, and imagination.

This is the bridge to the next unit. The present chapter asks why and when an organism samples beyond its current policy. The next unit asks how experience becomes the maps and memories from which routes, scenes, and possible futures are later constructed.

41.6 Stability and exploration are complementary control states

Chapter 38 described the problem of keeping a valid goal effective across distraction while remaining able to update when evidence changes the situation. Exploration is not the opposite of that capacity. It is one of the ways updating occurs.

A novel event can be an irrelevant distractor, evidence that the current policy has failed, or an opportunity worth sampling. Those cases cannot be classified from novelty alone. A flash in the periphery should usually not pull a thirsty animal off a reliable route to water. A newly fallen tree blocking that route must reorganize behavior. The controller has to determine whether the present event changes the problem.

Directed exploration can itself require strong goal maintenance. A scientist who systematically varies one condition, a child testing how a mechanism works, or an animal checking several branches of a route is not surrendering to capture. Each must resist immediate rewards and unrelated novelty long enough to complete an information-gathering policy. Random exploration likewise need not be uncontrolled distraction. Increasing variability can be a regulated response to uncertainty or environmental change.

Recordings from monkeys show what such a change of control state can look like. Ebitz and colleagues recorded from frontal eye field neurons while animals shifted between exploiting a reliably rewarded target and exploring alternatives. During exploration, spatially selective choice-predictive activity weakened and population dynamics associated with the chosen target emerged later. At the same time, behavior and neural activity became more sensitive to new reward outcomes [@ebitzetal2018exploration]. The established sensorimotor mapping loosened while learning increased.

That reconfiguration is not simply a failure of prefrontal control. It is a change in what control is doing. A stable mapping is useful when the environment is reliable. The same stability becomes a liability after reward contingencies change. In monkeys performing reversal learning, environmental volatility increased the influence of recent outcomes and altered outcome- and choice-related activity in orbitofrontal and dorsolateral prefrontal cortex [@massietal2018volatility]. The nervous system does not need one mechanism for persistence and another for flexibility. It needs recurrent circuits whose stability changes with evidence about the world.

The relationship can therefore be stated more precisely than a contest between dorsolateral persistence and exploratory capture:

Adaptive control must neither surrender to every novelty nor preserve a policy after the world that justified it has changed.

Exploration is controlled relaxation, variation, or replacement of a current policy when the expected return from information exceeds the cost of sampling. Exploitation is controlled protection of a policy whose estimated consequences remain good enough. Neither state is intrinsically superior. The problem is the transition between them.

41.7 The machinery is distributed and evolutionarily old

No single region decides whether an organism should explore. Different forms of exploration require different information, actions, timescales, and costs. The relevant machinery spans frontal cortex, posterior cortex, hippocampal formation, basal ganglia, amygdala, hypothalamus, neuromodulatory systems, diencephalon, midbrain, and brainstem.

41.7.1 Frontal systems keep alternatives and horizons effective

Human imaging first drew attention to the frontal pole. In a dynamic gambling task, exploratory decisions produced greater activity in frontopolar cortex and intraparietal sulcus, whereas activity in striatum and ventromedial prefrontal cortex more closely followed value-guided exploitation [@dawetal2006exploration]. Another experiment found that lateral frontopolar activity tracked evidence favoring a foregone alternative and changed its coupling with other regions when participants switched [@boorman2009alternatives].

These findings identify contributions, not an exploration center. The strongest causal evidence is narrower. Inhibiting right frontopolar cortex with transcranial magnetic stimulation reduced directed exploration in the Horizon task while leaving random exploration intact [@zajkowski2017frontopolar]. Frontopolar cortex therefore helps information acquire control when uncertain alternatives must be compared across a future horizon. The same result argues against assigning all exploration to that territory.

Medial frontal cortex enters when continuing a current course must be compared with searching elsewhere. In a human foraging task, dorsal anterior cingulate activity tracked variables related to leaving a current option and searching an environment whose average richness varied [@kollingetal2012foraging]. A subsequent experiment that separated those variables found that activity attributed to foraging value was better explained by choice difficulty [@shenhavetal2014difficulty]. The disagreement is instructive. Medial frontal activity participates when effort, uncertainty, switching, and control demand must be integrated. One fMRI contrast cannot turn that participation into a dedicated foraging computation.

Lateral frontal populations maintain task rules, current evidence, and the horizon over which information can pay. They also participate in the reconfiguration described above. Their role is not simply to suppress novelty. They help determine which variable—current reward, uncertainty, an alternative policy, or a long-range goal—continues to organize action.

41.7.2 Value, learning, and action selection

Exploration already contains a value problem. The organism must compare immediate reward with the possible benefit of information, subtract energetic and opportunity costs, and incorporate current bodily state. Orbital and ventromedial prefrontal systems contribute state- and outcome-specific value signals; they do not wait for an exploration system to deliver a completed map.

In monkeys choosing among familiar and newly introduced options, orbitofrontal neurons encoded chosen stimulus identity, outcomes, estimated value, and an exploration bonus associated with the possible future value of sampling a novel option [@costaaverbeck2020orbitofrontal]. Ventral striatum and amygdala also represented different aspects of immediate exploitative value and future exploratory value [@costaetal2019subcortical]. Exploration is therefore embedded in the same cortical–subcortical systems that learn consequences and select actions.

Dopamine makes the point especially clearly because its contribution changes with projection, target, state, and task. The macaque information experiment showed dopamine signals related to advance information about reward [@brombergmartinhikosaka2009information]. Blocking dopamine transport increased monkeys’ preference for newly introduced options by increasing their initial estimated value, without producing a general increase in choice noise or learning rate [@costaetal2014dopamine]. In hungry mice, by contrast, increased investigation of a novel object depended on suppressed dopamine signaling in the tail of the striatum [@kamathetal2025hunger]. There is no single dopamine signal that makes the unknown attractive. Through anatomically distinct circuits, dopamine systems can alter learning, salience, risk assessment, action vigor, and the value assigned to novelty.

Basal-ganglia output can also regulate the breadth of action sampling. During associative learning in monkeys, lower firing in a subset of internal globus pallidus neurons predicted early exploratory responses, whereas higher firing accompanied later exploitation of the learned response [@shethetal2011basalganglia]. This fits the broader architecture developed in Chapter 30. Basal ganglia do not generate one psychological faculty. They alter which actions gain access to downstream controllers and how tightly alternatives are constrained.

Baseline and task-evoked pupil diameter have also changed as participants moved between sustained task engagement and lower-utility states interpreted within an explore–exploit framework [@gilzenratetal2010pupil]. Such findings are consistent with a contribution from neuromodulatory systems, including locus coeruleus–noradrenergic projections, that can alter neural gain and the stability of ongoing processing. Pupil diameter is influenced by many processes and is not a direct assay of one nucleus. The important point is that exploration involves changes in the operating state of distributed circuits, not only a comparison made at the frontal pole.

41.7.3 Older circuits investigate a world

The ability to orient toward and investigate uncertainty is far older than human granular prefrontal cortex. Tectal and brainstem systems turn eyes, head, and body toward sudden or biologically important events. Hypothalamic state signals change what risks are acceptable. Diencephalic and midbrain circuits organize sustained investigation.

The zona incerta provides a particularly clear comparative example. In mice, inhibitory neurons in medial zona incerta are necessary for sustained investigation of novel objects and conspecifics. They receive excitatory input from prelimbic cortex and promote deep investigation partly by inhibiting periaqueductal gray [@ahmadlouetal2021investigatory]. In macaques, a projection from temporal cortex to zona incerta contributes causally to novelty seeking [@ogasawaraetal2022novelty]. These circuits differ in detail, but both demonstrate that investigatory behavior is implemented through cortical–subcortical loops rather than invented by a uniquely human frontal territory.

Human frontal elaboration changes the scale of the problem. It allows uncertain alternatives to remain effective across long delays, permits several branches of a plan to be compared, and lets symbolic or socially transmitted information guide where exploration is directed. A person can explore a hypothesis, a legal precedent, an archive, or another person’s testimony without physically roaming a landscape. The ancient operation—sample beyond the currently favored policy—has been extended across language, institutions, and imagined futures.

41.8 Exploration as an allostatic investment

The scientific distinctions now permit the chapter’s larger claim to be stated more strongly and more precisely. Exploration is not an allostatic drive separate from hunger, thirst, and other regulation. Chapter 37 already showed that those systems are predictive: bodily needs can be anticipated, and their expected future consequences can alter present value. Allostasis is not a faculty that belongs to one motive. It is regulation organized in advance of an expected demand.

Exploration becomes allostatic when present sampling improves the organism’s capacity to regulate later. The organism pays now in energy, time, exposure, and foregone reward. The return can be a shorter route to water, an alternative refuge, a more accurate estimate of danger, a new action–outcome relation, a reliable partner, or evidence that the current policy is obsolete.

This book uses future controllability for that return. The phrase does not mean a feeling of mastery. It means the number, accessibility, reliability, and estimated consequences of the actions through which a later disturbance can be corrected. An animal with one known route to water is vulnerable to one obstruction. An animal that has sampled several routes, learned their costs, and discovered another source possesses greater future controllability even while it is not thirsty.

The investment has identifiable boundary conditions. Information is most useful when the organism is likely to encounter the environment again, can retain or generalize what it learns, and faces enough stability for the information to remain valid. Exploration loses value when the setting is one-shot, when change is so rapid that yesterday’s sample misleads, or when present risk is catastrophic. Novelty can also be engineered to capture attention without improving later control. An exploratory system can waste time, enter a trap, or continue sampling after the useful uncertainty has been resolved.

These failures do not weaken the functional claim. They define it. Selection does not produce a homunculus that knows why information will matter. It produces state-sensitive mechanisms that vary behavior, investigate novelty, value uncertainty, and learn from the result. Such mechanisms are favored when their average future return exceeds their present cost. The animal need not represent that evolutionary reason. Directed exploration can still represent the local fact that information gained now will improve choices later.

Exploration is an allostatic investment when present information acquisition enlarges the set of reliable actions available to future regulation.

This formulation preserves what was strongest in the original idea. Wandering when full can stock knowledge before need becomes urgent. It no longer requires a unitary curiosity drive, a satiety switch, or a frontal region that decides when to indulge the unknown. It describes an organism-level function implemented by several policies and many interacting circuits.

41.9 Coda: a represented future still has to be mobilized

Exploration can make an option available. Valuation can make it desirable. Prospective bodily change can make its consequences felt, and maintained control can keep it effective after the cue disappears. None of those operations guarantees that action will begin.

A person may describe a sensible plan, recognize its value, and retain every component needed to carry it out. The plan can remain only something said. Between an available future and an initiated course of action lies another control problem: mobilizing effort, selecting a starting action, sustaining engagement, and reinitiating after interruption.

The syndromes of apathy, abulia, and akinetic mutism expose that gap. They do not reveal one reservoir of will. They reveal what happens when distributed systems linking expected value, bodily preparation, effort, basal-ganglia selection, medial frontal control, thalamus, and motor release no longer bring an available goal into action.

The next chapter, Tomorrow Never Comes, follows the future to that threshold.

Established findings. Organisms balance exploitation of known options with several distinguishable forms of exploration. Human behavior separates directed information seeking from exploration produced by increased choice variability, and both change with the number of future decisions in which information can be used. Novelty and uncertainty can influence exploration through different mechanisms. Bodily state can narrow, redirect, or increase exploration rather than simply turning it off. Frontopolar, lateral frontal, medial frontal, orbital, striatal, amygdalar, hippocampal, hypothalamic, diencephalic, midbrain, brainstem, and neuromodulatory systems make distinguishable contributions. Investigatory circuitry is evolutionarily older than human frontal pole.

Working synthesis. Exploration can function as an allostatic investment in future controllability. By paying present costs to acquire information, an organism can enlarge the number, accessibility, and reliability of the actions through which later disturbances may be corrected. This is an organism-level functional proposal. It does not identify a unitary curiosity drive, treat satiety as an exploration switch, or localize exploration to one cortical center.

Open questions. It remains uncertain how bodily states redistribute different forms of exploration across natural settings; how novelty, uncertainty, volatility, risk, and future horizon are integrated; how directed and random strategies interact during extended behavior; and how information acquired through physical, social, and symbolic exploration is consolidated into maps and models. The neuromodulatory mechanisms are also unresolved. Dopamine and noradrenaline act through heterogeneous projections, and the same transmitter can have different effects in different targets and states. Finally, an information-acquiring policy can improve future control or become a route for distraction, compulsion, and engineered capture. Explaining that boundary is part of explaining exploration itself.