37  A Common Currency

Bodily Need, Allostatic Prediction, and Value-Guided Choice

37.1 Value begins in a regulated body

Chapter 35 showed that consequences alter the competition among actions, and that the effect of a consequence depends on what the organism currently needs. A food can lose its control after selective satiety. A cue for concentrated salt can acquire motivational force when sodium depletion makes its predicted outcome useful. The identity of the outcome remains the same while its current value changes.

A further problem appears when several outcomes are available at once. A foraging animal may have access to food in one direction, water in another, shade nearby, shelter farther away, and unexplored terrain beyond. These possibilities differ in substance, location, delay, effort, and risk. They nevertheless compete for the same body and for the same interval of action. The animal cannot eat, drink, cool itself, hide, and investigate at the same moment. Several predicted consequences must become comparable enough for one course of action to gain control.

This problem is often introduced through economic choice, but economics is the late and culturally elaborated case. The older problem is regulatory. A vertebrate can be hungry, dehydrated, overheated, injured, threatened, and uncertain about its surroundings at the same time. Each condition changes the consequences of possible actions. Water gains value as osmotic and volume deficits grow. Shade gains value before body temperature reaches a dangerous level. A familiar shelter becomes more valuable when predation risk rises. Exploration becomes more valuable when urgent deficits are quiet and information acquired now can improve later control.

Value is therefore not a substance contained in the object. It is a relation between an expected outcome and an organism in a particular state. Michel Cabanac used the term alliesthesia for the change in pleasantness produced by a change in internal state: warmth can be welcome when cold and oppressive when hot; a food can be attractive when hungry and unappealing after satiety [@cabanac1971pleasure]. The sensory event need not change. What changes is the event’s significance for regulation.

The same relation can be written compactly as \(V(o \mid s)\) rather than \(V(o)\): the value of outcome \(o\) conditional on state \(s\). The relevant state includes more than the present concentrations of glucose, sodium, or water. It includes recent experience, current context, expected effort, available alternatives, social commitments, and predictions about what the body and environment will be like when the outcome arrives. Homeostasis makes a current regulatory error matter. Allostasis makes an expected future error matter before it occurs.

Need-sensitive circuits already display this anticipatory organization. AgRP neurons associated with hunger and subfornical-organ neurons associated with thirst carry aversive motivational signals, and cues predicting food can rapidly reduce AgRP activity before nutrients have entered the circulation [@betleyetal2015need]. More recent work shows the complementary effect: learned contextual information predicting an impending period without food can rapidly activate AgRP neurons before the future energy deficit has developed [@walkeretal2026anticipating]. These hypothalamic signals do not perform every comparison described in this chapter. They alter the state in which cortical, striatal, limbic, and brainstem systems evaluate what the available outcomes would accomplish.

Not every human value can be reduced to correction of one physiological variable. Money, promises, professional goals, social approval, and knowledge depend on learning, language, institutions, and models of futures that may be years away. Their value is not a disguised calorie count. They are nevertheless evaluated by an embodied organism. They alter expected access to resources, safety, affiliation, obligation, information, and future courses of action. The biological root of subjective value is not a single homeostatic set point. It is the significance of a predicted consequence for this organism’s control of its present and anticipated states.

Economic choice is a special case of an older regulatory problem: several predicted consequences compete for control of one body.

The terms surrounding value should remain distinct. An outcome is an event that an action or cue predicts. A reinforcer is defined by its effect on learning or behavior. Subjective value is the current weight assigned to a predicted or obtained outcome. Pleasure is a hedonic response during experience. Motivation concerns the mobilization and persistence of action. Salience concerns processing priority. A distant water source can be highly valued before it is tasted; a cue can motivate pursuit without being pleasurable itself; a painful stimulus can be intensely salient while carrying strongly negative value. These variables interact, but none can substitute for all the others.

37.2 Comparable enough to choose

A regulated body gives value a direction, but it does not by itself solve the comparison. Food may correct an energy deficit, water an osmotic deficit, shade a thermal load, and information an uncertainty that could become costly later. The nervous system must transform these consequences into a format in which each can influence the same decision.

A common currency is a domain-general decision format in which dissimilar expected outcomes become commensurable enough to guide one choice. The definition does not require a literal neural coin, one global meter, or one cortical location into which every attribute disappears. It requires signals whose magnitude has a comparable behavioral meaning across different options: greater positive values favor pursuit, more negative values favor avoidance, and the relative values help determine which action gains control.

Commensurability is not identity erasure. A controller that retained only a scalar labeled 7 would not know whether the seven referred to water, food, warmth, or social contact. It could not selectively reduce the value of one food after satiety, revise one prediction after a contingency change, or calculate how an outcome would interact with a later need. Adaptive choice therefore uses at least two kinds of information in parallel:

Outcome identity specifies what is expected to happen. Decision value specifies how strongly that expected outcome should influence this choice now.

Delay, uncertainty, and effort must also remain available. Water ten seconds away and water an hour away are the same kind of outcome but not the same option. A high-value food behind a dangerous barrier may lose to a modest food already within reach. A course that improves future controllability may deserve present effort even though it does not correct the strongest current deficit. Decision value can summarize these dimensions for a particular comparison, but the richer representation must remain accessible so that the summary can be recomputed when the state or task changes.

A mathematical utility function makes a related but narrower point. If choices are sufficiently orderly, their ranking can be represented by assigning numbers to the alternatives. That representation describes the ordering; it does not identify the neural algorithm that produced it. A utility function therefore does not prove that every attribute is converted into one explicit number in one brain region. Neural evidence is required to determine where domain-general value signals occur, what information remains outcome-specific, and how comparison changes action.

37.3 An exchange rate in orbitofrontal cortex

The clearest single-neuron evidence for a common decision format comes from experiments in which monkeys chose between two juices offered in varying quantities. The crucial behavioral measure was obtained before interpreting the neurons. By changing the amount of each juice, Padoa-Schioppa and Assad identified the indifference point at which the monkey chose either offer about equally often. If one unit of juice A was chosen as often as three units of juice B, then the animal’s revealed exchange rate was approximately \(1A = 3B\). Physical quantity and subjective value could now be separated [@padoaschioppaassad2006].

The exchange rate was not read directly from taste receptors. It was inferred from repeated choices. It also belonged to that animal under those task conditions. Another monkey, or the same monkey in another state, could reveal a different ratio. The experiment therefore supplied a behavioral scale against which neural activity could be tested.

Recordings in orbitofrontal cortex identified three prominent response types. Offer-value cells encoded the value of one particular offered juice. A cell associated with juice A changed its firing as the amount of A changed, with quantity expressed in units shaped by the monkey’s preference. These cells were generally good-specific: an offer-value A cell did not simply emit the same response whenever either juice had equivalent value. Chosen-value cells encoded the value of the option that the monkey selected in a format that generalized across the two juices. Chosen-juice cells encoded which juice was chosen. The original article called the latter response taste; later work commonly calls it chosen juice.

This division is more informative than a population containing only one abstract number. Offer-value activity preserved the identity of each candidate outcome while placing its quantity on a preference-weighted scale. Chosen-value activity represented the value of the selected offer across identities. Chosen-juice activity preserved the identity of the choice. The population therefore contained both the inputs to comparison and a representation of its result.

The offer locations varied, allowing value and chosen identity to be separated from the direction of the saccade used to report the choice. The principal signals were consequently not simple visual-location or motor-command signals. In this task, the OFC represented goods and their values before the choice had been reduced to a particular movement.

Later perturbation studies moved the argument beyond correlation. Low-current electrical stimulation altered the subjective value assigned to individual offers and biased choices in the predicted direction. Stronger stimulation disrupted valuation and comparison and increased choice variability [@ballestaetal2020causal]. A subsequent experiment used weak stimulation during the period in which the second offer became available. Choice became more variable even though the estimated offer values did not shift, showing a selective disruption of value comparison [@ballestaetal2022comparison].

The result supports a direct conclusion. In this juice-choice task, OFC populations encode the values and identities of offered goods, contribute causally to their comparison, and represent the value and identity of the selected outcome. Other regions can participate in the same decision, and other kinds of choice can weight action-based representations more strongly. Neither fact weakens the causal result obtained in this task. The OFC is not merely reporting a choice made elsewhere, and the value signal is not itself a hidden chooser. Comparison is a population process that transforms several represented alternatives into an emerging choice.

37.4 Shared and specific value codes in humans

Single-neuron recording provides precise evidence within one kind of choice. Human imaging asks whether related signals recur when the outcomes differ more substantially. Chib and colleagues measured real purchasing decisions involving foods, nonfood consumer items, and monetary gambles. Activity in an overlapping region of ventromedial prefrontal cortex covaried with each participant’s valuation across all three categories [@chib2009common]. Money is useful in this design because it has no fixed taste, texture, or consummatory response. Its value is learned from the future outcomes for which it can be exchanged.

A coordinate-based meta-analysis of 206 fMRI studies found a recurring pattern across a much larger literature. Positive relationships between modeled subjective value and BOLD activity clustered especially in vmPFC and anterior ventral striatum. The pattern appeared during choice and outcome receipt and for both primary and monetary outcomes [@bartraetal2013valuation]. This convergence establishes that domain-general value relationships recur in these regions. It does not establish that precisely the same neurons encode every outcome or that BOLD amplitude is a universal linear meter.

Pattern-based studies reveal the information that regional overlap alone cannot resolve. McNamee and colleagues found category-independent value information in medial prefrontal cortex together with category-dependent value patterns in more ventral medial orbitofrontal regions [@mcnameeetal2013category]. Howard and colleagues independently manipulated the identity and value of predicted food odors. OFC patterns carried identity-specific value information, whereas vmPFC carried a more identity-general value signal [@howardetal2015identity]. In macaques, large-scale recordings have also shown complementary emphases within frontal circuits: OFC populations represented outcome flavor especially strongly, whereas ventrolateral prefrontal populations emphasized outcome probability; preference influenced both representations [@stollrudebeck2024preferences].

These findings replace the simplest version of the currency metaphor with a stronger account. Domain-general and outcome-specific value codes coexist. A common signal allows dissimilar alternatives to exert comparable influence on choice. Specific signals preserve the sensory, bodily, temporal, and causal information required to predict consequences and revalue them selectively. The substance of the outcome does not wash out. It remains available while a more general decision variable is constructed.

The ventral striatum is part of this valuation system, not an action mechanism waiting for cortex to provide a price. Striatal populations carry learned value, prediction-error, motivational, state, and action-related information. OFC and vmPFC interact with the striatum through recurrent loops in which predicted outcomes, current state, learned policies, and action competition continually affect one another. The evidence for domain-general vmPFC signals therefore does not restore a serial architecture in which cortex values and striatum merely obeys.

37.5 The body changes the value of the expected outcome

The common-currency experiments establish comparison. Homeostatic revaluation reveals what must be preserved for comparison to remain adaptive. When the body changes, the nervous system must alter the value of the relevant outcome without erasing what that outcome is.

37.5.1 Satiety changes one outcome, not value in general

Sensory-specific satiety provides the cleanest example. Eating one food to satiety reduces the desire for that food more than for other foods. The effect is selective because the controller retains outcome identity. It does not merely lower a global reward dial.

In monkeys, orbitofrontal neurons that responded to the odor or sight of a particular food reduced their responses after the animal ate that food to satiety. Responses to other foods were relatively preserved [@critchleyrolls1996satiety]. The physical odor and image remained the same. The internal state changed their current regulatory significance, and the neuronal response changed selectively with it.

A human conditioning experiment carried the same logic into prediction. Participants learned that two arbitrary visual cues predicted two different food odors. They then consumed one associated food to satiety. When the cues were presented again without the odors, responses in OFC and amygdala decreased for the cue predicting the devalued outcome while responses to the other cue were maintained [@gottfriedetal2003devaluation]. The neural effect concerned an expected outcome, not merely a taste currently in the mouth.

Causal stimulation evidence strengthens this conclusion. Howard and colleagues used connectivity-guided theta-burst stimulation to disrupt a human OFC network before selective devaluation. Participants still showed devaluation of the food odors themselves, but they continued to choose visual cues predicting the sated outcome. The perturbation did not abolish preference in general. It disrupted the use of an outcome-specific model to infer what the cue was worth after the body had changed [@howardetal2020stimulation].

A scalar value without identity could not support this selectivity. The system must know that this cue predicts this food, that the food has just been consumed, and that its current value has declined while another outcome remains useful. Outcome identity and current value are separate enough to be recombined.

37.5.2 A need can revalue a cue before new experience

Physiological state can also increase value immediately. In the sodium-appetite experiment introduced in Chapter 35, rats had learned a cue predicting a highly concentrated salt solution that was ordinarily unpleasant. Sodium depletion later transformed the cue’s motivational force. On its first presentation in the depleted state—before the animal had tasted the salt under that new condition—ventral pallidal neurons responded strongly to the salt cue [@tindelletal2009dynamic]. The nervous system combined a stored outcome identity with a novel bodily state and recomputed what the predicted consequence now meant.

This is not ordinary reinforcement learning from a newly rewarding taste. No new salt experience was required before the cue changed. The learned relation cue predicts concentrated salt remained available, and the depleted state changed the value assigned to that predicted outcome. The experiment shows why a cached value cannot be the entire representation. Cached values are efficient summaries of earlier experience; flexible control also needs a model that preserves what consequence is expected.

37.5.3 Allostasis values correction before the error arrives

Homeostatic examples are sometimes described as though value rises only after a deficit has already developed. Predictive regulation works earlier. Food-related cues rapidly inhibit hunger-related AgRP neurons before ingestion, and the reduction can help assign preference to cues associated with relief of the aversive need state [@betleyetal2015need]. Conversely, contexts predicting future fasting can activate AgRP neurons before the energy shortfall occurs [@walkeretal2026anticipating]. The regulatory system incorporates evidence about what is about to happen.

The same principle extends beyond feeding. A route toward shade can gain value while core temperature remains within its viable range because the animal predicts continued heat load. A water source can gain value before severe dehydration because the route is long and the cost of waiting is high. Shelter can gain value before the predator appears because wind, odor, or time of day predicts risk. The controller is not assigning worth only to outcomes that remove current discomfort. It is assigning worth to outcomes expected to preserve control margins across future states.

Allostatic valuation therefore depends on two predictions: what outcome an action will produce, and what state the organism is likely to occupy when that outcome arrives. The same cup of water can have low value after drinking, high value during dehydration, and high anticipatory value before a long crossing through heat. Value remains conditional on state, but the relevant state may be predicted rather than present.

37.6 Value belongs to an outcome in a state

The state conditioning value is not only bodily. The same visible cue can predict different outcomes in different contexts, and the relevant context may be hidden. Recent history, location, a rule, another agent’s behavior, or an unobserved transition can determine which situation currently applies. A controller must infer that latent state before it can assign the correct value.

Human OFC activity contains information about such hidden task states. In the experiment by Schuck and colleagues, the currently visible event did not fully specify the task state; the correct representation depended on recent transitions and relationships among events. Multivoxel patterns in OFC represented the latent states needed to organize behavior [@schuck2016map]. This result broadens the role of orbital cortex beyond a table of prices. The region participates in representing the situation in which a price must be computed.

Rodent sensory preconditioning demonstrates why that state representation matters. Rats first experienced a neutral relation in which cue A predicted cue B. Later, cue B was paired with food. At test, cue A could support responding only if the animal inferred the chain \(A \rightarrow B \rightarrow food\); A had never been reinforced directly. Temporary OFC inactivation at test abolished responding to A while sparing responding to the directly reinforced B cue [@jonesetal2012inferred]. OFC was required when value had to be derived from an associative model, not when a previously cached value was sufficient.

Several operations that are often grouped under flexible value updating should remain distinct. In devaluation, the outcome remains the same while its current desirability changes. In reversal, the mapping between a cue and its outcome changes. In latent-state inference, the visible cue is ambiguous because hidden context determines what it means. In model-based inference, the current value must be computed from learned relations rather than retrieved as one stored number. These operations recruit overlapping frontal, striatal, amygdalar, hippocampal, thalamic, and sensory systems, but they solve different control problems.

The firm conclusion is that OFC is especially important when the organism must determine which state it occupies and use an outcome-specific model to infer what a predicted consequence is worth now. That function explains why OFC can encode economic value in one task, hidden state in another, and selective devaluation in a third. Value is not an independent substance produced by the region. It is computed from a representation of the current situation and the consequence expected within it.

37.7 Choice emerges from a distributed loop

No single region contains all of the information required for value-guided choice. Sensory and temporal cortical systems represent the properties and identity of possible outcomes. Hippocampal systems contribute context, relations, and remembered transitions. Hypothalamic and brainstem systems monitor and regulate the body; insular systems represent interoceptive and gustatory consequences; amygdalar circuits attach affective and learned significance to predictive cues. OFC is strongly positioned to represent outcome identity, latent state, and inferred current value. vmPFC often carries an integrated, domain-general decision variable. Ventral striatum and dopamine-related circuits contribute learned values, prediction errors, motivation, and policy change. Lateral and medial frontal systems incorporate rules, effort, uncertainty, and the cost of control. Basal-ganglia, thalamic, premotor, and motor circuits determine which policy gains effective access to the body.

These are regional biases within recurrent loops, not sealed modules in a production line. Hypothalamic state changes cortical and striatal valuation. OFC predictions alter striatal competition. Striatal learning changes which outcomes and actions are expected. Attention changes which attributes enter comparison. Action changes the environment and the body, producing feedback that updates both value and state. The cortex–striatum relation is therefore not a handoff in which cortex calculates a number and the basal ganglia execute its command.

A decision value biases competition; it does not choose by itself. The choice also depends on how alternatives are represented, which attributes are attended, what latent state is inferred, how uncertainty and delay are treated, which actions are available, and how rapidly the resulting competition is resolved. A larger value can favor one course without becoming a little agent that selects it.

The common-currency idea remains useful once placed inside this loop. It identifies a real computational requirement and a real class of neural signals. Unlike consequences must become comparable enough to affect one action. The comparison succeeds because domain-general decision variables are constructed alongside representations that retain what each outcome is, what state currently applies, and how the body is expected to change.

37.8 Coda: a value must remain in control

Valuation solves only the first moment of a longer problem. A selected outcome may disappear from view while several intermediate actions are completed. Its value may be temporarily weaker than the attraction of a new cue. A plan may require passing an immediately available reward in order to obtain the outcome that ranked higher when the decision was made.

The nervous system must therefore do more than compute what matters. It must keep the selected consequence effective after the offer has vanished, protect it from distraction, and revise it when new evidence genuinely changes the situation. Stability without revision produces perseveration. Revision without stability produces capture by whatever is most salient now.

That is the problem of Holding the Line. Valuation asks which expected future should govern action. The next chapter asks how that future continues to govern as the present changes.

Established findings. Subjective value is conditional on bodily and task state rather than fixed in the outcome. OFC populations in the monkey juice-choice task encode offer value, chosen value, and chosen identity, and perturbing OFC alters valuation and comparison. Human vmPFC and anterior ventral striatum repeatedly show domain-general relationships with subjective value. Domain-general signals coexist with category- and identity-specific representations. Satiety and physiological need selectively revalue predicted outcomes while preserving their identity. OFC contributes to latent-state representation and to the use of associative models when current value must be inferred. Value-guided choice is produced by distributed recurrent loops linking bodily regulation, outcome prediction, learning, comparison, and action selection.

Open questions. It remains to determine whether every class of choice uses the same domain-general format or whether several partially distinct currencies dominate in different circuits and contexts. The cellular organization of common and identity-specific value codes in humans is unresolved. Good-based and action-based representations can both influence choice, and their relative causal weight across ecological, economic, social, habitual, and explicitly planned decisions remains an active problem. The precise division of labor among OFC, vmPFC, ventrolateral and medial frontal cortex, amygdala, and striatum also changes with species, task design, and measurement scale. These questions concern the implementation and scope of valuation, not whether bodily state, predicted outcome, and learned consequence jointly govern value-guided behavior.