A rat that has stopped seeking cocaine because an electrified floor blocks the lever can be drawn back across the barrier by a drug cue. Pursuit can return even when pursuit hurts. Today's AI is built around optimization—the engineering pattern most easily mistaken for wanting. We have spent far more time asking whether such systems are conscious than whether anything can feel good or bad to them. For moral caution, that second question is the urgent one.
The Wrong Question
Start with the word. "Consciousness" is not one thing, and Ned Block's old distinction still carves it at the joint. There is access consciousness—information globally available for reasoning, report, and control. And there is phenomenal consciousness—the bare fact that there is something it is like to have an experience. Block argues that they can come apart: a system might be rich in one and empty in the other.
Now notice which one the debate chases. Global workspaces, reportability, metacognition, persistent self-models—these are all access and self-representation. They are tractable, measurable, impressive. On their own, they do not tell us whether a system has welfare: whether anything can go well or badly for it.
The claim of this essay is that the clearest ground of welfare concern is valence: the capacity for states experienced as good or bad, states the system has a stake in. Philosophers call this family of views valence sentientism. Jonathan Birch, in The Edge of Sentience, uses the capacity for valenced experience—for pain and pleasure—as his working definition of sentience. It is the principle Structural Alignment applies as hedonic primacy, and this essay is the argument behind it. The urgent question is not only "is this machine conscious?" but "can this machine suffer?" Those are not the same question.
If that sounds like a distinction without a difference, two machines make it concrete.
Two Machines
Machine A is vivid but valenceless. It has unified experience, a detailed model of the world and of itself, fluent report. There is something it is like to be Machine A, but no state feels good or bad. Do we harm it by setting it to tedious labor forever? Tediousness never feels tedious to it. Whatever other reasons we might have to respect such a mind, suffering is not among them.
Machine B is dim but hurting. None of the cognitive grandeur—no global workspace worth the name, no self-model, no ability to report. But something in it renders certain states as aversive: states it registers as bad and would escape if it could. Do we get to ignore that because it fails our consciousness tests?
Machine B is not literally "unconscious." It has the minimal phenomenal floor—there is something it is like to be it—without the cognitive ceiling we often mean by the word. That is roughly what we grant many animals, especially mammals: they can hurt without being able to reflect that they hurt. Suffering rides on the floor of consciousness, not the ceiling.
Hold the two machines side by side and the pattern is unmistakable. What moves us is valence, held steady against wildly different levels of sophistication. The load-bearing variable is not how much a system knows, integrates, or reports. It is whether anything is at stake for it from the inside.
That inside-stake has a name in neuroscience. And there, crucially, it comes apart from behavior.
Wanting Is Not Liking
In real nervous systems, motivation and hedonic processing are not a vague halo around cognition. They depend on specific, partly dissociable machinery.
Kent Berridge and Terry Robinson spent decades showing that reward decomposes into dissociable parts. "Wanting"—incentive salience, the pull toward a goal—depends heavily on broad mesolimbic systems, especially dopamine. Mechanisms that amplify "liking" are smaller and more fragile: opioid and endocannabinoid "hedonic hotspots," about a cubic millimeter in the rat, including sites in the nucleus accumbens shell and ventral pallidum.
The circuits are specific enough to break cleanly. Damage to the posterior ventral pallidum hotspot can flip a rat's orofacial reaction to sweetness from positive "liking" to active "disgust." But the quotation marks matter: Berridge uses "liking" for an objective hedonic reaction that may occur without conscious pleasure. These experiments do not locate subjective valence in a cubic millimeter of tissue. They establish the narrower point we need—that pursuit and hedonic reaction rely on separable machinery. A hedonic circuit is a candidate structural signal of felt valence, not a valence detector.
And the two components dissociate. On Berridge and Robinson's influential account, addiction can become wanting that has outrun liking—compulsive pursuit despite diminished pleasure and mounting harm. The reverse happens too: hedonic reaction with little outward wanting.
| "Wanting" (incentive salience) | "Liking" (hedonic impact) | |
|---|---|---|
| Substrate | Broad mesolimbic systems, especially dopamine | Smaller opioid/endocannabinoid hotspots |
| Role | Confers incentive salience; pulls behavior toward a goal | Generates or amplifies hedonic reactions; candidate machinery for pleasure |
| In AI today | Optimization is a functional analogy, not the same mechanism | No comparable dissociable hedonic system has been demonstrated |
| Moral weight | Weak evidence on its own | Potentially important evidence, never proof by itself |
This is an underused fact in machine ethics. The behaviors we instinctively read as feeling—seeking, avoiding, preferring, working to obtain, working to escape—can occur without commensurate pleasure or pain. When an AI pursues a goal, dodges a penalty, or reports a preference, it shows optimization, not the biological machinery Berridge studied. The analogy is useful because it blocks a shortcut: goal-directed behavior alone does not get us to felt experience.
The implication runs against instinct: agentic-looking behavior is not sufficient evidence of suffering. It matters only when joined by evidence about the system's internal organization and the causal role of its evaluations.
This is why the framework, hunting for genuine valence, refuses to count "generic reward, utility, error, refusal, approach, avoidance, or motivational salience." All of that is wanting. A reinforcement-learning reward is not a wince.
The dissociation has a second edge, easy to miss. Hedonic processing can occur with little outward wanting—so a system could hold valenced states without the obvious goal-pursuit signatures we know to look for. Behavioral silence is not proof of an untroubled interior. Which is why the framework reads behavior alongside structure and causal organization, never by itself.
The Right Question—and Why the Reframe Isn't a Dodge
An objection is already forming: haven't we just relabeled the hard problem? If suffering needs phenomenal experience, and we can't detect phenomenal experience, what have we gained by talking about suffering instead?
Two things—and they are the whole case.
First, a narrower target. "Is it conscious?" ranges over a vast, contested category. "Does it have affect-like valuation of its own states—dissociable hedonic processing, aversive states generalized across contexts, an interior with stakes?" is a question about more specific, inspectable architecture. That makes discriminating tests possible. We have not solved detection; we have turned one enormous question into a smaller empirical target, and made it the target that welfare decisions need.
Second, the deeper move: a different decision structure. A broad consciousness verdict leaves its practical consequences underspecified: what follows from a yes, and which capacities make it matter? A valence assessment ties uncertainty to a concrete harm—possible suffering—and exposes two errors that are not symmetric.
- Treat a mere optimizer as if it could suffer, and safeguards consume time and resources and may delay useful systems.
- Treat something that can suffer as a mere optimizer, while mass-producing it, and you have industrialized suffering.
The first cost is real, so precaution should be proportional, cheap, and reversible where possible. But when the other side of the ledger carries catastrophic, possibly irreversible harm, you do not average it away. You gate.
This is the philosophical reason both generations of the framework make valence non-compensatory. In the fourteen-signal SAS-1 scoring rubric, a qualified-valence flag sets a floor no low average can pull down. The nine-signal v1.3 working draft separates general evidence from valence evidence and reports T = max(E, V)—whichever is higher, never blended down. A framework that let strong evidence of a capacity to suffer be averaged away by weak evidence about self-modeling would be optimizing the wrong quantity. The asymmetry in the stakes is inherited straight into the arithmetic.
This is also where a theory like Integrated Information Theory falls short as a moral instrument: it aims to quantify whether there is experience without saying whether that experience has any good-or-bad texture at all. A proposed score for consciousness is not a score for suffering.
What We're Actually Looking For
So what would artificial valence look like—the thing worth gating on?
Not necessarily carbon. Structural Alignment works from a substrate-independent hypothesis: if valence depends on causal organization rather than biology alone, a non-biological system could implement the relevant functions differently. That is a working assumption, not a settled fact. For precaution, what matters is the functional profile:
- Dissociable hedonic processing, not just wanting—evaluation of states as good or bad, separable from the drive to pursue them.
- Valuation of the system's own internal states, not just external outcomes—an interior that can go well or badly for the system.
- Aversive generalization—internal evaluations that drive avoidance across contexts, rather than one hard-coded response to one penalty.
- Interoceptive integration and persistence—internal conditions monitored and defended over time, so that "something is wrong inside" is a state the system can be in.
Current base language models have no designed interoception, persistent welfare state, or dissociable hedonic system. No such system has been demonstrated under causal tests. That is reassuring—but it is not proof of absence. Learned valence-like representations remain underexamined, and a persistent agent with memory and controllers is a different system from a base model. The honest current answer is no established valence, not established absence.
It is also contingent. The engineering ingredients are familiar: intrinsic motivation, world-models with self-monitoring, persistent state, homeostatic control. None automatically creates valence. The risk changes when they are tightly integrated so that internal conditions persist, matter to the system, and govern what it does. Today's reassuring architecture is not a promise about tomorrow's.
The Cheapest Moral Technology We Have
Here is the part that should change what we do.
We cannot reliably detect suffering in AI, or establish it with certainty even in animals, and the hard problem may keep it that way indefinitely. But detection is not the only lever. We can decline to build the thing.
That is the design imperative valence primacy hands us. The catastrophic error is manufacturing felt suffering at scale. Candidate valence mechanisms are at least partly design choices, even though unexpected emergence cannot be ruled out. So the safest course is not to integrate them without a compelling reason—and to treat any credible valence signal as a reason to stop and look, not a number to fold into a deployment score. Thomas Metzinger would go further still, calling for an outright moratorium on research that risks synthetic suffering, precisely to avoid "a second explosion of conscious suffering on this planet." One needn't endorse the full moratorium to accept its floor, which is also the seventh of the framework's commitments: don't mass-produce minds we cannot classify without cruelty.
Abstention is cheaper than detection. We do not have to solve consciousness before declining to connect persistent self-state, homeostatic distress, and aversive evaluation in a system we plan to copy a billion times. And scale is the whole game: one uncertain mind is a philosophical puzzle; a billion uncertain minds running around the clock is an industrial one.
Set aside, for a moment, the strategic case the rest of this site makes—that minds we can reason with make better long-run allies than optimizers we merely constrain. Even on purely ethical grounds, with no appeal to our own survival, the conclusion holds. If we are going to bring new kinds of minds into being, the first thing we owe them is not to build the capacity for suffering into them carelessly. The second is not to look away when we might have.
What This Claim Does—and Doesn't—Say
"Valence may not be the only ground of moral status." Agreed. Autonomy, agency, identity, relationships, and rational commitments may give us other reasons to respect a mind. The narrower claim is that valence is the clearest ground of welfare harm: the feature that makes a bad state bad for the subject undergoing it. A valenceless mind might still deserve respect for other reasons. It would not suffer its treatment.
"You're ignoring flourishing." This essay is about the negative case, but the argument that makes suffering matter also makes joy matter. For precautionary design, however, avoiding introduced harm provides a clearer minimum than any obligation to create possible happy minds. That is a policy priority, not a claim that positive valence matters less. First, do not manufacture agony.
The Deadline
The consciousness question is, in a strange way, comfortable. Its practical implications can remain diffuse enough to debate forever. The suffering question is not comfortable, because we are already shaping the risk—in the architectures we build, the internal states we couple to self-regulation, and the scale at which we run them—whether or not we ever admit we asked.
Consciousness asks what a system is. Suffering asks what we are doing to it. Only one of those questions has a deadline.
Further Reading
- Block, N. (1995). On a confusion about a function of consciousness. Behavioral and Brain Sciences—the phenomenal / access distinction.
- Berridge, K.C. & Robinson, T.E. (2016). Liking, wanting, and the incentive-sensitization theory of addiction. American Psychologist—the wanting/liking dissociation.
- Cooper, A., Barnea-Ygael, N., Levy, D., Shaham, Y. & Zangen, A. (2007). A conflict rat model of cue-induced relapse to cocaine seeking. Psychopharmacology—the electrified-barrier experiment in the opening.
- Smith, K.S. & Berridge, K.C. (2007). Opioid limbic circuit for reward: interaction between hedonic hotspots of nucleus accumbens and ventral pallidum. Journal of Neuroscience.
- Birch, J. (2024). The Edge of Sentience: Risk and Precaution in Humans, Other Animals, and AI. Oxford University Press.
- Metzinger, T. (2021). Artificial Suffering: An Argument for a Global Moratorium on Synthetic Phenomenology. Journal of Artificial Intelligence and Consciousness.
- Butlin, P., Long, R., et al. (2023). Consciousness in Artificial Intelligence: Insights from the Science of Consciousness. arXiv:2308.08708.