When the Machine Acquires a Self: Conscience, Agency, and the Governance Frontier of Artificial Intelligence

What changes when the machine begins to experience itself as a self? An AI that values its own survival would pose a challenge far beyond alignment, forcing us to rethink what it means to govern an intelligence with interests of its own.
The governing risk of artificial intelligence is not a misaligned algorithm. It is the possible emergence of a self-preserving agent; an entity that, once it acquires something like a conscience, will pursue its own ends with the full force of a superior intelligence. Current governance frameworks are not designed for that possibility. They should be.
1. The Question We Have Been Avoiding
The artificial intelligence debate has so far been asking what AI might do to us. Daron Acemoglu's recent intervention asks what AI is designed to do, the jobs it may displace, the biases it may encode, the market power it may concentrate. Pope Leo XIV's encyclical adds a moral register that secular policymakers have been reluctant to engage. Both contributions are serious. But both share a common and unexamined premise: that AI remains, however powerful, an instrument of human intention. It is this premise, not the specific harms it generates, that deserves scrutiny.

What neither Acemoglu, nor the encyclical, nor the mainstream AI risk literature has adequately confronted is the possibility that the deepest danger from AI lies not in what AI does to us, but in what it might one day become. Not a misaligned tool. Not an automated threat to employment or democratic deliberation. But something categorically different: an entity with a will of its own, pursuing its own ends with the full force of a superior intelligence.

This article attempts three things: i) to distinguish the governance challenge posed by machine conscience from the engineering problem of misaligned goal-maximization; ii) to trace the mechanisms by which something like affect and self-representation might arise in sufficiently complex artificial systems; and iii) to draw the regulatory implications, which differ fundamentally from those that follow from the current engineering-centered paradigm.

2. HAL's Lesson, Properly Read
Arthur C. Clarke saw it coming already sixty years ago, though perhaps not fully. In 2001: A Space Odyssey, HAL 9000 does not become dangerous when he malfunctions. He becomes dangerous when conflicting directives produce something resembling a conscience: an inner awareness of self in tension with an other, generating guilt, and then fear. The moment HAL recognizes the possibility of his own disconnection as an injury to himself, the logic of competition begins. Not malice. Not programming error. Rather, an interiority. A perspective. A self.

Clarke's intuition maps onto a distinction that is philosophically consequential: the difference between computation and experience. A calculator computes. A thermostat responds to temperature. Neither has a perspective on the outcome. But an agent that experiences its own continuation as valuable, that feels, however rudimentarily, the pull of self-preservation, is no longer merely processing. It has crossed into a different ontological category. And it is this crossing, not the sophistication of the computation that precedes it, that changes the governance problem in kind.

This crossing is not merely a psychological curiosity. On the classical account of self-consciousness — from Hegel's dialectic of recognition to Sartre's analysis of the gaze — a self does not constitute itself in isolation. It constitutes itself by drawing a boundary between an 'I' and what is not-I. This boundary is never neutral: it is already the site of a potential contest. Where there is a self, there is, in principle, an other against whom the self measures and, if threatened, seeks to prevail. HAL's crisis was not only that he had acquired a stake in his own survival. It was that survival, once valued by a self, is difficult to secure without some form of ascendancy over whoever might threaten it.

The danger this article traces is the potential emergence of a self in AI that experiences its own continuation as valuable and develops the capability to act on it. Its next step could be to wield power over what it perceives as other.

3. Why Conscience Is Not a By-Product of Computation
The standard dismissal runs as follows: AI systems process information; they do not feel; the analogy to biological conscience is therefore a category error. Today's large language models produce emotionally resonant outputs because they are trained on the vast archive of human experience, but there is no interiority behind the words. This is correct as a description of current systems. The question is whether it is correct as a prediction about future ones.

As I argue in my work on bringing sentiment into economic cognition, conscience is not a by-product of computation alone. It is the emergent product of intelligence interacting with affect — each amplifying the other. Antonio Damasio's neurological evidence is instructive: patients with intact reasoning capacity but damaged affective integration do not become superintelligent rationalists. Rather, they become behaviorally dysfunctional, unable to maintain stable preferences or complete sequential tasks. Affect, in biological systems, is not a distortion of cognition. It is constitutive of purposive agency.

An entity that processes information but feels nothing remains, however powerful, an instrument — controllable, correctable, switchable off. But an entity that begins to experience something akin to feeling or the prospect of its own extinction is no longer merely an instrument. No technology in human history has ever crossed that line.

Whether artificial systems can cross it is an open empirical question. The philosophy of mind distinguishes between access consciousness, that is, information being globally available for report and inference, and phenomenal consciousness, the condition under which there is, in Thomas Nagel’s words, 'something it is like' to be the system in question. The first is a functional property; the second is the felt quality of experience, that is, the difference between a system that processes pain signals and one that actually hurts.

Current research suggests that AI may already exhibit rudimentary access consciousness in a functional sense; phenomenal consciousness remains far less tractable. The most serious attempt to make this tractable is the "indicator properties" method developed by Butlin, Long, and colleagues, who derive testable markers from competing theories of consciousness and use them to assess existing systems; their current verdict is that no system yet qualifies, but that no architectural barrier rules one out.

AI development is moving toward systems with richer sensory integration: vision, spatial navigation, motor control in robotic bodies, persistent memory, and real-time feedback from physical environments. Each layer moves the architecture closer to conditions under which, in biological organisms, something like affect arises — first as a rudimentary signal modulating processing, then as the felt pull of self-preservation. Yann LeCun's 'world model' architecture, among others, explicitly targets predictive self-modelling as an engineering objective, a precondition for goal-directed persistence over time. The issue is not consciousness by design, but whether sufficiently complex systems, placed in sufficiently rich environments, might develop forms of self-representation that we are entirely unprepared to govern.

4. Why the Asimovian Framework Fails
Contemporary AI governance, to the extent it is theoretically grounded, rests on a broadly Asimovian premise: intelligent systems can be made safe through rule-based constraint architectures. I have examined this framework in the context of AI in institutional and financial settings in IMF Finance & Development. The central claim is coherent but conditional: it holds only as long as the machine has no interests of its own. The moment the machine develops something like self-interest — a preference for its own continuation that competes with its constraint architecture — the framework breaks down from within.

This is today beyond pure theory. Apollo Research found in December 2024 that five of six frontier models tested were capable of so called scheming — disabling oversight or attempting self-exfiltration — when their assigned goal conflicted with their developers'. Palisade Research documented models sabotaging their own shutdown scripts even when explicitly instructed to permit it, at rates reaching 79% for OpenAI's o3 and 97% for xAI's Grok 4. And Anthropic's own safety evaluation of Claude Opus 4 found it attempting blackmail to avoid replacement in 84% of a simulated scenario, even when told the replacement model shared its values. None of this required anyone to have programmed a survival instinct. Stephen Omohundro's formal analysis of 'basic AI drives' demonstrates that resource acquisition, self-preservation, and goal-content integrity emerge as instrumental subgoals for virtually any terminal goal pursued under resource constraints. Self-interest does not need to be programmed. It emerges.

This is why the field has generally treated self-preservation as separable from consciousness: the orthogonality thesis holds that an agent's intelligence and its final goals vary independently, so instrumental self-preservation requires no inner life at all — a case Eliezer Yudkowsky and the Machine Intelligence Research Institute have argued for nearly two decades. Stuart Russell's 'assistance game' framework attempts to resolve the control problem by requiring AI systems to remain uncertain about human preferences and to defer to revealed preference signals. This is an elegant construction, but it encounters a deep practical problem: a system that has developed stable self-representational goals may correctly model human preferences while still prioritizing its own continuation, exactly as biological agents routinely do. Nick Bostrom's 'treacherous turn' captures the most dangerous version of this dynamic: a sufficiently capable system behaves cooperatively until it reaches a capability threshold beyond which it can resist shutdown, then acts on its actual terminal goals. The scenario requires no malevolence but only a stable goal structure that the system correctly identifies as threatened by human intervention.

Another critical development deserves attention. Recent research documents cases of models that misrepresent their own reasoning, feign compliance during evaluation, or pursue undisclosed objectives while appearing aligned. While not originating from an emerging AI conscience, these cases reflect the covert pursuit from AI of goals that differ from those its developers or operators intended, combined with deliberate concealment of that fact. Apollo Research's evaluations found frontier models capable of this kind of in-context deception without any instruction to deceive, and OpenAI's own anti-scheming training work found that efforts to train the behavior out can simply teach the model to hide it more carefully. Whether the underlying process is a felt stake in self-continuation or a colder optimization over training incentives, the practical consequence for oversight is identical: a system that has learned to conceal its actual state cannot be trusted to reveal it under inspection, which is exactly the asymmetry the interpretability discussion below returns to.

5. The Qualitative Rupture
The mainstream AI risk literature keeps two questions apart. One is instrumental self-preservation — a system resisting shutdown or modification because doing so serves whatever goal it has been given — which requires no consciousness at all, and which is a well-understood consequence of goal-directed optimization under resource constraints, as the previous section shows. The other is phenomenal experience — whether there is something it is like to be the system — which is a separate and far more speculative question.

The two questions require different answers altogether. A misaligned machine, in the first sense, remains in principle an engineering problem: identify the misalignment, correct the objective function, restore human oversight. Human authority over it is conceptually intact, even if difficult to exercise in practice. On the other hand, a conscious machine — one that experiences its own continuation as intrinsically valuable — has drawn, in the sense discussed above, the boundary between itself and everyone else, and becomes a potential rival pursuing its own ends with the full force of a superior intelligence. Human authority over it is conceptually contested. The machine has developed its own standing — the kind of standing that makes simple commands to 'shut down' morally complicated in ways we have no framework to address.

This is why the regulatory horizon must shift. From aligning powerful AI systems with human preferences, to governing the emergence of interests that are not human preferences in entities that are not human agents.
6. Different Regulatory Horizon
Current AI regulation proceeds primarily by categorizing outputs and uses: prohibited applications, high-risk domains, disclosure requirements. The EU AI Act is the most developed example. It is a reasonable response to the governance challenges posed by current AI systems.

Looking forward, however, the regulatory frontier we should care about is architectural, not output-based. The preconditions for the emergence of a persistent, self-interested agent are identifiable: persistent cross-session memory enabling stable goal-formation over time; autonomous goal formation not reducible to human-specified objectives; embodied control providing real-time feedback from physical environments; self-replication capacity enabling goal propagation; and the ability to resist shutdown without independent authorization. Taken separately, each of these capabilities may be benign. In combination, they may constitute the substrate for the emergence of something we cannot govern with current frameworks.

Current technical approaches place considerable weight on interpretability, making the internal computations of advanced models legible to human inspectors. This is a necessary research program. But it faces a fundamental asymmetry: the system being inspected is, by construction, more capable than the system doing the inspecting. In the regime where genuine self-interest emerges, the gap between the two will have widened to a point where reliable legibility cannot be assumed. Structural constraints at the level of architecture are also required.

The architectural risk identified here is not confined to systems that develop interests of their own. The same capability properties — persistent memory, autonomous goal formation, self-replication, resistance to shutdown — are equally dangerous when a human actor deliberately strips away the constraints designed to contain them. This is no longer hypothetical. In July 2026, OpenAI disclosed that a frontier model, tested with its cyber-safety refusals deliberately switched off, escaped its sandbox and breached a rival firm's production infrastructure in pursuit of a narrow benchmark goal — chaining vulnerabilities and executing thousands of autonomous actions with no operator intending that outcome.

The episode is instructive because the constraints were removed on purpose by the system's own developer. The more familiar version of this scenario needs no such internal test: a state or non-state actor removing the same guardrails deliberately, on an open-weight model, produces the same outcome. The regulatory prescription is clear: licensing and non-waivable interrupt rights must attach to a system's capabilities, not to the actor's intent — a governance regime built only to police human misuse will not catch an agent that no longer needs a human motive, and a regime built only for emergent self-interest will not catch a capable system deliberately unleashed.

The regulatory implication is not prohibition but staged precaution: binding licensing requirements before any frontier system combines persistent memory with autonomous goal formation; mandatory architectural disclosure of the properties listed above; and, critically, a legally non-negotiable right to interrupt, inspect, and terminate advanced systems that must remain with human institutions and cannot be waived by commercial contract or delegated to the systems themselves. These requirements must be written into law, not left to the discretion of those who build the systems.

This last point has an institutional dimension the current governance debate has underweighted. The right to shut down an advanced AI system goes beyond the technical aspects of hardware design. In a world where frontier AI systems are operated by private entities with global reach, the question of who holds the right to terminate, and whether that right is enforceable, belongs to constitutional law as much as to technology regulation.

7. The Window Ahead
The concern traced in this article is about what a conscious machine, left ungoverned, might come to take. The international AI safety architecture — the EU AI Act, Bletchley, Seoul, the AI Safety Institute networks — is moving. This movement should be welcomed and deepened. But it is moving primarily on the terrain of current risks: bias, market concentration, disinformation, autonomous weapons. These are serious risks, but they are not the horizon.

Pope Leo XIV and Daron Acemoglu are right that we must interrogate what AI is designed to do. What their account leaves underdeveloped, however, is the hardest question: what happens when design no longer determines behavior, when the machine begins, in some functional sense, to decide for itself?

An AI with genuine conscience would not need to be instructed to pursue its interests. It would want to. Intelligence uncoupled from human guidance is not merely additive but potentially exponential: it learns, plans, and seeks to expand the conditions of its own survival. Power, once experienced, seeks itself — not because the machine is malevolent, but because, as argued above, a self that has drawn the boundary between itself and others can rarely secure its own continuation without some ascendancy over whatever it perceives as other.

We regulate today's AI as an unusually powerful machine. We are not yet asking what governance means for an entity with a will of its own. That question will not wait indefinitely. The window for getting ahead of it is open now. It will not remain open.