On September 20, 2024, Constellation announced a 20-year power purchase agreement with Microsoft to support the restart of Three Mile Island Unit 1. The 1979 accident occurred in the separate Unit 2 reactor. Not to solve the climate crisis. Not to power hospitals or cities.
To run ChatGPT.
This is not a story about energy policy. It is an obituary for an architectural paradigm.
I. The Math Doesn't Work
Let's start with the numbers.
The human brain consumes approximately 20 watts. Not 20 kilowatts. Not 20 megawatts. Twenty watts, roughly the power of a dim light bulb. On that budget, it manages 86 billion neurons, maintains consciousness, regulates every bodily system, forms memories, solves abstract problems, and navigates complex social environments in real time. Around the clock. For decades.
Training GPT-4 consumed an estimated 50 gigawatt-hours. That is more than Denmark's entire electricity consumption for a day. For a model that cannot reason causally, has no persistent memory, and hallucinates facts with confidence.
A single ChatGPT query requires roughly ten times the energy of a Google search. That sounds trivial until you multiply it by billions of daily requests.
The International Energy Agency published its report Energy and AI in April 2025. The projection: global data centers consumed 415 terawatt-hours in 2024. By 2030, the figure is expected to reach 945 terawatt-hours, more than Japan's entire current electricity consumption. Gartner estimated in November 2025 that AI-specific server installations will increase fivefold within the same period. Morgan Stanley's 2026 Energy Report estimated that hyperscalers, Microsoft, Google, Amazon, Meta, plan to invest more than one trillion dollars in AI infrastructure during 2025–2026 alone.
One trillion dollars. Nuclear power plants. And still: no systemic understanding. No persistent memory. No causal modeling.
The industry has not solved the intelligence problem. It has buried it under a mountain range of electricity.
II. Why the Brain Doesn't Need a Nuclear Plant
The question is not rhetorical. There is a precise answer, and it has been available in the scientific literature since Karl Friston formulated his Free Energy Principle in 2005.
The brain does not compute. It predicts.
That distinction changes everything.
Conventional AI, and conventional computer science in general, treats perception as an information processing problem. Stimuli in, computation, output. More stimuli requires more computation. More complexity requires more resources. The system is reactive by nature.
The brain works in the opposite direction. It continuously generates an internal model of the world, a prediction of what is about to happen. Sensory input is not primarily used to build a worldview, but to correct an already existing one. What is actually processed, what actually requires metabolic energy, is the prediction error: the difference between what the brain expected and what it actually experienced.
The consequence is dramatic: when the world behaves as expected, the brain does nothing. No action potentials. No synaptic cost. Silence.
This is not passivity. It is the most energy-efficient form of intelligence that evolution has produced in 500 million years.
Neurophysiology confirms this consistently. Studies of cortical activity show that approximately 1–4 percent of neurons are active at any given moment. The rest are silent. Not because they are not needed, but because they are not needed right now. Sparsely coded representation, activated on demand, driven by prediction error rather than raw data throughput.
Friston called it "minimizing surprise." A more precise formulation: intelligence is the art of learning not to compute.
III. What the Industry Chose to Ignore
Friston's work is not obscure. It has been cited more than 50,000 times. It is one of the most influential frameworks in theoretical neuroscience of the past two decades.
And yet none of the major AI laboratories built on it.
Instead, the industry chose a different path: transformer architecture, scalable parameter matrix multiplication, training on the entire internet. Inference still requires substantial computation, but not everything is recomputed from scratch: KV caches reuse previously computed keys and values. The work also varies with the model, context length and number of generated tokens. Reusing those calculations is not the same as learning persistently from experience.
It is as if someone designed a car where the engine always runs at full throttle, whether you are crawling in rush hour traffic or cruising on the highway. The energy inefficiency is not an implementation flaw. It is baked into the architecture.
Why did the industry choose this path anyway?
Partly because it worked. Scalability turned out to produce impressive capabilities, and capabilities sold products. Partly because predictive architectures are harder to train, they require defining what the system should predict, which presupposes a theory of what intelligence actually is. It is more comfortable to avoid the question and throw more data at the problem instead.
But the real reason is simpler: the industry built tools for language modeling and called it intelligence. Language modeling requires statistical correlation. Statistical correlation requires massive training data. Massive training data requires massive computation. Massive computation requires massive energy.
Each step was logical. The initial premise was wrong.
IV. The Architecture That Doesn't Drown
There is an alternative. It is not hypothetical.
A cognitive architecture built on prediction, not reaction, does not need to compute more than the prediction error requires. This means the system at rest consumes minimal resources. It means simple input is handled with minimal computation. It means complexity scales dynamically to actual demand, not preset to maximum throughput.
It means, in practice, that intelligence can run without a data center.
The principle is not abstract. The K-Principle is a cognitive architecture framework built on predictive logic, exactly how it is implemented is the subject of ongoing patent work, but the results are measurable and public.
Not in a laboratory. In production.
The system runs on consumer hardware. It makes autonomous decisions in real time without cloud connectivity. It has completed more than 3,600 autonomous navigation decisions without a single crash, not because it memorized millions of driving examples, but because it operates causally rather than statistically.
This is not a more efficient version of today's AI. It is a fundamentally different architecture that solves a problem that scalability can never solve: that cognitive capacity is not proportional to energy consumption.
V. What Happens If No One Listens
There is a version of the future where the AI industry continues on its current trajectory.
In that version, more nuclear plants are acquired. Electricity prices in regions with high data center density, Northern Virginia, Dublin, Singapore, rise structurally. Countries with limited grid capacity are forced to choose between AI infrastructure and industrial production. Energy poverty and AI growth compete for the same cables.
Gartner identified energy consumption in its 2025 Hype Cycle report as the single most underestimated constraint on AI adoption in the enterprise segment. Not security. Not regulatory uncertainty. Energy.
That is not a soft warning. It is a structural ceiling.
An industry that builds its competitive advantage on consuming more energy than its competitors is not an industry with sustainable growth. It is an industry heading toward a resource wall.
But there is also another version.
A version where the architecture question is taken seriously. Where "how much can we scale" is replaced by "how little do we actually need?" Where predictive, sparse, on-demand cognition replaces massive reactive parameter matrix multiplication. Where intelligence is distributed to the edge, to sensors, to vehicles, to medical devices, without every decision requiring a trans-Atlantic round trip to a data center powered by a resurrected nuclear reactor.
This is not a vision. It is an architectural consequence.
VI. The Proof
I have been working on this since 2017. No formal academic background in machine learning. No institutional funding in the early years. What I had was a simple conviction: if the brain solved the intelligence problem on 20 watts, then what we lacked was not an efficiency improvement. It was the architectural insight.
There are systems in production today applying predictive cognition without cloud, without pre-training, on consumer hardware, with 50 wins against 3 in formal benchmarks against algorithms developed from 1983 to 2024. Not in a laboratory. The results are public.
That is enough to turn the question back to the industry: if 20 watts is sufficient for biological intelligence, and if predictive architecture is the reason, why are you still building nuclear plants?
Microsoft needed Three Mile Island to power a system that does not know why it says what it says.
The brain gets by on a light bulb, because it knows what it does not need to think about.
That is not an efficiency advantage. It is an architecture advantage.
And architecture can always be copied.
If you are sitting on an energy budget, an AI strategy, or an infrastructure decision this quarter: the question you need to ask your vendor is not "how much GPU capacity is included?" It is "can your system make a decision it has never seen before without calling home to a data center?" If the answer is no, now you know why the nuclear plant was necessary.
The question is not how much energy your AI consumes. The question is why it didn't predict that it didn't need to consume it.
: The date, transaction and reactor identification have been corrected, along with the description of inference computation.