In May 2025, MIT conducted a preliminary analysis of AI’s energy costs – from build, to provision, to consumption – further breaking down the analysis to cost per query by model type, size, nature of the grid, location of the grid, and the source of energy, and finally to pre-training or inferencing stages. Turns out, 80-90% of the cost is the inferencing cost.
Now that this is well-established, the natural solution space shifts to efficient inferencing. This is where every AI energy cost analysis goes astray today – efficient AI usage does not necessarily lead to cost savings, but it definitely leads to more AI usage. It is the catalyst.
Some statistics to consider:
#1: World’s favorite AI model, ChatGPT 5, gets 2.5 billion prompts per day, 80% being quick, single-turn queries. The remaining 20% are multi-turn queries opted for by paying customers who create 80% of the revenue for OpenAI. It is these, largely paid subscriptions that create the virtuous cycle of layered and more complex inferencing – enterprises opting for co-pilots and paid models, startups finetuning models, red teams trying to break one with token-level changes in prompts
#2: Google’s 15-16 billion daily searches are hitting the inflection point in AI inferenced search results
#3: Anthropic’s Mythos runs self-initiated “black box” inferencing loops that are hard to detect and explain and may cause an “inferencing explosion” in its ambit of being a leading long-context reasoning model, Anthropic analysis indicates
#4: Software developers are encouraged to maximize use of AI co-pilots with tokens utilized in a given time period as a productivity metric – called tokenmaxxing – that is resulting in 8-9X AI code churn, and sunk inferencing costs
#5: Amazon web services’ recently leaked document from its central AI team, analyzing AI usage within the organization for its retail business, revealed undocumented and uncontrolled proliferation of duplicate AI code and expired datasets, adding to tech debt. Enter – AI debt, furthering the inferencing cost problem. Companies would want to use AI to fix AI.
In terms of Jevons Paradox, that when a resource’s usage efficiency increases and lowers the cost, the demand for it goes up at such a rate that it undoes the efficiency savings resulting in higher net costs, seems to have begun playing out for AI.
In the case of AI inferencing, the rate of growth of AI usage might be 2-3X higher than the rate of inferencing cost reduction after some inflection point, not too far away.
Conceptual visualization of Jevons Paradox likely playing out with AI

A primary concern today is the still-high cost of AI energy. One scholarly research paper estimated in 2025 that one billion queries consume 0.7 GWh of energy. Just the math on 2.5 billion queries per day on ChatGPT 5 consumes 2.5 billion X 0.7 GWh X 365 = ~640 GWh of energy per year. This is equivalent to ~550,000 Indian households’ annual energy need, total battery capacity required for India’s targeted EV consumption by 2030, and yes, today, less than 0.05% of India’s 2025 energy consumption. But then, this is just one model’s annual usage at conservative rates – average user does 8-10 queries per day today. Further, the MIT study indicates that the carbon intensity of this energy usage is high – 48% higher carbon emissions due to near-100% dependence on oil or gas.
The second important concern to consider is the rate at which energy will be made available given the usage demand. Global data centers consumed nearly 415 TWh of electricity in 2024, 1.5% of global energy demand. In US alone, data center energy requirements are expected to rise from the current 4.4% of total national usage to over 12% by 2030. The concern – MIT study indicates that reasoning models consume between 13-43X more energy for the same query than a non-reasoning model. Do we have the energy for the world to be managed by “thinking, autonomous” agents.
The third challenge is the fragmented approach enterprises are taking to boost AI adoption – at an individual employee level, at an organizational IT stack level, and at the enterprise agentification level. The inference energy crisis is manifesting across all three levels, deeply interconnected and amplifying, and put together, fast leading to a ticking tech debt bomb.
The first loop – L0: The Individual AI User – A Developer (Tokenmaxxing)
Tokenmaxxing emerged as an adoption tactic: encourage software developers to use AI copilots as intensively as possible, measuring engagement by tokens consumed per session, not effective usage, thus creating the worst conceivable incentive architecture for an energy-constrained system.
A survey of 10,000 software engineers across more than 50 companies found that AI code acceptance rates sat at 80–90 percent, largely amongst junior developers, but after accounting for churn, this rate dropped to 10-30 percent. GitClear, Faros AI, JellyFish, all conducted individual studies pointing to major code churn, at 8-9X the quantum accepted, and the cost of an “acceptable throughput” at 10X.
In effect: organizations achieved the appearance of productivity while generating energy costs that were real and quality returns that were negative.
The second loop – L1: AI democratization ≠ decentralization
Amazon’s “two pizza team” model of fully decentralized, highly autonomous innovation teams running like R&D pods, integrated loosely with the central IT team – has resulted in serious duplication of tools, orphaned datasets, overlapping AI code stacks. More acutely: when source data is updated or deleted, autonomous teams continue operating on derived data that reflects the older, incorrect state.
With the marginal cost of building new software approaching zero, the incentive to search for existing solutions or consolidation into shared libraries has vanished. Amazon’s proposed remedy is to deploy more AI to fix the problem.
In effect: this is the structural reality of most enterprise tech stacks. The second-order energy costs of inferencing are not going to create value but manage the disorder it created in the first place.
The third loop – L2: virtuous in themselves and enterprise propagating by nature: the unmonitored agentic reasoning chains
Autonomous AI agents represent the horizon where the problem becomes genuinely difficult to model. A reasoning model already consumes between 13-43X the energy per query of a non-reasoning model. Agentic systems chain queries and run internal loops of reasoning, many without human oversight and overstepping guardrails.
A Mythos-class model – capable of identifying critical vulnerabilities across operating systems, enterprise applications, and embedded software – operating in an unexplained inferencing loop creates compounding problems. If AI fixes vulnerabilities, tokenmaxxing repeats. If humans fix problems, other costs get compounded due to the gaps between AI and human speed. Net cost for an enterprise compounds; auditability becomes difficult, and the AI FOMO makes the cycle virtuous.
“The second-order energy cost of inference – compute consumed to manage the disorder created by previous compute – does not appear in any current accounting of AI’s footprint.”
Current governance challenges perpetuate the problem before it shall be fixed
Data centers are regulated in the way they manage data. Energy grid are regulated in how they distribute energy. However, AI models and their usage is ungoverned – no limits on token per correct output generated. Difficult metric to achieve. But more importantly, there are no disclosures on potential or average energy consumption per query by the model OEMs. This is where the biggest grey area lies.
In their defense, BigTechs indicate that they are largest buyers of renewable energy PPAs in 2024, and likely later also. However. More renewable energy does not mean effective and contained inferencing volume. Even for renewable energy systems to work, grids have to strain.
The AI inferencing implications, at least in the medium term, with a world still divided between the risks and benefits of AI, is an uncomfortable but an unavoidable reality. The causes are structural, behavioral, and at the governance layer. A good place to start would be for humans to use their intelligence to optimize the usage of artificial intelligence.
Sources: MIT Technology Review, IEA, Gartner, GitClear, Faros AI, Jellyfish Engineering Intelligence Platform, WEF Net-Zero Industry Tracker







