A vendor can call an AI platform “renewable-powered,” “efficient,” or “low-carbon” and still leave a health system unable to verify whether the deployed tool is environmentally sustainable. The evidence appraisal starts there: the environmental impact of AI data centers is significant and growing; the most cited global projection is credible but scenario-dependent; per-model energy, water, and carbon claims are usually third-party estimates rather than provider disclosures; and confident vendor-level sustainability claims are weak unless operational data are made auditable.
That distinction matters for health systems because AI adoption decisions are no longer judged only on clinical performance, cost, privacy, and workflow fit. A hospital considering AI-at-scale may be asked whether it is shifting environmental burden to electricity systems, water-stressed regions, local air quality, or hardware waste streams that sit far outside the procurement committee’s room.

What the strongest system-level evidence supports
The International Energy Agency’s 2026 analysis is the most useful starting point because it separates today’s measured system burden from alternative 2030 scenarios. The IEA estimates that global data centers consumed about 415 TWh of electricity in 2024, roughly 1.5% of global electricity consumption, and projects that demand will rise to about 945 TWh by 2030 in its Base Case.[1]
That is not a trivial load. It is also not a single fixed future. The IEA’s 2030 sensitivity cases range from about 700 TWh in a Headwinds scenario to more than 1,700 TWh in a Lift-Off scenario, a spread large enough to change procurement narratives, grid-planning assumptions, and emissions forecasts.[1] When a deck turns the 2030 number into a settled fact, it has already stepped beyond the evidence.
| Claim type | Evidence status | What can be concluded |
|---|---|---|
| Global data-center electricity demand is rising | Supported by IEA baseline and scenarios | The footprint is material and projected to grow, but the exact 2030 level is scenario-dependent |
| AI is becoming a larger share of data-center power | Supported at system level by IEA estimates | AI-specific demand is increasingly important for data-center planning |
| A named commercial model has a known per-query footprint | Mostly unsupported by provider disclosure | Published estimates depend on inferred hardware, utilization, cooling, and deployment assumptions |
| A vendor’s AI tool is environmentally sustainable | Usually not substantiated without operational data | The claim needs location-specific energy, emissions, water, and hardware evidence |
The IEA also estimates that data centers account for about 0.5% of global CO₂ emissions today and could rise to about 1% by 2030, making the sector one of the relatively few where emissions are projected to increase.[1] This is the point at which “electricity demand” and “emissions” must be kept separate. A data center can draw more power while reporting lower emissions if its electricity supply is decarbonizing or if it relies on contractual clean-energy instruments. Conversely, rapid load growth in a fossil-heavy grid can make an efficient facility environmentally consequential.
The IEA itself warns that historical data-center electricity estimates can be widely divergent because there are no common definitions and no mandatory reporting.[1] For an appraisal, that caveat is not a footnote. If the baseline is uncertain, then model-specific precision built on top of that baseline should be treated cautiously.
AI’s share of the load is increasing
The environmental impact of AI data centers is not just a story about more data centers in general. The IEA estimates that AI’s share of data-center power demand rose from about 9% in 2022 to about 29% in 2025, and projects it could reach about 46% by 2030.[1] That trajectory is why AI-specific governance belongs in the conversation, even when the facility is owned by a cloud provider rather than a hospital.
Still, the denominator matters. A health system buying an AI documentation tool, imaging model, coding assistant, or analytics platform usually does not know which data center will serve the workload, what share of capacity it will use, whether inference is batched, which accelerators are involved, or how cooling is handled. Without those facts, the buyer can discuss sector-level risk but cannot verify the tool’s actual footprint.
A peer-reviewed reconciliation by Ren and colleagues helps explain why public estimates diverge: assumptions about model size, hardware, utilization, workload mix, carbon intensity, and water accounting can produce order-of-magnitude differences in estimated LLM impacts.[2] The useful lesson is not that every estimate is useless. It is that an estimate without its assumptions is not evidence a procurement committee can audit.
Training is visible, but inference carries much of the operating burden
Public discussion often lingers on the energy used to train a large model. That is understandable: training runs are large, discrete, and easier to dramatize. But for deployed AI tools, the recurring burden comes from inference — the day-to-day use of the model after training.

UNU-INWEH and Jegham and colleagues report that 80% to 90% of total AI energy demand comes from inference rather than training.[3][4] For health systems, this shifts the relevant question from “How expensive was it to train the foundation model?” to “What happens when this tool is called thousands or millions of times across clinical, administrative, and patient-facing workflows?”
That distinction also changes what evidence a vendor should provide. A one-time training disclosure does not establish the environmental impact of a deployed AI service. The relevant operating evidence includes inference volume, model routing, accelerator type, utilization, latency choices, caching, data-center location, cooling method, and grid carbon intensity.
Per-query numbers are useful only when their assumptions travel with them
Per-query comparisons are tempting because they make an abstract infrastructure burden feel concrete. Jegham and colleagues estimate that a short GPT-4o query consumes 0.42 Wh, about 40% more than a Google search, and that scaling to 700 million queries per day would equal the electricity use of about 35,000 US homes.[4] The same study estimates GPT-4o annual water consumption at about 1.3 million to 1.6 million kL, described as equivalent to annual drinking water for roughly 1.2 million people.[4]
Those figures are worth reading, but not as provider-certified facts. They are third-party estimates that rely on inferred hardware configurations, utilization assumptions, workload choices, and water-accounting methods. A short prompt, a long reasoning task, a small routed model, and a large general-purpose model do not have the same footprint.
The spread can be large. Jegham and colleagues estimate that DeepSeek-R1/o3 long prompts consume more than 33 Wh, more than 70 times the 0.45 Wh estimate for GPT-4.1 nano.[4] That difference is a warning against generic “AI query” claims. The unit of analysis has to match the deployed task.

For procurement, the problem is not that academics estimate. Estimation is necessary when providers do not disclose operational metrics. The problem is when a vendor turns an estimate, or a generic efficiency statement, into a sustainability claim about a particular product deployment.
Water, air pollution, land, and e-waste are not captured by electricity claims
Electricity is the easiest environmental metric to foreground, but it is not the only burden. A 2025 Nature Sustainability study of US AI servers projected annual emissions of 24 to 44 Mt CO₂e and water consumption of 731 million to 1,125 million m³ by 2030.[5] The study also modeled mitigation potential, estimating that best practices could reduce projected carbon and water impacts by 73% and 86%, respectively.[5]
That mitigation finding matters. It means the environmental impact is not simply an unavoidable byproduct of using AI. Facility location, electricity procurement, hardware utilization, cooling choices, scheduling, and operational practices can change the outcome. But the study focuses on US AI servers and assumes current spatial distribution patterns persist, so it should not be stretched into a universal global forecast.
Water claims need the same discipline. An EESI summary reports that about two-thirds of data centers built since 2022 are in water-stressed regions, that Google reported more than 5 billion gallons of water use in 2023, and that 31% came from medium- or high-water-scarcity watersheds.[6] EESI also reports that 80% of water withdrawn for cooling evaporates permanently.[6] Withdrawal, consumption, and evaporative loss are not interchangeable terms; a vendor that does not distinguish them is asking the buyer to accept an imprecise environmental claim.
Local air pollution is narrower but important. Harvard Chan coverage of work by Cork and Dominici reported that PM2.5 from on-site gas turbines at data centers drove about 90% of modeled air-pollution health impacts, and that one Virginia facility was associated with an estimated $53 million to $99 million per year in health damages, including 3.4 to 6.5 additional premature deaths per year.[7] This should not be read as proof that every data center imposes the same health burden. It is a facility-specific estimate using specialized methods, and its relevance depends on local generation, backup power, emissions controls, nearby populations, and exposure pathways.
The less visible infrastructure burdens are also unevenly distributed. UNU reporting estimates that AI infrastructure could require more than 14,500 km² of land and generate 2.5 million tonnes of e-waste per year by 2030.[3] Lincoln Institute reporting notes that more than 90% of AI-specialized computing is concentrated in two countries, the United States and China, while more than 150 nations lack domestic AI infrastructure.[8] These facts do not assign the same burden to every AI deployment, but they make “cloud-based” a poor synonym for “impact-free.”
What would substantiate a vendor sustainability claim?
A credible claim does not need perfect certainty. It does need traceable evidence. For a health system, the appraisal question is not whether the vendor can use the word “sustainable” in a brochure. It is whether the claim survives the same basic sequence used for other AI assertions: identify the claim, name the evidence source, separate observed data from projections, expose assumptions, and state what can and cannot be concluded.
| Procurement question | Why it matters | Evidence to request |
|---|---|---|
| Which model or models serve the contracted workflow? | Different model sizes and routing choices can change inference energy substantially | Model identity, routing logic, fallback model use, and expected task mix |
| Where will inference run? | Electricity carbon intensity, water stress, and local air impacts are location-dependent | Data-center region, grid mix, cooling method, and water-source information |
| What is measured versus estimated? | Provider disclosures and third-party estimates carry different evidentiary weight | Measured energy use, measured water use where available, and assumptions for estimates |
| How will usage scale? | Inference burden rises with deployment volume and workflow integration | Expected call volume, peak load, batching, caching, and monitoring plan |
| What reductions are operational, not just contractual? | Renewable claims may not show actual load timing or local impact | Time-matched electricity data, efficiency practices, utilization targets, and cooling optimization |
The weakest form of evidence is a broad sustainability phrase attached to an undisclosed cloud deployment. A stronger form is a third-party estimate with transparent assumptions. Stronger still is provider-disclosed operational data tied to the contracted workload, location, time period, and measurement method. The current public evidence base rarely reaches that last level for commercial AI models.
That evidence gap should change the language used in governance documents. A committee can reasonably say that the vendor uses infrastructure with renewable-energy commitments, that projected sector impacts are material, or that independent estimates suggest a range of possible per-query footprints. It should be much more cautious about saying that a specific AI tool is environmentally sustainable unless the vendor discloses model-specific and location-specific operational data.
Evidence appraisal judgment
The best available evidence supports concern, not fatalism. AI data centers already represent a meaningful and growing electricity load; AI is becoming a larger share of data-center demand; inference appears to dominate operational energy use; and water, local air pollution, land, and e-waste burdens can be material depending on facility location and operating choices.[1][3][4][5][6][7][8]
For a health system evaluating an AI vendor, the defensible conclusion is narrow: mitigation appears possible, but a claim that a particular AI product is sustainable remains unproven unless it is backed by transparent, workload-specific, location-specific evidence. Global projections establish scale, and third-party benchmarks illustrate plausible ranges; neither substitutes for auditable operational disclosure.
References
- Energy demand from AI. International Energy Agency.
- Reconciling the contrasting narratives on the environmental impact of large language models. Scientific Reports. 2024.
- AI’s environmental footprint. UN News. June 2026.
- The Environmental Impact of Large Language Models: Energy, Water, and Carbon Emissions. arXiv. 2025.
- AI servers could annually emit 24–44 Mt CO₂e and consume 731–1,125 million m³ water by 2030. Nature Sustainability. 2025.
- Data Centers and Water Consumption. Environmental and Energy Study Institute. 2025.
- Analyzing air pollution, health, economic risks from AI data centers. Harvard T.H. Chan School of Public Health. April 2026.
- The Land and Water Impacts of Data Centers. Lincoln Institute of Land Policy.