Evidence Verdict
When a vendor, chatbot, or financial wellness platform says it can estimate retirement healthcare costs, the first question is not whether the number looks reasonable. It is whether the model has been validated against actual retiree healthcare expenditures in the population where it is being used.
On that standard, the evidence does not support relying on AI tools for individual retirement healthcare cost projections. This is not because AI has been disproven for the task. It is because the necessary validation evidence has not been reported: no peer-reviewed study identified for this appraisal validates an AI tool’s individual-level retirement healthcare cost projections against real expenditure outcomes.
| Appraisal Item | Finding |
|---|---|
| Category | Evidence appraisal of retirement healthcare cost projection tools |
| Clinical or financial decision use | Individual retirement healthcare cost planning; employer financial wellness and benefits evaluation |
| Evidence verdict | Insufficient evidence for individual cost projection reliability |
| Peer-reviewed validation against actual retiree expenditures | Not reported in the materials reviewed |
| External validation of proprietary AI outputs | Not reported |
| FDA or medical-device regulatory status | Not applicable to ordinary retirement planning tools; not a substitute for validation |
| Risk of bias | Not supportable for individual projection accuracy because validation studies are absent |
| Evidence confidence | Very low for individual projection reliability |

The practical issue is familiar to anyone who has reviewed a benefits tool after the sales meeting. A projection can look individualized because it accepts age, state, assets, diagnoses, or family history. That does not mean it has been calibrated, externally tested, or shown to help people make better healthcare-spending decisions in retirement.
What Would Count as Validation
For the question “do AI tools help with retirement healthcare cost planning,” a useful validation study would need more than a plausible estimate. It would compare individual-level predictions with actual healthcare spending over time, specify the population studied, disclose the outcome being predicted, report calibration and error, and describe how the model performs for people whose costs are not close to the average.
The outcome also matters. A tool might estimate Medicare premiums, out-of-pocket medical spending, long-term care exposure, or savings needed to reach a probability target. Those are different endpoints. A model that approximates a broad actuarial benchmark has not thereby demonstrated that it can predict one household’s future spending.
Long-term care is the failure mode that should make evaluators slow down. Many consumer-facing retirement estimates exclude it, while many households assume it is covered by Medicare. If a tool does not surface that distinction, it can preserve the user’s largest false premise while still producing a polished answer.
The Baseline Estimates Are Imperfect, but They Are Named
The conventional retirement healthcare estimates are not personal predictions. Their value is that they make assumptions visible enough to argue with. Fidelity’s 2025 estimate put retiree healthcare costs at $172,500 for a 65-year-old couple, excluding long-term care.[1] Milliman’s 2025 Retiree Health Cost Index produced a much wider range, from $128,000 to $313,000 depending on the retirement pathway, and also excluded long-term care and most dental and vision expenses.[2] EBRI’s January 2024 report, using 2023 data, projected that a couple would need $351,000 in savings for Medicare-related expenses under a Medicare-only assumption.[3]
Those figures disagree because they are measuring different planning constructs. That is not a defect by itself. It is a warning label. A benefits committee looking at an AI-generated “personalized” number should ask whether the tool is estimating the same thing as Fidelity, Milliman, or EBRI, or whether it has silently changed the endpoint.
| Source | Reported Estimate | Important Exclusions or Assumptions |
|---|---|---|
| Fidelity 2025 | $172,500 for a 65-year-old couple | Excludes long-term care |
| Milliman 2025 | $128,000 to $313,000 depending on retirement pathway | Excludes long-term care and most dental and vision expenses |
| EBRI January 2024 | $351,000 for a couple | Assumes Medicare only; based on 2023 data |
EBRI’s figure also needs a timing caveat: the report used 2023 data, and Medicare Part D’s $2,000 out-of-pocket cap took effect in 2025.[3] That does not make the estimate useless. It shows why any tool producing a current projection needs to state which benefit rules, inflation assumptions, and coverage categories it has actually modeled.
Personalized AI Claims Raise a Different Evidence Question
A generic actuarial estimate asks, in effect, “What does a defined type of household need under stated assumptions?” A personalized AI tool asks a more demanding question: “What will this household likely spend?” The second claim requires stronger evidence, not weaker evidence, because the user is more likely to treat the output as personally actionable.
That is where the evidence thins out. Proprietary model claims can describe large datasets, machine-learning methods, or individualized projections, but those claims are not the same as peer-reviewed validation. Without an external study, evaluators cannot tell whether the model is well calibrated, whether it overfits its training data, which populations are underrepresented, or how often it misses the cases that matter most financially.
This distinction is especially important in employer-sponsored financial wellness. If a tool is used only to prompt conversations, its evidentiary burden is lower. If the same output is placed in front of employees as a retirement healthcare estimate, the procurement question changes. Someone will eventually have to explain why the organization trusted that number.
Waterlily: A Dedicated Tool, but Still an Unvalidated One
Waterlily deserves separate treatment because it is not a general-purpose chatbot improvising an answer. It is presented as a dedicated tool for long-term care planning. In a 2025 Kiplinger test, Waterlily projected a $1.48 million long-term care liability for a 60-year-old reporter and $1.72 million for his wife.[4]
The scale of those numbers is useful in one sense: it forces long-term care into the planning conversation rather than treating it as a footnote. But the projection itself remains difficult to appraise. Waterlily claims more than 500 million data points and a 50,000-family training dataset, according to the Kiplinger account.[4] Those are vendor-disclosed inputs, not independent validation results.
The missing pieces are the pieces an evaluator needs most. What expenditure outcome was predicted? Over what horizon? How were long-term care utilization and intensity defined? Was the model tested outside the development dataset? How wide were the prediction intervals? Did it perform differently by income, geography, sex, race, family structure, or existing disability? The public description does not answer those questions.
None of that proves Waterlily is wrong. It means the public evidence does not let a reviewer distinguish a well-calibrated projection engine from a compelling proprietary estimate. For a procurement file, that is not a small distinction.
Chatbot Tests Show Instability, Not Purpose-Built Failure
General-purpose chatbots should not be judged as if they were actuarial retirement healthcare calculators. ChatGPT, Claude, and Perplexity were not built primarily to estimate Medicare spending, long-term care exposure, or retiree out-of-pocket trajectories. Still, the way they behave matters because employees and plan participants can reach them without procurement, training, or guardrails.
In a 2025 CBS News test, ChatGPT, Claude, and Perplexity all walked back their initial retirement healthcare conclusions when prompted to extend the planning horizon or include long-term care. Claude’s answer changed from “tight but doable” to “meaningfully underfunded.”[5] That pattern is not a validation study, but it is operationally important: the answer changed when the user introduced one of the most consequential omissions.
AARP’s 2025 test of ChatGPT as a retirement planner found inconsistent outputs and reported that long-term care costs did not surface unprompted.[6] That is the wrong direction of failure for a planning assistant. A user who already knows to ask about long-term care is not the user most at risk.
The Medicare misconception makes this more than a prompt-engineering nuisance. A Nationwide 2025 survey reported that 58% of Americans incorrectly believe Medicare covers long-term care.[7] CBS reported that the general-purpose chatbots it tested did not correct the long-term care issue unless explicitly asked.[5] A system can sound helpful while leaving the central planning error untouched.
Behavioral Risk Arrives Before Validation
The evidence problem would be less urgent if users treated AI answers as casual brainstorming. The available behavior signal points the other way. Rehmann, citing Credit Karma and Empower 2026 data, reported that 85% of AI financial-advice users took action based on the advice, while more than half reported poor outcomes.[8]
That data point does not prove AI retirement healthcare projections caused harm. It does show why a wellness-program evaluator should not assume that a disclaimer will keep an output in the realm of education. If the tool gives a number, many users will treat the number as a decision input.
System-Wide AI Cost Arguments Do Not Validate Planning Tools
Broader healthcare AI cost literature is sometimes pulled into conversations about retirement planning tools, but it answers a different question. A 2023 NBER working paper estimated that AI could generate 5% to 10% system-wide healthcare savings.[9] In July 2026, NEJM Catalyst published an argument that AI may be more likely to increase total costs under fee-for-service incentives.[10]
Neither source validates a consumer or employer-facing model that predicts one retiree’s future healthcare expenses. A system-wide cost-savings projection is not an individual prediction study. A payment-incentive argument is not a calibration analysis. Both are relevant to the broader uncertainty around AI and healthcare costs, but neither closes the evidence gap for retirement healthcare planning tools.
How to Read an AI Retirement Healthcare Estimate in Procurement
For an employer benefits team or retirement-plan fiduciary, the useful review question is not “Does the tool use AI?” It is “What claim are we allowing this tool to make?” A scenario explorer that helps employees ask better questions about Medicare, supplemental coverage, caregiving, and long-term care is different from a tool that implies it can estimate an individual household’s future costs.
- Ask the vendor to identify the exact outcome predicted: premiums, out-of-pocket medical costs, long-term care costs, total healthcare spending, or savings needed for a probability target.
- Ask whether the model has been externally validated against actual retiree expenditure outcomes, not only benchmarked against published averages.
- Ask for calibration, error, and subgroup performance, including how the model handles high-cost tails.
- Ask whether long-term care is included, excluded, or handled as a separate scenario, and whether the tool corrects the Medicare coverage misconception without waiting for the user to ask.
- Ask how outputs are presented to users: education, scenario planning, advisor discussion, or an individualized projection.
A tool can be operationally useful before it is validated for individual projection. It may organize inputs, prompt a discussion with an advisor, make long-term care visible, or help users compare assumptions. Those uses should be labeled as scenario exploration, not validated individual cost prediction.
ClinicalMind cannot support relying on current AI tools for individual retirement healthcare cost projections until peer-reviewed validation against real expenditure outcomes exists. The evidence available now supports a narrower use: AI may help users ask better planning questions, but it has not been shown to answer the individual cost question reliably.
References
- 2025 Retiree Health Care Cost Estimate, Fidelity Newsroom, 2025.
- Milliman Retiree Health Cost Index, Milliman, 2025.
- Projected Savings Medicare Beneficiaries Need for Health Expenses, EBRI, January 2024.
- I Tried a New AI Tool to Answer One of the Hardest Retirement Questions We All Face, Kiplinger, 2025.
- More people are using AI for retirement planning, but how accurate is it?, CBS News, 2025.
- Could ChatGPT Be Your Retirement Planner?, AARP, 2025.
- Medicare Long-Term Care Misconceptions, Rethinking65, 2025.
- AI Used Widely (Though Not Always Wisely) for Retirement Planning, Rehmann, 2026.
- The Potential Impact of Artificial Intelligence on Healthcare Spending, National Bureau of Economic Research, 2023.
- AI and the Cost of Health Care, NEJM Catalyst, July 2026.