Post-lung transplant recovery and AI monitoring meet at a difficult clinical boundary: the period after discharge is data-rich, physiologically unstable, and still often managed through episodic contact. A patient may be doing well by symptoms while spirometry drifts, tacrolimus exposure varies, CT changes accumulate, donor-derived or molecular signals shift, and a connected device quietly records a trend that no one has yet acted on. The clinical question is not whether an algorithm can find a pattern. It is whether that pattern can safely inform the next decision in a transplant program that already has bronchoscopy slots, drug toxicity, infection risk, rejection risk, and distance from the center competing for attention.

That distinction matters because the evidence now spans very different tools. Some models predict primary graft dysfunction in the early postoperative window. Others try to detect chronic lung allograft dysfunction, classify rejection biology, estimate tacrolimus troughs, or forecast longer-term survival. Remote patient monitoring programs add another layer: less algorithmically sophisticated in many cases, but closer to the practical work of catching deterioration before it becomes an avoidable readmission.

Translucent lungs surrounded by digital data streams, waveform lines, and icons for CT imaging, spirometry, and molecular analysis

The Surveillance Problem Is Bigger Than a Single Signal

Lung transplant follow-up is not one monitoring problem. Primary graft dysfunction, acute rejection, infection, tacrolimus toxicity or underexposure, chronic lung allograft dysfunction, and survival risk emerge on different time horizons and from different data streams. The useful AI question therefore changes by task. A PGD model may help with early risk stratification. A CLAD model may support earlier diagnostic evaluation. A tacrolimus model sits close to dose adjustment, which raises the evidentiary bar because an inaccurate estimate can quickly become a clinical action.

A 2026 narrative review of AI in lung transplantation is useful because it keeps these applications in one transplant-specific frame rather than treating “AI monitoring” as a single maturity category. The review describes machine learning models for organ allocation, post-transplant complications, CLAD detection, molecular diagnostics, immunosuppression support, and survival prediction, while also noting that much of the evidence remains retrospective or single-center rather than prospectively validated across transplant programs.[1]

Clinical taskData modalityReported performance signalMost plausible clinical use
Severe primary graft dysfunction predictionClinical variables; metabolomics; CT imagingAUC 0.82 for a clinical-variable random forest; AUC 0.90 for volatile organic compound metabolomics; AUC 0.74 for a 3D ResNet on CTEarly postoperative risk stratification and closer surveillance
CLAD or bronchiolitis obliterans syndrome detectionCT texture, eNose breath analysis, airway microbiome featuresAUC 0.85 for CT texture analysis of bronchiolitis obliterans syndrome; AUC 0.91 for eNose CLAD detectionAdjunctive evaluation when lung function trends or symptoms raise concern
Molecular rejection classificationGene expression from biopsy tissueMolecular TCMR scores associated with graft loss risk independent of histologyDiagnostic adjunct when histology and clinical suspicion do not fully align
Tacrolimus trough predictionPast doses, route, prior trough levelsPredictions within 30% of measured concentrations 88.5% of the time in 117 lung transplant patientsMedication exposure forecasting, not autonomous dosing
Long-term survival predictionDonor-recipient matching featuresRandom survival forest iAUC 0.879 versus 0.658 for Cox regressionRisk attribution and follow-up planning, pending portability testing

Early Complication Prediction: PGD Models Have a Clear Target, But Not Yet a Clear Workflow

Primary graft dysfunction is one of the cleaner targets for early post-transplant prediction because the window is narrow and the consequence of deterioration is immediate. In the 2026 review, a clinical-variable random forest model for severe PGD reached an AUC of 0.82, while a volatile organic compound metabolomics approach reached an AUC of 0.90. A CT-based 3D ResNet model reported a lower AUC of 0.74.[1] These are not interchangeable tools. They ask different operational questions: can the team use variables already present in care, does a metabolomics workflow exist at the required speed, and does CT add enough signal to justify imaging-dependent prediction?

The clinical-variable model is attractive because it resembles the data environment transplant teams already work in. The metabolomics result is stronger numerically, but performance alone does not answer whether the assay can be returned quickly enough to change respiratory support, ICU surveillance intensity, or escalation thresholds. The CT model may be useful where imaging is already obtained, but a model that depends on CT cannot be treated as a continuous surveillance layer unless imaging timing, acquisition, and interpretation are standardized enough for the model to travel.

This is where retrospective AUCs can overpromise if they are read too quickly. A PGD risk score does not need to “manage” the patient to be clinically useful. It might simply flag a subgroup for closer multidisciplinary review, earlier respiratory therapy attention, or lower tolerance for ambiguous deterioration. But the study design still needs to show which action follows the alert, how often the alert is wrong, and whether the additional vigilance improves outcomes rather than only increasing alarms.

CLAD Detection Is a Better Example of Multimodal Surveillance

Chronic lung allograft dysfunction is the kind of problem that makes single-signal monitoring feel inadequate. Spirometry trends matter, but they do not capture the whole biology. CT texture, exhaled breath signatures, airway microbiome patterns, symptoms, and pathology may each carry partial information. In the lung transplant AI review, CT texture analysis for bronchiolitis obliterans syndrome reached an AUC of 0.85, and eNose breath analysis reached an AUC of 0.91 for CLAD detection. Airway microbiome signatures were also described as having discriminative value.[1]

Stylized lungs connected to nodes representing PGD prediction, CLAD detection, molecular diagnostics, tacrolimus dosing, and remote patient monitoring

For CLAD, the appeal is not merely earlier labeling. It is earlier recognition of a trajectory that might otherwise be normalized until the next clinic visit or pulmonary function test confirms what the patient has already been living through. CT texture analysis could make radiographic change more reproducible. Breath analysis could offer a lower-burden signal if it proves stable across devices and centers. Microbiome features may help characterize risk or phenotype, but they also introduce familiar questions about sampling, antibiotics, infection, sequencing methods, and local ecology.

The unresolved issue is not whether these modalities are interesting. They are. The issue is whether a transplant clinic can use them to decide what happens next: repeat spirometry sooner, obtain CT, schedule bronchoscopy, adjust immunosuppression, investigate infection, or keep watching. CLAD tools will be judged less by their isolated discrimination statistics than by whether they clarify those branching decisions without flooding teams with ambiguous risk.

Molecular Diagnostics Add a Different Kind of AI

Not all AI in post-transplant surveillance looks like a dashboard prediction score. Molecular diagnostics bring machine learning into tissue interpretation. The Molecular Microscope Diagnostic System applies unsupervised machine learning to biopsy gene-expression data to classify T cell-mediated rejection and antibody-mediated rejection; molecular TCMR scores have been reported to predict graft loss risk independently of histology.[1]

That is clinically important because histology can be noisy, and transplant teams regularly manage cases where pathology, symptoms, imaging, donor-specific antibody status, infection workup, and lung function do not tell the same story. A molecular classifier does not remove the need for clinicopathologic judgment. It may, however, give the team another structured signal when the biopsy appearance is equivocal or when the consequences of undertreating rejection and overtreating infection risk are both serious.

The governance question is sharper here than in many prediction papers. If molecular results alter treatment intensity, then assay reproducibility, turnaround time, interpretability, and integration into pathology review are not secondary implementation details. They are part of the clinical evidence.

Tacrolimus Prediction Sits Closest to a Daily Decision

Tacrolimus modeling deserves careful language because it is easy for a prediction tool to be mistaken for a dosing system. In a 2025 study summarized in the lung transplant AI review and indexed in PubMed, an LSTM time-series model used only three inputs: past tacrolimus doses, route, and previous trough levels. Across 117 lung transplant patients, it predicted concentrations within 30% of actual measured troughs 88.5% of the time. The same evidence summary notes that more than 50% of patients are out of range early after surgery, and SHAP analysis identified dose-concentration history as the key driver.[1][2]

That is a meaningful result because early post-transplant tacrolimus management is repetitive, high-stakes, and imperfect. A model that stays close to measured troughs using limited inputs could help clinicians anticipate exposure before the next lab result, identify patients whose concentrations are behaving unexpectedly, or prioritize pharmacist review. It is also exactly the kind of tool that should not drift into autonomous recommendation without a prospective dose-intervention trial.

The difference is practical. A trough prediction can say, in effect, “this patient’s next level may be outside the expected range.” A dosing recommendation says, “change the medication this way.” The second statement carries more liability, more need for guardrails, and more dependence on local protocols, interacting medications, renal function, infection status, enteral absorption, route changes, and clinician review. The LSTM evidence supports continued evaluation of tacrolimus forecasting; it does not establish a closed-loop immunosuppression system.

Survival Models May Help Stratify Risk, If They Survive Center-to-Center Variation

Long-term survival prediction is less immediately actionable than tacrolimus forecasting, but it can still influence follow-up intensity, counseling, and research enrichment. A 2023 JAMA Network Open study cited in the 2026 review found that random survival forests using donor-recipient matching features outperformed Cox regression, with an iAUC of 0.879 versus 0.658. SHAP methods were used to support patient-level risk attribution.[1][3]

The attraction of survival modeling is individualized risk structure. The limitation is that transplant survival is shaped by center practices as well as donor and recipient features. Listing behavior, donor acceptance, induction and maintenance immunosuppression, infection prophylaxis, surveillance bronchoscopy patterns, rehabilitation resources, and thresholds for hospitalization can all influence outcomes. A survival model trained in one data environment may remain statistically impressive while becoming operationally brittle elsewhere.

Remote Monitoring Shows What Happens When Signals Reach a Care Team

Remote patient monitoring belongs in the same conversation, but not in the same evidence bucket. Bluetooth spirometers, pulse oximeters, scales, symptom surveys, and dashboards are usually digital health implementation tools first. They may include alerting logic or triage analytics, but that is different from validating an adaptive machine learning model for rejection or drug exposure. The distinction matters because RPM can be clinically useful even when the “AI” component is modest.

Institutional reports from lung transplant RPM programs are still worth attention because they describe workflow consequences rather than only model performance. Keck Medicine of USC reported a Bluetooth-based program comparing 28 monitored patients with 28 historical controls, with 44% fewer readmissions, 54% fewer hospital days, and an estimated cost reduction of about $132,000 per patient-year.[4] Mayo Clinic reported an RPM experience involving 116 patients who lived a median 234 miles from the transplant center; approximately 470 alerts occurred during the study period, and about one in four triggered a care change.[5]

Those figures are compelling because they point to the operational unit that matters: an alert reviewed by someone who can change care. But they should be read as institutional reporting unless and until the underlying primary articles are fully available for independent appraisal. Historical controls, center-specific workflows, staffing models, and patient selection can all affect readmission results. A program that works because a highly engaged transplant team watches the queue closely may not reproduce in a center that adds devices without adding review capacity.

The broader digital health literature gives some support to the direction of effect. A 2026 meta-analysis of 11 randomized controlled trials including 1,187 patients found lower readmissions with digital health interventions, with an odds ratio of 0.59, equivalent to a 41% reduction. It also reported improvement in SF-36 quality-of-life scores by a mean difference of 3.52, lower HADS anxiety and depression scores by 1.77 and 1.54 points respectively, and improved medication adherence with an odds ratio of 2.48. The exercise-adherence subgroup was much smaller, with two studies and 97 patients.[6]

For a broader view of how RPM evidence is developing across specialties, see Clinical Evidence for AI in Remote Patient Monitoring. Lung transplant recovery is a narrower use case than general cardiometabolic monitoring, but it makes the same implementation lesson harder to ignore: data transmission is not care unless a staffed, governed workflow decides what to do with the data.

What Is Closest to Deployment?

Deployment readiness is not the same as algorithmic performance. A transplant program evaluating these tools needs to ask what decision the tool informs, how fast the output arrives, whether clinicians can verify or override it, and whether the model has been tested in a population that resembles its own patients.

  • Most immediately usable: RPM-supported follow-up, when a center has clear alert thresholds, review staffing, escalation pathways, and documentation practices.
  • Promising for defined pilots: PGD and CLAD risk models that are limited to adjunctive risk stratification rather than autonomous decisions.
  • Clinically important but governance-heavy: molecular diagnostics that influence rejection classification and treatment intensity.
  • Close to daily management but not dosing-autonomous: tacrolimus trough prediction models that support review rather than prescribe dose changes.
  • Useful for planning and research: survival models with explainability features, provided they are validated across centers and time periods.

The common barrier is portability. Lung transplant centers differ in EHR structure, immunosuppression regimens, surveillance bronchoscopy schedules, laboratory timing, antimicrobial practices, device availability, and thresholds for readmission. These differences are not nuisance variables; they are part of the care system the model is trying to enter. A tool trained on one center’s behavior may partly learn that center’s protocol rather than a generalizable transplant signal.

The Regulatory and Governance Gap Is Now a Clinical Issue

Continuous-learning software as a medical device is often discussed as a regulatory abstraction, but in transplant surveillance it becomes a bedside problem. If a model updates with new data, the transplant team needs to know when its behavior changes, whether performance is monitored by subgroup, who approves updates, how false alerts and missed events are reviewed, and whether the model remains aligned with current protocols.

This is especially important for tools near treatment decisions. A CLAD risk alert may lead to diagnostic testing. A molecular rejection classifier may change immunosuppression. A tacrolimus prediction may influence how urgently a pharmacist or physician reviews a dose. These are not consumer wellness nudges. They are clinical decision-support functions in a high-consequence specialty, and they need audit trails, escalation rules, and accountability.

The field does not need reflexive pessimism. Retrospective AI work can identify signals worth testing, and some of these signals are strong enough to justify prospective pilots. But a pilot should be honest about its unit of evaluation. It is not only testing a model. It is testing whether a model, a data feed, a clinician review process, and a center-specific intervention pathway produce better care than current surveillance.

A Measured Readiness Judgment

AI-powered post-lung transplant surveillance is clinically meaningful enough to track and pilot in defined workflows. The strongest near-term opportunities are complication risk stratification, adjunctive CLAD and rejection detection, tacrolimus exposure forecasting, and RPM-supported follow-up for patients whose deterioration might otherwise surface late. The evidence is not mature enough to justify broad claims of an AI-managed transplant clinic.

Broader adoption still depends on prospective multi-center validation, portability across center-specific data and protocols, transparent governance for model updates, and clear regulatory treatment when these tools function as software as a medical device. Until then, the right posture is neither dismissal nor enthusiasm. It is disciplined deployment: choose the clinical decision first, then decide whether the model is good enough to sit beside it.

References

  1. Transforming lung transplantation with artificial intelligence: a narrative review from organ allocation to post-transplant management. PMC. 2026.
  2. Predicting tacrolimus concentrations in lung transplant recipients using machine learning. The Journal of Heart and Lung Transplantation. 2025.
  3. Machine Learning–Based Prediction of Long-term Survival After Lung Transplant. JAMA Network Open. 2023.
  4. Remote patient monitoring reduces hospital readmissions for lung transplant patients. Keck Medicine of USC.
  5. Mayo Clinic study shows remote patient monitoring can improve care for lung transplant patients. Mayo Clinic.
  6. Effects of digital health interventions in patients following lung transplantation: a systematic review and meta-analysis. Journal of Cardiothoracic Surgery. 2026.