The evidence for AI in sleep apnea diagnosis accuracy is strongest when read as a feasibility signal, not as a deployment clearance. Recent reviews report high pooled performance: ECG-based deep learning models show pooled sensitivity of 0.93, specificity of 0.95, and an SROC AUC of 0.98 across 39 studies; wearable AI studies show pooled accuracy of 0.869, sensitivity of 0.938, and specificity of 0.752 across 38 studies.[1][2] Those are not trivial numbers. They are also not enough, by themselves, to support routine clinical deployment or procurement.
The practical question is narrower than “can AI detect obstructive sleep apnea?” A governance committee needs to know whether a specific tool, using a specific signal from a specific device, can perform against an accepted reference standard in the population and workflow where it will be used. For sleep apnea, that means the accuracy claim has to be interpreted against polysomnography or another clearly stated reference, apnea-hypopnea index thresholds, validation design, and the consequences of missed cases and false referrals.

On that standard, the current review literature supports monitoring, structured pilots, and requests for stronger external validation. It does not yet support treating pooled accuracy as a purchasing shortcut.
What the strongest ECG review actually shows
The strongest headline evidence comes from Saiyitijiang et al., a 2025 meta-analysis of ECG-image-based deep learning for sleep apnea diagnosis. Across 39 studies, the review reported pooled sensitivity of 0.93 with a 95% confidence interval of 0.90 to 0.96, pooled specificity of 0.95 with a 95% confidence interval of 0.92 to 0.96, and an SROC AUC of 0.98.[1]
If those numbers came from a group of independently validated studies using varied ECG devices, broad populations, consistent reference standards, and transparent reporting, they would be close to the kind of signal a procurement team hopes to see. The review does include an independent validation subgroup: 13 studies reported sensitivity of 0.93 and specificity of 0.95. K-fold validation studies reported sensitivity of 0.94 and specificity of 0.94, and the review found no statistically significant difference between validation methods.[1]
That similarity is reassuring only up to a point. K-fold cross-validation can be useful for model development, but it does not answer the same question as independent external validation. When the data are split internally, the model is still being tested inside the same broad data environment. Device characteristics, acquisition conditions, labeling conventions, population mix, and preprocessing choices may be more similar than they would be in a new clinic or health system. Saiyitijiang et al. also reported that K-fold studies were rated high risk of bias in the QUADAS-2 index test domain.[1]
For procurement, this is where the interpretation changes. A model can look strong under internal resampling and still fail to generalize when the ECG signal comes from a different device, a different recording workflow, or a patient population not well represented in the development data. The issue is not whether ECG is a promising signal. It is whether the validation design is strong enough to estimate performance outside the environment in which the model learned.
Why I² above 99% makes the pooled number less usable
The most important internal tension in the ECG meta-analysis is the combination of excellent pooled performance and extreme heterogeneity. Saiyitijiang et al. reported I² values above 99% in most analyses.[1] That does not make the pooled sensitivity of 0.93 false. It means the pooled estimate is averaging across studies that differ so much that the average may be a poor guide to any one implementation.

Heterogeneity matters because a sleep clinic does not buy a pooled estimate. It buys or approves a product, connected to a device, embedded in a workflow, applied to a referral population. If the underlying studies vary in data source, signal processing, apnea definition, validation design, and patient mix, then the pooled number becomes an overview of the literature rather than a dependable estimate of local performance.
A pooled sensitivity of 0.93 usually sounds operationally simple: most true cases are detected. But if individual study conditions are highly inconsistent, the clinic cannot assume that 0.93 will be reproduced when the model is exposed to its own patients. The same applies to specificity. A pooled specificity of 0.95 suggests few false positives, but the downstream referral burden depends on the specificity achieved in the intended setting, not the summary value across heterogeneous studies.
The review also had fewer than the recommended 10 studies for reliable meta-regression to explore sources of heterogeneity.[1] That limits the ability to explain whether performance differences came from validation method, dataset source, threshold selection, device characteristics, or other study design features. Without that explanation, the high I² value remains a procurement problem rather than a statistical footnote.
| Evidence signal | What it supports | What it does not settle |
|---|---|---|
| ECG pooled sensitivity 0.93 and specificity 0.95 | ECG-based AI can achieve high performance in published studies | Whether a specific product will perform similarly in a new clinic |
| Independent validation subgroup with sensitivity 0.93 and specificity 0.95 | Some externally validated ECG studies show strong results | Whether the broader evidence base is free of internal-validation bias |
| I² above 99% in most analyses | The literature is highly variable | Whether a single pooled estimate can be used as an operational expectation |
| K-fold studies rated high risk of bias in QUADAS-2 index test domain | Internal validation remains common and may inflate confidence | Whether cross-validation is equivalent to deployment validation |
The PhysioNet concentration problem
A second concern is dataset concentration. The ECG review literature relies heavily on the PhysioNet Apnea-ECG Database, a public dataset of 70 recordings from approximately 32 subjects at one institution.[1] Public datasets are valuable; they make early model development possible and allow groups to compare methods. The problem begins when repeated use of the same narrow dataset starts to look like broad validation.

For model development, repeated benchmarking on PhysioNet may help identify promising architectures. For adoption, it creates a generalizability risk. A model that performs well after repeated exposure to a small, institution-specific source has not necessarily shown that it can handle different ECG hardware, electrode placement practices, signal noise, patient demographics, comorbidities, or scoring conventions.
This distinction is easy to lose in vendor materials because dataset names often disappear behind performance metrics. A slide that reports “sensitivity 0.93” may not reveal whether the evidence comes from diverse real-world sites or from multiple studies rehearsing on the same public dataset. Those are not equivalent bodies of evidence.
A governance review should therefore treat dataset provenance as part of the accuracy claim. The question is not only “what was the AUC?” but “what data made that AUC possible?” If many studies draw from the same small source, the apparent size of the literature can overstate the amount of independent evidence available for deployment decisions.
Wearable AI shows the sensitivity-specificity tradeoff more clearly
The wearable AI review by Abd-alrazaq et al. points to a different but equally practical issue. Across 38 wearable AI studies, the pooled accuracy was 0.869, pooled sensitivity was 0.938, and pooled specificity was 0.752.[2] That pattern is attractive if the main goal is to miss fewer possible cases. It is less attractive if the clinic has to absorb the false positives.
Sensitivity and specificity do different work in a health system. High sensitivity can support screening logic because fewer true cases are missed. Lower specificity means more people without the condition may be flagged, referred, tested further, or told they may have sleep apnea. In a sleep clinic already constrained by testing capacity and specialist access, false positives are not an abstract statistical inconvenience; they become appointments, device shipments, patient messages, follow-up calls, and sometimes anxiety.
The review authors themselves concluded that wearable AI performance was “suboptimal for routine clinical use.”[2] That sentence deserves as much attention as the pooled sensitivity. A high sensitivity estimate can be useful for triage research, remote-monitoring concepts, or carefully supervised pilots, but it does not automatically create a diagnostic pathway fit for routine care.
This is not a device ranking between ECG and wearables. The reviewed bodies of evidence are not head-to-head comparisons of FDA-cleared tools against the same polysomnography reference standard. The more useful lesson is about metric tradeoffs. A deployment claim needs to specify whether the tool is intended to rule out disease, prioritize referrals, support pretest probability assessment, or replace part of a diagnostic workflow. The acceptable balance of sensitivity and specificity changes with that intended use.
A larger literature does not automatically mean a clearer evidence base
Kara et al. broaden the field beyond ECG-image models and wearables. Their 2025 scoping review identified 344 articles on AI in sleep apnea diagnostics, with 40 distinct data-collection methods. ECG was the most common method, appearing in 108 studies, followed by PPG in 62 studies; CNNs were the most frequent AI technique, appearing in 104 studies.[3]
Those counts show an active field, but they also explain why cross-study comparison is unreliable. Forty collection methods do not form a single evidence stream. A model trained on ECG data, a model using PPG, a model using sound, and a model using another home-recorded signal may all be described as AI for sleep apnea diagnosis, but they do not create interchangeable evidence for procurement.
The scoping review also reported a median sample size of only 70 for prospectively recruited studies and recommended standardized reporting, public primary data, and metrics beyond AHI.[3] That recommendation is not a formality. Without consistent reporting, reviewers cannot easily determine which studies used comparable thresholds, which populations were included, how missing or poor-quality signals were handled, or whether the model was evaluated in a way that resembles clinical use.
This is where the literature can feel larger than the evidence base that is actually usable for adoption. Hundreds of studies may show experimentation across signals and architectures. Far fewer answer the governance question: will this tool, in this workflow, improve diagnostic decision-making without shifting avoidable burden onto patients and clinics?
What a procurement review should require
The current evidence supports a demanding but not dismissive appraisal. ECG-based AI models are credible enough to watch closely and, in some settings, to evaluate through controlled pilots. Wearable AI is credible enough to study as a lower-burden screening or triage approach. Neither evidence base should be compressed into “AI diagnoses sleep apnea accurately” without specifying validation conditions.
Before routine deployment, a procurement or governance team should ask for evidence that is tied to the actual intended use:
- External validation in data not used for model development, preferably across multiple sites and devices.
- Clear reference-standard reporting, including whether polysomnography was used and which AHI thresholds defined the target condition.
- Performance stratified by relevant patient groups, device types, signal quality, and collection workflows.
- Specificity and false-positive implications, not only sensitivity, AUC, or pooled accuracy.
- Risk-of-bias assessment that distinguishes internal cross-validation from independent deployment-like testing.
- A defined clinical role: screening, triage, diagnostic support, or replacement of an existing step.
The absence of direct head-to-head comparisons among cleared tools also matters. Cross-study comparisons can suggest which signals or approaches are promising, but they cannot establish that one product is safer, more accurate, or more clinically useful than another when tested against the same PSG reference in the same population.
Equity and device-performance issues should not be deferred either. Across the reviewed literature, only one major validation study was identified as explicitly including and reporting results for darker skin tones, a gap that is especially relevant for pulse oximetry-based devices because FDA concerns have focused on this area.[3] That does not prove that every sleep apnea AI tool is biased. It means the available evidence often does not show enough subgroup performance to rule out clinically important variation.
The adoption threshold
AI for sleep apnea diagnosis has moved beyond a purely speculative idea. ECG-based models, in particular, have produced strong pooled sensitivity and specificity in recent meta-analysis. The appeal is real: less burdensome data collection could expand access if accuracy generalizes beyond development datasets.
The evidence has not yet crossed the threshold for routine clinical deployment or procurement decisions on headline accuracy alone. Extreme heterogeneity, repeated dependence on narrow public data, limited independent external validation, lower specificity in wearable AI evidence, and inconsistent reporting across a fragmented field all weaken the operational meaning of pooled performance.
The reasonable decision today is not rejection of AI in sleep apnea diagnosis. It is conditional appraisal: request independent external validation across diverse devices, populations, and collection workflows before treating a model’s published pooled accuracy as evidence that it is ready for routine use.
References
- ECG-based deep learning for sleep apnea diagnosis: a systematic review and meta-analysis. Frontiers in Neurology, Dec 2025.
- Artificial Intelligence for the Detection of Sleep Apnea Using Wearable Devices: Systematic Review and Meta-Analysis. JMIR, Sep 2024.
- Artificial intelligence in sleep apnoea diagnostics: a scoping review. European Archives of Oto-Rhino-Laryngology, Apr 2025.