
Evidence signal before any procurement decision
If the claim on the slide is “AI Lyme test, 90%+ accuracy,” the first question is not whether the number is impressive. It is what denominator produced it: 30 blinded specimens, 63 held-out patients, 118 samples against healthy controls, or a conference abstract that has not yet become a peer-reviewed validation paper.
- Use: informational evidence appraisal, not clinical advice or a diagnostic recommendation for an individual patient.
- Reviewer credential placeholder: clinical AI / laboratory medicine reviewer credential to be inserted before publication.
- Regulatory status: FDA has cleared conventional Lyme serology indications, including 2019 modified two-tier testing indications for existing assays; those clearances should not be read as clearance for an AI/ML Lyme diagnostic model. FDA described Lyme disease as affecting more than 476,000 people in the U.S. each year in that clearance announcement. [1]
- AI/ML Lyme diagnostics reviewed here: no FDA-cleared AI/ML Lyme diagnostic was identified in the reviewed materials. Treat any AI Lyme diagnostic claim as research-stage unless the vendor supplies a specific FDA 510(k), De Novo, or PMA record for the exact product and intended use.
- Overall evidence verdict: strongest signals are in serology point-of-care and host-response classifiers; serum metabolomics is technically interesting but weakened by healthy-control validation; rash-image AI is the least dependable evidence cluster; treatment-side AI evidence supports discovery and observational hypothesis generation, not clinical efficacy.
The clinical problem is real. Early Lyme disease can present before serology behaves cleanly, and conventional testing is not a perfect gold standard. NASEM’s 2025 report describes standard two-tier testing sensitivity at initial diagnosis as roughly 30%, and persistent symptoms after Lyme disease as affecting 5% to 36% of patients six or more months after treatment. [2] That does not make every AI workaround credible. It means the validation problem is harder than a healthy-control accuracy figure makes it look.
Where the evidence is actually sitting
The most useful way to read the current literature is not “AI versus standard testing.” It is by modality, because the bias structure changes. A blood-based classifier trained to find host-response signal in seronegative early disease has a different evidence problem from a rash-image model trained on curated internet images. A point-of-care serology platform has a different regulatory path from a metabolomics signature protected by a patent. The headline accuracy number hides those differences.

| Modality | Best reported performance in reviewed materials | Validation set and comparator | Regulatory status in reviewed materials | Main evidence concern |
|---|---|---|---|---|
| Host-response gene expression, 31-gene Lyme classifier | 90% sensitivity, 100% specificity, and 95.2% accuracy on an independent 63-sample test set; positive in 85.7% of seronegative early-Lyme patients. [3] | Single 263-sample study; held-out test set, with particular interest in early Lyme where antibody tests can be negative. [3] | No FDA clearance identified for this AI classifier. | Clinically meaningful target, but no external replication at another institution and no prospective deployment evidence. |
| xVFA single-tier serology point-of-care approach | 95.5% sensitivity and 100% specificity on a blinded 30-sample Lyme Disease Biobank validation subset; 88% accuracy on a 32-sample CDC panel; 100% agreement with standard two-tier testing and 96.6% agreement with modified two-tier testing. [4] | Model trained on 40 samples; blinded validation subset contained 30 samples; CDC panel contained 32 samples. [4] | No FDA clearance identified; UCLA coverage described clinical availability as potentially years away. [5] | Striking numbers, but small training and validation sets, no sample-size calculation, code withheld under material transfer terms pending IP, patent and Biopeptides Corp affiliations. [4] |
| Serum metabolomics with SSVM classifier | 98.13% balanced success rate, 96.25% sensitivity, and 100% specificity on 118 sequestered samples. [6] | Validation against healthy controls, not febrile, rheumatologic, dermatologic, or other infection mimics. [6] | No FDA clearance identified for this classifier. | Healthy-control contrast makes discrimination easier than the intended clinical setting; authors reported patent protection. [6] |
| Erythema migrans rash imaging, Hopkins/APL work | 2019 report: 86.53% accuracy using 1,834 curated online images plus 116 clinical images from 63 patients; 2020 report: 71.58% accuracy for EM-vs-other-skin-lesion and 94.23% for EM-vs-normal, with 88.55% clinical sensitivity. [7][8] | Mix of curated online images and limited clinical-image material; image labels and population spectrum remain central concerns. [7][8] | No FDA clearance identified for a rash-image Lyme AI tool. | Weakest modality cluster: plausible adjunctive use, but vulnerable to dataset curation, spectrum bias, and unstable high-accuracy claims. |
| Retracted rash-image high-accuracy papers | One paper claimed 98.97% accuracy and was followed by a 2025 retraction notice; another article carried a 97% claim and is marked retracted. [9][10] | Not a usable evidence base for procurement or clinical claims. [9][10] | No FDA clearance identified. | Retracted literature should not be laundered into marketing claims or secondary slide decks. |
| ADLM 2025 decision-tree panel / ACES Diagnostics | Press-release and conference-abstract claim: greater than 90% sensitivity and specificity, and greater than 90% early-case detection compared with 27% for the standard method. [11] | Reported n=123 Lyme samples and 197 uninfected samples; peer-reviewed validation paper not available in reviewed materials. [11] | No FDA clearance identified; end-2026 commercialization was described as a target, not an achieved regulatory fact. [11] | Useful to track, but conference and press-release evidence is not equivalent to independent peer-reviewed validation. |
The early-Lyme validation trap
Early Lyme disease is exactly where better diagnostics would matter. It is also where validation is least forgiving. If a reference standard misses many early infections, a new classifier may look “wrong” when it is detecting true disease that antibody-based testing has not yet captured. But the reverse is also possible: a classifier can look strong because the comparison group is too clean, too healthy, or too unlike the patients who actually arrive in clinic.
That is why a governance review should avoid a lazy formulation: “AI beat the gold standard.” In early Lyme, the standard is not gold enough to settle the matter by itself. The right question is whether the study used a defensible case definition, enrolled the right mimics, separated training from validation, and then repeated the result outside the developer’s environment.
The 31-gene classifier is interesting because it addresses that failure point directly. In the Communications Medicine study, the classifier was positive in 85.7% of seronegative early-Lyme patients, and the independent test set performance was 90% sensitivity, 100% specificity, and 95.2% accuracy. The same paper notes the clinical limitation of rash-dependent recognition: erythema migrans is seen only 60% to 70% of the time. [3]
That is a better target than automating something clinicians already do comfortably. But it is still one study. The cohort size was 263 samples, the independent test set was 63 samples, and the reviewed materials do not include external replication at another institution. [3] A hospital committee should treat the result as a serious translational signal, not as a deployable diagnostic replacement.
xVFA: the most procurement-tempting numbers, and the smallest denominator
The xVFA study has the kind of result that moves quickly through vendor conversations: a simpler, single-tier, point-of-care-style platform with 95.5% sensitivity and 100% specificity on a blinded Lyme Disease Biobank validation subset. It also reported 88% accuracy on a CDC panel, 100% agreement with standard two-tier testing, and 96.6% agreement with modified two-tier testing. [4]
Those figures deserve attention because they are attached to a real clinical bottleneck: testing that is easier to run and potentially more useful early. But the study’s structure should slow any adoption discussion. The model was trained on 40 samples. The blinded validation subset contained 30 samples. The CDC panel contained 32 samples. The paper reported no sample-size calculation, and the code was not publicly released; access was constrained by material transfer terms while intellectual-property issues were pending. The paper also disclosed patents and Biopeptides Corp affiliations. [4]
None of those facts makes the assay wrong. They make the confidence interval, the reproducibility question, and the conflict-management plan central. A point-of-care Lyme test does not become clinically proven because a small blinded subset looks excellent. It becomes credible after the result survives larger, prespecified, externally replicated cohorts that include the patients who cause diagnostic uncertainty.
The translational timeline also matters. UCLA’s own coverage described the technology as potentially taking years to reach the clinic. [5] That is a useful institutional reality check when a pitch deck compresses research prototype, clinical validation, manufacturing, regulatory clearance, and procurement readiness into one phrase.
Metabolomics looks strong until the control group changes
The serum-metabolomics classifier reported a 98.13% balanced success rate, 96.25% sensitivity, and 100% specificity on 118 sequestered samples. [6] A laboratory medicine reader should notice both halves of that sentence: the performance is high, and the validation set is not large.
The deeper issue is not only sample size. The classifier was trained and tested against healthy controls. [6] That is an easier problem than the real one, where Lyme may need to be distinguished from viral syndromes, other tick-borne infections, autoimmune disease, nonspecific febrile illness, cellulitis, dermatitis, and the large category of patients whose symptoms do not line up neatly on the first visit.
A healthy-control design can be appropriate for early discovery. It is not enough for a diagnostic claim aimed at clinics. Before procurement, the relevant test is not “Lyme versus healthy.” It is “Lyme versus the conditions a clinician would plausibly confuse with Lyme at the time the decision is made.”
Rash-image AI is not useless, but it is the least dependable evidence cluster
Computer vision has an obvious appeal in Lyme disease because erythema migrans can make the diagnosis clinically straightforward when it is classic and recognized. The problem is that rashes are photographed under uneven lighting, at different disease stages, across skin tones, with variable labels, and against a long list of dermatologic mimics. A high image-classification number can dissolve quickly when the model leaves its image library.
The Hopkins/APL work should not be dismissed together with the worst of the rash-image literature. The 2019 report used 1,834 curated online images plus 116 clinical images from 63 patients and reported 86.53% accuracy. [7] The 2020 work reported 71.58% accuracy for erythema migrans versus other skin lesions, 94.23% for erythema migrans versus normal skin, and 88.55% clinical sensitivity. [8] That is at least aimed at the right visual problem.
But this remains a weaker modality for clinical adoption than blood-based early-disease classifiers. The evidence base is still limited, and the distance between curated-image performance and front-door diagnostic performance is large. It should be treated as possible triage or educational support, not as a replacement for clinical assessment and laboratory reasoning.
The retracted high-accuracy rash papers are a separate warning. One paper associated with a 98.97% accuracy claim was followed by a 2025 retraction notice, and another Lyme image-classification article carrying a 97% claim is marked retracted. [9][10] Those claims should not appear in evidence packets except as examples of why source verification matters.
Conference claims belong in a tracking file, not in the validation column
The ADLM 2025 ACES Diagnostics decision-tree panel belongs in the evidence map, but in a separate category. The press release reported greater than 90% sensitivity and specificity, greater than 90% detection of early cases compared with 27% for the standard method, and a dataset of 123 Lyme and 197 uninfected samples. It also described end-2026 commercialization as a target. [11]
That is enough to watch. It is not enough to buy. A conference abstract or press release does not provide the full protocol, case definitions, prespecified analysis plan, subgroup behavior, missing-data handling, or independent replication needed for a clinical AI governance decision.
FDA status: keep conventional test clearance separate from AI clearance
The clean regulatory statement is narrower than many marketing conversations imply. FDA has cleared conventional Lyme testing indications, including 2019 indications that allowed certain existing assays to be used in a modified two-tier testing algorithm. [1] That is not the same thing as FDA clearance of an AI model for Lyme diagnosis.
For the AI diagnostics reviewed here—the 31-gene classifier, xVFA classifier, serum-metabolomics classifier, rash-image models, and ADLM decision-tree panel—no FDA-cleared AI/ML diagnostic status was identified in the reviewed source set. A procurement file should still require a fresh FDA 510(k), De Novo, and PMA database check for the exact product name, manufacturer, intended use, and version. The absence of a clearance in an article review is not a substitute for regulatory due diligence.
This distinction is familiar across clinical AI. A strong retrospective study can be useful without being cleared, and a cleared device can still require local validation. ClinicalMind has covered similar evidence-versus-regulatory separation in PICTURE AI and in its discussion of on-device AI validation gaps. The same discipline is needed here.
Treatment evidence is much thinner than diagnosis evidence
The treatment side should not be stretched to match the diagnostic literature. There is no trial-level evidence in the reviewed materials that an AI tool improves Lyme treatment outcomes. The available work is observational, definitional, or preclinical.
The MyLymeData machine-learning study analyzed 2,162 self-reported unwell participants and 131 well participants. High response was associated with antibiotic use, longer treatment duration, and tick-focused clinician oversight. The cohort was self-selected and 85% female, and the study design was observational. [12] That can generate hypotheses about patient-reported outcomes and care patterns. It cannot establish that an AI-guided treatment strategy works.
NASEM’s 2025 infection-associated chronic illness report is useful for setting the uncertainty baseline. It reported that 5% to 36% of patients may have persistent symptoms six or more months after Lyme treatment, and described machine learning for infection-associated chronic illnesses as still in preliminary stages of validation. [2] That is a careful way to frame the field: important patient burden, unresolved mechanisms, and early analytic tools.
The Tufts AI antibiotic-discovery work is even earlier in the clinical chain. Researchers used AI-enabled screening across 60,000 compounds, identified several hundred active candidates, and had no candidates in clinical trials at the time described. [13] That is preclinical hypothesis generation, not evidence that an AI-discovered Lyme treatment improves patient outcomes.
What to ask before accepting an AI Lyme claim
A governance committee does not need to resolve every controversy in Lyme diagnostics to make a sound procurement decision. It needs to force the claim into the right evidence category.
- What exact intended use is being claimed: diagnosis, triage, adjunctive interpretation, antibiotic discovery, treatment selection, or outcome prediction?
- How many patients or samples were used for training, tuning, internal validation, and external validation?
- Was the validation cohort independent of the developers and collected at another institution?
- How were early-Lyme cases defined when standard serology was negative?
- Were controls healthy volunteers, or were they look-alike patients with febrile, dermatologic, rheumatologic, neurologic, or other tick-borne disease presentations?
- Were performance results prespecified, and are confidence intervals reported around sensitivity, specificity, and accuracy?
- Is the code, model, assay protocol, or feature set accessible for independent evaluation, or restricted by material transfer, patent, or commercial terms?
- What patents, equity, licensing arrangements, or developer affiliations exist?
- What exact FDA record supports the claim, if any? Ask for product name, manufacturer, submission number, clearance date, intended use, and version.
- Does the evidence show analytical separation, clinical diagnostic validity, clinical utility, or improved treatment outcomes? Those are not the same claim.
The procurement bottom line is straightforward. AI in Lyme disease diagnosis has promising research signals, especially where it tries to detect early disease that serology can miss. It does not yet have enough independently replicated, clinically representative evidence to support broad claims that 90%+ accuracy equals proven diagnostic performance. On treatment, the reviewed AI evidence remains discovery-stage or observational, not proof of improved patient outcomes.
References
- FDA Clears New Indications for Existing Lyme Disease Tests That May Help Streamline Diagnoses, U.S. Food and Drug Administration, July 29, 2019.
- Infection-Associated Chronic Illnesses: Toward a Consensus Definition and Better Care, National Academies Press / NCBI Bookshelf, 2025.
- A diagnostic classifier for gene expression-based identification of early Lyme disease, Communications Medicine, 2022.
- Single-tier testing with a multiplexed vertical flow assay for serodiagnosis of Lyme disease, Nature Communications, 2024.
- Lyme disease early detection could get a boost from simpler, faster testing technology, UCLA Broad Stem Cell Research Center.
- A targeted metabolomics approach to the diagnosis of Lyme disease, Scientific Reports, 2022.
- An Artificial Intelligence and Deep Learning Based Approach Can Improve the Early Detection of the Lyme Disease Rash, Johns Hopkins Lyme Disease Research Center.
- Automated Detection of Erythema Migrans in Early Lyme Disease and Other Confounding Skin Lesions via Deep Learning, Johns Hopkins Lyme Disease Research Center.
- Retraction: A novel framework for diagnosis of Lyme disease based on deep learning and image processing, Skin Research and Technology, November 2025.
- An Automated Method for Identifying Lyme Disease Using Deep Learning, Computational Intelligence and Neuroscience, 2022.
- AI Poised to Revolutionize Lyme Disease Testing and Treatment, Association for Diagnostics & Laboratory Medicine, July 2025.
- Treatment Response in Patients With Persistent Lyme Disease Symptoms: A Machine Learning Approach, Healthcare, 2020.
- Harnessing the Power of AI to Fight Lyme Disease, Tufts University School of Medicine, 2026.