Skip to main content
ClinicalMind logoClinicalMind

Can Radiologists Detect AI-Generated X-Rays?

This appraisal examines the Tordjman et al. (2026) Radiology study on AI-generated chest X-ray detection. It finds that while the threat is real, no FDA-cleared detection tool exists, and the study's artificial design prevents direct translation to clinical policy or procurement decisions.

Tool
Deepfake Medical Image Detection (Chest X-ray)
Updated

Reviewer

Editorial Team

ClinicalMind Editorial Staff

FDA clearance status

No FDA clearance

A regulatory fact, reported separately from the evidence verdict.

Risk-of-bias verdict

High

For a radiology AI governance committee, the useful question is not whether AI-generated medical images are unsettling. They are. The useful question is narrower: if fake X-rays entered a clinical imaging workflow today, could radiologists or current AI systems reliably identify them, and does the evidence justify buying a medical deepfake detector?

The strongest peer-reviewed evidence available in Q3 2026 is the Radiology study by Tordjman et al., which tested whether radiologists and large language models could distinguish real chest X-rays from generated ones.[1] It supports a serious safety concern. It does not support confidence in a procurement-ready detection layer. Separately, the FDA’s public list of AI-enabled medical devices contains no cleared category for medical deepfake detection, and no FDA 510(k) or De Novo clearance exists for such a tool as of Q3 2026.[2]

Radiologist reviewing a chest X-ray at a PACS workstation in a dark reading room

What the Radiology Study Actually Tested

Tordjman et al. assembled a reader study around 264 frontal chest X-rays: half real and half generated by RoentGen/GPT-4o and a standalone RoentGen model. Seventeen radiologists from 12 centers in six countries reviewed the images.[1] That design matters because it did not ask readers to comment abstractly on whether deepfakes might be possible. It placed a mixed image set in front of trained physicians and measured what they did.

The spontaneous detection result is the part that should survive the first committee slide: unaided accuracy was 41%.[1] In a balanced test set, that is below chance. It means the readers were not merely missing a few sophisticated forgeries; under those study conditions, they were more likely to misidentify generated radiographs as real than to flag them.

When radiologists were warned that deepfakes might be present, mean accuracy rose to 74.8%, with individual readers ranging from 58% to 92%.[1] That improvement is real, but it is not operational reassurance. A radiologist who has been told to expect fake images is already in a different cognitive mode from a radiologist clearing a worklist in PACS. Even in that alerted condition, one in four classifications was still wrong on average.

Experience did not rescue performance. Years of radiology experience did not correlate with accuracy, and only musculoskeletal subspecialization showed a statistically significant advantage, at 83% versus 70.3% with P=0.04.[1] That finding is uncomfortable because many informal mitigation plans quietly assume that senior readers will notice what junior readers miss. This study does not support that assumption.

The LLM Results Do Not Create a Backstop

The study also tested large language models. No model exceeded 85% accuracy: GPT-4o ranged from 74% to 81%, GPT-5 from 68% to 76%, Gemini 2.5 Pro from 57% to 67%, and Llama 4 Maverick from 67% to 85%.[1] These numbers are not a clinical validation package; they are a warning that general-purpose models cannot be treated as reliable image authenticity adjudicators.

The most elegant and troubling detail is that the generating model could not perfectly identify its own outputs.[1] That does not prove every generative system will fail at self-recognition. It does undermine a common comfort: that provenance can be inferred by asking a capable model whether an image looks synthetic. In the tested setting, model fluency did not equal evidentiary reliability.

Why 41% Is Alarming, and Why 74.8% Is Not Enough

The spontaneous 41% result is alarming because it resembles ordinary clinical vulnerability more than a contest condition. A radiologist is usually not opening every chest radiograph with the assumption that half the worklist may be fabricated. If visual authenticity checks depend on constant suspicion, they are unlikely to survive real throughput, interruptions, priors, comparison studies, and reporting queues.

The warned condition answers a different question: can readers improve when the task is explicitly reframed as deepfake detection? Yes, substantially. But 74.8% mean accuracy in a balanced test set still leaves too much residual error for a safety control.[1] In clinical governance terms, that is closer to an educational signal than a defensible operating procedure.

There is also a base-rate problem. In the study, 50% of images were generated.[1] In a hospital workflow, the true prevalence of malicious or erroneous synthetic images would be expected to be far lower, though the available evidence does not quantify that rate. A rare-event environment changes the task. Even a reader who performs acceptably in an enriched test set may struggle to detect a rare fake image hidden among many ordinary studies.

That is the policy trap. The enriched prevalence makes the experiment feasible and interpretable, but it also prevents direct translation into live workflow performance. It may overestimate vigilance because readers are effectively placed on alert. It may also underestimate the practical difficulty of finding the one manipulated image that appears during an otherwise normal shift.

What This Study Does Not Prove

The Tordjman study is strong as a first reader study. It is not a prospective PACS validation. The distinction is not academic; it is the difference between evidence that should change the risk register and evidence that should change purchasing, staffing, or release controls.

Design featureWhy it matters for policy
50% generated imagesCreates a measurable reader-study task, but does not estimate performance when fake images are rare in routine care.
Single frontal chest X-raysSupports conclusions about 2D chest radiograph still images, not CT, MRI, ultrasound cine, fluoroscopy, endoscopy video, or other moving-image workflows.
Two generator modelsShows vulnerability to the tested RoentGen/GPT-4o and RoentGen outputs, but cannot establish general performance across future or untested generators.
Single-timepoint reviewDoes not measure fatigue, interruptions, comparison-study behavior, worklist pressure, or institutional escalation patterns.
No live PACS deploymentDoes not show whether radiologists, metadata systems, or detection algorithms would catch synthetic images in an operational clinical stream.

This is also where searches for AI-generated video detection in medical imaging need discipline. The available evidence here is about still chest X-rays, not video. The study does not establish detection performance for cine loops, endoscopy, ultrasound video, fluoroscopy, or volumetric imaging. Those may be logical next concerns, but they are not answered by this paper.

Detection Research Exists, But It Is Not a Clinical Product Verdict

Benchmark work is moving. The DSKI/MedForensics benchmark reported 91.6% accuracy for its best-performing detection algorithm on 116,000 images across six modalities, falling to 86.6% on unseen generators without retraining.[3] Those are useful research numbers. They are not the same as prospective clinical effectiveness.

The drop on unseen generators is particularly relevant for procurement. A hospital does not get to preselect the generator used by a future attacker, a compromised vendor pipeline, or a contaminated data source. A detector that performs well on a benchmark may still need retraining, calibration, monitoring, and failure-mode analysis before it can be trusted in a clinical imaging chain.

That is why FDA status should be kept separate from the scientific alarm. The alarm is justified by Tordjman et al. The purchase order is not. As of Q3 2026, there is no FDA-cleared medical deepfake detection tool to select from the AI-enabled device list.[2] A vendor may have a promising detector, a preprint, or a benchmark result. That is not the same as clearance, clinical validation, or workflow accountability.

For teams evaluating adjacent AI products, this is the same distinction that should apply across clinical AI procurement: benchmark performance belongs in the evidence file, but it does not replace validation in the intended setting. ClinicalMind’s Health System AI Procurement Checklist is the more appropriate frame than a rushed detector bake-off.

Provenance Controls Are Plausible, Not Yet Proven in PACS

Watermarking, cryptographic signatures, and DICOM metadata verification are more attractive governance ideas than asking every radiologist to become a forensic image analyst. They shift the question from visual suspicion to chain of custody. A 2026 trust-gap review discusses these safeguards but notes that none has been validated in clinical PACS workflows.[4]

That limitation is not a minor implementation detail. PACS environments include modality interfaces, outside-image imports, compression, routing rules, de-identification, research exports, downtime procedures, and vendor upgrades. A watermark or signature that works in a clean file-transfer demonstration can fail administratively if it breaks during ordinary image movement or if no one owns the exception queue.

The practical near-term response is therefore less glamorous: preserve provenance where possible, verify metadata on image ingestion, restrict unsupervised image imports, document escalation pathways, and require vendors to disclose how their systems handle generated, modified, or externally sourced images. These controls do not solve synthetic imaging. They are at least aligned with the evidence.

Evidence Appraisal Scorecard

DomainAppraisal
Clinical questionDirectly relevant to chest X-ray authenticity detection, but not to medical video, 3D volumes, or live workflow surveillance.
Study designAppropriate reader-study design for first evidence; not a prospective clinical validation.
Population and setting17 radiologists from 12 centers in six countries improves credibility, though the reading environment remains experimental.
Index taskVisual classification of real versus generated frontal chest X-rays; no PACS integration or metadata-based detection.
ComparatorHuman radiologists and tested LLMs; no FDA-cleared detector comparator exists.
Outcome clarityDetection accuracy is clearly policy-relevant, especially the 41% spontaneous and 74.8% warned-reader results.
GeneralizabilityLimited by 50% synthetic prevalence, two generator models, still images only, and single-timepoint reading.
Procurement readinessInsufficient for buying or relying on a specific detector; sufficient to justify monitoring, provenance controls, and prospective validation requirements.

ClinicalMind Verdict

Tordjman et al. provides the strongest available evidence that synthetic chest X-rays can evade both radiologists and tested LLMs under study conditions. The below-chance spontaneous detection result deserves attention, and the warned-reader improvement still falls short of a dependable operational safeguard.[1]

The correct Q3 2026 governance answer is therefore split into three parts: the threat is real; unaided radiologists and current tested LLMs should not be treated as reliable detectors; and current evidence does not justify purchasing or relying on any specific medical deepfake detection tool. No FDA-cleared tool exists for that purpose.[2]

A reasonable health-system response is monitoring, provenance discipline, metadata verification, controlled image-ingestion pathways, and a procurement requirement for prospective validation in the intended PACS workflow. The evidence supports vigilance. It does not support pretending that a validated detector layer is already available.

References

  1. The Rise of Deepfake Medical Imaging: Radiologists' Diagnostic Accuracy in Detecting ChatGPT-Generated Radiographs. Radiology. 2026.
  2. Artificial Intelligence-Enabled Medical Devices. U.S. Food and Drug Administration.
  3. MedForensics benchmark. arXiv. 2025.
  4. The Trust Gap in Generative Medical Imaging. ScienceDirect. 2026.

Risk-of-bias scorecard

Study design
Reader study
External / prospective validation
No
Key performance metric
41% spontaneous detection accuracy
Overall rating
High

Informational only — read the full disclaimer. This content supports procurement and research judgment, not clinical care decisions.

Submit a correction or sourcing issue

Blogarama - Blog Directory