The practical question behind ai copyright settlements and medical ai development is no longer whether copyright lawyers are worried about generative AI. It is whether a medical AI team can prove, at the point of diligence, procurement, or FDA review, where its training data came from and what rights traveled with it.

That proof used to be treated as two separate chores. Counsel reviewed licenses, scraping history, vendor warranties, and indemnities. Regulatory affairs assembled the training-data story for the device file: data sources, intended-use fit, annotation, demographic coverage, bias controls, and performance evidence. In 2026, those files are starting to collapse into one another. A dataset that cannot survive a copyright chain-of-custody review is also harder to defend as a regulated training asset. A dataset that cannot be described clearly enough for FDA-facing documentation will not become less risky when an investor or hospital asks who had the right to share it.

Legal and medical regulatory streams converging on a data provenance shield

The Bartz v. Anthropic settlement matters to medical AI developers for exactly that reason. The headline number is large, but the operational lesson is sharper: courts and counterparties are beginning to care not just about what a model does with data, but how the developer acquired and retained the material that made training possible.

Bartz draws the line medical AI teams should already be drawing internally

In Bartz v. Anthropic, the court distinguished between training on lawfully acquired works and the use or storage of pirated copies. Norton Rose Fulbright’s 2026 litigation update describes the ruling as finding that training on lawfully acquired works could qualify as transformative fair use, while the retention of pirated copies from shadow-library sources such as Library Genesis created a separate infringement problem.[1]

That distinction is more useful to medical AI companies than a broad claim that “AI training is legal” or “AI training is infringement.” The question moves upstream: was the input lawfully acquired, licensed, transferred, de-identified, annotated, and retained under terms the developer can still prove?

The settlement was reported at $1.5 billion, involving roughly 500,000 works, with statutory damages exposure of up to $150,000 per work.[1][2] NPR described it as the largest copyright recovery in history.[3] The Authors Guild noted that the settlement received preliminary approval in September 2025 and that a final approval hearing was set for May 14, 2026, meaning details of the final payout structure could still adjust.[2]

For a clinical AI developer, the point is not that a medical imaging startup has the same exposure profile as a general-purpose language model company. It usually does not. The defensible statement is closer to: we obtained this dataset from this source, under this agreement, with these permissions, after these transformations, using these annotations, with this retention policy, and with these exclusions.

Kadrey v. Meta, as discussed in the same litigation context, reinforces the same point: source legality matters when courts analyze AI training and fair use.[1] That does not turn every clinical dataset into a copyright minefield. It does make undocumented acquisition history a business risk, especially when the training corpus includes medical images, pathology slides, journal content, textbooks, proprietary reports, curated annotations, or third-party benchmark datasets.

The litigation environment has also hardened. An independent tracker counted 87 active U.S. copyright lawsuits against AI companies as of March 5, 2026, up from about 30 at the end of 2024.[4] That count is not an official court-docket compilation, and it should not be treated as a precise government statistic. It is still a useful signal: copyright risk has moved from theoretical memo language into active litigation and settlement behavior.

The FDA pressure arrives from a different door

FDA is not enforcing copyright law. Its concern is whether an AI-enabled medical device software function is safe, effective, and supported by data appropriate to the intended use. But the documentation FDA is moving toward asks many of the same factual questions that a copyright and data-rights review asks.

FDA’s January 2025 draft guidance on AI-enabled device software functions is not final binding law. Public comments were due April 7, 2025, and the final version may differ. Still, the draft is a strong indicator of the regulatory record manufacturers should expect to build. Fenwick’s analysis of the draft guidance notes FDA’s expectations around training data source documentation, annotation process details, bias control evidence, and model-card-style information in premarket submissions.[5]

Those expectations are not cosmetic. Training data source documentation forces the developer to identify where data came from and why it is suitable for the device’s intended population and clinical use. Annotation process details force the developer to explain who labeled the data, what instructions they followed, how disagreements were resolved, and how label quality was controlled. Bias control evidence forces the developer to show that performance was not built on a dataset that systematically underrepresents or misrepresents relevant patient groups. Model-card-style disclosures push the team toward a more structured account of model purpose, data, limitations, and performance.[5]

A company can try to maintain one version of the truth for FDA and another for counsel, but that usually fails under pressure. The regulatory submission says the model was trained on a multi-institutional dataset. The contract file shows only one institution granted machine-learning training rights. The data science notebook includes a public dataset added late in development. The annotation vendor agreement is silent on downstream commercial use. The premarket record describes demographic balancing, but the raw data inventory cannot show which records were excluded and why. None of those gaps begins as a philosophical dispute about AI. They begin as ordinary provenance failures.

What provenance has to prove

“Data provenance” is often used as a soft governance phrase. In medical AI development, it needs to become a traceable evidence package. At minimum, it should answer six questions without depending on the memory of the engineer who assembled the first training run.

  • What data was used: the source systems, public datasets, institutional extracts, vendor-supplied data, synthetic or augmented data, and any externally sourced reference material.
  • How it was obtained: the agreement, permission, public-use terms, data-use approval, institutional review pathway, or other basis for access.
  • Who had authority to share or license it: the hospital, registry, publisher, vendor, researcher, data collaborative, or other rights holder, with any limits on sublicensing or commercial training.
  • How it was transformed: de-identification, normalization, segmentation, filtering, augmentation, deduplication, linkage, format conversion, and split assignment.
  • How it was annotated: annotator qualifications, labeling instructions, review hierarchy, adjudication method, quality checks, and version history.
  • Where it went: which model versions used it, which experiments included it, whether copies were retained, whether excluded data was deleted, and whether downstream fine-tuning or validation reused it.

The same answers serve different reviewers. Counsel wants the rights trail. FDA reviewers want to understand the development and validation data behind a device function. Investors want to know whether a core training asset can survive diligence. A health system governance committee wants assurance that a vendor’s model was not built on data practices the institution would not tolerate internally. For the health system side of that governance problem, the same operational gap appears in broader shadow-AI controls, not just in regulated device development; ClinicalMind’s discussion of healthcare AI governance frameworks is a useful adjacent lens.

The hard part is that provenance is easiest to document before training begins and hardest to reconstruct after model performance looks promising. Once a prototype works, no one wants to remove a dataset, rerun an experiment, or admit that the best-performing version depended on a source with unclear rights. But that is exactly when the file begins to matter. A weak data record turns an engineering success into a regulatory and commercial cleanup project.

Medical AI teams need precision here. Treating all clinical data as copyrighted overstates the risk and leads to poor controls. Treating all health data as free training material understates the risk and invites exactly the kind of source problem that Bartz made visible.

Comparison of lower-risk raw clinical data and higher-risk curated medical AI datasets

WIPO’s March 2026 analysis makes the useful distinction: raw clinical data generally is not protected by copyright, while healthcare AI training datasets may be protected when they involve creative selection or arrangement.[6] That distinction matters because many medical AI assets are not just raw measurements. They are selected, cleaned, labeled, sequenced, enriched, and packaged for a purpose.

Data typeCopyright concernPractical provenance focus
Raw clinical measurements, lab values, and many de-identified EHR fieldsGenerally lower copyright concern when treated as factual data, though privacy, contract, and health-data rules still matterAccess authority, de-identification process, permitted use, patient population, transformations, and retention
Radiology images, pathology slides, dermatology photographs, waveform displays, and other clinical mediaHigher concern when image creation, compilation, storage terms, or dataset packaging create rights questionsInstitutional ownership or license, consent or authorization pathway where relevant, vendor/platform terms, and downstream training rights
Published research, medical textbooks, guideline content, and journal figuresHigher copyright concern because the expressive work itself may be protectedPublisher or author license, permitted machine-learning use, extraction method, and whether copies are retained
Curated benchmark datasets, registries, labeled image collections, and expert annotationsPotential concern where selection, arrangement, or annotation reflects protectable authorship or contractual controlDataset terms, annotator agreements, rights to labels, sublicensing limits, version history, and commercial-use restrictions
Synthetic or augmented data derived from clinical source dataDepends on the rights and restrictions attached to the source data and the transformation processSource lineage, generation method, validation, whether synthetic outputs can reconstruct protected or restricted source material

The low-copyright-risk category is not a low-compliance category. De-identified EHR data, lab values, and raw observations still raise privacy, contract, institutional review, security, and data-use concerns. The point is narrower: copyright analysis should not be pasted over the entire clinical data environment as if a sodium value and a textbook chapter raise the same issue.

Conversely, a dataset can be medically familiar and still legally complicated. A pathology archive may include images created inside a health system, annotations by external specialists, a storage vendor’s platform terms, research-use limitations, and later commercial fine-tuning by a startup. A public medical imaging dataset may be easy to download but still require close review of its terms, contributor permissions, and permitted downstream use. “Public” is not the same as “cleared for commercial medical AI training.”

The overlap is not abstract. It appears in the rows of the same dataset inventory.

QuestionCopyright and data-rights useFDA-facing use
Source identityShows who supplied the material and who may hold rightsShows what data population supports training and validation
Acquisition basisSupports license, permission, public-use, or contractual analysisSupports credibility of the development record and data governance process
Permitted usesDetermines whether model training, commercial deployment, sublicensing, or retention is allowedHelps define whether the submitted model was developed under controlled and reproducible conditions
Annotation lineageIdentifies ownership or license issues in labels and derived datasetsSupports label quality, clinical validity, and bias-control claims
Transformations and exclusionsShows whether restricted materials were removed, modified, or retainedExplains preprocessing, representativeness, bias mitigation, and reproducibility
Model-version linkageIdentifies which models may be affected by a disputed sourceConnects training data to performance evidence and change management

This is why provenance should not live only in a legal folder or only in a regulatory submission workspace. A license clause saying data may be used for “research” but not commercial product development affects both legal clearance and the validity of the development pathway. An annotation vendor agreement that fails to assign or license label rights affects both copyright posture and the ability to explain the supervised learning process. A late-stage dataset swap affects both rights analysis and the evidence trail supporting the submitted model.

The teams that struggle most are often not careless; they are fragmented. Data science tracks experiments in notebooks and object stores. Clinical operations tracks chart-review protocols. Legal tracks executed agreements. Regulatory tracks the device file. Security tracks data movement. Procurement tracks vendors. No single person can answer the simple diligence question: which exact data sources trained this model version, under which permissions, and with which exclusions?

Interoperability work can help, but it does not solve rights provenance by itself. FHIR-based exchange, TEFCA participation, and cleaner clinical data infrastructure can make health data more usable, but permission, purpose limitation, and downstream training rights still need their own record. That distinction is important for teams building on modern health-data infrastructure, as discussed in ClinicalMind’s analysis of TEFCA, FHIR, and clinical AI.

The licensing norm is emerging, but healthcare will not copy entertainment

The U.S. Copyright Office’s 2025 Part 3 report called for scalable licensing mechanisms for copyrighted works used in AI training.[7] Voluntary licensing examples involving Disney and OpenAI, UMG and Udio, and WMG and Suno are signs of an emerging commercial norm around negotiated access to valuable training material.[7]

Those examples should not be imported too neatly into clinical AI. A catalog of songs or films is not the same as a longitudinal oncology dataset, a radiology archive, or a pathology-slide collection with clinical metadata. Healthcare data carries privacy obligations, institutional duties, consent and authorization questions, research restrictions, and patient-trust concerns that entertainment licensing does not need to solve.

Still, the licensing trend has one direct implication for medical AI: “available” is becoming a weaker defense than “authorized.” If a developer trains on a public repository, a scraped article set, or a dataset inherited from an academic collaborator, the operational question is not merely whether the files could be accessed. It is whether the developer can show permission for the specific use that occurred.

The development-stage problem: provenance must start before the model is good

Provenance controls are often postponed because early research feels exploratory. That is understandable and dangerous. The training set that begins as a sandbox frequently becomes the basis for a pre-submission meeting, a seed-round diligence packet, a pilot with a health system, or a 510(k), De Novo, or PMA strategy. By then, the team may no longer know which files were temporary, which records were excluded, or which public dataset was used only in an early baseline but still influenced architecture decisions.

A more defensible approach is to treat datasets as controlled development assets from the beginning. That does not mean every experiment needs a full legal memo. It means every dataset entering the environment should receive a stable identifier, source record, rights status, permitted-use note, retention rule, and linkage to model versions. If the team later learns that a source cannot be used commercially, it should be possible to identify the affected experiments without interviewing half the engineering staff.

Clinical examples make the point concrete. In cancer screening AI, questions about training data are inseparable from questions about real-world performance, target population, and regulatory status; ClinicalMind’s review of AI in cancer screening shows why evidence quality depends heavily on the data behind the model. In primary care, where ambient documentation, triage, risk scoring, and administrative automation may draw from mixed clinical and textual sources, the evidence gap is just as tied to data sourcing choices; see ClinicalMind’s discussion of AI in primary care.

The same discipline matters after launch. A model updated with new site data, new annotations, or new third-party reference material may change both its regulatory record and its rights posture. Postmarket learning, fine-tuning, and performance monitoring all need provenance continuity. Otherwise, a company can clear its original training set and still lose control of the data chain during maintenance.

What to do in Q3 2026

The practical move for 2026 is not to create separate copyright, FDA, bias, and procurement provenance systems. That produces duplicate records, inconsistent answers, and late-stage surprises. The stronger move is to build one training-data provenance system that can output different views for different reviewers.

For legal review, the system should show acquisition basis, license terms, rights holder, permitted uses, restrictions, and deletion obligations. For FDA-facing work, it should show data sources, development versus validation splits, preprocessing, annotation process, representativeness, bias controls, and model-version linkage. For procurement and health system governance, it should show whether the vendor can explain the data chain in plain terms and whether the institution is being asked to accept undisclosed training-data risk.

The copyright cases will continue to develop. The final contours of fair use for AI training are not settled across every context. But waiting for perfect legal certainty is not a serious operating plan. Bartz has already made source legality a central practical issue. FDA has already signaled that training-data documentation belongs in the premarket record. WIPO has already clarified why raw clinical data and curated healthcare AI datasets require different copyright analysis.[1][5][6]

Medical AI companies with defensible data chains reduce legal exposure and make their products easier to regulate, easier to diligence, easier for health systems to procure, and easier for clinicians and patients to trust.

References

  1. AI in litigation series: An update on AI copyright cases in 2026, Norton Rose Fulbright
  2. What Authors Need to Know About the Anthropic Settlement, Authors Guild
  3. Anthropic settlement authors copyright AI, NPR, September 5, 2025
  4. Latest U.S. Map of Copyright Suits v. AI Companies: Total 87, Mar. 5, 2026, ChatGPT Is Eating the World, March 5, 2026
  5. FDA Issues Draft Guidances on AI in Medical Devices, Drug Development: What Manufacturers and Sponsors Need to Know, Fenwick
  6. Intellectual property rights in healthcare-related AI, WIPO, March 2026
  7. AI Copyright Lawsuit Developments 2025, Copyright Alliance, 2025