AI Training Data Copyright Litigation: What Enterprises Should Learn from NYT v. OpenAI
NYT v. OpenAI centers on whether training large language models on copyrighted news content, and generating near-verbatim outputs, constitutes infringement or protected fair use. The case remains unresolved, but it and related litigation are already reshaping how enterprises should evaluate AI vendor contracts, indemnification terms, and technical controls for output review and training data provenance.
What NYT v. OpenAI Alleges
The New York Times filed suit against OpenAI and Microsoft in the U.S. District Court for the Southern District of New York in December 2023. The complaint alleges that OpenAI used millions of Times articles without authorization to train GPT models, and that ChatGPT has reproduced near-verbatim excerpts of Times content in its outputs. In some instances, the complaint alleges the model generated fabricated content and misattributed it to the Times, raising separate concerns about output accuracy alongside copyright exposure.
OpenAI's defense rests substantially on fair use, arguing that training on copyrighted text is a transformative use that does not substitute for the original works in the market. As of the most recent publicly available filings, the case has not reached a final ruling on fair use or damages. Discovery disputes, including questions about data retention and log preservation, remain active. The case's outcome, whenever it arrives, will not resolve every open question in AI copyright law, but it is likely to influence how courts evaluate similar claims against other model providers.
Fair Use and the Training-vs-Output Distinction
Across NYT v. OpenAI and related cases, courts are increasingly treating two types of claims as legally distinct: claims about ingesting copyrighted material during training, and claims about a model reproducing that material in its outputs. This distinction matters because fair use analysis depends on the specific use at issue, not just the presence of copyrighted material somewhere in a system's pipeline.
Authors Guild v. OpenAI raises similar training data claims on behalf of fiction authors and is being litigated in parallel to NYT v. OpenAI. In Thomson Reuters v. ROSS Intelligence, a Delaware federal court rejected a fair use defense in 2025 regarding the use of Westlaw headnotes to train a legal research tool, an early signal that courts will not uniformly accept transformative use arguments for training data. Getty Images' parallel litigation against Stability AI in the UK and US has narrowed some claims through preliminary rulings, again separating training-time ingestion issues from output-level reproduction issues.
These rulings do not predict how NYT v. OpenAI will resolve, since the facts, models, and content types differ. But together they show that fair use outcomes are fact-specific and inconsistent across jurisdictions, which is itself a risk enterprises should factor into vendor evaluation.
Why This Matters for Enterprise AI Buyers
Most enterprises deploying commercial AI models are not parties to this litigation, but they inherit exposure through the models and vendors they use. Three categories of risk are relevant:
- Vendor liability: if a court finds a provider's training practices infringing, enterprises using that provider's models may face uncertainty about continued access, licensing terms, or required changes to deployed models.
- Retraining or model-deprecation obligations: an adverse ruling could require a vendor to retrain or withdraw a model, with downstream effects on enterprise applications built on that model.
- Output usage risk: even if training practices are found lawful, an enterprise's own use of AI-generated outputs that closely resemble copyrighted source material carries independent infringement risk.
None of these risks require a plaintiff win to materialize as a business concern. The uncertainty itself affects how enterprises should structure vendor contracts, monitor model outputs, and document decisions about which models and configurations are approved for production use.
Contractual Protections to Evaluate
Given this uncertainty, governance and procurement teams should review vendor agreements for the following:
- Indemnification scope: which infringement claims are covered, and what exclusions apply.
- Guardrail configuration requirements: whether indemnification coverage is contingent on using the vendor's default or recommended guardrail settings.
- Licensing and continued access terms: what happens if a court requires a vendor to change licensing terms or restrict access to a model.
- Model deprecation and retraining clauses: how the vendor will notify and support customers if a model must be retrained or withdrawn as a result of litigation.
Technical and Governance Controls to Reduce Exposure
Contractual protections work best alongside operational controls. Governance teams should maintain:
- An inventory of AI vendors and models in production, with indemnification terms mapped to each.
- Documented confirmation that default guardrail configurations required for indemnification are actually in use, along with a record of any deviations.
- Policy enforcement and audit logging that record how models were configured and used.
- Monitoring of model and tool usage to support review of outputs that may closely resemble copyrighted source material.
Practical Next Steps for AI Governance Teams
Enterprises do not need to wait for litigation to resolve before acting. Governance teams can start by mapping which AI vendors and models are in production and identifying the indemnification scope and exclusions attached to each. Where models rely on default guardrail configurations for indemnification, governance teams should confirm those configurations are actually in use and document any deviations.
Runtime governance capabilities, including policy enforcement, audit logging, and monitoring of model and tool usage, can support these efforts by giving enterprises a documented record of how models were configured and used when questions about copyright exposure arise. This does not eliminate legal risk tied to unresolved litigation, but it strengthens an enterprise's position to demonstrate diligence and respond to vendor or regulatory inquiries as case law develops.
Document how AI models are configured and used
As training data litigation continues to unfold, enterprises benefit from a documented record of model configuration, usage, and output review. Runtime governance provides that operational foundation.
Explore runtime governance