The fundamental accusation is simple: AI companies hoovered up copyrighted books, articles, images, code, and music without asking permission or paying a dime. OpenAI, Meta, Stability AI, Anthropic, and others built multi-billion-dollar businesses by training their models on works that took humans years to create. Now the lawsuits are piling up faster than a ChatGPT response, and the legal system is scrambling to figure out whether this constitutes the biggest copyright heist in history or falls under fair use. Getty Images fired the opening salvo in January 2023, suing Stability AI in both US and UK courts for allegedly copying 12 million photographs from its database to train the image generator Stable Diffusion. The smoking gun? Stability's model occasionally spits out images with Getty's watermark still visible, mangled but recognizable. Getty isn't asking for a licensing deal anymore, they want statutory damages that could reach into the billions. In the Southern District of New York (case 1:23-cv-00135), Judge Katherine Polk Failla is weighing whether Stability's scraping constitutes transformative fair use or blatant infringement. As of May 2024, the case survived Stability's motion to dismiss, meaning discovery will force the company to reveal exactly what training data it used and how. The Authors Guild coordinated class-action lawsuits that read like a who's-who of American literature. Sarah Silverman, John Grisham, Michael Chabon, George R.R. Martin, and Jodi Picoult all sued OpenAI and Meta (cases 3:23-cv-03416 and 3:23-cv-03223 in Northern District of California). Their complaint is elegant: when you prompt ChatGPT or LLaMA (Large Language Model Meta AI) to summarize their books, the models produce accurate plot details, character arcs, and even writing style mimicry, proof that the full copyrighted texts were ingested during training. OpenAI's defense rests on two pillars: fair use (the training is transformative, doesn't replace the original market) and the argument that learning from copyrighted works is what humans do every day. Judge Vince Chhabria hasn't bought it wholesale. In February 2024, he allowed core copyright claims to proceed while dismissing some derivative claims around Digital Millennium Copyright Act (DMCA) violations. The music industry entered the fray with a vengeance. The Recording Industry Association of America (RIAA), representing Universal Music Group (UMG), Sony Music Entertainment, and Warner Music Group, sued AI music generators Suno and Udio in June 2024 (cases 1:24-cv-05410 and 1:24-cv-05311, District of Massachusetts). Their evidence is damning: they prompted the AI with specific lyrics or style cues and got output that replicated chunks of Mariah Carey, ABBA, and James Brown recordings. Suno CEO Mikey Shulman admitted in a BBC interview that the company trained on copyrighted music but argues the output is transformative. The RIAA disagrees, they're seeking $150,000 per willfully infringed work, and they've documented over 400 songs in the complaint alone. How do AI companies actually manage this legal minefield? The strategies fall into four buckets:

  1. Fair Use Defense: Claim that training AI is transformative use under Section 107 of the Copyright Act. Point to Google Books (Authors Guild v. Google, 2015) where scanning millions of books for search was deemed fair use.
  1. No Access Equals No Infringement: Argue that once trained, the model doesn't "store" copyrighted works, so outputs aren't copies. This worked partially in GitHub Copilot litigation, where Judge Jon S. Tigar noted in November 2023 that plaintiffs struggled to prove direct copying.
  1. Licensing Partnerships: OpenAI signed deals with Associated Press, Axel Springer, and Financial Times in 2023-2024 to license content. These agreements typically pay $5-10 million annually per major publisher, a fraction of what training data is actually worth but enough to create a "good actor" narrative.
  1. Blame the Open Internet: Companies argue they scraped publicly available data, which courts have historically treated differently than paywalled content. The Computer Fraud and Abuse Act (CFAA) typically doesn't apply to public web scraping (hiQ Labs v. LinkedIn, 9th Circuit 2022).

But here's where it gets wild: researchers at Carnegie Mellon University and the Allen Institute for AI published "Scalable Extraction of Training Data from (Production) Language Models" in November 2023. They demonstrated that GPT-3.5 could be prompted to regurgitate verbatim passages from Harry Potter, confirming that models memorize substantial copyrighted chunks, not just statistical patterns. The European Union's AI Act, which entered into force in August 2024, requires AI developers to publish detailed summaries of copyrighted training data. Article 53 specifically mandates transparency about what content was used, giving rights holders a roadmap to sue. Japan offers a contrasting approach: their Copyright Act Article 30-4, amended in 2019, explicitly allows AI training on copyrighted works for non-expressive use. This is why Stability AI moved substantial operations to Tokyo in 2023. However, Japanese courts still prohibit outputs that reproduce "the expression" of original works, creating a murky line between learning and copying that mirrors the US fair use debate.