AI systems trained on millions of creator works—music, images, text, books—without artists' consent or knowledge. This wasn't an accident; it was the deliberate foundation of generative AI's explosive growth, and the legal and ethical reckoning is just beginning.
How AI training without creator consent became industry standard
Most major AI models were built by systematically downloading internet content, treating creator work as freely available material for building commercial products. Companies like OpenAI, Meta, Stability AI, and Google collected vast datasets—including copyrighted music, images, books, and articles—without seeking permission. Creators found their work in training datasets through luck or third-party investigation tools. They received no payment, no notice, and no opportunity to opt out.
The scope is striking. The Atlantic reported on music incorporated into AI training datasets. Getty Images sued Stability AI for using approximately 12 million copyrighted photographs. Music publisher Kudos Records alleged that Suno trained its AI music generator on copyrighted recordings without consent. And when Anthropic settled its authors' lawsuit, the settlement covered reportedly substantial numbers of copyrighted works.
This is the first post in our series exploring creator rights in the AI era. The issue isn't just money—it's consent, transparency, and who controls the narrative of creativity itself.
Why AI training without creator consent happened in the first place
AI developers faced a choice: obtain permission from millions of creators and work out licensing agreements, or train their systems on whatever content they could access online. They chose the latter path. Large language models and generative AI systems need enormous datasets—billions of text samples, hundreds of millions of images, millions of songs. Securing permission from each creator would dramatically slow development and increase costs significantly.
Instead, companies deployed automated tools to gather content from across the internet. Platforms justified this under "fair use"—the legal principle permitting limited use of copyrighted material for purposes that transform it. However, legal justifiability differs from ethical responsibility. Fair use doesn't guarantee no harm occurred; it means a court might rule in your favor if challenged.
Creators received no notification because revealing this process would generate opposition. Transparency would invite public criticism. Seeking permission would result in many refusals.
The SZA case: When a major artist discovered her work was used in AI training
In 2023, SZA publicly discussed concerns about her music appearing in AI training datasets. She highlighted the consent and transparency crisis at the heart of AI training. She did not authorize this. She was not informed. She has no straightforward path to having her music removed.
SZA's response centered on an important distinction: this extends beyond financial compensation. She framed it as fundamentally about consent and fair treatment, emphasizing how Black artists and creators remain especially vulnerable, given their history of seeing their work exploited by music and publishing industries. AI magnified these existing vulnerabilities. While the scraping affected all creators indiscriminately, the consequences were not equally distributed—artists from marginalized communities, who often lack legal resources and industry connections, faced the most severe losses.
The contradiction: copyright protection exists specifically to safeguard creators. Yet AI developers exploited legal ambiguity to weaken that protection at massive scale.
Copyright stripping and deliberate obfuscation: Key legal concerns
Multiple lawsuits have alleged that AI companies did more than simply use copyrighted works—they actively removed or obscured copyright ownership information. This distinction carries legal weight. Fair use typically applies when copyrighted material is used in a transformative manner while keeping attribution visible. Removing or erasing copyright metadata indicates awareness of potential wrongdoing and weakens fair-use arguments.
This conduct has appeared across several cases and points to intentional concealment meant to prevent proper attribution and creator compensation.
Consent versus compensation: Why this distinction matters
The fundamental issue in AI training without consent centers on power and transparency, not merely financial reward. A creator might agree to payment for AI training use. However, none were given that opportunity.
Genuine consent requires:
- Genuine choice, not automatic participation
- Clear information about what will be used and for what purpose
- Agency in determining how work appears in training systems
Payment is significant—creators deserve fair compensation for their work. Yet consent must come first. You cannot ethically justify using someone's work without permission simply because you can compensate them later.
FAQ
What datasets included creator content?
Significant datasets include Common Crawl (which gathered billions of web pages), Book Corpus and comparable text collections, and various image databases used for image generation systems. Complete dataset compositions remain unclear; companies typically do not publicly reveal full details of their training sources. Researchers using reverse-engineering techniques have begun identifying specific works in training data, but most companies do not practice comprehensive public disclosure.
Can creators remove their work from AI training data?
Currently, no formal process exists. Identifying your work in a dataset does not automatically grant you removal rights. Several companies have introduced voluntary artist opt-out options (Stability AI offers one example), though these are discretionary rather than legally mandated. Courts and legislators are working to establish removal rights, but creators currently have limited options for action.
Is this legal?
The legal status remains uncertain. AI companies contend that fair use permits training on lawfully accessed data. Court rulings to date have shown mixed results—some courts have upheld fair use for transformative training applications, while others have raised concerns about uses that substitute for the original market. Settlements and ongoing cases have not yet settled the broader question of whether fair use applies to AI training on copyrighted materials.
What legal cases are advancing creator protections?
The New York Times filed suit against OpenAI; Sarah Silverman and additional authors brought cases against Meta and OpenAI. These cases focus on training without authorization and intellectual property infringement, and their outcomes could establish new boundaries for what fair use means in AI development.
What's being done to prevent this going forward?
Lawmakers in several regions have drafted regulatory proposals. Multiple jurisdictions are exploring requirements that would demand explicit permission before training AI on protected works or personal information. As of mid-2025, most proposals remain under discussion or in early legislative consideration.
Where this leaves creators in 2025
Companies maintain legal advantages, yet momentum is shifting. Ongoing cases continue, legislative action is accelerating, and creators are developing better tools to identify and document when their work has been used without permission.
Yet the core problem—using creator work to train AI without consent or transparency—remains legally unresolved. This moment represents an inflection point for the industry. The question isn't whether scraping occurred; it's whether future AI training will require creator permission.
Next in this series: How creators are fighting back—from detection tools to legal strategies to collective action.
References
- The Atlantic: How AI Scraped Music
- Getty Images Sues Stability AI
- The New York Times v. OpenAI
- Sarah Silverman and Authors v. Meta and OpenAI
