AI & Copyright
Concerns about attribution, consent, and artists' livelihoods are legitimate. But the specific claim that AI image generation is "theft" or "plagiarism" rests on technical and legal misconceptions.
The Greatest Art Heist in History: How Generative AI Steals from Artists
- Models use stochastic latent-diffusion, not image databases — they don't store or "look up" images.
- Style is not protected under copyright law (17 U.S.C. §102(b)).
- Licensed, consent-based training already exists (Shutterstock–OpenAI, Adobe Firefly).
- The U.S. Copyright Office has not declared dataset training infringing; fair-use analysis is ongoing.
- Job-loss figures are speculative — no peer-reviewed study shows 90% displacement.
of prompts show any unintentional memorization — not systemic copying
the inference compute of pixel-space diffusion, and falling each generation
How latent diffusion actually works
The "reverse-engineers captions from memorized images" framing misunderstands the architecture. Models like Stable Diffusion and DALL-E 2 run an iterative denoising diffusion process in the latent space of a variational autoencoder. Sampling is stochastic, not deterministic; weights don't store JPEGs. The unCLIP architecture generates an image-embedding first, then decodes it — not a lookup or "photocopy."
Source: High-Resolution Image Synthesis with Latent Diffusion Models — Rombach et al.
On the environmental argument
Critiques often cherry-pick training costs without a baseline (artist workstations, renders, logistics) and ignore that recent latent-diffusion models need roughly 1/10 the inference FLOPs of pixel-space diffusion, with each generation more efficient than the last.
Conclusion
The "theft" framing relies on technical misconceptions, legal overreach, and flawed analogies to physical property. Legitimate concerns about attribution, consent, and economic transition are real — but better addressed through targeted policy (licensing, opt-out mechanisms, artist-centric platforms) than sweeping moral condemnation.
Consent-based AI already exists — and it’s growing
The debate usually stops at “theft.” But you don’t have to choose between AI and artists’ consent: a growing set of tools train only on licensed, opted-in data — and every time one gets used, that becomes more of the norm.
A nonprofit that certifies generative-AI models which don’t use any creator’s work without a license. Look for its certification — or ask a tool’s makers whether they have it.
A 10B text-to-image model trained exclusively on ~80M licensed, copyright-safe images from Freepik’s own catalogue — the first open model at this scale built on legally-cleared data. Free to run or self-host.
Others are moving the same way — Adobe Firefly (Adobe Stock + licensed/public-domain) and the contributor-compensated models from Shutterstock and Getty. The point worth holding onto: licensing and consent are workable today, not a fantasy for “later.”
Part of the AI Problems Index · see the Risk Atlas and Environmental Impact.