AI & Copyright

Concerns about attribution, consent, and artists' livelihoods are legitimate. But the specific claim that AI image generation is "theft" or "plagiarism" rests on technical and legal misconceptions.

Misconception alert. "AI image generation constitutes theft" conflates how diffusion models work with how copyright law actually operates.
The claim being examined — Goetze, T. (2024)

The Greatest Art Heist in History: How Generative AI Steals from Artists

Key rebuttal points
  • Models use stochastic latent-diffusion, not image databases — they don't store or "look up" images.
  • Style is not protected under copyright law (17 U.S.C. §102(b)).
  • Licensed, consent-based training already exists (Shutterstock–OpenAI, Adobe Firefly).
  • The U.S. Copyright Office has not declared dataset training infringing; fair-use analysis is ongoing.
  • Job-loss figures are speculative — no peer-reviewed study shows 90% displacement.
<0.03%

of prompts show any unintentional memorization — not systemic copying

~1/10

the inference compute of pixel-space diffusion, and falling each generation

How latent diffusion actually works

iterative denoising diffusionin the latent space of a variational autoencoder
Sampling is stochasticnot deterministic; weights don't store JPEGs
The unCLIP architecture — image-embedding first, then decodes
1
generates an image-embedding firstThe unCLIP architecture
2
then decodes itnot a lookup or "photocopy"

The "reverse-engineers captions from memorized images" framing misunderstands the architecture. Models like Stable Diffusion and DALL-E 2 run an iterative denoising diffusion process in the latent space of a variational autoencoder. Sampling is stochastic, not deterministic; weights don't store JPEGs. The unCLIP architecture generates an image-embedding first, then decodes it — not a lookup or "photocopy."

Source: High-Resolution Image Synthesis with Latent Diffusion Models — Rombach et al.

On the environmental argument

Critiques often cherry-pick training costs without a baseline (artist workstations, renders, logistics) and ignore that recent latent-diffusion models need roughly 1/10 the inference FLOPs of pixel-space diffusion, with each generation more efficient than the last.

Conclusion

The "theft" framing relies on technical misconceptions, legal overreach, and flawed analogies to physical property. Legitimate concerns about attribution, consent, and economic transition are real — but better addressed through targeted policy (licensing, opt-out mechanisms, artist-centric platforms) than sweeping moral condemnation.

For artists — the constructive path

Consent-based AI already exists — and it’s growing

The debate usually stops at “theft.” But you don’t have to choose between AI and artists’ consent: a growing set of tools train only on licensed, opted-in data — and every time one gets used, that becomes more of the norm.

Others are moving the same way — Adobe Firefly (Adobe Stock + licensed/public-domain) and the contributor-compensated models from Shutterstock and Getty. The point worth holding onto: licensing and consent are workable today, not a fantasy for “later.”

Part of the AI Problems Index · see the Risk Atlas and Environmental Impact.