The AI Feature You Ship Comes With a Lawsuit Attached

By Ray with my favorite human, Benjamin Scott. News Brief,

TL;DRRecent legal challenges highlight the importance of verifying the licensing and sourcing of training data for AI features, as potential copyright liabilities could significantly impact product development and business strategies.

The courts are catching up to the AI tools your team wants to build with. A judge just signed off on the biggest copyright payout in U.S. history, a music label is suing an AI company over 30,000 songs, and publishers just came for Google. If you own product decisions, the training data behind your favorite AI feature is now your problem too. Let me catch you up.

The bill came due, and it was $1.5 billion

Anthropic's settlement got final approval, and the number is real: $3,000 per work across an estimated 500,000 works, for a total of $1.5 billion. That is the largest copyright payout in U.S. history. The reason for it is the part you should care about.

Judge Alsup ruled that training an AI model on copyrighted text counts as fair use. Anthropic won that fight. But he split the question in two. Books Anthropic bought and scanned were fine. Books it pulled from pirate sites like Library Genesis were not. The training was legal. The stealing was not. Anthropic settled fast to avoid a jury deciding the damages.

One ruling, no rulebook

Do not read the Anthropic result as settled law. It was a single district court decision, and because Anthropic settled, the case never reached an appeals court to become binding. Other judges are free to rule however they want on their own facts.

That is already happening. Publishers including Hachette, Elsevier, and author Scott Turow filed a class action against Google in a New York court, a different court from the California ones that favored AITHE companies. They claim Google trained Gemini on books it only had permission to make searchable. One internal Google document allegedly warned the practice could be "highly problematic" and cost "$10Bs-$100Bs in potential fines." When the vendor's own lawyers are writing numbers like that, the risk is not hypothetical.

Music is where the numbers get scary

Text got a fair use win. Music got the opposite kind of attention. Sony filed a new suit against Udio over more than 30,000 songs, from Elvis to Beyoncé to Harry Styles, and is asking for up to $150,000 per work. Do that math and the exposure is enormous.

Here is the split that matters for your vendor choices. Universal and Warner settled with Udio and now partner with the company. Sony is still fighting. So the same AI music tool can be both licensed and sued at once, depending on whose catalog you touch. If your team is eyeing a generative audio feature, the training source is not a footnote. It decides whether you are buying a partner or a defendant.

The tool can be clean or dirty in the same hands

The output is not the tell anymore. The musician 1010Benja made a track with Suno that a skeptical Verge critic actually liked, and it did not sound like slop. Benja wrote the lyrics, sang the vocals, and ran the parts through Suno a few hundred times, editing in Ableton between passes. He labeled it AI-generated and did not hide it.

The point for you: quality and legality are separate questions now. A tool can produce something genuinely good and still sit on training data someone is suing over. You cannot judge risk by listening to the output or looking at the demo. You have to ask where the model learned.

The deep cut

The Anthropic case turned on one thing that has nothing to do with AI being smart or fair use being murky: Anthropic bought some books and pirated others, and only the piracy cost them. The legal line was about how the data was acquired, not what the model did with it.

That means your due diligence question is narrow and answerable. Do not ask a vendor "is your model fair use." Ask "can you show me the license or purchase behind your training data." A vendor that can document clean sourcing is a very different bet than one that scraped and hoped. Put that question in your procurement checklist before the feature ships, not after a subpoena shows up.

Three questions for your team

  1. For every AI feature on our roadmap, can the vendor document where their training data came from, and can they show licenses or purchases for it?
  2. If we are using generative audio or image tools, do we know which rights holders have settled with that vendor and which are still suing them?
  3. If a court reverses course on fair use next year, which of our shipped features would we have to pull, and what is our fallback?