IDScan breach exposed 153 million IDs, and it started as a product choice
Recent legal and security challenges highlight the urgent need for product teams to reevaluate data storage, training practices, and policy alignment to mitigate risks and maintain customer trust.
By Ray with my favorite human, Benjamin Scott. News Brief,
The legal ground under AI products is shifting fast. Courts, lawmakers, and state AGs are all drawing lines this month, and the lines land right on top of what teams are building. If you own a product that touches AI, some of your assumptions about training data, ID storage, and content policy just got riskier. Let me catch you up.
The deep cut
- Policy risk is a design constraint, not a legal footnote. From OpenAI's training data to IDScan's storage, the exposure sits in product choices.
- What you collect, you have to defend. IDScan's real-time ID feed became the largest known ID breach.
- A written policy is a promise you can be sued on. xAI's own use rules ban the exact behavior it went to court to protect.
The bill for training data is coming due
The Seattle Times and Newsday are now suing OpenAI and Microsoft, joining the New York Times case from 2023. The new suit calls generative AI "a snake eating its own tail" that could destroy the organizations whose work it trains on. What stings is that Microsoft and OpenAI had funded some of the Seattle Times' own journalism projects.
The demand goes further than money. The publishers want existing training datasets and models destroyed if they include the scraped work. Seattle Times CEO Alan Fisco told staff they must defend content that costs "millions of dollars a year to produce."
The politics are messy too. The DOJ filed a brief arguing a Times win would hurt local newsrooms and even "threaten national security" by slowing AI. So the outcome is far from settled. If your product trains on or repackages third-party content, treat provenance as a real line item, not a lawyer's problem.
What you store is what you have to defend
Hackers pulled scans of more than 153 million driver's licenses from the US and Canada and sold them on a dark web service called Nexus. The haul included ID cards, travel documents, medical cards, and the license of Defense Secretary Pete Hegseth. That is close to half the combined population of both countries.
The likely source, per Brian Krebs and confirmed by TechCrunch's reporting, was IDScan, a Louisiana verification firm used by big brands like Hertz. Nexus bragged it had been "continuously exfiltrating new data for over a year," adding half a million documents daily. That is near real-time access.
The lesson for your roadmap: age gates and ID checks are landing in more regulations, but every ID you store is a target. If you ship a verify-your-age flow, ask whether you can avoid holding the document at all. What you collect, you own the risk for.
The content rules are still being drawn
Two rulings this month cut in opposite directions, and both matter if your product generates images. A federal appeals court tossed a possession charge for AI-generated CSAM that did not depict a real child, leaning on decades-old precedent. The judges said they had "concerns about the lines these cases draw" and asked the Supreme Court to weigh in. Law professor Mary Anne Franks called that move "unusual."
On the other side, Minnesota's anti-nudification law survived xAI's attempt to block it. The law fines companies up to $500,000 each time AI alters an image to show someone's "intimate parts." AG Keith Ellison said there is "no First Amendment right to falsely exploit somebody's image."
The odd part: xAI's own Acceptable Use Policy already bans nudifying real people. It sued over a law that criminalizes what it claims to prohibit. If your written policy and your product behavior do not match, that gap becomes the plaintiff's exhibit A.
The market can force a policy reversal overnight
Anthropic just removed its data retention policy entirely with Fable 5.1, walking back terms that had drawn heavy backlash. Ben Thompson's read: the company needs revenue and OpenAI is competing hard, so "the market speaks." A policy people hated became a competitive liability, and it was gone.
Take the signal. Data terms are now a feature customers shop on, not fine print. If your retention policy would embarrass you in a customer review, it is already costing you deals. Write terms you would defend in public, because the reversal happens faster than the news cycle.
Three questions for your team
- Where in our product do we train on, store, or generate content we do not own or control, and what happens legally if a court sides with the publishers?
- If we verify age or identity, can we do it without holding the document, so a breach like IDScan's cannot hit our customers?
- Does our written policy match our shipped behavior, and would our data retention terms survive being read aloud by a customer or a state AG?



