Policy

Cara Art Platform Scraper Now Helps Build Anti-Scraping Tool

After harvesting 12 million images from the artist portfolio site, a remorseful data scraper is collaborating with its founder on protective technology.

Omega Editorial· August 28, 2026· 3 min read

From antagonist to ally

A student who scraped 12 million artworks from Cara—an artist portfolio platform designed to resist AI training—has apologized and is now collaborating with the site's founder on a tool to help artists track when their work appears in AI datasets.

The scraper, who goes by "Heft," posted a 12-terabyte archive of essentially Cara's entire public image library to Reddit on August 13, calling it "a fun project" that cost less than $10. The incident triggered two additional copycat scrapes within days, spiking server costs and prompting hundreds of artists to delete their portfolios, according to Cara founder Jingna Zhang, as first reported by WIRED.

Zhang, a photographer who launched Cara in early 2023 with volunteers, built the platform specifically for the 1.5 million artists who oppose unauthorized use of their work in AI model training. The site filters out AI-generated images and integrates Glaze, a tool that masks artistic style from scrapers.

Why it matters

The episode illustrates a fundamental tension in AI development: even platforms built explicitly to protect creators from data harvesting remain technically vulnerable to scraping. The collaboration between Zhang and Heft represents an unusual attempt to build practical defenses in the absence of clear legal protections, while highlighting how easily accessible artist portfolios remain to anyone with basic technical skills and minimal resources.

The aftermath and response

Heft initially posted his dataset to the subreddit r/DefendingAIArt but says he "made a foolish decision to attempt to ragebait" and was "carried away by trolling in the comments." After seeing artists report panic attacks and watching them delete entire portfolios, he reached out to Zhang directly.

"He felt very bad to see how hurt people were," Zhang said. "So he decided to help us."

A second scraper pulled 8.5 million links and metadata from Cara and uploaded them to Hugging Face, which refused to remove the URLs, stating that "no copies of the artworks are hosted here." A third scraper obtained 123,000 images along with user bios containing personal information, sharing everything on Academic Torrents.

Zhang launched a GoFundMe seeking $120,000 for legal fees to explore cyber and copyright law strategies. As of late August, she had raised over $100,000.

Building Lantern

Heft now works in Cara's Discord server as a troubleshooter, explaining why proposed anti-scraping measures won't work. "He's just helping us, you know, taking the time to explain" structural vulnerabilities, Zhang said, noting he can "break through in like a few minutes, literally."

Because Heft believes "no site can be made truly 'unscrapable,'" he and Zhang are building Lantern, an open-source tool that creates a "one-way fingerprint" for images without storing them. The system scans new public AI image datasets and notifies artists when their work appears, providing links so they can request removal or file takedown notices.

Heft notes that while 12 million images "is not a lot to train an image model," and major AI labs likely train on much larger web scrapes from organizations like LAION, the tool aims to give artists awareness and some measure of control.

Zhang hopes the scraping incidents will draw policymaker attention beyond copyright issues. "Getting attacked by AI is just going to become so, so commonplace," she said. "And most of the time, the attacker won't say sorry—let alone try to set things right."

Details were first reported by Miles Klee at WIRED.

#ai training data#web scraping#artist rights#copyright#cara platform#data protection

This is an original analysis by the Omega editorial team. Source reporting: WIRED.

Want systems like this working for your business?

Book a Call

More in Policy

Policy· 3 min read

Wisconsin Congressman Uses AI-Generated Attack Video in Tight Race

Rep. Derrick Van Orden's synthetic deepfake of opponent Rebecca Cooke highlights gaps in social media regulation as states struggle with disclosure laws.

Via AI Watch · Aug 28, 2026
Policy· 3 min read

EU AI Act enforcement begins with transparency requirements

Major tech companies are adopting watermarking and disclosure standards as Brussels tests whether its sweeping rulebook can work in practice.

Via AI Watch · Aug 28, 2026
Policy· 3 min read

Judge Blocks Pentagon Ban on Anthropic Over AI Safety Stance

Federal court rules Defense Department illegally retaliated against Claude maker for refusing military surveillance and autonomous weapons use.

Via AI Watch · Aug 28, 2026