← All posts

Millions for a comment: how tech giants buy AI training data

Reddit APIs, OpenAI media deals, the Spirit Airlines archive, and Anthropic’s book scans — why model training now starts with a check, not the open web.

Millions for a comment: how tech giants buy AI training data
Contents

In brief

A Habr essay tracks how forum posts, newsrooms, and even used paper books turned into priced assets.

Think of oil: while it “just flows,” nobody meters barrels; when the plant runs dry, buyers pay for access.

The path runs from Reddit’s paid API and OpenAI’s media licenses to Google’s Spirit Airlines bankruptcy auction and Anthropic’s buy-scan-destroy book program.

What happened

The large-language-model boom is usually dated to 2022 and ChatGPT on GPT-3.5.

Training needs other people’s text, code, images, and chat — without a human trail there is nothing to learn from. The author uses a simple picture: a new bakery “does not exist” for a recommender until someone opens its map page, leaves a review, and a rating. Information appears only with an action.

When a company lacks a corpus, a well-funded player does not wait for users to upload PDFs; it buys access. Forums are especially valuable here — long archives of questions and answers already filtered by real conversation.

Reddit was an early loud signal. In April 2023 it shut free API access; two months later the API was paid.

The iOS client Apollo shut down — its developer put API cost at roughly $20 million a year. Early in 2024 Reddit struck a deal with Google of about $60 million a year to license content for AI training; by IPO time, similar agreements already exceeded $200 million.

In spring 2024 OpenAI announced partnerships almost weekly: Stack Overflow (which had just cut about 28% of staff amid traffic decline), Reddit again, then News Corp for more than $250 million over five years.

Next came Dotdash Meredith for at least $16 million, Axel Springer at about $13 million a year for up to three years, and the Financial Times at $5–10 million a year. Wiley closed two deals at $21 million and $23 million with unnamed buyers; Informa took a $10 million starter fee from Microsoft. Shutterstock said AI licensing brought about $104 million in 2023.

An Epoch AI forecast from 2024 warned that publicly available human-made content as training fuel would start shrinking sharply from 2026 and be nearly exhausted by 2032.

Against that backdrop, on 2 May 2026 the failed Spirit Airlines became another kind of feedstock. Google won a bankruptcy auction for corporate data at $10 million — on the order of 100 million emails and about 500 million Microsoft Teams messages plus internal documents; the company says it does not receive customer personal data.

Separately, Anthropic. In July 2026 a San Francisco federal court approved a $1.5 billion settlement of a publishers’ and authors’ suit over books taken from torrents — about $3,000 per work named in the deal.

From the same case came Project Panama: since early 2024 Anthropic has bought physical books, scanned them for Claude training, and destroyed the originals — millions of used volumes. A clean chain of title became more expensive than a grey downloadable archive.

Why it matters

The market is no longer only arguing about “can we scrape the open web?”

Paid APIs, press licenses, and bankruptcy archives point to a new normal: provenance is an asset, like iron or coal at a plant. When the open stream thins, buyers purchase what is not public — company mail, internal databases, scanned books with a clear chain of title.

The Epoch AI forecast and the Spirit Airlines deal read as one story: public text is finite, and corporate archives are the next layer. Law and courts catch up not with slogans but with numbers in settlement agreements.

For builders this changes how you talk about your own product. If you hoard logs, tickets, or document dumps “for our future model,” that is no longer a free byproduct — it is an asset with legal and commercial weight.

Conversely, someone else’s corpus without a license increasingly ends not in a grey zone but in a billion-dollar bill, as in the Anthropic case. Provenance and quality start to cost more than raw volume.

In practice

This is a map of other people’s deals, not a price list for your dataset. Read it as a risk chart if you train or fine-tune, buy corpora, or simply store user content.

  1. Record provenance first: license, contract, export source — before a corpus enters training.
  2. Do not plan only on “whatever the open API still exposes”: Reddit showed how fast access becomes paid.
  3. Separate public web, licensed media, and internal corporate archives — each has its own legal regime.
  4. If you keep user text “for a future model,” describe consent and purpose up front — or the asset becomes a lawsuit.
  5. Watch court frames like Anthropic’s: “we downloaded torrents but did not train on them” no longer reads as a safe public excuse.

Takeaway

In a few years a forum comment, a newspaper article, and a stack of used books stopped being “just content” — buyers pay tens and hundreds of millions for them, and copyright breaches can already cost billions.

The loop closed: everyday digital traces became feedstock without which the language-model plant does not run. The next “meaningless” comment you leave is, in effect, an asset giants will write checks for.

Comments

Loading comments…