One dollar thirty-two: the economics of training data
Follow the value chain from public writing and data annotation to licensed archives, and examine who is paid at each step.

In 2021 and 2022, OpenAI paid an outsourcing firm called Sama to have Kenyan workers read and label the worst text on the internet, so that a model could be taught to refuse to produce it. The contracts were worth about $200,000 and involved around three dozen people.
The workers took home between $1.32 and $2.00 an hour. The most junior were on a basic salary of 21,000 Kenyan shillings a month, about $170. OpenAI paid Sama $12.50 an hour for that labour — between six and nine times what the person doing it received.

The gap compares the price paid to the contractor with workers’ take-home pay. It is not a profit-margin calculation: those measures cover different costs and obligations. It also leaves another question unanswered—what arrangements, if any, applied to the people whose writing supplied the original text?
This piece compares several stages of the data economy using published figures. The examples involve different datasets, employers and years; they do not trace one sentence through a single verified supply chain.
Step one: somebody writes it, for nothing

Start with what a corpus — the body of text a model is trained on — is actually made of, because almost everyone guesses wrong. The Pile, an 825-gigabyte corpus that trained a generation of open models, is 18.11% general web crawl, 14.40% PubMed Central, 12.07% a collection of books, 10.01% pages linked from Reddit, 8.96% ArXiv, 7.59% GitHub, 6.12% American court records and 5.13% Stack Exchange. Wikipedia — the source most people assume dominates — is 1.53% of it.
The proportions differ elsewhere but the shape does not. GPT-3 drew 410 billion of its 499 billion tokens from a filtered web crawl, weighted at 60% of the training mix, with Wikipedia at 3%. LLaMA was 67% Common Crawl, 15% C4, 4.5% GitHub and 4.5% Wikipedia — though Wikipedia was passed over 2.45 times, against 1.10 for the crawl, because it is better written.
And when the Washington Post and the Allen Institute took apart Google’s C4 corpus, they found 15.7 million domains. The largest single contributor was Google’s own patent search, followed by Wikipedia and a subscription ebook service. The top thousand domains accounted for only 8% of the tokens — which is the number that matters, because it means the corpus is not a few big sites. It is the long tail of the web, in bulk, and there is no plausible mechanism by which fifteen million domains negotiate anything.
The presence of material in a training corpus does not itself establish a payment to its original authors. The datasets described above combine sources with different licences and acquisition histories; they cannot support a blanket statement that every contributor was unpaid or never gave permission.
Step two: the well runs dry

Here is what happened next, measured directly from the source rather than taken from anyone's chart. I counted the questions created on Stack Overflow in each calendar month, from the forum's own public record.
In January 2014, at the peak, people asked 188,520 questions. In January 2023, two months after ChatGPT was released, 96,684. In January 2024, 47,581. In January 2025, 18,193. In January 2026, 3,711.
In July 2026, 1,346. That is 99.3% below the peak, and 98.8% below the month ChatGPT launched.

Stack Exchange accounts for 5.13% of The Pile. The decline in public questions is consistent with changes in how developers seek help, including private AI tools, but this time series alone cannot attribute the decline to ChatGPT or show which training sets used each contribution.
Stack Overflow announced a partnership with OpenAI in May 2024. The agreement illustrates a platform licensing access to a body of user contributions. It does not by itself establish a change to the licence of every existing answer or identify payments to individual contributors.
Copying does not remove the existing archive. The separate question is whether incentives to produce new public answers are weakening. Question counts measure activity on this forum, not the supply or quality of all training data.
Step three: somebody labels it
Raw text can be used for pre-training after collection and processing. Human annotation, ranking, rewriting and expert evaluation contribute to other stages, including alignment and assessment. The rates below concern these different kinds of work.
At the bottom: Venezuelan crowdworkers on Appen, Hive Micro and Spare5 averaged "a little more than 90 cents an hour" in 2022, and by mid-2018 made up around three-quarters of those platforms' workforces. The Kenyan annotators above took home $1.32 to $2.00.
At the top, and recently: Mercor pays its expert contractors — doctors, lawyers, PhDs writing evaluation data — an average of over $85 an hour, more than $1.5 million a day in total, and passes 60 to 70% of its top-line revenue straight through to them. Approved annotators at Surge are on 30 to 40 cents a working minute, which is $18 to $24 an hour.
| Who | Rate | As stated | Year |
|---|---|---|---|
| Venezuelan crowdworkers on Appen, Hive Micro, Spare5 | about $0.90 | per hour | 2022 |
| Sama annotators working on OpenAI contracts, Kenya | $1.32–$2.00 take-home | per hour | 2021–22 |
| The same annotators, most junior | 21,000 Kenyan shillings, about $170 | per month | 2021–22 |
| What OpenAI paid Sama for that hour | $12.50 | per hour | 2021–22 |
| Approved annotators at Surge | 30–40 cents | per working minute | 2026 |
| Expert contractors at Mercor | over $85 on average | per hour | 2025 |
| Shutterstock contributors | $0.0069 median | per image, per payout round | 2023 |
| The Anthropic settlement class if you were pirated | about $3,000 | per book | 2026 |
The endpoints span very different jobs, countries, years and payment definitions. Toxic-content classification and expert evaluation both contribute to model development, but the rate comparison is not a measure of wage inflation or a like-for-like comparison of working conditions. It also does not account for unpaid waiting time, benefits, fees or local purchasing power.
The businesses in the middle are now large. Meta paid about $14.3 billion for 49% of Scale AI in June 2025, valuing it at $29 billion. Mercor raised at $10 billion four months later and reached $2 billion of annualised gross revenue by June 2026. And the one that did not adapt shows the other side: Appen’s revenue fell 14% to $234.3 million after Google terminated its contract.

Step four: somebody sells it back

The people who own an archive rather than having written it can charge for it, and a few do. Reddit’s audited accounts show its other-revenue line — content licensing plus direct sales — going from $15.2 million in 2023 to $114.7 million in 2024 to $140.0 million in 2025. Its agreement with Google was reported at roughly $60 million.
Set that against total 2025 revenue of $2.2 billion and it is 6.4% of the company: real money, not a transformation. And note who is not a party to it. Reddit sells conversations that Reddit's users wrote. These filings do not identify a distribution of this licensing revenue to the users whose conversations are included.
Where a platform has tried to share, the arithmetic is instructive. Shutterstock pays contributors a 20% average royalty on data-licensing revenue, proportionate to how much of their work is in the dataset sold. When photographers compared notes on one payout round, the median came to $0.0069 per image — seven-tenths of a cent. A thousand photographs earned about seven dollars.
This reported payout round illustrates how a payment can become small when allocated across many contributions. It is not an estimate of the value of every image, a universal compensation formula or a finding about what creators consider meaningful.

What the research literature adds, and what it does not
177 research abstracts on data annotation, crowdwork, digital labour and data provenance were read against one fixed checklist. 105 were about something else — "data labelling" collides with a large clinical and industrial literature about labelling things — leaving 72.
Among the 72 retained abstracts, the recorded classifications include 21 narrative reviews and 2 qualitative studies. Three met the checklist’s stronger-design criterion. That classification alone does not establish causal identification, and a pooled sample-size median across unlike studies is not a meaningful population estimate.
This selected literature describes annotation work from several perspectives. It is not an exhaustive review of the economics of data. The prices below combine investigations, filings, company information and a forum activity series; each has its own population and measurement limits.
The case that none of this matters much longer
There is a serious argument that parts of the market described above may change as synthetic data improves, and it comes from the researchers who first raised the alarm about running out of data.
Epoch AI estimated in 2024 that the total effective stock of human-generated public text is on the order of 300 trillion tokens, and that it would be fully used somewhere between 2026 and 2032. That finding launched a thousand pieces about the data wall. Three months later the same group modelled the constraints on scaling to 2030 and concluded that "the most binding constraints are power and chip availability" — not data.
And the substitution is already measurable. NVIDIA reports that in building one of its large models, "over 98% of data used in our model alignment process is synthetically generated". That is alignment rather than pre-training, and the distinction matters — the base model was trained on a separate pre-training dataset — but alignment is precisely the stage that the annotation industry exists to serve.
One possible outcome is that synthetic data substitutes for some routine annotation while demand for expert evaluation continues. The NVIDIA result describes one alignment pipeline; it does not measure employment effects across the industry. Whether workers lose hours, move into different tasks or receive higher rates requires separate evidence.
What it all comes to

Prism’s full-quarter series contains 5,046 articles mentioning training data, 424 mentioning data labelling and 1,090 mentioning data annotation. The searches overlap and have different breadth. They show the language used in recorded coverage, not a direct comparison of the importance of material and labour.
The examples reveal several separate payment systems: compensation for annotation, rates for specialist evaluation, licensing revenue recognised by platforms and distributions to contributors. Combining them into one chain would imply links the sources have not established. Keeping them separate makes the remaining questions clearer: who supplied the input, what rights were licensed, what work was performed, and who received the payment?
A platform’s licensing revenue does not automatically tell us how its contributors benefit. Equally, a small payout in one distribution round does not establish that all creators receive the same rate. Contract terms and allocation rules are part of the evidence still needed.
The reported Shutterstock payout illustrates the difficulty of dividing a payment across many contributions. It does not settle the merits of collective licensing, negotiated rates, different allocation rules or other compensation arrangements.
For further company discovery, explore artificial-intelligence companies in India. The directory does not establish whether each company buys, creates or uses labelled data.
The measurement in this article that will still matter in ten years is not the wage gap. It is 1,346.
How this was put together. The question-volume series is our own, counted per calendar month from the public record of the forum itself rather than taken from anyone's chart. Every wage figure carries the year and the currency its source stated, and none has been converted, because a wage converted across four years and three countries is a number with no meaning. Corpus compositions are quoted from the papers describing them; company figures come from audited filings where one exists and are labelled as reported estimates where one does not. Market-size forecasts were left out: the ones available are vendor and analyst projections whose base figures are smaller than the disclosed revenues of the companies they claim to size. Fliar Prism, a research tool, was used to find the coverage and the studies and to measure the news series, over its collections on 2026-08-27; those counts run to 1 June 2026 and are floors. 72 research abstracts on data annotation and digital labour were read against one fixed checklist. Quotations were checked against the pages they came from.
Sources and further reading
Sources are also linked next to the claims they support. The original research edition is dated 27 Aug 2026.
- The contracts were worth about $200,000time.com
- The Pilear5iv.labs.arxiv.org
- GPT-3ar5iv.labs.arxiv.org
- LLaMAar5iv.labs.arxiv.org
- took apart Google’s C4 corpuswww.washingtonpost.com
- Google’s own patent searchgigazine.net
- top thousand domains accounted for only 8% of the tokenskschaul.com
- Stack Overflow announced a partnership with OpenAI in May 2024stackoverflow.co
- Venezuelan crowdworkerswww.technologyreview.com
- Mercortechcrunch.com
- passes 60 to 70% of its top-line revenuesacra.com
- Shutterstock contributorspetapixel.com
- The Anthropic settlement classauthorsguild.org
- Meta paid about $14.3 billion for 49% of Scale AItechcrunch.com
- Appen’s revenue fell 14%www.fool.com.au
- Reddit’s audited accountswww.sec.gov
- agreement with Google was reported at roughly $60 millionwww.cbsnews.com
- Shutterstock pays contributors a 20% average royaltysubmit.shutterstock.com
- Epoch AI estimated in 2024epoch.ai
- the same group modelled the constraints on scaling to 2030epoch.ai
- "over 98% of data used in our model alignment process is synthetically generated"arxiv.org