Non-recurring: publishers and the AI licensing economy
What public filings and reported deals reveal about the money publishers receive when they license their archives to AI companies.

On 30 July 2026 Informa, which owns the academic publisher Taylor & Francis, published its results for the first half of the year. Its Academic Markets division is presented with and without non-recurring data contracts. The comparison matters because a division's growth rate is not the group's growth rate.
Underlying revenue growth, excluding non-recurring data contracts: up 5.4%.
Underlying revenue, unadjusted for those contracts: down 4.4%.

The gap between those two lines is money that artificial-intelligence companies paid for access to an academic archive. In 2024 it was more than $75 million — enough that Informa's own annual report describes the levels as "exceptional" and notes the group's earnings growth was 16% "absent FX and non-recurring data contracts". Two years later the same money is the reason a division looks like it shrank.
The comparison exposes a central question in the licensing economy: how much of an announced payment turns into repeatable revenue? One company’s accounting does not describe the whole market, but it gives a concrete place to start.
What has actually been paid

There are, on the most careful public count, about 150 agreements between news publishers and AI companies. The Tow Center for Digital Journalism tracks them along with the lawsuits and the grants: 207 items in total, of which 151 are deals.
Almost none of them has a disclosed value.
The table distinguishes amounts disclosed by a party from values reported by journalists. Two entries are labelled disclosed; one of those is a grant rather than a licence. This is a reviewed sample, not a census of every AI agreement.
| Parties | Announced | Reported value | Basis |
|---|---|---|---|
| Informa / Taylor & Francis – Microsoft | May 2024 | $10m initial plus three annual payments | disclosed, in a regulatory announcement |
| OpenAI + Microsoft – Lenfest Institute | Oct 2024 | $10m — a grant, not a licence | disclosed |
| News Corp – OpenAI | May 2024 | $250m over five years | reported estimate |
| News Corp – Meta | Mar 2026 | up to $50m a year | reported estimate |
| New York Times – Amazon | May 2025 | $20–25m a year | reported estimate |
| Reddit – Google | Feb 2024 | $60m a year | reported estimate |
| Dotdash Meredith – OpenAI | May 2024 | at least $16m a year | reported estimate, from a parent's filing |
| Axel Springer – OpenAI | Dec 2023 | $25–30m over three years | reported estimate, and contested |
| Associated Press, Financial Times, Le Monde, Guardian, Washington Post, Condé Nast, Hearst, Vox, The Atlantic, TIME, Axios, Schibsted, Gannett, Getty, Nine and about thirty more | 2023–2026 | not disclosed | — |
Most headline prices in this table are reported estimates rather than separately disclosed contract values. That makes them useful context, but they cannot be treated as an audited total for the market. Differing estimates of the same agreement should not be added together.
The first finding is a disclosure gap: much of the pricing of journalism for AI remains outside public accounts. A company can disclose licensing revenue without disclosing the value or terms of an individual agreement.
Where it does appear in the accounts

A handful of companies do report it, and their filings are the only hard evidence of what this business is. The clearest is Shutterstock, which sells images rather than journalism but breaks out the relevant line every quarter.
In 2025 its Data, Distribution and Services revenue was $203.3 million, 21% of the company's total, up 16% on the year. On that basis it is one of the great success stories of the licensing era.
The quarterly pattern changed from growth in Q3 2025 to declines in Q4 and early 2026. Q1 2026 Data, Distribution and Services revenue fell 47%, to $21.0 million. The company attributes fluctuations in this line partly to the timing of metadata-licence delivery. That is evidence of timing-sensitive revenue, not proof that an archive can be sold only once.
Shutterstock also discloses about $35.3 million of contracted but unsatisfied data-related obligations, to be recognised over five years. A straight-line division is not company guidance: actual recognition depends on the contracts and delivery schedules.
The pattern repeats wherever anyone discloses. Thomson Reuters booked $18 million of generative-AI content licensing in a single strong quarter at the end of 2023, and by 2025 was explaining that revenues were partly offset by "lower generative AI related transactional content licensing revenue". Reddit's other revenue — which includes content licensing and other direct sales — reached $140.0 million in 2025, a real number and 6.4% of the company; but of the obligations remaining on the books at mid-2026, only a third fall in 2027.
CuriosityStream’s Q2 2026 filing presents a different mix: licensing contributed about 61% of revenue. This includes traditional content licensing, AI-related arrangements and barter. The same quarter included $5.3 million in barter licence fees, where content was exchanged for content. The licensing share cannot therefore be read as the share of cash revenue from AI.
And then there is News Corp, whose chief executive has been the most quotable advocate of getting paid. Its FY2026 annual report mentions "higher content licensing revenues" twice, folded into circulation and subscription lines, and breaks out no AI licensing figure at all. The company with the largest reported deal in the industry does not report it.

What it comes to per article
Two litigation-related figures offer a different comparison from licence prices. Two numbers exist that put a price on a single piece of work, and both come from litigation rather than from a negotiation.
In its claim against Perplexity, Yomiuri Shimbun asks for ¥16,500 per article across 119,467 articles. And the Anthropic settlement, finally approved on 20 July 2026, values a book at $3,000 — a $1.5 billion fund, a settlement concerning pirated books, not a market price for lawfully licensed training material.
The other way to size it is against the business. Reporting on the Amazon–New York Times arrangement noted that the payments equal almost 1% of the Times' total revenue. The same arithmetic on Dotdash Meredith's deal — $16 million against revenues over $1.5 billion — gives about 1%. Axel Springer's, well under 1%.
A payment of around one per cent of revenue can still matter. These comparisons alone do not establish recurring income, permanent transfer of rights or the profitability of a licence.
Who is not in the room

About 150 deals exist. There are several thousand news publishers.
The Reuters Institute asked publishers directly. In its 2025 survey of 326 leaders across 51 countries, the majority said they had no deal and did not expect to be offered one, and 72% wanted collective bargaining. In the 2026 survey only 20% expected significant income from AI licensing — and another 20% expected none at all.
The surveys indicate uneven access to licensing income. They do not establish that every small, local, independent or non-English publisher has been excluded, or identify why each publisher lacks a deal.
The refusal, and what my own measurement found

The publishers who cannot sell can still refuse, and increasingly do. A live count of 1,154 news publishers run by an open-source project found, on the day this was written, that 626 of them — 54.2% — instruct OpenAI, Google's AI crawler or Common Crawl to stay out. Among the top 100 UK and US news sites, 79% block AI training bots. In February 2024 the equivalent figure for 150 top sites was 48%, and the researchers noted then that none of the sites they examined had reversed course.
The Guardian blocked OpenAI in September 2023 and announced an OpenAI partnership in February 2025. This shows that a publisher’s relationship can change; it does not establish that every blocking decision is a negotiating tactic.
My own measurement of the wider web puts a number on how precisely that bid is now being made. In a sample of 4,762 sites with a readable robots.txt, drawn from the 15.6 million domains in the web collection in Prism and fetched on 27 August 2026, 576 sites refuse a crawler that collects training data while allowing every user-request crawler included in the measurement. And of the 708 sites that shut out GPTBot, the number that also shut out Google's search crawler is 0.
The pattern is consistent with distinguishing search discovery from automated reuse. Robots.txt records a site operator’s instructions; this measurement does not establish whether each crawler complied or determine the legal effect of those instructions.
The case for taking the money
It is stronger than the sceptics allow.
The public accounts show that data licensing can be material to some businesses. Shutterstock’s 2025 Data, Distribution and Services line accounted for roughly a fifth of its revenue, although the category is broader than training-data licensing alone. Whether that income recurs depends on the contract, delivery obligations and recognition schedule.
There is a separate distribution question: does a licensing agreement bring qualified visits or subscribers as well as a payment? A useful comparison would track licensed and unlicensed publishers with similar size and audiences before and after an agreement. An observed referral difference alone cannot identify the effect of the deal.
The legal record also needs careful scope. In Bartz v. Anthropic, the district court distinguished the training use before it from the acquisition and retention of pirated books. The final settlement approved in July 2026 concerns past acquisition and copying of listed works; future conduct and output claims remain outside its release. Neither the opinion nor the settlement supplies a blanket answer for every publisher, dataset or use.

Prism's measured full-quarter news series contains 1,973 articles mentioning copyright lawsuits and 5,046 mentioning training data. These keyword searches overlap and differ in breadth. Their relative volume cannot establish the legal strength of a claim or the size of the licensing market.
Why the market stalled

The most economical explanation of everything above came from a Nieman Lab prediction at the end of 2025. As long as Google's search crawler and its AI-training crawler function as a single system, it argued, the market for licensing journalism into AI models is effectively stalled. And the follow-up: "Every other AI company sees the same incentive structure. Why pay for something your largest competitor gets for free?""
This is one proposed explanation for the bargaining problem. Disclosure practices, contract timing, publisher scale and the uses licensed also matter. The evidence here does not isolate a single cause of the market’s development.
The filings show a mixed market. Some payments are non-recurring or timing-sensitive; some businesses report continuing AI-related revenue. An agreement count, a headline deal estimate and a recognised revenue line answer different questions.
For related company discovery, explore digital publishing platform companies in India. Directory inclusion does not establish a company’s archive size, licensing eligibility or exposure to AI competition.
The useful next measurement is repeatable revenue: what a publisher recognises from licensing, whether it recurs, and what rights and delivery obligations sit behind it.
How this was put together. Every money figure above is labelled by where it comes from — a filing, a company statement, or a value reported by journalists — and the distinction is kept in the table rather than smoothed away, because the table separates disclosed values from reported estimates, and identifies the disclosed grant separately. Charts plot filed figures only. The robots.txt measurement is our own, run over Fliar Prism's web collection: 6,643 domains sampled and fetched on 27 August 2026, of which 4,762 returned a readable file; every share quoted is of those. Prism was also used to find the coverage and the studies and to measure the news series, over its collections on 2026-08-27; those counts run to 1 June 2026 and are floors. Quotations were checked against the pages they came from.
Sources and further reading
Sources are also linked next to the claims they support. The original research edition is dated 27 Aug 2026.
- more than $75 millionwww.informa.com
- The Tow Center for Digital Journalism tracks themtow.cjr.org
- $203.3 million, 21% of the company's totalinvestor.shutterstock.com
- Q1 2026 Data, Distribution and Services revenue fell 47%, to $21.0 millioninvestor.shutterstock.com
- about $35.3 million of contracted but unsatisfied data-related obligationswww.sec.gov
- $18 million of generative-AI content licensingwww.sec.gov
- "lower generative AI related transactional content licensing revenue"www.sec.gov
- other revenue — which includes content licensing and other direct saleswww.sec.gov
- CuriosityStream’s Q2 2026 filingwww.sec.gov
- FY2026 annual reportwww.sec.gov
- the Anthropic settlement, finally approved on 20 July 2026storage.courtlistener.com
- A live count of 1,154 news publisherspalewi.re
- Among the top 100 UK and US news siteswww.buzzstream.com
- blocked OpenAI in September 2023www.theguardian.com
- announced an OpenAI partnership in February 2025www.theguardian.com
- Bartz v. Anthropicstorage.courtlistener.com
- final settlement approved in July 2026authorsguild.org
- the market for licensing journalism into AI models is effectively stalledweb.archive.org