Publishers Keen To Continue Controlling Knowledge

A computer circuit board with a brain on it

​Last week, OpenAI announced open access for 100,000 STEM researchers at 500 universities around the world. But when a medical scientist at one of the ten Aus universities that made the cut filled in the form, he did not qualify. Problem was, his publications are in rated journals, while OpenAI wants a track-record in pre-prints.

It looks like a pawing jab in the world championship bout for open-ish access to research. And both contenders are prize-fighters, in it for the money. The world champion is the giant for-profit publishers who are keen on LLMs trained on their journals. Just so long as the AI and the articles it analyses are behind their paywalls, where they continue to monetise academics’ work that they do not have to pay for.

The contender is the AI developers, who want all the world’s research accessed by their models, without their having paid for it either. And even as publishers switch from pay to read to pay to publish journals, they are extending their business models with their own custom AIS and data analytics of their own content.

It is a fight neither wants to draw, but whatever result, the loser will continue to be researchers who do not have access to every publisher’s AI and who use open domain versions of research. Them and the taxpayers who (occasionally) fund them.

The OA for all research to knowledge argument does not work in engineering, chemistry and materials sciences where work with industry on specific problems/product development means scientists do not want their work to be available everywhere. Meanwhile medical research behind paywalls restricts more work on an article’s findings to subscribers of the journal. It also prevents curious researchers replicating experiments to see if they get the same results.

Still, across the vast steppe of scholarship, OA things could be worse. The EU forced publishers to adjust content pricing in part from pay to read to pay to publish (“article processing charges”). OA articles from big five publisher journals are now estimated (by Google Gemini) to represent 25% to 50% of the whole. And that is a huge amount of knowledge for LLMs to process.

Except, publishers want to stop them scraping free to read research; arguing that means LLMs can be “infinite substitution machines,” that take copyright content and change it not-much, so it can stand as a literal alternative to the original. And the models know how to make credible their simulacrum of a scholarly paper, because they have learnt from so many of them. “Journal articles are uniquely useful to LLM development,” the Association of American Publishers argues in a US court case against Meta, “they form a highly curated, professional, trusted, and authoritative system of the expression of scientific research.”

The publishers apparently hate the idea of anybody getting research of great community value for nothing – excepting them.

Share:

Facebook
Twitter
Pinterest
LinkedIn

Sign Up for Our Newsletter

Subscribe to us to always stay in touch with us and get latest news, insights, jobs and events!!