Kyle Siler
This is a large-scale survey study that quantitatively identifies the current status and social disparities of Large Language Model (LLM) adoption in academia by analyzing the texts of 7.3 million scholarly articles.
There was a lack of large-scale empirical evidence on how Large Language Models (LLMs) are changing the production methods of academic research and how their adoption is distributed within the academic community.
The study analyzed 7.3 million articles from four major publishers between 2020-2025. It developed a set of 228 focal words exhibiting sharp post-2022 frequency increases consistent with LLM output to score each article's level of LLM influence. Difference-in-differences (DID) models were used to analyze variations in LLM-associated language use across regions, institutional ranks, publishers, disciplines, and journal tiers.
By 2025, an estimated 57% of published articles showed evidence of LLM influence, a significant increase from 12% in 2023. Adoption rates varied distinctly by economic development, proximity to English, institutional prestige, publisher type, and academic field, with lower-ranked institutions or for-profit publishers showing higher rates. This study demonstrates that LLM adoption in academic writing is pervasive but socially stratified, contributing to the understanding of the essential social dynamics for governing the evolving relationship between AI and academic knowledge production.