When people seek guidance from science, it’s often in search of imaginative possibilities with almost-boring reliability. And one evening in 1861, surrounded by the arena in the lecture hall of London’s Royal Institution, a 30 year old physicist changed how everyone sees the world by turning out the lights.

When James Clerk Maxwell projected the first color photograph onto a screen in front of an admiring crowd of scholars, it was one of the most elegant examples of the power of basic science. Maxwell had combined cutting-edge theories of biology and psychology with novel work on the physics of light to calculate (mathematically) how to create almost any color with red, blue, and green filters. The only problem? Despite the successful demo, scientists tried and failed for 100 years to reproduce Maxwell’s photo, and it wasn’t until after his death 30 years later that color photography was actually invented (Klein et al 2022).

In two new papers, my co-authors and I show that computational social science has failed in similar ways, and that in the rush to adopt AI, the problem is about to become much worse. In both cases, part of the problem is something that 19th century scientists struggled with too: attempts by corporations to control science.

Commercial Determinants of Replication Infeasibility

When a group of scholars announced computational social science (CSS) as a new field in 2009, they hoped that widespread data collection on human behavior by tech firms and governments would unlock important insights about humanity. To promote this new field, scientists asked readers to imagine an alternative world where self-interested tech firms would anoint “a privileged set of academic researchers” who use private data to “produce papers that cannot be critiqued or replicated” (Lazer et al 2009).

In a new paper with Cassidy Waldrip and David Lazer, we look back on CSS articles in the most prominent journals and find evidence for this dystopic outcome in the first two decades of the field. In an analysis of all 187 articles involving tech platforms in Science, Nature, and PNAS from 2004 into 2025, only 26% of those studies could be repeated today. The reason? Corporate control over data access.

Only 26% of high profile computational social science studies can be repeated today. The reason? Corporate control over data access.

This paper makes two contributions beyond this simple and devastating count: first, it applies a wider discussion in science and especially public health about how corporations put their thumb on the scale of public knowledge. This point is illustrated very clearly when we compare studies that involved commercial entities and studies that involved Wikipedia, which is a nonprofit. Zero of the papers that required special permission can be repeated; 32% of studies using corporate APIs are repeatable. And Wikipedia? 60% of studies using Wikipedia can be repeated, with the barrier to replication in the remaining 40% often being a proprietary dataset that the authors added independently of Wikipedia.

Counts and percentages of articles that were replicable in the population of all social and behavioral articles involving technology platforms published in Nature, Science, and PNAS between 2004 and 2025.

I expect some people will read different stories from this chart. Some will celebrate over a hundred studies that would not have been possible without careful negotiation by extraordinarily persistent-scholars inside and outside of companies. The problem? Unlike scientists who tried and failed to replicate Maxwell’s color photography, it’s not possible to find out if many of these findings would repeat.

I hope this paper is a wake-up call for scientists – especially early career scholars – who are deciding where to put their energy. While data from industry collaborations can seem like a fast-track to great science, scholars who care about repeatable knowledge should consider their options:

Proprietary LLMs impede scientific transparency and reproducibility

In the rush to apply generative AI to every part of the scientific process, social scientists are rushing headlong into a repeat of the failures of computational social science. That’s what Killian McLoughlin, Alexis Palmer, Molly Crockett, and I argue in a new article in Nature Reviews Psychology.

When scientists use closed-weight, proprietary models (such as OpenAI’s GPT, Anthropic’s Claude and Google’s Gemini), they make the scientific process subject to the whims of corporations that aren’t incentivized to prioritize science. The resulting analyses often fail to meet fundamental scientific standards of transparency and reproducibility.

I first ran into this problem in 2017 when Jigsaw, a Google-related thinktank, launched “Perspective API” as a prototype in open source content moderation algorithms. During an early preview, I was excited by the conversation it could start and worried about the risks of using it for science. The system’s definitions had no scientific backing, no reliability measures, and it was constantly changing without versioning. A decade later, as Google makes plans to shut down the API, a group of scientists have admitted that “the tool was never fit for the central role it came to play, with consequences for validity, reliability, and reproducibility” (Hartmann et al 2026). Thousands of papers are affected by this house of cards.

LLMs make the scientific process subject to the whims of corporations that aren’t incentivized to prioritize science

The same thing is happening to research involving LLMs. As Alexis has shown in a study on using LLMs for classification tasks: the same prompt applied to the same data at two different times can yield different results. In some cases, inexplicable changes in model behavior can change the findings of a paper. Ouch.

Why does this happen? We compare science at moments of technical upheaval to a “panicked gold rush” where people feel like they must compete with others to win an early-adopter competition. Just like the California mining industry in the 19th century, this rush creates externalities – pollution and ecological damage to science that can take decades to untangle (Isenberg 2026).

What can we do about this? Solving the problems of proprietary LLMs will require similar work on regulation, economics, ethics, and technology production as computational social science. But hopefully we’re smarter and wiser this time around? Some of the recommendations in our article include:

  • Develop and use open models on hardware you control (we offer some advice on this)
  • Create reproducible frameworks for LLM research (Killian has a stealth project in this area!)
  • Standardize reporting practices – projects like GuideLLM are already working to create those standards

Where to from here?

Speaking for myself, I truly hope these articles help spur a constructive if challenging conversation about how to sustain scientific inquiry at a time when the rush to embrace AI has a flavor of panic to it. Compared to 2024, last year saw a 33% decline of faculty positions in computer science. Given that AI is one the greatest growth areas in CS and elsewhere, I can see why people would feel extreme pressure to move fast and break things. Conversations about the reliability of knowledge can get personal very quickly, devolving into debates about what counts as science, moral purity, and what strategies should get someone a job.

I don’t think that categorical imperatives are useful here. But those of us who want something to work more than once or twice clearly have some work to do, and I hope these papers help folks find others who are working on solutions.

When I was a student in the humanities, one of the most riveting lectures I attended was a talk by a poetry lecturer about the difference between mastery and understanding. When Maxwell made his ground-breaking if unrepeatable experiment, it was a technical failure on a sound theoretical basis. Maxwell’s success was incomplete because he wanted to understand the world, not just master it. The companies that seek mastery often have a relationship of convenience with science, and their grinding momentum will lead them to disrupt the search for understanding if that’s what it takes to follow their incentives.

When people turn to scientists for answers, it’s often out of a yearning to understand the world, make sense of our place in it, and change things for the better. Repeatable science is only one tool for achieving those ends, but when we invoke it, we need to deliver on that promise. I hope these two articles prompt creative efforts to advance that search in the presence of the commercial forces that make science less reliable.

References

I am grateful to Daniel Gasiencia, who made the header photo available under a Creative Commons License.