Do researchers actually reuse publicly available data and code? For years, this question sat at the heart of open science advocacy, yet answers were largely anecdotal. Despite strong policy signals and growing expectations around data sharing, there was no reliable, scalable way to measure whether shared datasets were being reused—or what impact that reuse had on subsequent research.
That has now changed. Developed by DataSeer in partnership with their Open Science Indicator partner PLOS (with input from the broader Open Science community) and the Michael J. Fox Foundation (MJFF), the new AI technology enables large-scale, accurate detection of data reuse across the scholarly literature. The Data Reuse LLM represents a major step forward in making data a measurable, first-class research output, whose influence can be tracked beyond the paper in which it first appeared.
Why detecting data reuse is so difficult
One of the core arguments for open research data is that reuse accelerates discovery, reduces duplication, and amplifies the return on research investment. But these benefits depend on researchers actually finding, understanding, and reusing shared datasets. Historically, whether — and how often — that happens has been almost impossible to determine.
Traditional approaches to measuring reuse rely on formal data citations, typically DOIs from known repositories appearing in reference lists. In practice, this captures only a fraction of real-world reuse. Researchers often reference data indirectly: by citing accession numbers, including URLs in the text, naming repositories without identifiers, or describing reused datasets narratively. In other cases, data may be shared informally or privately between research groups. As a result, most data reuse has remained invisible to conventional citation-based metrics.
For MJFF, overcoming this invisibility was a priority. After piloting progressive open science policies through the Aligning Science Across Parkinson’s (ASAP) initiative, MJFF began integrating these practices across its broader research network. Since 2022, MJFF has championed open science ecosystem-wide, reinforcing its commitment to accelerate Parkinson’s research by maximizing reuse and downstream impact. But without robust measurement, it was difficult to assess whether those policies were translating into meaningful changes in research behavior and outcomes. MJFF partnered with DataSeer to address this gap.
A new way to see data reuse that was always there
The result is the Data Reuse LLM: an LLM-based system designed to identify whether data and code in a research article were newly generated or reused, and – critically – to trace reused datasets back to their original sources. Unlike traditional methods, the model does not depend on the presence of formal data citations. Instead, it analyzes the full text of articles to detect reuse even when references are incomplete, indirect, or entirely unstructured.
This new approach allows reuse to be identified across a wide range of citation practices, including accession numbers, repository names, URLs, and descriptive mentions in the text. Rather than requiring perfect compliance with data citation standards, the system infers open science actions from information authors already provide.

Applied to a corpus of 6,000 MJFF-funded articles published between 2005 and 2025, the Data Reuse LLM revealed a much richer picture of reuse than was previously visible. Evidence of dataset reuse was found in 43% of articles—far higher than the roughly 2% typically detected using identifier-based approaches alone. The model successfully identified 79% of reuse instances, compared with 59% using traditional methods.
These results suggest that data reuse has been more widespread than previously captured by conventional citation-based metrics.
Measuring data impact beyond the paper
Beyond detecting reuse, the Data Reuse LLM makes it possible to assess research impact in a fundamentally new way. By tracing how datasets are reused across successive publications, the system can quantify the influence of data itself—not just the articles that first report it.
For example, in the MJFF-funded article “Cross species systems biology discovers glial DDR2, STOM, and KANK2 as therapeutic targets in progressive supranuclear palsy,” the model identified six generated datasets (the first row of red dots in the figure below). Two of these datasets went on to be reused in 16 subsequent studies (second and third row of blue dots) within the MJFF corpus, with one dataset proving especially influential.

Patterns like this reveal an important asymmetry: a small number of datasets account for a disproportionate share of downstream impact. In this case, four of the six initially generated datasets were not reused within the corpus, while one dataset was reused over ten times. These dynamics are largely invisible when impact is measured solely at the article level.
From measurement to insight
Measuring data reuse at scale provides concrete evidence for the effectiveness of open science policies and investments. In the MJFF corpus, year-by-year analysis shows clear improvements following the foundation’s recent policy push. Between 2022 and 2024, the proportion of articles reusing data increased from 72% to 86%, sharing of newly generated data rose from 43% to 55%, and code sharing increased from 24% to 38%.

But the implications go beyond evaluation. Reliable reuse metrics raise new questions: What distinguishes datasets that are widely reused from those that are not? Can early signals of reuse help funders identify especially high-impact outputs? How might these insights inform study design, data management practices, or funding strategy?
As the Data Reuse LLM continues to evolve, DataSeer expects further improvements in sensitivity, accuracy, and transparency. Ultimately, this technology has the potential to support routine, large-scale assessments of data impact across disciplines — bringing long-overdue visibility to how shared data drives scientific progress.




