Scientific breakthroughs are built on insights from good data, but sometimes they are hidden beneath an ocean of too much of it. This is precisely the problem the California Institute for Regenerative Medicine (CIRM) has decided to tackle by approving a new infrastructure program.
CIRM recently approved investments in the creation of tools that will help make sense of the vast and complex biological datasets generated during institute-funded research. The program is aimed at developing new software to integrate regenerative medicine data and extract new insights from it.
“Our mission drives us forward,” noted Jenna Bayram, CIRM Senior Scientific Officer, during a board meeting. “We are developing tools and strategies to maximize the value of CIRM’s investments, advancing science and creating new ways to work with data.”
The new program allocates 10 million dollars for 15–20 grants of 500 thousand dollars each. Each project will address a specific data integration challenge. Some tools will link different types of data, while others will improve or support widely used open-source software adapted for the needs of regenerative medicine. A portion of the projects will update existing solutions to ensure they meet modern computing and artificial intelligence capabilities.
The Data Science and Software Engineering Awards are aimed at creating open-source tools that integrate various types of biomedical data. Connecting this data in a user-friendly way is intended by CIRM to accelerate the identification and validation of new drug targets and biomarkers for treating diseases.
This is about creating an infrastructure that will allow researchers to work with datasets more quickly. Particular attention is paid to software capable of translating the diversity of data generated in regenerative medicine research and clinical trials into an easily exchangeable format.
Several years ago, the NIH launched a similar initiative—the Biomedical Data Translator program—to integrate petabytes of biomedical data from federal research. CIRM is also building on its own previous investments in data sharing, including the Data Explorer tool, which already contains more than 800 CIRM datasets and makes them more accessible to researchers.
The problem of making sense of complex data has long faced government research. As far back as 2014, former NIH Director Francis Collins warned Congress that the growing wave of data could overwhelm researchers. Since then, the scale and complexity of data have only grown. CIRM projects and clinical trials have already produced more than 1200 datasets, intensifying the need for software tools to connect them.
Some tools already combine different types of biomedical data; however, many open-source systems require updates to keep pace with modern computing and AI. New or improved software better meets the needs of regenerative medicine researchers. The goal of integrating different types of data is to accelerate research and maximize CIRM's initial investments, Bayram emphasized.
Open, transparent, reusable, and extensible software can provide the infrastructure needed to overcome these barriers. CIRM requires applicants to think not only about creating a tool but also about its deployment, use, support, and improvement over time. The institute is looking for projects with broad impact—tools that work across different systems, data types, and research contexts.
Data science evolves rapidly, and even popular tools need regular updates. This investment helps researchers keep pace with progress in computing and AI, as well as the constantly changing landscape of new research data. The open-source solution developers who receive support will be able to turn fragmented data into meaningful discoveries.

