The NSCORE Data Lake will be a shared computational environment where working groups can bring together the data needed to address their research questions. Rather than serving as a comprehensive database or permanent repository, it will support the temporary staging, organization, and analysis of selected datasets from established repositories, published sources, and participating researchers.
The contents of the Data Lake will develop in response to the needs of NSCORE projects. Depending on the question, a working group might use organismal trait measurements, molecular or behavioral data, images or recordings, biodiversity records, or environmental observations. These examples illustrate the range of data that could be used; they do not represent a fixed collection that NSCORE promises to maintain.
NSCORE will develop tools to help researchers document, find, connect, and prepare these varied datasets for analysis while tracking their sources, transformations, and access conditions. When projects conclude, eligible data products will be deposited in appropriate public repositories rather than permanently housed in the Data Lake.
This project is supported by the U.S. National Science Foundation under Cooperative Agreement 2438843. Any opinions, findings and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the National Science Foundation.