Approach

Two obstacles commonly stand between organismal biologists and broad comparative work: the effort required to find, combine, and clean disparate datasets, and the scarcity of accessible computational expertise to analyze them well. NSCORE addresses both through a single scientific environment that helps working groups be successful in their goals:

  • Source: Stage datasets from established repositories, published databases, and researchers' own collections into a common data lake.
  • Integrate & Curate: Use AI, including large language models, to resolve structural, syntactic, and semantic problems in tabular data (including metadata), and to extract measurable traits from images, audio, and video. Curation happens on demand rather than on upload, when a dataset is needed for analysis.
  • Search: Find and rank the datasets and attributes most likely to improve a given model or answer a given question.
  • Analyze: Apply causal inference, machine learning, generative AI, agentic AI, and explainable AI to linked datasets, with tooling designed for biologists rather than machine learning specialists.
  • Govern: Track provenance, lineage, and access policy for both raw and derived data so that analyses are open, reproducible, and consistent with the terms data providers set.
Logo

 

This project is supported by the U.S. National Science Foundation under Cooperative Agreement 2438843. Any opinions, findings and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the National Science Foundation.