← Back to NASA Technology Projects
AMP: An Automated Metadata Pipeline (AMP)
Completed
TRL 5 (started at 3, targeting 5)
Description
A core function of the AIST Analytic Center Framework is to facilitate research and analysis that uses the full spectrum of data products available in archives hosting relevant, publicly available data. Key to this is making data "FAIR" (findable, accessible, interoperable, and reusable), not just for humans, but for automated systems. Effective data discovery services and fully automated, machine-driven transactions require metadata that can be understood by both humans and machines. But such metadata are uncommon. More commonly, metadata records are inadequately contextualized, incomplete, or simply do not exist. When they do exist, they often lack the semantic underpinnings to make them meaningful. The goal of the Automated Metadata Pipeline (AMP) project is to (1) Develop a fully-automated metadata pipeline that integrates machine learning and ontologies to generate syntactically and semantically consistent metadata records that advance FAIR objectives and support Earth science research for a diverse group of stakeholders ranging from scientists to policy makers, and (2) Demonstrate the application of AMP-enhanced data to automate and substantially improve the use and reuse of NASA-hosted data in environmental/Earth systems models in the ARtificial Intelligence for Ecosystem Services (ARIES; Villa et al. 2014) platform, a distributed network of ecosystem services and Earth science models and data that relies on semantics to assemble network-available data and model components into ecosystem services models, built on demand and optimized for the context of application. ARIES is a context-aware modeling system that, given adequate, semantically-grounded metadata, is capable of finding and assessing the suitability of candidate data sets for use with a particular model, and automatically establishing linkages between data sources and model components, performing a variety of mediation and pre-processing tasks to integrate heterogeneous data sets. The AMP project aims to use machine learning techniques to auto-generate semantically consistent, variable-level metadata records for NASA data products and, in collaboration with the ARIES developer and user communities, demonstrate their value in supporting scientific research. In so doing, we hope to achieve several objectives: • Address usability and scalability issues for data providers and metadata curators in connection with tools for generating robust variable-level metadata records; • Improve the semantic interoperability of target NASA data products by linking concepts in AMP-generated metadata records to terms from well-established, external vocabularies such as the Environment Ontology (ENVO), thereby taking advantage of existing term mappings between ENVO and ontologies developed by other communities of practice, including NASA's Semantic Web for Earth and Environment Technology (SWEET) Ontology; • Demonstrate the benefits of semantically interoperable, FAIR data across communities of practice. AMP will provide a platform for auto-generating robust, FAIR-promoting, semantically consistent metadata records using neural nets to assign variables to ontological classes in the AMP Ontology. Our approach recognizes that there is a wealth of information contained within the data itself, which can be exploited to generate accurate and consistent metadata. We will work with the Goddard Space Flight Center's (GSFC) Earth Science (GES) Data and Information Services Center (DISC), which will provide access to data via its OPENDAP servers for training and testing the AMP pipeline, and for use by the ARIES platform. AMP advances the Analytic Center Framework objectives of allowing seamless integration of new and user-supplied components and data; increasing research capabilities and speed; handling large volumes of data efficiently; and providing novel data discovery tools.
Benefits
Advance Earth system science knowledge through the identification, development, and demonstration of innovative information systems technologies
Details
| Technology area | Software, Modeling, Simulation, and Information Processing > Information Processing and Artificial Intelligence |
| Program | Advanced Information Systems Technology (AIST) |
| Lead organization | LINGUA LOGICA LLC, Denver, CO |
| Start date | 2019-12-01 |
| End date | 2022-05-31 |
Project contacts
Listed on TechPort itself — the most direct way to ask about this specific project.
How to get involved
This is early/mid-stage (TRL 5) — the most realistic path in is NASA SBIR/STTR, which funds small businesses and research institutions to develop technology aligned with NASA's needs (equity-free, phased funding). Check whether a current SBIR/STTR solicitation topic overlaps with this project's technology area, or contact the project directly (above) to ask.
None of these are guaranteed paths for this specific project — TechPort itself doesn't have an "apply" button. Reaching out to the contact(s) above with a specific question is usually the fastest way to find out what's actually open.