← Back to NASA Technology Projects
Reproducible Containers for Advancing Process-oriented Collaborative Analytics
Completed
TRL 6
Description
For science to reliably support new discoveries, its results must be reproducible. This has proved to be a severe challenge. Lack of reproducibility significantly impacts collaborative analytics, which are essential to rapidly advance process-oriented model diagnostics (PMD). As scientists move from performance-oriented metrics toward process-oriented metrics of models, routine tasks require diagnostics on analytic pipelines. These diagnostics help to understand biases, identify errors, and assess processes within the modeling and analysis framework that lead to a metric. To conduct diagnostics, scientists refer to the "same pipeline". Referring to the same software and data, however, becomes contentious---scientists iteratively tune and train pipelines with parameters, changing model and analysis settings. Often reproducibility in terms of sharing a common analytics pipeline, and methodically comparing against different datasets cannot be achieved. While tracking logs, provenance, and sufficient statistics are used, these methods remain disjoint from analysis files and only provide post hoc reproducibility. We believe a critical impediment to conducting reproducible science is the lack of software and data packaging methods which methodically encapsulate content and record associated lineage. Without such methods, deciding a common reference pipeline and scaling-up collaborative analytics, particularly in operational settings, becomes a challenge. Container technology, such as Docker, Singularity, provides content encapsulation and improves software portability and is being used for conducting reproducible science. Containers are useful for well-established and documented analysis pipelines, but, in our experience, the technology has a steep learning curve and significant overhead of use, especially for iterative, diagnostic methods. This proposal aims to establish reproducible scientific containers that are easy-to-use and are lightweight. Reproducible containers will transparently encapsulate complex, data-intensive, process-oriented model analytics, will be easy and efficient to share between collaborators, and will enable reproducibility in heterogeneous environments. Reproducible containers, developed by the PI so-far, rely on reference executions of an application to automatically containerize all necessary and sufficient dependencies associated with the application. They record application provenance and enable repeatability in different environments. Such containers have met with considerable success, demonstrating a lightweight alternative to regular containers for computational experiments conducted by individual geoscientists in the domains of Solid Earth, Hydrology, and Space Science. However, in their current form, containers, reproducible or otherwise (such as Docker), are not data-savvy---they are oblivious to spatio-temporal semantics of data and either include all data used by an application or exclude it entirely. When all data is included, containers become bloated; alternatively when excluded, they cause network contention at the virtual file system. The target outcome of this project is to develop reproducible containers that are data-savvy---that is to retain their original properties of automatic containerization, provenance tracking, and repeatability guarantees, but provide ease of operation with spatio-temporal scientific data, and are efficient to share and repeat even when an application uses a large amount of data. This outcome will be achieved by (i) developing an I/O-efficient data observation layer within the container, and (ii) including spatio-temporal data harmonization methods when containers encapsulate heterogeneous datasets (ii) applying data-savvy, reproducible containers to process-oriented precipitation feature (PF) diagnostics, and (iv) finally assessing how diagnostics improve with the use of data-savvy, provenance-tracking reproducible container.
Benefits
Advance Earth system science knowledge through the Identification, develop, and demonstrate innovative information systems technologies
Details
| Technology area | Software, Modeling, Simulation, and Information Processing > Other Software, Modeling, Simulation, and Information Processing |
| Program | Advanced Information Systems Technology (AIST) |
| Lead organization | DePaul University, Chicago, IL |
| Start date | 2022-08-08 |
| End date | 2025-08-07 |
Project contacts
Listed on TechPort itself — the most direct way to ask about this specific project.
How to get involved
This is early/mid-stage (TRL 6) — the most realistic path in is NASA SBIR/STTR, which funds small businesses and research institutions to develop technology aligned with NASA's needs (equity-free, phased funding). Check whether a current SBIR/STTR solicitation topic overlaps with this project's technology area, or contact the project directly (above) to ask.
None of these are guaranteed paths for this specific project — TechPort itself doesn't have an "apply" button. Reaching out to the contact(s) above with a specific question is usually the fastest way to find out what's actually open.