← Back to NASA Technology Projects

Causal And Reinforcement Learning (CARL) for Concepts for Ocean Worlds Life Detection Technology (COLDTech) (CARL)

Completed

Description

Exploring Ocean Worlds requires adaptation to obstacles never seen before by the system, while also conducting science operations. One solution to this problem is utilizing reinforcement learning to learn strategies to reach a goal. Although reinforcement learning can learn powerful strategies from simulation, it lacks the capability to reason about causal relationships observed in the environment, which are needed by a system to determine the source of an unanticipated problem and to adapt to circumvent the cause. Lockheed Martin (LM) has extensive experience in Deep Reinforcement Learning (DRL), having applied that technique to driving an autonomous vehicle, flying aircraft, guiding groups of missiles, and controlling a Barrett WAM arm. These applications span a wide range from low-level control to higher level strategies, in some cases combining them by using hierarchical learning. LM has extensive experience going from simulator to testbed to real world applications using Robot Operating System (ROS) and has developed deep learning techniques that enable real-time processing on low-size, weight and power (SWaP) hardware with path to flight. Our past work has demonstrated that DRL is an effective tool to develop agent behaviors for a system, without explicit a priori knowledge of the environment and its physical properties, and we will show how augmenting DRL with causal models extends its capabilities to diagnose and handle faults in the Ocean Worlds Autonomy Testbed for Exploration Research & Simulation (OceanWATERS) simulator. Therefore, LM proposes CARL, a neuro-symbolic approach using causal modeling and DRL techniques to perform higher level planning using a hierarchical reinforcement learning structure. The goal and objective of the proposed work is to apply LM’s CARL algorithm to detect, isolate, identify and classify faults even if it has never encountered them before and modify its plans to circumvent the root causes of failures to assure mission success and maximize science return. As part of the development, Jet Propulsion Laboratory and LM collaborate to develop fault-tolerant autonomous solutions for COLDArm in OceanWATERS that can specifically learn causal graphs of the arm and its environment (i.e., to detect faults) and behaviors that utilize a learned causal model that can execute science objectives with minimal interruptions (i.e., circumvent faults) and human intervention. Using CARL with COLDArm, the mission will be robust to known unknowns and will provide higher level planning to handle unknown unknowns. Given the time delay, it is necessary for a system on a remote world to be able to operate autonomously, even in the presence of known and unknown unknowns. For the former, parameters of the physical world that may not be fully known prior to landing (e.g., soil particulate size or surface friction), we will train the algorithm over a wide range of possible values, so that the agent does not overfit to one (possibly incorrect) value. Furthermore, we propose to train the agent to explicitly predict the values of these unobserved parameters based on the responses of the system to its actions. Thus, predicting these parameters will lead the agent to learn a model with a causal relationship between these parameters and the desired actions. We will incorporate expert knowledge in the causal model based on the physics of the environment, which we postulate will allow the system to more easily handle faults due to unknown unknowns in OceanWATERS. For instance, if the system captures a fault from the camera hardware whether internal or external, the causal models can be used to deduce the source of the fault and synthesize behavior to work around the fault. By incorporating causal reasoning and training a DRL agent in OceanWATERS that utilizes COLDArm, CARL addresses NASA’s exploration goals for future Ocean Worlds missions, dealing with both known and unknown unknowns.

Benefits

Developing Instrument or spacecraft technology to improve measurements for future planetary science missions

Details

Technology areaAutonomous Systems > Situational and Self-Awareness Technologies
ProgramConcepts for Ocean Worlds Life Detection Technology (COLDTech)
Lead organizationLockheed Martin Inc., Palo Alto, CA
Start date2021-06-01
End date2023-05-31

Project contacts

Listed on TechPort itself — the most direct way to ask about this specific project.

How to get involved

This is a mature technology (TRL 7+) — the realistic path in is usually NASA's Technology Transfer Program: licensing an existing NASA patent, or a Space Act Agreement to use NASA facilities/expertise directly. NASA also runs a startup licensing program with no upfront fee for companies formed to commercialize a specific NASA technology.

None of these are guaranteed paths for this specific project — TechPort itself doesn't have an "apply" button. Reaching out to the contact(s) above with a specific question is usually the fastest way to find out what's actually open.