This National Science Foundation (NSF) Project Grant award, under the Computer and Information Science and Engineering (CISE) Program (CFDA 47.070), provides $439,425 to the University of Utah to enable the safe deployment of learning-enabled systems that can robustly learn objectives from human feedback despite uncertainty. The key objectives are to: 1) develop approaches that provide high-confidence bounds on policy performance when learning from human input, 2) create tests to verify with high-confidence that learned rewards and policies are correct, and 3) develop techniques to ensure robustness against reward misidentification and misgeneralization during policy optimization. This 3-year project, running from August 2024 to July 2027, aims to advance capabilities for safe and reliable human-AI alignment, with potential applications in domestic robots, recommendation systems, self-driving cars, and large language models.
Generated 1/28/25, 8:34 AM