This four-year, $400,000 project grant from the National Science Foundation's Computer and Information Science and Engineering program aims to develop new techniques for 3D scene understanding. Specifically, the Stanford University researchers will explore implicit neural representations to model scene structure and details from raw images and videos. They will integrate findings into course development and partner with organizations to teach artificial intelligence, computer vision, and graphics to underrepresented students. Key innovations include new scene representations that bridge neural and physical models; methods to infer representations from data in a self-supervised manner; and leveraging the representation for downstream tasks like interactive design and scene synthesis. The work supports the NSF's mission to advance computing and information science through investigator-initiated research and education.