This $300,000 Project Grant award from the National Science Foundation's (NSF) Computer and Information Science and Engineering (CISE) program (CFDA 47.070) aims to enhance computer vision capabilities by developing deep learning models that can better recognize and process the configural relationships between object parts. The research at Temple University will focus on teaching computers to focus on the critical components that define an object, rather than being distracted by background elements. This will be achieved by modifying transformer-based vision models to limit their attention to the most relevant image patches, and by developing a novel graph-transformer architecture for holistic shape understanding. The project has broad applications in fields like medical imaging, augmented reality, and robotics. The research will also involve students from high school to doctoral programs, with a focus on including underrepresented groups.