This three-year, $399,918 project grant from the National Science Foundation's Computer and Information Science and Engineering program aims to develop flexible deep speech synthesis through gestural modeling. Awarded to Carnegie Mellon University on October 1, 2021 and set to conclude on September 30, 2024, the grant supports collaborative research between Carnegie Mellon and the University of California, Berkeley. Specifically, the universities will explicitly introduce mechanistic aspects of human speech production into acoustic modeling for text-to-speech synthesis. This interdisciplinary work seeks to advance flexible synthetic speech capabilities through a deeper understanding of the human gestural system underlying natural speech. The involvement of Berkeley further aims to develop and apply knowledge of linguistics and phonetics to strengthen the gestural basis of the acoustic models.
Generated 1/6/24, 4:26 PM