Project Grant 2230817

Award Date 8/15/22
Completion Date 7/31/26
Dollars Obligated $150K
Funding Federal Agency
National Science Foundation
Federal Grant Program
47.070
Assistance Type
Project Grant
Place of Performance
Urbana, IL 61801, USA
Similar Awards
This $449,275 federal Project Grant award from the National Science Foundation's (NSF) Social, Behavioral, and Economic Sciences (CFDA 47.075) program supports research to advance the scientific understanding of human language processing and change. The project, led by the University of California, Davis (UC Davis), combines expertise from linguistics, psychology, and computational approaches to investigate how language users integrate general language constraints and verb-specific preferences...
The National Science Foundation (NSF) awarded a $174,967 EAGER (Early-concept Grants for Exploratory Research) Project Grant under the Computer and Information Science and Engineering (CISE) program (CFDA 47.070) to the Trustees of Indiana University. This 2-year project, running from October 1, 2024 to September 30, 2026, aims to study how well machine language tools can grasp cultural nuances to enhance cross-cultural communication. The project will focus on developing a new knowledge base...
This Project Grant award, provided by the National Science Foundation's (NSF) Social, Behavioral, and Economic Sciences (SBE) program (CFDA 47.075), supports research into the factors that influence how humans speak and understand language. The $407,738 award to Vanderbilt University will investigate the impact of an individual's daily language exposure, cognitive abilities, and reading habits on their preferences and performance when processing complex English sentence structures. The project...
The National Science Foundation awarded a $299,569 EAGER (EArly-concept Grants for Exploratory Research) Project Grant to the University of Rochester to develop an approach that leverages large language models and artificial intelligence to accelerate the discovery of earth-abundant, active, and selective catalysts for the reverse water-gas shift reaction. The project aims to demonstrate that language-based representations can be universally applied to materials discovery processes expressed...
This Project Grant award from the National Science Foundation's Social, Behavioral, and Economic Sciences (CFDA 47.075) program provides $204,593 to The MGH Institute Of Health Professions, Inc. to develop a "challenge suite" - a dataset and evaluation tool - to systematically compare mathematical theories of how humans learn language through unsupervised pattern detection. The goal is to assemble a comprehensive set of experiments on adult statistical word segmentation that any viable...
This Project Grant award from the National Science Foundation's (NSF) Social, Behavioral, and Economic Sciences (SBE) program (CFDA 47.075) provides $160,000 in funding to support an early career scientist developing a framework for process-level models of human language comprehension. The goal is to bridge the gap between human language behavior and computational models by analyzing incremental language processing behavior, particularly in response to imperfect input, and implementing models to...
This Project Grant award of $334,731 from the National Science Foundation's (NSF) Computer and Information Science and Engineering (CISE) program (CFDA 47.070) will support research at George Mason University (GMU) to develop language technologies for data-scarce languages. The key objectives are to: Build datasets to explore the meta-linguistic capabilities of large language models. Leverage grammars to build large language models for new languages, similar to how humans learn second languages....
This Project Grant award from the National Science Foundation (NSF) under the Social, Behavioral, and Economic Sciences (CFDA 47.075) program provides $104,114 to Haverford College to conduct research on the mechanisms underlying human language learning. The research aims to better understand how people construct mental representations of words and learn the morphological and phonological patterns that govern language. The project will develop interpretable computational models and algorithms...
This National Science Foundation project grant award of $116,990 supports research into deconstructing wordlikeness judgments through interconnected experimental and computational studies. Funded under the Social, Behavioral, and Economic Sciences program (CFDA 47.075), the award will provide research training and mentoring opportunities for graduate students in modern experimental and computational methods related to phonotactic knowledge and language sciences. Specifically, the City University...
This Project Grant award of $340,217.00 from the National Science Foundation (NSF) Computer and Information Science and Engineering (CISE) program (CFDA 47.070) will fund the development of innovative tools based on AI large language models (LLMs) to help individuals with limited majority-language proficiency improve their language abilities for effective information seeking. The project, led by the University of Maryland, College Park, will empower users to design their own personalized digital...

EAGER: BUILDING IDIOMATICITY INTO NATURAL LANGUAGE PROCESSING -IDIOMATIC EXPRESSIONS ARE AN ESSENTIAL COMPONENT OF EVERYDAY LANGUAGE USE AND THE HALLMARK OF NATIVE LANGUAGE ABILITY. CONSIDER THE PHRASE THROW AWAY; PROFICIENT SPEAKERS CAN EFFORTLESSLY UNDERSTAND THAT THE PHRASE TAKES A FIGURATIVE MEANING IN ?BRITAIN THREW AWAY ALL THE ACHIEVEMENTS OF THE LAST DECADE.? AND A LITERAL SENSE IN ?HE THREW AWAY HIS CIGARETTE AND BURIED HIS HEAD IN HIS ARMS.? THIS EARLY GRANT FOR EXPLORATORY RESEARCH (EAGER) WILL BUILD A HIGH-QUALITY DATASET FOR COMPUTERS TO UNDERSTAND THE DIFFERENCES BETWEEN FIGURATIVE AND LITERAL SENSES OF THESE EXPRESSIONS IN GENERAL ENGLISH TEXT. THE MAIN NOVELTY OF THIS PROJECT WILL BE IN COLLECTING A LARGE CLASS OF IDIOMATIC EXPRESSIONS AND SENTENCES CONTAINING THEM TO LET COMPUTERS LEARN THE INHERENT VARIABILITY BETWEEN A VARIETY OF IDIOMATIC PHRASES. COLLECTING MANY SENTENCES WITH PHRASES THAT HAVE A FIGURATIVE AND LITERAL MEANING WILL PERMIT COMPUTERS BETTER UNDERSTAND THE NUANCES WITH WHICH THESE EXPRESSIONS ARE USED IN EVERYDAY CONVERSATIONS AND WRITING. BEYOND UNDERSTANDING THEM, THE COLLECTED. EXAMPLES WILL HELP COMPUTERS USE THESE EXPRESSIONS LIKE NATIVE SPEAKERS DO WHEN AUTOMATICALLY WRITING TEXT AND EVEN SUGGEST APPROPRIATE EXPRESSIONS IN SPECIFIC CONTEXTS. THIS EAGER PROJECT IS ESSENTIALLY INTERDISCIPLINARY SPANNING THE AREAS OF LINGUISTICS AND COMPUTATION AND WILL INVESTIGATE NOVEL PARADIGMS FOR NATURAL LANGUAGE PROCESSING THAT ARE IDIOMATICITY-AWARE. AS SUCH, IT WILL HAVE TWO RESEARCH AIMS: (1) CREATING A HIGH-QUALITY DATASET OF PHRASAL VERBS ANNOTATED WITH THEIR CONTEXT-SPECIFIC SENSES AND THEIR LITERAL/FIGURATIVE EQUIVALENT FORMS, AND (2) TESTING THE PERFORMANCE OF STATE-OF-THE-ART IDIOMATICITY-AWARE ALGORITHMS. BECAUSE IDIOMATIC EXPRESSIONS VARY WIDELY IN FORM AND STRUCTURE, THE FOCUS ON PHRASAL VERBS (ALSO KNOWN AS VERB-PARTICLE CONSTRUCTIONS) IN THE CONTEXT OF THE EXPLORATORY PROJECT WILL PERMIT STUDYING A VERY FREQUENT CLASS OF IDIOMATIC EXPRESSIONS THAT ARE SYNTACTICALLY DIFFERENT FROM THOSE IN CURRENTLY AVAILABLE DATASETS. THE PRIMARY RISK OF THIS PROJECT STEMS FROM ITS EXPLORATORY NATURE OF CREATING LARGE CORPORA WITH SUFFICIENT COVERAGE FOR LANGUAGE MODEL TRAINING. GIVEN THEIR PREVALENCE IN NATURAL LANGUAGE, THE DATASET OF PHRASAL VERBS IN ENGLISH WILL SUPPLEMENT AVAILABLE DATASETS ON IDIOMATIC EXPRESSIONS IN TERMS OF THEIR VARIETY. MOREOVER, THEIR FIGURATIVE AND LITERAL AMBIGUITY IN CONTEXT (APART FROM THEIR POLYSEMY) WILL PERMIT A DIVERSE LOOK AT THE PHENOMENON OF NON-COMPOSITIONALITY THAT CHARACTERIZES IDIOMATIC EXPRESSIONS. THUS, THE DATASET WILL SERVE AS A TRAINING AND TEST BED FOR ALGORITHMS THAT DETECT, INTERPRET, AND GENERATE A BROAD CLASS OF IDIOMATIC EXPRESSIONS. THIS EFFORT WILL LEAD TO NEW NATURAL LANGUAGE PROCESSING ALGORITHMS FOR ACCURATE INTERPRETATION AND GENERATION OF IDIOMATIC EXPRESSIONS TOWARDS A MORE HUMAN-LIKE LANGUAGE PROCESSING ABILITY IN MACHINES. THIS AWARD REFLECTS NSF'S STATUTORY MISSION AND HAS BEEN DEEMED WORTHY OF SUPPORT THROUGH EVALUATION USING THE FOUNDATION'S INTELLECTUAL MERIT AND BROADER IMPACTS REVIEW CRITERIA.

Posted 8/1/22, 12:00 AM