Project Grant 2504353

Award Date 10/1/25
Completion Date 9/30/28
Dollars Obligated $675K
Federal Grant Program
47.070
Assistance Type
Project Grant
Place of Performance
Pittsburgh, PA 15213, USA
Similar Awards
This Project Grant award from the National Science Foundation (NSF) Computer and Information Science and Engineering (CISE) program (CFDA 47.070) provides $225,000 to Colorado State University to improve the code generation capabilities of large language models (LLMs). The project aims to develop novel program analysis techniques that can provide meaningful feedback to LLMs during training and decoding, enabling them to generate higher-quality, semantically-aware code. Key research activities...
This $100,000 Project Grant award from the National Science Foundation (NSF) Computer and Information Science and Engineering (CFDA 47.070) program aims to develop a holistic benchmarking infrastructure for evaluating large language models used in software engineering. The key activities include: Conducting surveys and interviews with the software engineering and machine learning research communities to gather requirements and understand barriers in evaluating large language models for code....
This National Science Foundation (NSF) award under the Computer and Information Science and Engineering (CISE) Federal Grant Program (CFDA 47.070) provides $295,169 to Carnegie Mellon University (CMU) to develop an experimental approach that leverages large language models (LLMs) and rule-based symbolic AI to generate and verify high-performance math kernels. The project aims to address the challenge of ensuring the correctness of code snippets produced by generative AI systems, which lack the...
This $875,000 Project Grant was awarded by the National Science Foundation (NSF) under the Computer and Information Science and Engineering (CISE) Federal Grant Program (CFDA 47.070) to the University of California, San Diego (UCSD). The project aims to develop new techniques for aligning large language models (LLMs) with formal specifications in order to generate high-quality computer code that provably matches user intent. Specifically, the project will: (1) develop grammar-aligned decoding...
This National Science Foundation (NSF) Computer and Information Science and Engineering (CISE) Federal Grant Program (CFDA 47.070) award to Purdue University will fund a $402,777 research project titled "CAREER: REGULARIZING LARGE LANGUAGE MODELS FOR SAFE AND RELIABLE PROGRAM GENERATION." The project aims to advance the understanding of limitations in using large language models (LLMs) for program generation and develop methods to enhance the correctness, safety, and robustness of...
This $450,000 Project Grant award from the National Science Foundation's (NSF) Computer and Information Science and Engineering (CISE) program will support research at the University of California, Riverside (UC Riverside) to explore how large language models (LLMs) can assist with program analysis. The project aims to investigate strategies for integrating LLMs with existing software analysis tools to improve their accuracy and scalability in detecting bugs and vulnerabilities in complex...
Carnegie Mellon University was awarded a $500,000 Project Grant from the National Science Foundation under the Computer and Information Science and Engineering federal grant program (CFDA 47.070). The university will deliver research and education services focused on developing programming language techniques to make Bayesian inference algorithms more usable, accessible and reliable for challenging data science problems. Specifically, the university aims to apply static analysis, type systems...
This $100,000 Project Grant was awarded by the National Science Foundation's (NSF) Computer and Information Science and Engineering (CISE) program (CFDA 47.070). The grant supports the development of an open, community-driven evaluation infrastructure to assess the safety risks of large language models (LLMs). The project aims to: (i) conduct surveys and interviews with experts to identify critical safety concerns and evaluation gaps; (ii) organize a workshop to refine the evaluation...
This National Science Foundation (NSF) $900,000 "COLLABORATIVE RESEARCH: SHF: MEDIUM: NATURAL LANGUAGE MODELS WITH EXECUTION DATA FOR SOFTWARE TESTING" Project Grant, awarded under the Computer and Information Science and Engineering program (CFDA 47.070), aims to develop natural language processing (NLP) models to simplify the development and maintenance of software tests. Key objectives include test generation, completion, update, and migration across programming languages, targeting...
This Project Grant award of $360,000.00 from the National Science Foundation (NSF) Computer and Information Science and Engineering (CISE) program (CFDA 47.070) to the Massachusetts Institute of Technology (MIT) aims to develop new methods for efficient, architecture-aware algorithms for large language models (LLMs). The goals are to make existing LLM applications more efficient, enable new applications, and broaden access to this transformative AI technology. The key focus areas include: (1)...

This Project Grant award of $675,000 was provided by the National Science Foundation (NSF) under the Computer and Information Science and Engineering (CFDA 47.070) program. The award aims to improve the ability of large language models (LLMs) to generate high-quality source code by deeply integrating program analysis techniques into the LLM training, code generation, and evaluation processes. The key products and services to be delivered include:

  1. Developing novel quantitative program analysis techniques, such as symbolic execution and Bayesian analysis, to provide meaningful metrics for evaluating the quality of LLM-generated code.
  2. Using these program quality metrics to train a differentiable reward model that can assess the quality of partial or complete code generated by LLMs. This reward model will then be used to fine-tune the LLMs for improved code generation.
  3. Leveraging the reward model and similarity-based program ranking techniques to constrain and prune the LLM's decoding tree during code generation, further enhancing the quality of the output.
  4. Creating new semantics-guided benchmarks consisting of realistic coding tasks for training and evaluating the code generation capabilities of LLMs.

The award was made to Carnegie Mellon University and is planned to run from October 1, 2025, through September 30, 2028. No subawards are planned under this project.

Generated 8/5/25, 6:26 AM