This Project Grant award of $675,000 was provided by the National Science Foundation (NSF) under the Computer and Information Science and Engineering (CFDA 47.070) program. The award aims to improve the ability of large language models (LLMs) to generate high-quality source code by deeply integrating program analysis techniques into the LLM training, code generation, and evaluation processes. The key products and services to be delivered include:
- Developing novel quantitative program analysis techniques, such as symbolic execution and Bayesian analysis, to provide meaningful metrics for evaluating the quality of LLM-generated code.
- Using these program quality metrics to train a differentiable reward model that can assess the quality of partial or complete code generated by LLMs. This reward model will then be used to fine-tune the LLMs for improved code generation.
- Leveraging the reward model and similarity-based program ranking techniques to constrain and prune the LLM's decoding tree during code generation, further enhancing the quality of the output.
- Creating new semantics-guided benchmarks consisting of realistic coding tasks for training and evaluating the code generation capabilities of LLMs.
The award was made to Carnegie Mellon University and is planned to run from October 1, 2025, through September 30, 2028. No subawards are planned under this project.
Generated 8/5/25, 6:26 AM