Project Grant 2339216
- This National Science Foundation (NSF) Division of Information and Intelligent Systems Project Grant aims to advance "few-round active learning" algorithms and enable more efficient training of supervised machine learning models. The $300,000 award to Virginia Polytechnic Institute & State University (Virginia Tech), running from August 1, 2023 to July 31, 2026, will support research to: 1) develop methods for quantifying the utility of unlabeled data for active learning tasks, and...
- This EAGER (Early-concept Grants for Exploratory Research) project grant, awarded by the National Science Foundation's (NSF) Computer and Information Science and Engineering (CISE) program (CFDA 47.070), aims to investigate the use of ensemble model diversity as a secondary optimization signal for improving machine learning model performance. The project, awarded to the University of Massachusetts Boston, will receive $299,883 over a 2-year period from January 1, 2025 to December 31, 2026. The...
- This Project Grant award of $450,000 from the National Science Foundation's Engineering program (CFDA 47.041) will support research on bi-level optimization for hierarchical machine learning problems. The award to the Regents of the University of Minnesota, conducting the work through their Office of Sponsored Projects Administration, aims to develop new approaches for modeling, analyzing, and innovating on a wide array of emerging machine learning applications using bi-level optimization...
- This $250,000 Project Grant award from the National Science Foundation's (NSF) Technology, Innovation, and Partnerships (TIP) program (CFDA 47.084) will fund a study on the ethical and financial trade-offs of machine learning (ML) training methods used to develop large language models (LLMs) like ChatGPT. The study, conducted by Georgia Tech Research Corp, will: 1) Identify current practices among researchers for training LLMs using large datasets; 2) Examine the trade-offs of different data...
- This three-year National Science Foundation project grant of $300,000 will fund research to advance trustworthy machine learning through bi-level optimization. The grantee, the University of California, Santa Barbara, will develop new algorithms and computational methods to achieve robust and fair deep learning. Specifically, the project will create a bi-level optimization framework for robust learning, defenses against adversarial examples and distribution shifts, and a full-stack robustness...
- This $349,427 CAREER grant from the National Science Foundation's (NSF) Computer and Information Science and Engineering (CISE) program supports research at The Leland Stanford Junior University (Stanford University) to develop machine learning (ML) techniques to improve the performance of discrete optimization algorithms. The project aims to address challenges in efficiently solving complex combinatorial optimization problems, such as those encountered in supply chain logistics and...
- This three-year, $300,000 project grant from the National Science Foundation's Division of Information and Intelligent Systems, under the Computer and Information Science and Engineering program (CFDA 47.070), will support the development of new algorithms and computational methods for trustworthy machine learning via bi-level optimization. The grantee, Michigan State University, will advance the theoretical understanding and practical implementation of robust and fair deep learning....
- This $431,250 Project Grant award from the National Science Foundation's (NSF) Computer and Information Science and Engineering (CISE) program supports research by Carnegie Mellon University to develop robust machine learning (ML) systems that can operate reliably in complex, real-world environments. The 5-year project (8/1/2025 - 7/31/2030) aims to bridge theoretical analysis and practical experimentation to address the brittleness of current ML models, which can fail unexpectedly under...
- This $400,000 Project Grant awarded by the National Science Foundation (NSF) under the Computer and Information Science and Engineering (CFDA 47.070) program supports the development of a new framework and tools for advancing data-centric artificial intelligence (AI) through generative approaches to feature space reconstruction. The project aims to transform the traditional way of constructing feature spaces by using deep generative learning instead of manual or classical discrete search...
- This $236,099 federal Project Grant award from the National Science Foundation's (NSF) Computer and Information Science and Engineering (CISE) program is supporting research to develop robust optimization and machine learning algorithms capable of handling dynamic and uncertain data environments. The research aims to advance optimization techniques for fundamental supervised learning tasks, yielding computationally and data-efficient algorithms with provable error guarantees. This work will...
CAREER: MITIGATING THE LACK OF LABELED TRAINING DATA IN MACHINE LEARNING BASED ON MULTI-LEVEL OPTIMIZATION -MACHINE LEARNING HAS DEMONSTRATED GREAT SUCCESS IN NUMEROUS APPLICATIONS SUCH AS AUTONOMOUS DRIVING, EARLY DETECTION OF DISEASES, DRUG DESIGN, ETC. THE ACCURACY OF MACHINE LEARNING MODELS HIGHLY DEPENDS ON THE ACCESSIBILITY OF LARGE-SCALE, HUMAN-LABELED TRAINING DATA. HOWEVER, SUCH DATA IS OFTEN VERY CHALLENGING TO ACQUIRE IN SPECIALIZED DOMAINS SUCH AS HEALTHCARE, LEGISLATION, ENVIRONMENTAL SCIENCES DUE TO THE HIGH COSTS INVOLVED IN OBTAINING HIGH-GRADE HUMAN LABELS AND DATA PRIVACY CONCERNS. THIS PROJECT WILL ADVANCE SCIENCE BY PROVIDING ALGORITHMS, SOFTWARE, AND SYSTEMS THAT CAN AUTOMATICALLY GENERATE HIGH-QUALITY LABELED DATA TO MITIGATE THE LACK OF LABELED TRAINING DATA IN SPECIFIC DOMAINS AND AND ALLOW TRAINING OF HIGHLY ACCURATE MACHINE LEARNING MODELS. THE PROJECT WILL SIGNIFICANTLY BROADEN THE APPLICABILITY OF MACHINE LEARNING ACROSS VARIOUS APPLICATION AREAS BY LOWERING DATA BARRIERS AND WILL SUBSTANTIALLY REDUCE THE LABOR COSTS OF MANUAL DATA ANNOTATION. FOR EXAMPLE, IT WILL PROMOTE SCIENTIFIC DISCOVERY IN STRUCTURAL BIOLOGY AND HIGH-ENERGY PHYSICS AND STREAMLINE ENGINEERING DESIGN IN WIRELESS COMMUNICATION. IT WILL FACILITATE THE EARLY DETECTION OF SEPSIS, LUNG CANCER, PARKINSON'S DISEASE, AND SLEEP APNEA, IMPROVING PATIENT OUTCOMES AND QUALITY OF LIFE. APPLIED TO COMPOUND DESIGN AND CEMENT PRODUCTION, THE DEVELOPED TECHNOLOGIES HAVE THE POTENTIAL TO EXPEDITE DRUG DISCOVERY AND REDUCE ENERGY CONSUMPTION. TO ACHIEVE THE GOAL OF CREATING HIGH-QUALITY LABELED TRAINING DATA, THIS PROJECT WILL DEVELOP THREE COMPLEMENTARY PARADIGMS OF NOVEL APPROACHES BASED ON MULTI-LEVEL OPTIMIZATION AND LARGE LANGUAGE MODELS, FOR: 1) END-TO-END GENERATION OF LABELED DATA; 2) ANNOTATION OF UNLABELED DATA; AND, 3) EXAMPLE-SPECIFIC ADAPTATION/SELECTION OF LABELED SOURCE DATA, RESPECTIVELY. FIRST, THE PROPOSED DATA GENERATION METHODS WILL LEVERAGE THE WORST-CASE AND CLASS-SPECIFIC PERFORMANCE OF DOWNSTREAM MODELS TO PROVIDE END-TO-END AND FINE-GRAINED GUIDANCE FOR GENERATING DATA (WITH COMPLEX LABELS) THAT IS TAILORED TO IMPROVE THE ACCURACY AND ROBUSTNESS OF DOWNSTREAM MODELS, AND TO PROMOTE BALANCED PERFORMANCE ACROSS DIFFERENT CLASSES. SECOND, THE PROPOSED DATA ANNOTATION METHODS WILL LEVERAGE AN END-TO-END MECHANISM THAT CAPITALIZES ON LARGE LANGUAGE MODELS, A SEQUENCE OF VERIFICATION PROCEDURES, AND AVAILABLE SIDE INFORMATION TO MAXIMIZE THE ACCURACY OF GENERATED LABELS. THIRD, THE PROPOSED ADAPTATION/SELECTION METHODS WILL DISTINGUISH BETWEEN SOURCE EXAMPLES THAT ARE INSIDE OR OUTSIDE OF A TARGET DOMAIN AND SUBSEQUENTLY DETERMINE AN EXAMPLE-SPECIFIC ADAPTATION/SELECTION ACTION END-TO-END TO ENSURE OPTIMAL USE OF SOURCE DATA. IN ADDITION, THE PROPOSED NOVEL OPTIMIZATION ALGORITHMS AND DISTRIBUTED SYSTEMS WILL EFFECTIVELY TACKLE NEW CHALLENGES RELATED TO MULTI-LEVEL OPTIMIZATION, INCLUDING NON-DIFFERENTIABILITY, INCOMPATIBILITY WITH THE OPTIMIZERS OF LARGE LANGUAGE MODELS, AND SCALABILITY. THIS PROJECT REPRESENTS THE FIRST ONE SYSTEMATICALLY LEVERAGING MULTI-LEVEL OPTIMIZATION TO CREATE LABELED DATA, EFFECTIVELY ADDRESSING A FUNDAMENTAL KNOWLEDGE GAP THAT EXISTING METHODS OFTEN LACK CAPABILITIES TO PERFORM END-TO-END EXECUTION OF MULTIPLE LEARNING STAGES AND THEREFORE FALL SHORT IN TAILORING GENERATED DATA TO IMPROVE DOWNSTREAM MODELS? PERFORMANCE. ANOTHER SIGNIFICANT INNOVATION OF THIS PROJECT IS ITS EFFECTIVE HARNESSING OF LARGE LANGUAGE MODELS FOR DATA ANNOTATION, WHICH WILL SUBSTANTIALLY REDUCE THE COSTS OF MANUAL LABELING. THIS AWARD REFLECTS NSF'S STATUTORY MISSION AND HAS BEEN DEEMED WORTHY OF SUPPORT THROUGH EVALUATION USING THE FOUNDATION'S INTELLECTUAL MERIT AND BROADER IMPACTS REVIEW CRITERIA.- SUBAWARDS ARE NOT PLANNED FOR THIS AWARD.
Mod # | Description | ReasonForModification | Federal Obligation | Date |
|---|---|---|---|---|
| Not listed | $80.0k | 8/18/25 | ||
| Not listed | $125.0k | 8/4/25 | ||
| Not listed | $165.0k | 4/17/24 |