Project Grant 2403194
- This Project Grant award from the National Science Foundation (NSF) Computer and Information Science and Engineering (CISE) program (CFDA 47.070) provides $564,958 to Carnegie Mellon University from December 1, 2024 to November 30, 2027. The project aims to advance fundamental knowledge and innovate resource allocation algorithms to address the challenges posed by high variability and uncertainty in both demand and service in modern computing systems, especially those supporting machine learning...
- This $275,000 Project Grant award from the National Science Foundation's (NSF) Computer and Information Science and Engineering (CISE) program (CFDA 47.070) aims to develop and test new scheduling algorithms for improving the performance of modern database systems. The goal is to create models and policies that can efficiently allocate limited hardware resources, such as compute and memory, to handle a stream of parallelizable database queries with varying levels of parallelizability and service...
- This Project Grant from the National Science Foundation's Computer and Information Science and Engineering program provides $250,000 to the College of William and Mary to support research and education activities under the EASER framework from June 2022 through May 2025. The award will fund the development of an integrated paradigm for application execution on heterogeneous high-performance computing clusters. Key products and services include compiler-driven performance prediction models, an...
- This Project Grant from the National Science Foundation's $1.2 million Computer and Information Science and Engineering program will fund the development of the Machine Learning Accelerator Cohort Architecture at the University of Southern California from October 2022 to September 2025. The ML Accelerator Cohort Architecture aims to build a heterogeneous computing fabric to efficiently support diverse machine learning models, including novel approaches for privacy-preserving computation. Key...
- This National Science Foundation (NSF) Computer and Information Science and Engineering (CISE) Program award provides $599,386 to Northeastern University to conduct research on communication-aware algorithms for dynamic allocation of heterogeneous computing resources in distributed infrastructure systems. The key products delivered under this 3-year project grant include: Innovative solutions for scheduling large-scale, precedence-constrained computations such as Artificial Intelligence tasks...
- This National Science Foundation (NSF) Computer and Information Science and Engineering (CISE) Federal Grant Program (CFDA 47.070) Project Grant award in the amount of $325,000.00 was provided to Carnegie Mellon University to develop and test algorithms for scheduling a stream of parallelizable database queries. The goal is to create new scheduling policies that maximize the utilization of system resources such as compute and memory in order to reduce query latencies. The project targets...
- This $1.2 million Project Grant from the National Science Foundation's Computer and Information Science and Engineering program will fund research at the University of Pittsburgh to expedite machine learning applications on multi-GPU infrastructure. Specifically, the university will uncover and address architectural bottlenecks in deep neural network executions on multi-GPU systems. Researchers will redesign translation lookaside buffer hierarchies and page table walks to reduce address...
- The National Science Foundation (NSF) awarded a $600,000 Project Grant under the Computer and Information Science and Engineering (CISE) Federal Grant Program to The Ohio State University. The grant funds a 3-year project focused on developing a holistic management framework to handle bursty machine learning (ML) inference requests in data centers with latency guarantees and reduced capital expenses (CAPEX). The key components of the project include: 1) Co-locating ML inference and training...
- This Project Grant award from the National Science Foundation's (NSF) Computer and Information Science and Engineering (CISE) program (CFDA 47.070) provides $600,000 in funding to the University of Central Florida to develop software-hardware solutions that enable efficient execution of large AI foundation models, like those powering advanced AI applications, on smaller, resource-limited computer systems. The project aims to improve the efficiency, scalability, and resource utilization of...
- This $131,959 Project Grant awarded by the National Science Foundation's (NSF) Computer and Information Science and Engineering (CISE) program supports research conducted by Rutgers, The State University to develop new deep learning training methods that can efficiently scale to utilize high-performance computing (HPC) systems. The key goals are to: 1) Explore techniques like second-order information approximation, computation-communication tradeoffs, and data compression to enhance the speed...
COLLABORATIVE RESEARCH: CIF: MEDIUM: TOWARDS OPTIMAL SCHEDULING FOR PARALLELIZABLE MACHINE LEARNING TRAINING WORKLOADS -IN THE PAST FIFTEEN YEARS, THERE HAS BEEN EXPLOSIVE GROWTH IN THE NUMBER OF APPLICATIONS SEEKING TO LEVERAGE MACHINE LEARNING (ML) MODELS. BEFORE AN ML MODEL CAN BE DEPLOYED, THE MODEL MUST BE ?TRAINED? BY REPEATEDLY PROCESSING EXAMPLES. EACH EXAMPLE HELPS THE MODEL TO MAKE PROGRESSIVELY MORE ACCURATE PREDICTIONS, LEARNING HOW TO SOLVE THE PROBLEM AT HAND. UNFORTUNATELY, THIS TRAINING STEP REQUIRES A LARGE AMOUNT OF EXPENSIVE, HIGHLY-SPECIALIZED HARDWARE AND CAN TAKE SEVERAL HOURS TO COMPLETE. GIVEN LIMITED HARDWARE RESOURCES, IT IS NOT OBVIOUS HOW TO ALLOCATE (SHARE) THESE RESOURCES ACROSS A STREAM OF ML TRAINING JOBS. THE GOAL OF THIS PROJECT IS TO DEVELOP NEW RESOURCE ALLOCATION POLICIES THAT ALLOW US TO PRODUCE HIGHLY-ACCURATE ML MODELS, QUICKLY, AND WITH LIMITED RESOURCES. THE MAIN CHALLENGE IN THIS WORK IS THAT ML TRAINING JOBS PRESENT A NUMBER OF UNIQUE CHARACTERISTICS COMPARED TO OTHER COMPUTING WORKLOADS. ML TRAINING JOBS ARE HIGHLY PARALLELIZABLE, MEANING A SINGLE TRAINING JOB MIGHT RUN ACROSS MULTIPLE SERVERS. EACH JOB ALSO HAS A LARGE NUMBER OF CONFIGURATION OPTIONS THAT AFFECT HOW FAST THE MODEL LEARNS A PARTICULAR PROBLEM. THIS PROJECT USES MATHEMATICAL MODELING TO DEVELOP NEW ALLOCATION POLICIES SPECIFICALLY DESIGNED FOR ML TRAINING JOBS. THIS RESEARCH WILL BE ACCOMPANIED BY THE DEVELOPMENT OF NEW COURSES AND MENTORSHIP PROGRAMS DESIGNED TO RECRUIT STUDENTS FROM UNDERREPRESENTED BACKGROUNDS INTO RESEARCH ON THE MODELING OF COMPUTER SYSTEMS. MACHINE LEARNING (ML) MODELS ARE INCREASINGLY BEING DEPLOYED ACROSS A WIDE VARIETY OF APPLICATIONS. THE GROWTH IN ML MODELS HAS BEEN ACCOMPANIED BY THE DEVELOPMENT OF SPECIALIZED HARDWARE ACCELERATORS THAT HELP REDUCE THE TRAINING TIME FOR ML MODELS. HOWEVER, THERE HAS NOT BEEN A SIMILAR DEGREE OF SPECIALIZATION IN THE SCHEDULING ALGORITHMS USED TO TRAIN ML MODELS ON CLUSTERS OF SPECIALIZED HARDWARE. THE QUESTION OF HOW TO BEST SCHEDULE ML TRAINING JOBS IS NON-TRIVIAL. ML TRAINING JOBS PRESENT MANY UNIQUE CHARACTERISTICS COMPARED TO OTHER COMPUTING WORKLOADS. FIRST, ML TRAINING JOBS VARY IN THEIR DEGREE OF PARALLELIZABILITY (ABILITY TO SCALE OUT ACROSS SERVERS), WITH SOME JOBS BEING HIGHLY ELASTIC AND OTHERS BEING INELASTIC IN THEIR SCALABILITY. IT IS NOT CLEAR HOW TO ALLOCATE (SHARE) HARDWARE RESOURCES AMONG JOBS WITH SUCH DIFFERENT CHARACTERISTICS. TO MAKE THINGS MORE COMPLICATED, EACH TRAINING JOB IS PARAMETERIZED BY A SET OF TUNING PARAMETERS CALLED HYPERPARAMETERS. THE PARALLELIZABILITY OF AN ML TRAINING JOB VARIES OVER TIME AS THE JOB HYPERPARAMETERS CHANGE. HENCE, SYSTEMS WHICH EITHER RELY ON STATIC RESOURCE RESERVATIONS OR EXISTING HEURISTIC SCHEDULING POLICIES ARE POORLY SUITED FOR ML TRAINING JOBS. FINALLY, THE INHERENT WORK ASSOCIATED WITH AN ML JOB IS NOT A FIXED QUANTITY. INSTEAD, ML JOBS ARE GENERALLY TRAINED UNTIL THE MODEL MEETS A DESIRED LEVEL OF ACCURACY. HENCE, THE EXISTING SCHEDULING THEORY THAT FAVORS ?SHORT? JOBS DOES NOT APPLY IN THIS CASE. THIS PROPOSAL DEVELOPS A NEW THEORETIC FRAMEWORK FOR MODELING THE DYNAMICS OF ML JOBS RUNNING IN A SHARED CLUSTER. THIS FRAMEWORK WILL BE USED TO DEVELOP SCHEDULING AND ALLOCATION POLICIES THAT ARE SPECIALIZED TO ML TRAINING JOBS, AND TO PROVE PERFORMANCE GUARANTEES ON THESE POLICIES. THIS AWARD REFLECTS NSF'S STATUTORY MISSION AND HAS BEEN DEEMED WORTHY OF SUPPORT THROUGH EVALUATION USING THE FOUNDATION'S INTELLECTUAL MERIT AND BROADER IMPACTS REVIEW CRITERIA.- SUBAWARDS ARE NOT PLANNED FOR THIS AWARD.
Mod # | Description | ReasonForModification | Federal Obligation | Date |
|---|---|---|---|---|
| Not listed | $212.2k | 8/27/25 | ||
| Not listed | $406.0k | 6/3/24 |