Project Grant 2337943

Award Date 6/1/24
Completion Date 5/31/29
Dollars Obligated $172K
Federal Grant Program
47.049
Assistance Type
Project Grant
Place of Performance
Raleigh, NC 27695, USA
Similar Awards
This federal Project Grant award from the National Science Foundation's (NSF) Social, Behavioral, and Economic Sciences program (CFDA 47.075) provides $375,000.00 to Iowa State University of Science and Technology to develop statistical and machine learning tools for data integration and data fusion. The project aims to enhance the analysis of complex survey data with big data sources, as well as improve scientific conclusions drawn from multiple datasets. Key activities include research on mass...
This National Science Foundation (NSF) Project Grant award under the Mathematical and Physical Sciences program (CFDA 47.049) provides $250,000 to The Trustees of the University of Pennsylvania to develop advanced statistical methods for integrating and analyzing large-scale data from multiple sources, such as electronic health records and genomics data. The project aims to devise new data-driven algorithms with theoretical optimality guarantees for transfer learning, as well as adversarially...
This $320,502 Project Grant award from the National Science Foundation (NSF) Division of Information and Intelligent Systems is for a collaborative research project titled "Knowledge Discovery from Highly Heterogeneous, Sparse and Private Data in Biomedical Informatics." The research aims to mine healthcare data to identify patients likely to develop chronic conditions like type 2 diabetes and heart failure, and to develop models for opportunistic screening, particularly for...
This Project Grant from the National Science Foundation's National Center for Science and Engineering Statistics will fund the development of Bayesian statistical and machine learning methodologies tailored for complex survey and census data. Awarded $743,050 under the Social, Behavioral, and Economic Sciences program, the grant will support research at the University of Missouri from September 2022 through August 2025. The research aims to advance computational efficiency and expand...
This National Science Foundation project grant of $281,469 will fund research at the University of Colorado Denver from August 2022 through July 2025 to advance causal inference methods for heterogeneous data fusion. Specifically, the award will support developing new approaches to empirically estimate associations and perform causal inference when individual-level data cannot be obtained due to privacy or logistical constraints. The research aims to extend statistical methodologies to...
This $233,955 Project Grant award from the National Science Foundation's (NSF) Social, Behavioral, and Economic Sciences (CFDA 47.075) program will support a research project at The Washington University in St. Louis to develop improved methods for probabilistic data integration and record linkage without the use of unique identifiers. The central objectives are to create computationally efficient and accurate techniques for merging large datasets from multiple sources, which is a critical...
This three-year, $674,542 National Science Foundation project grant supports research at the University of California Santa Cruz to develop Bayesian statistical and machine learning methods for analyzing complex survey data from the federal statistical system. The grant falls under the NSF's Social, Behavioral, and Economic Sciences program (CFDA 47.075), which promotes basic research and education in these fields. Specifically, the investigators will extend existing models using data...
This $200,000 Project Grant award from the National Science Foundation (NSF) Division of Mathematical Sciences under the Mathematical and Physical Sciences program (CFDA 47.049) aims to develop novel feature selection techniques for supervised and unsupervised machine learning models. The research will focus on the "knockoff method" for identifying key predictive features while controlling false discoveries, incorporating microbiome data structures, handling missing values, and...
This Project Grant awarded by the National Science Foundation (NSF) under the Social, Behavioral, and Economic Sciences (CFDA 47.075) program provides $140,482 to develop a new methodology to distinguish between research findings that can be generalized across populations, places, and time, and those that cannot be generalized. The research aims to advance statistical meta-analysis by creating estimators and classification tools to identify predictable (generalizable) and unpredictable...
This Project Grant award from the National Science Foundation's (NSF) Computer and Information Science and Engineering (CISE) program, with a total funding of $299,996, aims to develop software for post-linkage data analysis. The project builds on record linkage methods to address potential errors and uncertainty in data analysis performed after merging multiple data sources. The software will be developed in popular data science programming languages and tailored to the needs of federal...

CAREER: NEW DATA INTEGRATION APPROACHES FOR EFFICIENT AND ROBUST META-ESTIMATION, MODEL FUSION AND TRANSFER LEARNING -STATISTICAL SCIENCE AIMS TO LEARN ABOUT NATURAL PHENOMENA BY DRAWING GENERALIZABLE CONCLUSIONS FROM AN AGGREGATE OF SIMILAR EXPERIMENTAL OBSERVATIONS. WITH THE RECENT ?BIG DATA? AND ?OPEN SCIENCE? REVOLUTIONS, SCIENTISTS HAVE SHIFTED THEIR FOCUS FROM AGGREGATING INDIVIDUAL OBSERVATIONS TO AGGREGATING MASSIVE PUBLICLY AVAILABLE DATASETS. THIS ENDEAVOR IS PREMISED ON THE HOPE OF IMPROVING THE ROBUSTNESS AND GENERALIZABILITY OF FINDINGS BY COMBINING INFORMATION FROM MULTIPLE DATASETS. FOR EXAMPLE, COMBINING DATA ON RARE DISEASE OUTCOMES ACROSS THE UNITED STATES CAN PAINT A MORE RELIABLE PICTURE THAN BASING CONCLUSIONS ONLY ON A SMALL NUMBER OF CASES IN ONE HOSPITAL. SIMILARLY, COMBINING DATA ON DISEASE RISK FACTORS ACROSS THE UNITED STATES CAN DISTINGUISH LOCAL FROM NATIONAL HEALTH TRENDS. TO DATE, STATISTICAL APPROACHES TO THESE DATA AGGREGATION OBJECTIVES HAVE BEEN LIMITED TO SIMPLE SETTINGS WITH LIMITED PRACTICAL UTILITY. IN RESPONSE TO THIS GAP, THIS PROJECT DEVELOPS NEW METHODS FOR AGGREGATING INFORMATION FROM MULTIPLE DATASETS IN THREE DISTINCT DATA INTEGRATION PROBLEMS GROUNDED IN SCIENTIFIC PRACTICE. THE DEVELOPED APPROACHES ARE INTUITIVE, PRINCIPLED AND ROBUST TO SUBSTANTIAL DIFFERENCES BETWEEN DATASETS, AND ARE BROADLY APPLICABLE IN MEDICAL, ECONOMIC AND SOCIAL SCIENCES, AMONG OTHERS. AMONG OTHER APPLICATIONS, THE PROJECT WILL DELIVER NEW TOOLS TO EXTRACT HEALTH INSIGHTS FROM LARGE ELECTRONIC HEALTH RECORDS DATABASES. THE PROJECT WILL SUPPORT UNDERGRADUATE AND GRADUATE STUDENT TRAINING, COURSE DEVELOPMENT, AND THE RECRUITMENT AND PROFESSIONAL MENTORING OF UNDER-REPRESENTED MINORITIES IN STATISTICS. FURTHER, THE PROJECT WILL IMPACT STEM EDUCATION THROUGH A DATA SCIENCE TEACHER TRAINING PROGRAM IN UNDERSERVED COMMUNITIES. THIS PROJECT DEVELOPS INTUITIVE, PRINCIPLED, ROBUST AND EFFICIENT METHODS IN THREE ESSENTIAL DATA INTEGRATION PROBLEMS: META-ANALYSIS, MODEL FUSION AND TRANSFER LEARNING. FIRST, THE PROJECT DELIVERS A SET OF META-ANALYSIS METHODS FOR PRIVACY-PRESERVING ONE-SHOT ESTIMATION AND INFERENCE USING A NEW NOTION OF DATASET SIMILARITY. THE PRIMARY NOVELTY IN THE APPROACH IS THE JOINT ESTIMATION OF BOTH DATASET-SPECIFIC PARAMETERS AND A COMBINED PARAMETER THAT BEARS SOME SIMILARITY TO THE CLASSIC META-ESTIMATOR. SECOND, THE PROJECT ESTABLISHES MODEL FUSION METHODS THAT LEARN THE CLUSTERING OF SIMILAR DATASETS. THE METHODS? UNIQUE FEATURE IS A MODEL FUSION THAT DIALS DATA INTEGRATION ALONG A SPECTRUM OF MORE TO LESS FUSION AND THEREBY DOES NOT FORCE MODEL PARAMETERS FROM CLUSTERED DATASETS TO BE EXACTLY EQUAL. THIRD, THE PROJECT DEVELOPS FLEXIBLE AND ROBUST TRANSFER LEARNING APPROACHES THAT LEVERAGE HISTORICAL INFORMATION FOR IMPROVED STATISTICAL EFFICIENCY IN A TARGET DATASET OF INTEREST. AN IMPORTANT ELEMENT OF THESE APPROACHES IS A FLEXIBLE SPECIFICATION OF THE TYPE OF MODELS FIT TO THE SOURCE DATASETS. ALL THREE SETS OF METHODS PLACE A PREMIUM ON INTERPRETABILITY, STATISTICAL EFFICIENCY AND ROBUSTNESS OF THE INFERENTIAL OUTPUT. THE PROJECT UNIFIES THE THREE SETS OF PROPOSED METHODS UNDER A FORMAL DATA INTEGRATION FRAMEWORK FORMULATED AROUND TWO AXIOMS OF DATA INTEGRATION. DATA INTEGRATION IDEAS PERVADE EVERY FIELD OF SCIENTIFIC STUDY IN WHICH DATA ARE COLLECTED, AND SO THE RESEARCH CONTRIBUTES TO SCIENTIFIC ENDEAVORS IN THE MEDICAL, ECONOMIC AND SOCIAL SCIENCES, AMONG OTHERS. THIS AWARD REFLECTS NSF'S STATUTORY MISSION AND HAS BEEN DEEMED WORTHY OF SUPPORT THROUGH EVALUATION USING THE FOUNDATION'S INTELLECTUAL MERIT AND BROADER IMPACTS REVIEW CRITERIA.- SUBAWARDS ARE NOT PLANNED FOR THIS AWARD.

Posted 1/30/24