About this file

This document is a draft Program Solicitation for the PRECISE-AI (Performance and Reliability Evaluation for Continuous modIfications and uSEability of AI) program, issued by the Advanced Research Projects Agency for Health (ARPA-H). The PRECISE-AI program aims to create self-correction techniques that maintain peak performance of predictive AI components across diverse clinical settings. Key areas of innovation include continuous monitoring, degradation detection, root cause analysis, self-correction, and improved clinician communication. PRECISE-AI will establish an open-source repository of tools to autonomously maintain the performance of clinical AI Decision Support Tools (AI-DSTs) and enhance their interpretability. The program plans to test these innovations in real-world settings to demonstrate measurable improvements in clinical decision-making. Multiple awards are anticipated, with an anticipated individual award structure of Other Transactions under the authority of 42 U.S.C. 290c(g)(1)(D). The program is structured in three phases over four years, with progression to subsequent phases dependent on performance against specified milestones and metrics. Proposals can address individual technical areas or combinations of areas related to automated ground truth extraction, degradation detection and self-correction, uncertainty quantification and clinician performance improvement, core data infrastructure, and independent verification and validation.

View the file

Other files for this federal contract opportunity

Other files attached to Performance and Reliability Evaluation for Continuous modIfications and uSEability of AI (PRECISE-AI), newest first.
File Type Posted
ARPA-H-SOL-25-113_DRAFT v2.pdf PDF
PRECISE-AI_ProposersDay Slides.pdf PDF
ARPA-H-SOL-25-113_Appendix A.docx DOCX document
OT Bundle of Attachments.zip ZIP file

On GovTribe

Work with this file on GovTribe

  • Download the original file
  • Contacts named in this file
  • Similar government files
  • Ask GovTribe AI about this file

Text version

This is a DRAFT Program Solicitation. This is not an invitation for solution summaries or proposals. Any such response will be disregarded. ARPA-H will not comment nor provide feedback on questions related to proposal technical approaches. Questions and Comments should be directed to PRECISEAI@arpa-h.gov.

This is a DRAFT Program Solicitation. This is not an invitation for solution summaries or proposals. Any such response will be disregarded. ARPA-H will not comment nor provide feedback on questions related to proposal technical approaches. Questions and Comments should be directed to PRECISEAI@arpa-h.gov.

PROGRAM SOLICITATION

PERFORMANCE AND RELIABILITY EVALUATION FOR CONTINUOUS

MODIFICATIONS AND USEABILITY OF AI (PRECISE-AI)

RESILIENT SYSTEMS OFFICE (RSO)

ADVANCED RESEARCH PROJECTS AGENCY FOR

HEALTH

ARPA-H-SOL-25-113

August 29, 2024

PROGRAM SOLICITIATION OVERVIEW INFORMATION

FEDERAL AGENCY NAME: Advanced Research Projects Agency for Health (ARPA-H)

SOLICITATION TITLE: Performance and Reliability Evaluation for Continuous modIfications and uSEability of AI (PRECISE-AI)

ANNOUNCEMENT TYPE: PROGRAM SOLICITATION (PS), Initial Announcement

SOLICITATION NUMBER: ARPA-H-SOL-25-113

Dates: (All times listed herein are Eastern Time)

• Proposers’ Day: October 2024 (specific date and location TBD)

• Program Solicitation Questions & Answers (Q&A) submission due date: TBA in final solicitation

• Solution Summary Due Date: TBA in final solicitation

• Proposal Due Date: TBA in final solicitation

CONCISE DESCRIPTION OF THE SOLICITATION:

The rapid advancement of artificial intelligence (AI) technologies is transforming healthcare by improving efficiencies, reducing costs, and enhancing health outcomes. This potential is evident with over 850 FDA-approved medical devices now incorporating AI functionalities, a tenfold increase from 2018 to 2023.

However, the ability to ensure the ongoing safety and efficacy of these AI systems has not kept pace. The conventional safety testing approach relies heavily on pre-market testing, assuming that these initial results will predict long-term performance. However, pre-market results often fail to account for variations in operational processes and patient demographics, leading to unpredictable post-market performance that currently requires manual oversight by vendors. Performance and Reliability Evaluation for Continuous modIfications and uSEability of AI (PRECISE-AI) aims to create a suite of self-correction techniques that make it possible to automatically maintain peak model performance of predictive AI components across diverse clinical settings. PRECISE-AI will advance novel approaches to optimally support clinician decision-making and scalably manage the performance of AI Decision Support Tools (AI-DSTs) after their commercial deployment. Key areas of innovation include continuous monitoring capabilities, degradation detection, root cause analysis, self-correction, and bidirectional communication with clinicians. This program will establish an open-source repository of tools to autonomously maintain the performance of clinical AI-DSTs while enhancing the interpretability and actionability of AI model outputs. The program will test these innovations in real-world settings to demonstrate measurable improvements in clinical decision-making. This program addresses the pressing need for continuous monitoring and updating of clinical AI models to ensure they remain effective and trustworthy over time.

ANTICIPATED INDIVIDUAL AWARD: Multiple awards are anticipated.

TYPES OF INSTRUMENTS THAT MAY BE AWARDED: Other Transactions awarded under the authority of 42 U.S.C. § 290c(g)(1)(D).

POINTS OF CONTACT (POC):

Technical Point of Contact: Berkman Sahiner, PRECISE-AI Program Manager, Resilient Systems Office

(RSO)

The PS Coordinator for this effort can be reached at: PRECISE-AI@arpa-h.gov.

ATTN: PRECISE-AI

Contents

PROGRAM SOLICITIATION OVERVIEW INFORMATION

1. PROGRAM INFORMATION

1.1. BACKGROUND

1.2. PROGRAM DESCRIPTION

1.3. PROGRAM SCOPE

1.4. PROGRAM STRUCTURE & TECHNICAL AREAS

1.1.1. Summary of Program Technical Areas

1.1.2. Program Structure

1.1.3. TA1: Automated Surrogate Ground Truth Label Extraction

1.1.4. TA2: Degradation Detection & Self-Correction

1.1.5. TA3: Quantify Uncertainty & Improve Clinician Performance

1.1.6. TA4: Core Data Infrastructure

1.1.7. TA5: Independent Verification and Validation

1.1.8. Program Milestones & Metrics

1.1.9. Proposal Scope

1.1.10. Common Requirements for All Proposals

1.1.11. Collaboration and Data Sharing

1.1.12. Open Software Standards

1.1.13. Commercial Transition Support

1.1.14. Equity Requirements

2. PS AWARD INFORMATION

3. ELIGIBILITY INFORMATION

3.1. ELIGIBLE APPLICANTS

3.2. PROHIBITION OF PERFORMER PARTICIPATION FROM FEDERALLY FUNDED

RESEARCH AND DEVELOPMENT CENTERS (FFRDCS) AND GOVERNMENT ENTITIES

3.3. NON-U.S. ORGANIZATIONS

3.4. ORGANIZATIONAL CONFLICTS OF INTEREST (OCI)

3.5. AGENCY SUPPLEMENTAL OCI POLICY

3.6. GOVERNMENT PROCEDURES

3.7. RESEARCH SECURITY DISCLOSURE

4. PROPOSAL AND SUBMISSION INFORMATION

4.1. GENERAL GUIDELINES

4.2. SOLUTION SUMMARY RESPONSES

4.3. PROPOSAL INSTRUCTIONS

4.3.1. Proposal Volume Templates

4.3.2. Model Other Transaction Agreement

4.4. PROPOSAL DUE DATE AND TIME

5. EVALUATION OF PROPOSALS

5.1. EVALUATION CRITERIA FOR AWARD

5.2. REVIEW AND SELECTION PROCESS

5.3. HANDLING OF COMPETITIon SENSITIVE INFORMATION

6. AWARDS

6.1. GENERAL GUIDELINES

6.2. NOTICES

6.2.1. Proposals

6.3. ADMINISTRATIVE AND NATIONAL POLICY REQUIREMENTS

6.3.1. System for Award Management (SAM) Registration and Universal Identifier Requirements ... 39

6.3.2. Controlled Unclassified Information (CUI) or Controlled Technical Information (CTI) on Non- DoD Information Systems

6.3.3. Intellectual Property (IP)

6.3.4. Human Subjects Research

6.3.5. Animal Subjects Research

6.4. ELECTRONIC INVOICING AND PAYMENTS

7. COMMUNICATIONS

1. PROGRAM INFORMATION

1.1. BACKGROUND

New artificial intelligence (AI) technologies are transforming healthcare, proffering improved efficiencies, reduced costs, and improved health outcomes. The market's interest is evident, with over 850 FDA-approved medical devices integrating AI functionalities — a tenfold increase from 2018 to 20231.

However, this growth has not been matched by scalable, automated capabilities to ensure the ongoing safety and efficacy after a vendor receives FDA clearance and after tools enter broad clinical use. Today’s safety testing approaches assume that pre-market studies predict post-market performance. These processes also presuppose that vendors have sufficient resources to detect and report deviations from pre-market performance, even when predictive AI tools are deployed across hundreds of clinical settings. Yet, evidence shows pre-market results often fail to predict the long-term performance of predictive AI tools due to variations in deployment contexts, including changes to operational processes and patient demographics 2,3.

Pre-market model performance studies typically rely on curated test data wherein each patient case is individually reviewed and manually assigned a “ground-truth label” (i.e., the most accurate estimate of the diagnosis given the available evidence). A ground truth label serves as the critical baseline for evaluating AI model performance. Once an AI tool is on the market, developers lack incentives to monitor and update their products, leading to sporadic updates and no mechanisms for hospitals to track product reliability.

The absence of ground truth labels in real-world clinical settings hinders hospitals from understanding AI tool performance in their environment putting patients and clinicians at increased risk. Current tools make it challenging for clinicians to detect performance deterioration of AI models, which can result from changes in clinical operations, IT infrastructure updates, and patient demographics. Performance may degrade differently across different clinical environments, and there is no established methodology to understand the root causes of the degradation and to take corrective action. These factors underscore the need for continual adaptation of AI systems to maintain their effectiveness in dynamic clinical environments.

National Health Impact: A Pathway to Improve Health Outcomes

The use of AI in clinical decision settings is rapidly expanding; however, there are currently no automated tools or mechanisms to monitor the performance and long-term safety of clinical AI Decision Support Tools (DSTs). Over time, clinical AI-DST performance can degrade without notifying clinicians. A study that evaluated 32 datasets from 4 industries found that a significant majority, 91%, of machine learning models experience a significant reduction in effectiveness as they age4. This degradation manifests as an 8-20% annual decrease in accuracy, or increased variability of errors,4,5 primarily due to shifts in the underlying data, which can erode trust in these systems.

Current post-market evaluation strategies are inadequate for addressing the issues faced by predictive AI models in clinical settings. Alarmingly, none of the existing clinical AI models undergo regular testing and

1 U.S. Food and Drug Administration. Artificial Intelligence and Machine Learning (AI/ML)-Enabled Medical Devices.

https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-and-machine-learning-aiml-enabled-medical-devices 2 Wong A, Cao J, Lyons PG, Dutta S, Major VJ, Ötles E, Singh K. Quantification of Sepsis Model Alerts in 24 US Hospitals Before and During the COVID-19 Pandemic. JAMA Netw Open. 2021 Nov 1;4(11):e2135286. doi:

10.1001/jamanetworkopen.2021.35286. PMID: 34797372; PMCID: PMC8605481 3 de Vries CF, Colosimo SJ, Staff RT, Dymiter JA, Yearsley J, Dinneen D, Boyle M, Harrison DJ, Anderson LA, Lip G; iCAIRD Radiology Collaboration. Impact of Different Mammography Systems on Artificial Intelligence Performance in Breast Cancer Screening. Radiol Artif Intell. 2023 Mar 22;5(3):e220146. doi: 10.1148/ryai.220146. PMID: 37293340; PMCID: PMC10245180.

4 Vela, D., Sharp. et al. Temporal quality degradation in AI models. Sci Rep. 2022. doi: 10.1038/s41598-022-15245-z 5 Yang J. et al. AI Gone Astray: Technical Supplement. arXiv preprint. 2022. doi: arXiv:2203.16452 updating during their clinical use to maintain accuracy. Only 44.7% of physicians stated that they would be willing to use AI-driven medicine6. Additionally, there are no robust mechanisms in place to continuously monitor and update these AI models during clinical operations, though preliminary trials are being conducted to explore potential solutions7. Currently, clinician intuition remains the primary method for detecting model degradation, which can result in harm before the issues are manually identified. Moreover, there are no clinically validated technical solutions yet available to fulfill the requirements of the 2023 Executive Order, which mandates the development of performance monitoring for medical AI models. The FDA is in the process of creating a regulatory pathway for automated post-market AI model updating to address these challenges8, but the industry still lacks robust, technically rigorous, scalable techniques for detecting model degradation and evaluating corrective actions.

1.2. PROGRAM DESCRIPTION

The PRECISE-AI program is soliciting proposals to create novel self-correction techniques that optimally maintain the peak performance of clinical AI decision support tools, both when operating independently and in combination with the clinicians who use them. PRECISE-AI aspires to create a future where predictive AI models continuously communicate with clinicians in a way that appropriately earns the clinician’s trust. Proposers should address critical challenges in the ability of AI Decision Support Tools (AI- DSTs) to actively monitor and maintain their optimal performance with consideration to local health systems, operational processes such as data acquisition, and patient characteristics, after their commercial deployment. Because such techniques are in their infancy, PRECISE-AI will advance the science behind continuous monitoring capabilities and move beyond monitoring into degradation detection, root cause analysis, self-correction, and bidirectional communication with clinicians. PRECISE-AI aims to create and validate robust AI degradation detection and auto-correction capabilities that provide the technical means to accomplish goals laid out in the AI Executive Order9, and the risk management frameworks created by organizations such as the FDA and NIST.

PRECISE-AI performers will create an open-source repository of tools that enable continuous monitoring and auto-correction of AI-DST. Proposers should focus specifically on AI-DSTs that make a diagnosis or prediction that is later validated by another test or clinical action. By automatically extracting surrogate ground truth label information from health records, automated monitoring techniques can continuously evaluate when AI decision support tools make accurate diagnoses and when they make mistakes. The term “surrogate ground truth label” above refers to an automatically extracted approximation to the ground truth label extracted using traditional, manual, and resource-intensive labeling. As an AI model ages, the percentage of mistakes increases due to changes in factors such as patient demographics or data acquisition techniques (technically termed as “dataset shifts”), affecting the accuracy of the AI tool.

PRECISE-AI will develop self-correction tools that tackle the effects of dataset shifts, as described next.

PRECISE-AI seeks novel proposals to analyze the monitoring data to suggest root causes for performance deterioration. The root cause analysis will be used to recommend self-corrective actions that enable the AI model to improve its performance. When performance degradation occurs, it will be important to communicate the heightened risk of medical errors to clinicians, the developers of the AI decision support

6 Tamori H, Yamashina H, Mukai M, Morii Y, Suzuki T, Ogasawara K. Acceptance of the Use of Artificial Intelligence in Medicine Among Japan's Doctors and the Public: A Questionnaire Survey. JMIR Hum Factors. 2022 Mar 16;9(1):e24680. doi:

10.2196/24680 7 https://www.fiercehealthcare.com/ai-and-machine-learning/epic-plans-launch-ai-validation-software-healthcare-organizations-test 8 https://www.fda.gov/regulatory-information/search-fda-guidance-documents/marketing-submission-recommendations-predetermined-change-control-plan-artificial 9 Exec. Order No. 14110 of Oct 30, 2023. https://www.federalregister.gov/documents/2023/11/01/2023- 24283/safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence tool, hospital administrators, and potentially the FDA. Proposers should describe novel mechanisms to streamline these communications and ensure that AI decision support tools not only generate trustworthy recommendations but also behave in a manner that earns stakeholder’s trust.

PRECISE-AI aims to improve the performance of clinicians that use AI-DSTs. Once the AI-DST monitoring and auto-correction tools are developed, it will be necessary to help transform these into meaningful insights for clinicians that enhance patient care. Proposers are encouraged to innovate new methods and tools for improving the transparency, interpretability, and actionability of AI model outputs to clinicians. Ultimately, researchers will test these tools in real world settings to demonstrate measurable improvements in clinical decision-making.

Cumulatively, PRECISE-AI seeks innovative proposals to create a large stepwise improvement over the current state of practice, where all patient groups are potentially exposed to AI-DSTs with degraded performance, resulting in possible patient harm. Proposers should endeavor to mitigate risk of patient harm beginning with patient one (by leveraging simulated data sets alongside real-world data) and precipitously decreasing risk for subsequent patients as additional real-world data is collected downstream of the AI- DST’s use. As a result, the first group of patients where the AI-DST shows a deviation from its expected performance would alert stakeholders to the problem, the likely cause, and potential solutions to mitigate risk to patients. In addition, uncertainty quantification and other bi-directional communication tools with clinicians will mitigate the risk for all patients.

1.3. PROGRAM SCOPE

The PRECISE-AI program seeks proposals for novel computational techniques that enable AI-DSTs to self-monitor and maintain high quality performance across diverse clinical contexts. AI-DSTs are decision support tools that are used in a context where they provide recommendations for clinicians, and the clinicians are the final decision-makers. Enhancing post-market model performance in an automated way will require research and development in three key areas: (1) automatic extraction of surrogate ground truth labels, (2) automated detection and correction of performance degradation, and (3) robust communication with clinicians. Finally, multi-institutional data infrastructure will accelerate the aggregation of insights from different hospitals and enable effective data collection, sharing, and analysis with vendors, clinicians, hospitals, and potentially the FDA.

Automated Extraction of Ground Truth

Continuous monitoring of clinical AI systems requires a continuous source of “ground truth” that can be used to determine when an AI system performs well and when it makes mistakes. A ground truth label is the most accurate diagnosis that can be made given the available evidence. Pre-market approaches to defining ground truth labels are resource-intensive, wherein each patient case is individually reviewed and assigned a label. Manual ground truth labeling generally relies on the consensus opinion of a panel of experts, and this approach does not scale.

Automated surrogate ground truth labeling requires research advances to extract the most relevant information from medical records and test results with high accuracy, and to approximate, as closely as possible, the traditional, resource-intensive ground truth extraction used in pre-market evaluation of AI- DSTs. A single data element, such as an isolated laboratory result or imaging finding, is often insufficient to accurately determine a patient's condition. Precise identification of the surrogate ground truth label might require the integration of multiple sources of information, such as radiology and pathology reports, clinical notes, biomarker assay results, ICD and SNOMED CT codes. This complexity demands the ability to ingest diverse data types from both structured and unstructured electronic health records (EHRs).

Additionally, data elements necessary for reliable surrogate ground truth labeling are stored in diverse formats across various database systems and information models within and between healthcare institutions.

This heterogeneity complicates the task of combining elements into a cohesive and accurate surrogate ground truth label. For instance, the terminology used in the radiology reports may vary among different institutions, and within an institution, the storage format of free-text data will be different from that of the ICD codes. Consequently, existing methods often rely on manual processes or ad hoc approaches that are not scalable or robust enough to handle the volume and variety of data encountered in real-world clinical environments. These challenges make it difficult to create a continuous stream of surrogate ground truth labels to enable the continuous validation of clinical AI-DST models. Therefore, PRECISE-AI seeks innovative research proposals that address the challenge of extracting surrogate ground truth labels across systems.

Clinical AI-DSTs are often developed around a specific input data type. To increase the breadth and generalizability of program innovations, PRECISE-AI efforts will be directed towards specific clinical use cases within the following clinical tracks: diagnosis through X-ray imaging, diagnosis through Computed Tomography (CT) imaging, clinical insights derived from patient Electronic Health Records, or alternative data input types.

Automated Detection and Correction of Performance Degradation Across Hospital Systems

When an AI model starts making more mistakes within a clinical context over time, it is called performance degradation. Existing performance monitoring approaches are manual and sporadic, limiting the ability to accurately detect performance degradation across clinical sites and hospital systems. PRECISE-AI seeks to develop novel approaches that provide a continuous and comprehensive view of model performance, making it possible to detect and respond to degradation as expeditiously as clinically possible. Additionally, these innovations would create mechanisms to quickly share performance data and insights among clinicians, hospital administrators, and developers, as well as hospital systems.

It will be important for performance degradation detection techniques to distinguish between normal variability in summary performance metrics such as sensitivity and specificity, and genuine performance degradation. Simulation capabilities are an example of an approach that could be used to test the behavior of AI models under various degradation scenarios at scale. High-precision and high-confidence detection of model degradation requires extensive analysis of degradation behaviors, which is not adequately supported by current methodologies. PRECISE-AI seeks proposals that address gaps in the timely and accurate detection of AI model degradation, to enable systemic improvements in the overall reliability and effectiveness of clinical AI systems.

Once performance degradation has been detected, root cause analysis is key to respond to the degradation.

PRECISE-AI seeks innovative techniques to empower hospitals and clinical sites by providing an automated analysis of the most likely root causes for a given performance degradation episode. The current lack of datasets hinders the ability to systematically study and address the underlying factors contributing to model degradation. Large-scale simulations will therefore likely play a key role in developing robust root cause analysis tools. At the same time, PRECISE-AI proposers should address the gap in real-world data availability for performance degradation and root cause analysis by continuously monitoring AI-DSTs. The developed methods will be validated and stress-tested with both simulated and real-world data, providing scientific and practical evidence into the performance of the developed root-cause analysis methods.

Research funded through the PRECISE-AI program will ultimately prototype novel machine learning and AI techniques to correct degraded model performance. Self-correction mechanisms for AI models carry significant risks, including the potential to inadvertently deteriorate overall performance or negatively impact specific subgroups. Therefore, research will focus on self-correction approaches that carefully balance the potential benefits of model updates against the risks, a process that is both complex and nuanced.

PRECISE-AI will advance research on reliable methods to enhance the effectiveness and safety of automated or semi-automated model updates. Methods of interest will recommend self-correcting actions that predictably improve AI-DST outputs, bolstering the overall reliability of clinical AI models.

Robust Communication with Clinicians

When tools fail to convey important nuances about the reliability of the AI outputs, it erodes clinician trust and limits the practical utility of these AI models. The task of quantifying model uncertainty varies significantly across different clinical use cases and existing methods are often not tailored to the specific needs of diverse clinical environments. PRECISE-AI performers will clinically evaluate novel approaches to convey model uncertainty and other supplementary model information to clinicians. Examples of supplementary information include AI explainability tools, capabilities enabling clinicians further interrogate the model (e.g., question-answering systems), or approaches to display of data from similar cases with ground truth. Because AI-DST capabilities are only as good as the decisions that clinicians make based on their outputs, PRECISE-AI performers will compare techniques for communicating model certainty and other supplementary model information to determine which ones give clinicians an appropriate level of trust in AI-DST outputs and ultimately improve clinician performance.

1.4. PROGRAM STRUCTURE & TECHNICAL AREAS

PRECISE-AI will advance auto-correction capabilities for predictive AI-DSTs. The program will develop a suite of tools and capabilities that continuously monitor the performance of predictive AI-DSTs in the clinic, detect when the performance deteriorates due to changes in a dynamic clinical environment, automatically suggest, and when appropriate, implement updates to the AI-DST models to counteract the effect of the changes in the environment, and establish improved and dynamic communication with the clinicians, resulting in improved clinical decision-making by clinicians.

1.1.1. Summary of Program Technical Areas

PRECISE-AI is comprised of five interconnected Technical Areas (TAs), outlined below.

• TA1 - Automated Surrogate Ground Truth Label Extraction: Automate the extraction of surrogate ground truth labels for individual patients across a diverse range of clinical use-cases. This will provide a robust foundation for assessing AI model performance continuously in TA2.

• TA2 - Degradation Detection & Self-Correction o TA2.1 - Continuous Degradation Detection Tools: Quickly and accurately differentiate between normal AI model performance variability and actual performance degradation based on surrogate ground truth labels extracted in TA1.

o TA2.2 - AI-Based Root-Cause-Analysis Tools: Develop AI-based tools that can automatically pinpoint the root causes of performance drops following alerts from degradation detection technologies developed in TA2.1.

o TA2.3 - AI Model Self-Correction Tools: Develop technologies enabling the AI models to automatically suggest and when appropriate, implement necessary updates to mitigate the issues detected in TA2.2.

• TA3 - Quantify Uncertainty & Improve Clinician Performance: Enhance clinician trust in AI-supported decision-making processes and improve their overall performance. This will be achieved by accurately quantifying and communicating uncertainty in AI model outputs and other supplementary information that aids in the clinician’s ability to appropriately trust AI models.

• TA4 - Core Data Infrastructure: Establish a core data infrastructure that supports interoperable data collection, sharing, analysis, and reporting among key stakeholders. During the program, this infrastructure will underpin the entire program, enabling effective coordination among performers.

After the program, aspects of TA4, such as reporting of performance degradations or identified root causes for degradations to stakeholders, are expected to continue so that the program’s effectiveness is maintained after the program’s lifecycle. Communication of analysis results may include public workshops and interactive performance dashboards, as well as mechanisms such as adverse event reporting and predetermined change control stipulations in regulatory submissions for communicating with regulators.

• TA5 - Independent Verification and Validation: Test and validate results from TA1, TA2, and TA3 to confirm performers’ progress. This will include validating the performers’ progress towards the metrics listed in Figure 3 below (Program Metrics) with independent data to ensure the generalization of the reported methods to new data and new clinical sites, as well as the verification and validation of simulation tools developed by other performers.

Each proposal responsive to this Solicitation can address a single Technical Area (TA) among TA1, TA2, TA3, TA4 and TA5, or a combination of TA1, TA2, and TA3. Proposals that address TA4 or TA5 cannot address any of the other TAs. Proposers to TA1 or TA2 must align their technologies to specific clinical tracks (Figure 4) with preference for technologies that can be generalized outside of their chosen track. Proposers to TA3 are expected to develop technologies that are largely generalizable across the program’s clinical tracks (i.e., X-ray imaging, diagnosis through Computed Tomography (CT) imaging, and clinical insights derived from patient Electronic Health Records). Proposals that apply to a combination of TA1, TA2, and TA3 (e.g., joint TA1/TA2 or joint TA1/TA2/TA3 proposals) may submit an alternative use case outside of the aforementioned clinical tracks.

Figure 1. Summary of Program Technical Areas

1.1.2. Program Structure

PRECISE-AI is a 4-year, 3-phase program that consists of five TAs. The program will advance technically rigorous techniques to continuously monitor and correct AI models used in real-world clinical decision settings. Performers will produce tool suites that include multiple innovative elements: automatic and continuous surrogate ground truth label extraction, continuous degradation detection, AI-based root-cause-analysis, model self-correction, and intuitive and actionable quantification of model outputs. Innovations throughout the program will be developed for use cases that fall within priority clinical tracks (i.e., X-ray imaging, diagnosis through Computed Tomography (CT) imaging, clinical insights derived from patient Electronic Health Records; see Figure 4 for more information). Progression to Phase II and to Phase III will depend on performance against milestones (Figure 2), metrics (Figure 3), and evaluation by the Independent

Verification and Validation (IV&V) performer.

PRECISE-AI will structure work across three phases as follows:

• Phase I (months 1-24): Performers will prototype their proposed approaches by providing computational models, component prototypes, and landscape reports. TA1 performers will collaborate with TA2-TA4 to generate use-case-specific surrogate ground truth labels across all sites.

• Phase II (months 25-36): Performers will refine prototypes, test in real-world clinical settings, and expand the breath of sites and use cases being covered for maximum impact.

• Phase III (months 37-48): Performers will integrate their component technologies into a commercial transition package, which can include any or all of the following options: an open-source tool suite that can be leveraged by multiple vendors to improve AI model performance; a self-monitoring medical device that will be submitted for FDA clearance; a performance monitoring capability that enables hospitals to monitor the performance of their AI-enabled medical devices; or a multi-organization capability that could assess patient safety across multiple hospital systems.

1.1.3. TA1: Automated Surrogate Ground Truth Label Extraction

TA1 proposals should describe an innovative, scalable, sustainable approach to automatically and continuously identify individual patient-level surrogate ground truth labels across a breadth of clinical use cases. For healthcare providers to evaluate whether AI tools are performing well in a site-specific manner, it is necessary to define surrogate ground truth labels that serve as the basis for comparison with the AI model output. TA1 aims to innovate scalable and automatic methods to derive surrogate ground truth labels by retrieving and analyzing disparate patient-level data sources to create surrogate ground truth labels.

For example, the purpose of a given clinical AI model may be to identify patients with pneumonia based on chest X-ray images. To evaluate how well the AI model is performing (done in TA2), TA1 proposers would establish a surrogate ground truth for each specific patient (i.e., what is the best estimate that a particular patient had pneumonia on the date of the chest X-ray image). This surrogate ground truth label should be triangulated from multiple data sources (e.g., ICD-9 or ICD-10 diagnostic codes, review of infection related symptoms, laboratory test results, prescribed medications, and radiology reports). TA1 proposers will also specify when the extracted surrogate ground truth labels are ready for use in TA2 tasks. For example, in the example above, this may be after the diagnostic codes, review of infection related symptoms, laboratory test results, prescribed medications, and radiology reports have been entered into the EHR. As described below, TA1 proposers will validate their surrogate ground truth extraction method, with the timeline specified in their proposal, by comparing it to traditional ground truth extraction methods (e.g., with a manual method and/or an expert panel and/or a longer timeline).

The surrogate ground truth labels that are produced by TA1 performers will serve as the basis for comparison when TA2 evaluates clinical AI model performance. The purpose of TA1 is not to develop AI- DSTs for the use cases selected by TA2, but to develop generalizable automated surrogate ground truth label extraction techniques to be used for the continuous monitoring and updating of AI-DST. To achieve this, TA1 proposals should focus on the following objectives:

Objective 1: Identify and ingest data elements that correlate with ground truth. TA1 performers will determine which disparate sources of healthcare information are necessary for establishing surrogate ground truth labels for each clinical use case, as outlined in Figure 4. TA1 will develop methods to extract these data elements from multiple sources in an interoperable manner to inform surrogate ground truth labeling.

TA1 performers will work with clinical partners and the developers of the AI-DSTs selected by TA2 to establish the baseline performance and to help track the temporal performance of TA2-selected AI-DSTs. To establish the initial performance baseline, performers may use historical (previously acquired) data from TA2 performers, other data repositories, and ground truth labels derived using traditional methods. During the program, TA1 performers will collaborate with hospitals or other clinical care sites to create algorithms that support the continuous extraction of data elements that provide patient-level surrogate ground truth labels. Additionally, TA1 performers will estimate the number of individual patient records required to reliably achieve the concordance levels targeted in Figure 3 between automatically extracted and traditionally defined ground truth labels. For instance, if a TA2 performer selects an AI-DST aligned with an X-ray based use case, the corresponding TA1 performer will need to estimate the number of patient records necessary for establishing statistically accurate and precise surrogate ground truth labels, and then extract relevant data from radiological imaging facilities, patient encounter notes, and other pertinent sources in the

EHR.

For each program use case selected by TA2 performers, TA1 performers will test and optimize NLP and multimodal data integration to consistently ingest data from different healthcare data infrastructures and information coding systems (e.g., ICD10, SNOMED CT, CPT codes, pharmacy and prescription codes, insurance claims data, clinicians' notes, clinical reports). TA1 proposals should describe how to combine the necessary breadth and numbers of disparate patient records from diverse EHR systems to establish a coherent, harmonized history of patient cases across multiple sites. Furthermore, they will extract metadata, such as relevant patient demographics, to enable sub-population comparisons.

TA1 performers will work with TA4 performers to deliver this data into TA4’s core data infrastructure using interoperable methods and a continuous integration continuous delivery (CICD) approach. TA1 performers will demonstrate the generalizability of their approach for surrogate ground truth label extraction across multiple use cases specified by TA2 performers and across multiple clinical sites in Phases II and III.

Objective 2: Develop methods to produce surrogate ground truth information automatically and efficiently. Once a set of data elements that correlate with the correct diagnosis have been identified and extracted, performers will establish scalable and automated methods to analyze the necessary data to continuously produce patient-specific (or case-specific) surrogate ground truth labels.

During the program, TA1 performers will develop surrogate ground truth label inference techniques that provide an accurate estimate of the diagnosis given the available evidence by integrating multi-modal data across contexts and time-points. Performers will leverage multiple approaches, including (i) extracting non-structured healthcare data from patient health records (e.g., separate test results, clinician notes); (ii) machine learning techniques that perform feature selection from a larger set of data elements; and (iii) complex methods for information fusion. TA1 performers will develop automated and scalable methods and tools to define surrogate ground truth labels, which can include, but are not limited to, rule-based systems, ML-based methods, and a combination of rule-based and ML-based methods. They will demonstrate the ability to continuously and automatically extract surrogate ground truth labels from different healthcare data.

Additionally, performers will demonstrate agreement between the automated and traditionally defined ground truth labels, using a representative and properly sized data set.

The primary deliverable of objective 2 will be the successful development of methodologies that enable continuous auto-extraction of relevant clinical data and generation of surrogate ground truth labels.

Objective 3: Engage hospital partners and create a diverse and representative dataset for program activities across all TAs. TA1 will work with multiple independently sourced hospital partners and create a diverse, representative longitudinal dataset that spans multiple institutions, patient demographics, coding systems, and interrelated clinical use cases. This dataset will be used for: training and testing TA1 surrogate ground truth inference algorithms; all activities for TA2 (including monitoring, root cause analysis, and self-correction); as well as communication and test activities for TA3.

Strong TA1 proposals will establish partnerships with hospitals and clinical sites that enables the program to evaluate AI-DST performance across a broad array of patients and environments, through data collected in real world practice for AI-DSTs at these sites. TA1 performers will be required to expand the number and diversity of clinical sites throughout the program (see Figure 3 for specific metrics). TA1 proposers will demonstrate that enough number of cases can be collected through their partnerships with these sites from important demographic subgroups (e.g., defined by ancestry, gender, age, geography) to make statistically meaningful comparisons of the performance of the AI-DST among these subgroups. The partnering hospitals or clinical sites should be varied in terms of their size, geographic location and demographics they serve, so that potential AI-DST performance differences among these sites can be measured.

During the program, TA1 performers will establish data use agreements with clinical sites to provide the necessary clinical data for the AI-DST use cases selected by TA2 performers. These use cases will fall into pre-defined clinical tracks, such as X-ray, CT, and EHR (Figure 4). The portfolio of clinical sites should allow for unbiased data collection in terms of patient demographics for the selected use cases.

Innovations from each TA will be tested across the array of hospital partners that have been assembled by TA1 performers. This real-world testing will enable improved AI performance monitoring and help demonstrate improved clinical decision-making.

TA1 proposals will address the following topics:

1. Describe the necessary input data types and quantify the number of patient records needed to enable surrogate ground truth label inference with statistical accuracy and precision for each use case.

2. Strong proposals will describe innovative, comprehensive surrogate ground truth labeling methods capable of spanning different health centers/clinics.

3. Strong proposals will describe generalizable methods that will be used for automatically extracting data from clinical reports, defining the surrogate ground truth labels, and explaining how the methods will be validated.

4. Strong proposals will limit the time horizon for data elements to be identified, so that model performance monitoring can be performed in as little time as possible after the acquisition of the input data for the AI-DST.

5. Describe supplementary metrics, if necessary, that may enable stakeholders to measure progress more accurately towards program goals.

6. Identify the clinical sites required in Phase I and provide evidence of their willingness to partner throughout the program. Strong proposals will identify more clinical sites than required in Phase I.

7. Include the historical percentage of patients at each site that belong to underrepresented populations, defined by ancestry, age, gender, geography, healthcare access and utilization.

Describe how the clinical site portfolio allows for unbiased data collection in terms of patient demographics for the selected use cases, comparing the demographics of patient sets to be collected to national averages to demonstrate the diversity of the overall patient populations and minimal data collection bias.

8. Strong proposals will select at least one site in Phase I and two sites in Phase II that have a preponderance of underrepresented populations and describe the approach to establish data use agreements with a growing number of clinical sites that incorporate these populations.

9. Provide a sharing plan that allows for interoperable data sharing of AI-DST input data, output data, and relevant metadata to program data repositories (e.g., TA4).

The following are out of scope for TA1 proposals:

• Training / developing new AI-DST models. Although it is permissible to use the data collected at the clinical sites to improve the AI-DST model performance, the improvement targeted in the overall program is with respect to performance degradations caused by dataset shifts in existing AI-DST models identified by TA2 performers.

1.1.4. TA2: Degradation Detection & Self-Correction

TA2 will develop AI auto-alert and auto-correction techniques to enable continuous end-to-end performance monitoring and self-corrective updating of clinical AI models by developing and combining capabilities for degradation detection, root cause analysis, and intelligent update implementation. TA2 performers will cover all three sub-TAs (2.1, 2.2, and 2.3), which are designed to seamlessly integrate, providing a comprehensive system that automatically identifies, analyzes, and corrects clinical AI model performance issues, thus maintaining the integrity and accuracy of AI applications in clinical settings. In their proposals, TA2 proposers should break down technical approach descriptions, cost estimates, and key personnel descriptions to the sub-TA level.

TA2.1: Continuous Degradation Detection Tools Strong TA2.1 proposals will describe innovative, scalable methods for accurate, timely, and granular detection of AI model degradation in clinical settings. Proposers are encouraged to develop a comprehensive framework for the continuous and automatic monitoring of key performance metrics, such as sensitivity, specificity, positive predictive value and predicted positive rate across various clinical sites. Advanced statistical analyses and machine learning techniques that distinguish between normal variability and genuine performance degradation are of interest. Relevant approaches include, but are not limited to, Statistical Process Control (SPC) charts10,11, Sequential Probability Ratio Test (SPRT)12 and Bayesian methods13.

Proposers are encouraged to address challenges around the unique statistical distribution of AI performance metrics, clinical factors for alarm threshold selection, allowable statistical variation, and patient volume.

Leveraging the metadata collected in TA1, TA2.1 proposers should describe a statistically robust approach to model degradation detection across clinically relevant subpopulations. TA2.1 degradation alerts should be more nuanced than a binary decision, so alerts may communicate that an AI-DST works well for one subgroup and not another. The approach should provide inputs to the root-cause analysis in TA2.2.

TA2.1 proposers can also consider how to detect changes in the statistical properties of input data to the AI model, such as variations in the preprocessing of input images. By combining detected changes at both the

10 Woodall, W. H. (2006). The Use of Control Charts in Healthcare and Public-Health Surveillance. Journal of Quality Technology, 38(2), 89–104. https://doi.org/10.1080/00224065.2006.11918593 11 Lowry, C. A., Woodall, W. H., Champ, C. W., & Rigdon, S. E. (1992). A Multivariate Exponentially Weighted Moving Average Control Chart. Technometrics, 34(1), 46–53. https://doi.org/10.1080/00401706.1992.10485232 12 Grigg OA, Farewell VT, Spiegelhalter DJ. Use of risk-adjusted CUSUM and RSPRT charts for monitoring in medical contexts. Stat Methods Med Res. 2003 Mar;12(2):147-70. doi: 10.1177/096228020301200205. PMID: 12665208.

13 Predictive Control Charts (PCC): A Bayesian approach in online monitoring of short runs. Journal of Quality Technology, 54(4), 367–391. https://doi.org/10.1080/00224065.2021.1916413 input and output of the model, it may be possible to improve performance degradation detection, resulting in methods with fewer false alarms and fast detection of true degradations. To optimize the methods for model performance degradation detection and to cover combinations of changes of the AI model input data and AI model output, performers can rely on not only clinical data but also simulation studies that use synthetic data. Proposers who adopt this approach are encouraged to develop innovative methods that use realistic synthetic data as well as clinically obtained data. Proposers who intend to use generative methods to obtain synthetic data are encouraged to consider using multiple fidelity metrics, such as the Frechet Inception Distance, to evaluate and ensure synthetic data quality.

TA2 performers will work with TA1 clinical sites to continuously and automatically monitor key performance metrics such as sensitivity, specificity, positive predictive value and predicted positive rate, enabling timely assessment of AI models across various clinical sites. TA2.1 will leverage scalable, automated surrogate ground truth labels extracted in TA1 to perform continuous and automated performance assessments. Furthermore, with the data-sharing tools implemented in TA4, continuous performance assessments can be conducted on an aggregated basis across multiple clinical sites.

TA2.1 proposals should describe a strategy to continuously monitor AI model performance at local and global scales, with few false alarms while ensuring quick and granular detection of deviations, particularly for specific patient subpopulations and clinical scenarios. TA2.1 proposals should explain how the team plans to compare local and site-aggregated performances and integrate sub-population monitoring features into their continuous monitoring activities, enabling the assessment of AI model performance across different patient demographics and clinical contexts. This will ensure that the AI models are effective for all patient populations and all clinical settings in which they are used.

Starting early in Phase I, TA2 performers will provide access to AI-DSTs for two use cases from different clinical tracks (Figure 4) to TA1 performers for testing. In Phase II, TA2 performers will provide access to AI-DSTs for two additional use cases to TA1 performers for continued testing.

TA2.2: AI-Based Root-Cause-Analysis Tools TA2.2 will develop AI-based root-cause-analysis tools for specified clinical use-cases. Performers will focus on identifying and testing potential root causes for AI model performance degradation across various use cases in the priority clinical tracks (Figure 4). Strong proposals will identify compelling strategies to identify root causes. Performers are encouraged to develop simulation frameworks or other approaches that mimic various data distribution changes, such as alterations in model inputs and ground truth labels14. Performers are encouraged to address common sources of dataset shifts including changes in patient demographics, data acquisition, storage and preprocessing, electronic health record management, and clinical definitions. In addition, proposers should describe how a diverse set of subject matter experts, including clinicians, AI developers, hospital administrators, and regulators, will contribute their insights to identify the set of potential root causes for performance degradation. TA2.2 proposers are encouraged to explain how large-scale simulations, or other approaches will enable testing of various types of changes to model inputs and ground truth labels to elucidate the effects of potential degradation scenarios and potential root causes.

Additionally, proposers should describe how simulations accurately reflect real-world scenarios, including the common sources of dataset shifts and the potential degradation root causes identified by subject matter experts.

Performers will adapt and extend various machine learning approaches to precisely attribute observed performance degradation…

This is the start of the file's text. The full file is on GovTribe.

File details come from the government source that posted it. Updated .