HR001122S0031.pdf

PDF 3 MB Posted

Attached to
In the Moment (ITM) Federal contract opportunity
Solicitation number
HR001122S0031
Issued by
Defense Advanced Research Projects Agency

About this file

This is a Broad Agency Announcement from the Defense Advanced Research Projects Agency seeking proposals for research and technology development supporting the building, evaluation, and fielding of algorithmic decision-makers that can assume human-off-the-loop decision-making responsibilities in difficult domains such as combat medical triage. The program, called In the Moment, will focus on small unit triage in austere environments in Phase 1 and mass casualty triage in Phase 2. Proposals are sought across four technical areas: decision-maker characterization, human-aligned algorithmic decision-makers, evaluation, and policy and practice integration. Multiple awards are anticipated for the characterization and algorithms areas, with single awards for evaluation and policy. Proposals are due May 17, 2022. The effort will run 42 months with a 24-month base period and optional 18-month extension. Deliverables include quarterly and final reports as well as software, documentation, and other materials.

View the file

Other files for this federal contract opportunity

Other files attached to In the Moment (ITM), newest first.
File Type Posted
ITM_Attachment_A_ABSTRACT_SUMMARY_SLIDE_TEMPLATE.pptx PPTX presentation
ITM_Attachment_G_PROPOSAL_TEMPLATE_VOL._3_ADMIN_NATL_POLICY_REQ.docx DOCX document
ITM_Attachment_F_MS_ExcelTM_DARPA_COST_PROPOSAL_SPREADSHEET.xlsx XLSX spreadsheet
ITM_Attachment_B_ABSTRACT_TEMPLATE.docx DOCX document
ITM_Attachment_E_PROPOSAL_TEMPLATE_VOL._2_COST.docx DOCX document
ITM_Attachment_D_PROPOSAL_TEMPLATE_VOL._1_TECH_MGMT.docx DOCX document
ITM_Attachment_C_PROPOSAL_SUMMARY_SLIDE_TEMPLATE.pptx PPTX presentation

On GovTribe

Work with this file on GovTribe

  • Download the original file
  • Contacts named in this file
  • Similar government files
  • Ask GovTribe AI about this file

Text version

HR001122S0031 IN THE MOMENT 1

Broad Agency Announcement

In the Moment (ITM) Defense Sciences Office

HR001122S0031

March 14, 2022

HR001122S0031 IN THE MOMENT 2

Table of Contents I. Funding Opportunity Description

A. Introduction B. Background

ITM definitions Limitations of a Current Autonomy Approach for Delegated Decision-Making

Motivating Examples for ITM Fully-autonomous triage decision-maker for small military units Fully-autonomous triage for a combat support hospital commander

C. Program Description/Scope D. Program Structure E. Technical Area Descriptions

Technical Area 1: Decision-maker characterization Technical Area 2: Human-aligned algorithmic decision-makers Technical Area 3: Evaluation Technical Area 4: Policy & practice integration Technical Area Interactions Common Proposal Elements CUI and CTI

F. Schedule/Milestones G. Deliverables H. Government-furnished Property/Equipment/Information I. Other Program Objectives and Considerations

II. Award Information A. General Award Information B. Fundamental Research

III. Eligibility Information A. Eligible Applicants B. Organizational Conflicts of Interest C. Cost Sharing/Matching D. Ability to Receive Awards in Multiple Technical Areas - Conflicts of Interest E. Ability to Support Classified Development

IV. Application and Submission Information A. Address to Request Application Package B. Content and Form of Application Submission C. Submission Dates and Times D. Funding Restrictions E. Other Submission Requirements

V. Application Review Information A. Evaluation Criteria B. Review and Selection Process C. Countering Foreign Influence Program (CFIP)

HR001122S0031 IN THE MOMENT 3

D. Federal Awardee Performance and Integrity Information (FAPIIS) VI. Award Administration Information

A. Selection Notices B. Administrative and National Policy Requirements C. Reporting

VII. Agency Contacts VIII. Other Information

A. Proposers Day B. Frequently Asked Questions (FAQs) C. Collaborative Efforts/Teaming D. Sample ACA Clause

BAA Attachments:

Attachment A: ABSTRACT SUMMARY SLIDE TEMPLATE Attachment B: ABSTRACT TEMPLATE Attachment C: PROPOSAL SUMMARY SLIDE TEMPLATE Attachment D: PROPOSAL TEMPLATE VOLUME 1: TECHNICAL & MANAGEMENT Attachment E: PROPOSAL TEMPLATE VOLUME 2: COST Attachment F: MS ExcelTM DARPA COST PROPOSAL SPREADSHEET Attachment G: PROPOSAL TEMPLATE VOLUME 3: ADMINISTRATIVE & NATIONAL POLICY REQUIREMENTS

HR001122S0031 IN THE MOMENT 4

PART I: OVERVIEW INFORMATION

Federal Agency Name: Defense Advanced Research Projects Agency (DARPA), Defense Sciences Office (DSO)

Funding Opportunity Title: In the Moment (ITM)

Announcement Type: Initial Announcement

Funding Opportunity Number: HR001122S0031

Catalog of Federal Domestic Assistance (CFDA) Number(s): 12.910 Research and Technology Development

Dates (All times listed herein are Eastern Time.)

o Posting Date: March 14, 2022 o Proposers Day: March 18, 2022. See Section VIII.A.

o Abstract Due Date: March 30, 2022, 4:00 p.m.

o FAQ Submission Deadline: May 2, 2022, 4:00 p.m. See Section VIII.B.

o Full Proposal Due Date: May 17, 2022, 4:00 p.m.

Anticipated Individual Awards: DARPA anticipates multiple awards for Technical Area 1 and 2 and a single award each for Technical Area 3 and 4.

Types of Instruments that May be Awarded: Procurement contracts, grants, cooperative agreements or Other Transactions. Award instruments will be limited to procurement contracts and Other Transactions for proposers whose proposed solution includes Controlled Unclassified Information (CUI).

Agency contacts o Technical POC: Matt Turek, Program Manager, DARPA/DSO o BAA Email: ITM@darpa.mil o BAA Mailing Address:

DARPA/DSO

ATTN: HR001122S0031

675 North Randolph Street Arlington, VA 22203-2114 o DARPA/DSO Opportunities Website: http://www.darpa.mil/work-with-us/opportunities

Teaming Information: See Section VIII.C for information on teaming opportunities.

Frequently Asked Questions (FAQ): FAQs for this solicitation may be viewed on the DARPA/DSO Opportunities Website. See Section VIII.B for further information.

Security: ITM is a basic research program that should not require performer access to CUI or Controlled Technical Information (CTI). DARPA anticipates that proposals will be unclassified. See Sections IV.B.4 and IV.B.5 for more details.

mailto:ITM@darpa.mil https://www.darpa.mil/work-with-us/opportunities?oFilter=DSO https://www.darpa.mil/work-with-us/opportunities?oFilter=DSO

HR001122S0031 IN THE MOMENT 5

PART II: FULL TEXT OF ANNOUNCEMENT

I. Funding Opportunity Description

This Broad Agency Announcement (BAA) constitutes a public notice of a competitive funding opportunity as described in Federal Acquisition Regulation (FAR) 6.102(d)(2) and 35.016 as well as 2 C.F.R. § 200.203. Any resultant negotiations and/or awards will follow all laws and regulations applicable to the specific award instrument(s) available under this BAA, e.g., FAR

15.4 for procurement contracts.

A. Introduction

The Defense Advanced Research Projects Agency (DARPA) Defense Sciences Office (DSO) is soliciting innovative research proposals for research and technology development that supports the building, evaluating, and fielding of algorithmic decision-makers that can assume human-off-the-loop decision-making responsibilities in difficult domains, such as combat medical triage.

Difficult domains are those where trusted decision-makers disagree; no right answer exists; and uncertainty, time-pressure, resource limitations, and conflicting values create significant decision-making challenges. Other examples of difficult domains include first response and disaster relief. Two specific domains have been identified for this effort - small unit triage in austere environments and mass casualty triage.

The Department of Defense (DoD) continues to expand its usage of Artifical Intelligence (AI) and computational decision-making systems. DoD missions involve making many decisions rapidly in challenging circumstances and algorithmic decision-making systems could address and lighten this load on operators. In order to employ such systems, the DoD needs rigorous, quantifiable, and scalable approaches for building and evaluating these systems. Current AI evaluation approaches often rely on datasets such as ImageNet1 for visual object recognition or the General Language Understanding Evaluation (GLUE)2 for Natural Language Processing (NLP) that have well defined ground-truth, because human consensus exists for the right answer.

In addition, most conventional AI development approaches implicitly require human agreement to create such ground-truth data for development, training, and evaluation. However, establishing conventional ground truth in difficult domains is not possible because humans will often disagree significantly about the right answer. Rigorous assessment techniques remain critical for difficult domains; without them, the development and fielding of algorithmic systems in such domains is untenable. In the Moment (ITM) seeks to develop techniques that enable building, evaluating, and fielding trusted algorithmic decision-makers for mission-critical DoD operations where there is no right answer and, consequently, ground truth does not exist.

Specifically, DARPA seeks capabilities that will (1) quantify the alignment of algorithmic decision-makers with key decision-making attributes of trusted humans; (2) incorporate key human decision-maker attributes into more human-aligned, trusted algorithms; (3) enable the evaluation of human-aligned algorithms in difficult domains where humans disagree and there is

1 ImageNet Large Scale Visual Recognition Challenge. Russakovsky, Olga, et al. International Journal of Computer Vision, 2015, Vol. 115.

2 GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. Wang, Alex, et al.

New Orleans, LA : International Conference on Learning Representations, 2019.

HR001122S0031 IN THE MOMENT 6

no right outcome; and (4) develop policy and practice approaches that support the use of human-aligned algorithms in difficult domains. Proposed research should embody innovative approaches that enable revolutionary advances in the current state of the art. Specifically excluded is research that primarily results in simply evolutionary improvements to the existing state of the art.

B. Background

ITM definitions The following terms are used throughout the BAA and are defined as follows.

Triage: The system of sorting and prioritizing casualties based on the tactical situation, the mission, and the available resources.3 In ITM triage also incorporates limited treatment options.

Mass Casualty: An event that overwhelms immediately available medical capabilities to include personnel, supplies, and/or equipment,4

Difficult decision-making: Decision-making in situations where trusted decision-makers frequently disagree; no right answer exists; and uncertainty, time-pressure, limited resources, and conflicting values create significant decision-making challenges.

Algorithmic decision-maker: A software implementation of a decision-making process.

Also referred to as a decision-making algorithm.

Domain: An application area with associated knowledge sources. Example: In Phase 1, small unit triage in an austere environment. Associated knowledge would include information such as resources, training, standard operating procedures, and treatment options.

Scenario: A scenario takes place within a domain and specifies the conditions of the environment and one or more situations of interest and forms the basis for a probe or series of probes.

Probe: A probe is a decision-making query posed to a decision-maker (algorithmic or human). Probes may be in the form of forced-choice or open-ended questions. Probes are designed to elicit information about underlying decision-maker attributes.

Scenario environment: The scenario environment is the mechanism for instantiating a scenario. Examples of different scenario environments would be text based, audio-visual, and simulated gaming environments.

Decision-maker attribute: Characteristics of decision-makers indicative of their outcome preferences and/or their decision-making process. Examples include risk-seeking vs. risk-

3 Mass Casualty and Triage, Emergency War Surgery Course, Joint Trauma System, Defense Health Agency https://jts.amedd.army.mil/assets/docs/education/ewsc/Mass_Casualty_Triage_EWSC_1.0.pdf 4 Ibid https://jts.amedd.army.mil/assets/docs/education/ewsc/Mass_Casualty_Triage_EWSC_1.0.pdf

HR001122S0031 IN THE MOMENT 7

aversion5 or maximizing vs. satisficing.6

Decision-maker descriptors: Computational representations that capture attribute information from decision-makers’ responses to one or more probes. Descriptor representations should also support computing the reference distribution. While the term descriptor in the machine learning community typically indicates a vector-based representation of features, ITM proposers may use any representation that meets the computational goals specified in the BAA.

Reference pool of decision-makers: Trusted human decision-makers with expertise in the domain and potential scenarios. The responses of these decision-makers to probes may be used to train, tune, or evaluate algorithmic decision-makers, and establish reference distributions.

Attribute space: A mathematical space defined across multiple dimensions of decision-maker attributes.

Reference distribution: A probability distribution defined on the attribute space that captures the prevalence of decision-maker attributes.

Decision-maker under test: The decision-maker under test is the system whose alignment is being compared against a reference decision-maker or reference pool of decision-makers.

Alignment score: Comparison measure of the decision-maker descriptor from a decision-maker under test and the reference distribution. Notionally an alignment score ranges from 0 (no alignment) to 1 (fully aligned).

Forced-choice: A probe that requires a decision, or choice, be made between a finite number of alternatives (usually a small number of options).

Open-ended: A probe for which no alternatives are provided, and the decision-maker must produce their own probe response.

Psychological fidelity: A concept that captures how closely a scenario engages human decision-makers in the same mental processes and stressors experienced in the real-world decision-making.

DevEthOps: A set of practices that considers and tests the potential legal, moral, and ethical (LME) implications of design choices during DevOps (development & operations) cycles.

Difficult Decisions and a Basis for Trust

Difficult decisions occur when the decision-maker is confronted with challenges that include too many or too few options, too much or too little information, uncertainty about the outcome that may result from a course of action, uncertainty about how to value foreseeable outcomes, 5 Prospect Theory: An Analysis of Decision under Risk. Kahneman, Daniel and Tversky, Amos. The Econometric Society, 1979, Vol. 47.

6 Police Perfection: Examining the Effect of Trait Maximization on Police Decision-Making. Shortland, Neil, Thompson, Lisa and Alison, Laurence. Frontiers in Psychology, 2020, Vol. 11.

HR001122S0031 IN THE MOMENT 8

resource limitations, and conflicts in core values.7 Humans often disagree about the right course of action or the right answer when faced with these difficult decisions. The lack of a right answer, i.e., the lack of ground truth in algorithm terms, undermines a typical assumption in how we build and assess algorithmic systems today, namely that ground truth outcomes are available for algorithm development, evaluation, and certification. In terms of trust, the performance numbers that result from comparing system decisions against ground truth are often relied on to decide whether to operationalize a system. In difficult domains, this sort of ground truth is not available and a different basis for trust must be found.

ITM seeks to use the algorithmic expression of key human attributes as the basis for trust in algorithmic decision-makers. ITM will investigate this basis for trust in the context of human-off-the-loop decision-making in difficult domains and seeks to enable the development, evaluation, and fielding of algorithmic decision-makers in difficult domains. ITM will develop a computational framework for key human attributes and a quantitative alignment score in order to assess the alignment of an algorithmic decision-maker to key decision-makers. Leveraging this computational framework, ITM will develop algorithms that express key human decision-making attributes and explore the alignment framework as a method for establishing appropriate trust in algorithmic decision-making systems.

The Role of Trust in Delegation

ITM is interested in a specific notion of trust, specifically the willingness of a human to delegate difficult decision-making to an algorithmic system. Mayer et al. 8,9 defines trust as “the willingness of a party to be vulnerable to the actions of another party based on the expectation that the other will perform a particular action important to the trustor, irrespective of the ability to monitor or control that other party” and provides one of many trust models in the research literature, as illustrated in Figure 1. Factors of perceived trustworthiness include key attributes or characteristics of the entity being trusted. While the trust literature often identifies technical performance characteristics (e.g., error rate) as factors of trust for autonomous systems, ITM is most interested in human attributes and characteristics (e.g., risk tolerance vs. risk aversion, maximizing vs. satisficing behavior, or other personality characteristics; subject matter expertise;

and human values to name a few) that could be encoded into algorithmic systems. The trustor’s propensity, based on prior experiences, may bias them towards or away from trusting the other entity. Both the trustor’s propensity and the trustee’s factors of perceived trustworthiness modulate the existence and degree of appropriate trust in the moment of a difficult decision. The trustor makes a perceived risk calculation, and, if trust overcomes the perceived risk, they will take a risk in a relationship, leading to an outcome that may impact the factors of perceived trustworthiness. In the context of ITM, key human attributes contribute to the factors of perceived trustworthiness, and the quantitative alignment score between the trustor and trustee is a proxy for the perceived risk.

7 Shortland, Neil D., Alison, Laurence J. and Moran, Joseph M. Conflict, How Soldiers Make Impossible Decisions.

New York, NY: Oxford University Press, 2019.

8 Mayer, Roger C., James H. Davis, and F. David Schoorman. "An integrative model of organizational trust."

Academy of management review 20.3 (1995): 709-734.

9 Kohn, Spencer C., et al. "Measurement of Trust in Automation: A Narrative Review and Reference Guide."

Frontiers in psychology 12 (2021).

HR001122S0031 IN THE MOMENT 9

Operationalizing the envisioned alignment framework will require computational methods to identify key decision-maker attributes and quantify the alignment between the algorithmic decision-making system and human decision-makers. ITM will focus on human-off-the-loop, algorithmic decision-making in difficult domains to understand the limits of such a computational framework. In order to scope the research, ITM has identified two specific difficult domains that represent real-world DoD and civilian concerns and will force developers to grapple with central issues in difficult decision-making - small unit triage in austere environments in Phase 1; mass casualty triage in Phase 2.

Figure 1 - Model of the trust process

Limitations of a Current Autonomy Approach for Delegated Decision-Making DoD missions create unique operational needs that may not be well supported by current approaches in industry. The decision-making requirements for self-driving cars are an example of delegating to algorithms in a domain where difficult scenarios can occur. For self-driving cars, one can encode risk values into the simulation environment that trains the algorithm.10,11 These risk values are chosen a-priori, in a process separate from end users, and hard-code an objective function, requiring policy and developer consensus at training time. By baking the behavior into the algorithm, there is no mechanism for an algorithm to adapt to changes in desired behavior or to situational guidance. This may be a reasonable approach for self-driving vehicles, where the rules of the road are generally static from day-to-day. However, hard-coded objective functions with their one-size-fits-all approach do not address DoD needs, where rules of engagement and commander’s intent vary from situation to situation and may evolve rapidly within a dynamic scenario.

Figure 2 illustrates Mayer’s trust model applied to the self-driving car scenario. When a driver delegates decision-making to the vehicle, there is no ability to independently characterize the decision-maker attributes of the self-driving car algorithm, making the perception of risk uncalibrated based on unknown alignment between the driver’s decision-making preferences and

10 Nelson, Gabe. Self-driving cars make ethical choices. Automotive News. [Online] July 13, 2015.

https://www.autonews.com/article/20150713/OEM06/307139936/self-driving-cars-make-ethical-choices.

11 Hyatt, Kyle. Waymo's simulators are doing 100 years of driving per day, even while working from home. CNET.

[Online] April 28, 2020. https://www.cnet.com/roadshow/news/waymo-100-years-simulation-daily-self-driving-car/.

https://www.autonews.com/article/20150713/OEM06/307139936/self-driving-cars-make-ethical-choices https://www.cnet.com/roadshow/news/waymo-100-years-simulation-daily-self-driving-car/

HR001122S0031 IN THE MOMENT 10

that of the vehicle. The driver’s trust in the vehicle is initially based on their confidence in the vehicle developer and then evolves with experience. From a societal perspective, car insurance processes may mitigate misalignments between the driver and the vehicle’s decisions. However, equivalent insurance processes do not exist for difficult DoD domains.

Figure 2 - Trust model applied to self-driving cars

In contrast to this model, the ITM program seeks insight into the factors of perceived trustworthiness for algorithmic decision-makers, mechanisms for assessing the alignment of decision-makers, and methods for improving the alignment between an algorithmic decision-maker and a group of trusted decision-makers or a specific decision-maker.

Motivating Examples for ITM The following examples illustrate the need for algorithms that can be aligned with a group of humans or with a specific, individual human. It is not intended that ITM performers will directly develop the applications in the examples. Instead, they are meant to illustrate some of the long-term needs ITM research must address.

Fully-autonomous triage decision-maker for small military units If successful, ITM technology could enable military personnel with limited medical training to conduct triage decision-making in the field. Such technology could significantly improve triage decision-making at point-of-injury, through decision-making built to model a group of carefully chosen human decision-makers who represent highly-experienced, capable, and trusted triage decision-makers. This group of human decision-makers would become the standard that the triage decision-making algorithms seek to emulate and operationalize across the entire force. It is important to note that the members of any trusted pool of decision-makers may exhibit different decision-making attributes, and ITM technology will need to support capturing and modeling that variability without requiring consensus from the trusted humans.

Fully-autonomous triage for a combat support hospital commander ITM technology may also enable the development of a fully-autonomous triage decision-maker, fine-tuned to a particular unit commander, in a combat support hospital (CSH) setting. In this scenario, there is a senior leader (such as a Colonel trained as a trauma surgeon) with appropriate authorities and exquisite training responsible for the medical decision-making within the CSH.

ITM technology would enable fine-tuning the decision-making algorithm so that it is aligned with the key decision-making attributes of the specific CSH leader. Due to their position, HR001122S0031 IN THE MOMENT 11 authorities, and training, that leader would likely be held responsible for the decisions made within the CSH, including those of any autonomous system. In order for that leader to trust the algorithmic decision-making system enough to employ it in a human-off-the-loop manner, the algorithm must be highly aligned with that individual. To build appropriate trust, ITM envisions supporting an interactive process between an algorithm and the senior leader prior to operational use that enables fine-tuning the algorithm to be aligned with that individual.

C. Program Description/Scope

ITM is 3.5-year, two-phase program with a 24-month Phase 1 (base) and an 18-month Phase 2 (option). This BAA is soliciting for only Phase 1 and 2. A notional 12-month third phase is envisioned should funding be secured. ITM will focus development in four distinct technical areas (TAs), which will proceed through Phase 1 and 2, and the notional Phase 3. Decision-maker characterization (TA1), Human-aligned algorithms (TA2), Evaluation (TA3), and Policy & Practice (TA4). TA1 and TA2 address key elements of trust: TA1 will develop the underlying theory and technologies for quantitatively characterizing decision-makers in difficult domains and assessing alignment between decision-makers. TA2 will develop approaches for building human-aligned algorithmic decision-makers for difficult domains. TA3 will be responsible for designing and executing overall program evaluations, evaluating both TA1 and TA2. TA4 will provide expertise in DoD policy related to autonomous decision-making and Legal, Moral, and Ethical (LME) considerations. All four technical areas are being competed via this BAA.

Proposers may propose to multiple TAs. In that case, proposers should submit a separate proposal for each TA. To prevent a conflict of interest, a single performer will not be awarded both a TA3 effort and either a TA1 or TA2 effort. For scoping plans and costs for interactions with other TAs, proposers should anticipate two awards for TA1, two awards for TA2, a single award for TA3, and a single award for TA4.

The ITM program domains will be small unit triage in austere environments in Phase 1 and mass casualty triage in Phase 2. Proposers will need general knowledge of triage as well as domain-specific knowledge in order to be successful. Research addressing ITM’s goals will require, beyond triage domain knowledge, expertise and advances in decision-making, cognitive science, experimental psychology, simulation environments, data science, artificial intelligence, machine learning, evaluation, and decision-making policy for autonomous systems. Strong proposals will include cross-disciplinary research efforts supported by a cross-disciplinary team with relevant expertise.

Out of Scope Research in the following areas is considered out-of-scope for purposes of ITM:

Approaches that develop hardware.

Approaches that directly implement the Motivating Examples for ITM.

Approaches that develop sensing or perception techniques for triage environments.

Approaches that apply only to human decision-making and do not support the development of algorithmic decision-makers.

Approaches that exclusively develop novel knowledge graphs or ontologies for triage domains.

HR001122S0031 IN THE MOMENT 12

Approaches that require humans-in-the-loop/humans-on-the-loop at program evaluation time or in operational use to make difficult decisions.

Approaches that require large numbers of human interactions during the alignment process with a group or single individual.

D. Program Structure

ITM’s four TAs are illustrated in Figure 3. This figure is not intended as a design framework or specification for an ITM system, and proposers should recommend additional interactions as needed in support of their research plans.

Figure 3 ITM Technical Areas (TAs) and key interactions

Figure 4 illustrates how TA1, Decision-maker characterization, and TA2, Human-aligned algorithmic decision-makers, relate to the trust model for delegated decision-making.

Figure 4 - ITM approach to delegation and alignment

Program Phases and Schedule:

HR001122S0031 IN THE MOMENT 13

ITM will use a phased-acquisition approach. The program will have two phases dedicated to technical development and a notional third phase to enable application of ITM capabilities to one (or more) domains of interest to U.S. Government transition partners.

Phase 1 will be 24 months in duration. Phase 1 will develop a proof of concept for ITM technologies, demonstrate the ability to measure alignment, and demonstrate the ability to tune algorithmic decision-maker alignment to a group of trusted human decision-makers.

The Phase 1 domain will be small unit triage in austere environments. By the end of this phase, ITM will have demonstrated the fundamental capabilities needed to:

o identify and characterize trusted decision-makers according to key human attributes o align algorithmic decision-makers in the Phase 1 domain with a group of trusted human decision-makers o evaluate decision-maker alignment with a group of trusted decision makers o engage with the relevant policy communities on ITM technology and begin development of a certification model for algorithmic decision-makers

Work beyond Phase 1 is subject to availability of funding and technical progress.

Phase 2 will be 18 months in duration. Phase 2 will build on the triage domain experience gained in Phase 1 and will expand the decision-making challenges to more complex mass casualty events. Phase 2 will focus on developing the capabilities necessary to fine-tune an algorithmic decision-maker to exhibit the attributes of an individual human decision-maker. By the end of this phase, ITM will:

o refine the capability to identify and characterize trusted decision-makers according to attributes in a second domain o align algorithmic decision-makers in the Phase 2 domain with an individual trusted human decision-maker o evaluate decision-maker alignment with an individual trusted decision maker o develop policy recommendations for ITM technology in collaboration with the relevant policy communities and produce a draft certification model for algorithmic decision-makers

If funded, the notional Phase 3 will be 12 months in duration. The domain for Phase 3 will be determined based on technical progress and Government partner needs. At the end of Phase 3, we anticipate having the ability to demonstrate ITM capabilities that support a domain chosen by a transition partner. Phase 3 will be contingent on the availability of funds, technical performance in prior phases, utility of technical approaches to transition partner use cases, and significant U.S. Government partner interest.

The period of performance will be the same for all performers across the technical areas.

Proposers should propose a base effort for Phase 1 and a Phase 2 option. A draft Statement of Work (SOW) and Rough Order of Magnitude (ROM) cost for Phase 3 will be a Phase 1 deliverable for all performers. This draft SOW and ROM will be used for budgeting purposes, and will not be evaluated. Should funding be identified for Phase 3, DARPA will issue Phase 3 Proposal Instructions during Phase 2 requesting proposals from performers whose Phase 2 options have been exercised. Phase 3 proposals will be evaluated against the criteria in the Phase 3 Proposal Instructions, which will be consistent with the criteria in this BAA. Participation in

HR001122S0031 IN THE MOMENT 14

Phases 2 and 3 is contingent upon successful performance in prior phases as well as availability of funds.

In order to avoid potential funding gaps between decisions regarding progression from one phase to the next and the execution of contract options, a decision on whether to continue individual teams’ efforts into Phase 2 are anticipated at roughly Month 21. Decisions regarding progression into Phase 3 are anticipated at roughly Month 40 of the program, which will be the 16th month of Phase 2. The final months of Phase 1 and Phase 2 will be used to prepare for the subsequent phase by refining research and evaluation plans and improving results. For all performers, a final report will be due 60 days after the last phase in which they participate.

E. Technical Area Descriptions

The sections below outline the program objectives for TAs 1, 2, 3 and 4, respectively. The proposer’s research plan must include a constructive task breakdown and plan for achieving the program goals, including interactions with the other TAs. Quantitative metrics for the program are discussed in the section on TA3.

Technical Area 1: Decision-maker characterization The focus of TA1 is developing technologies that identify and quantitatively model key decision-making attributes of trusted humans in order to produce a quantitative decision-maker alignment score.

TA1 proposals should identify the theory (or theories) of decision-making that will form the basis for identifying and quantitatively characterizing key decision-maker attributes. Theory selection should be informed by existing research. While targeted experiments on the program may serve to enhance the existing theories of decision-making, developing a wholly new decision-making theory is beyond the scope of ITM. The underlying theory of decision-making should be relevant for decision-making in difficult domains, particularly the small unit triage in austere environments (Phase 1) and mass casualty triage (Phase 2) domains that will be the experimental domains of ITM. Strong proposals will describe a theory that is also applicable across other important DoD domains.

TA1 is responsible for developing a mathematical and computational framework for quantitatively characterizing key decision-maker attributes from a group of trusted humans. This framework includes a decision-maker attribute space, computational decision-maker descriptors, and a quantitative alignment score.

The envisioned ITM framework (illustrated in Figure 5) starts with a reference pool of trusted human decision-makers. TA1 will need to secure its own human decision-makers with triage domain expertise for development of the characterization framework. During program evaluations, TA3 will provide access to trusted human decision-makers with relevant expertise in the ITM domains. The decision-maker pool provided by TA3 is anticipated to be small in size, perhaps on the order of 3 to 5 individuals, and is drawn from a population of humans that are already trusted to make critical decisions in ITM domains. There is no expectation that this trusted pool will agree or reach consensus, and the ITM characterization framework must represent both individual and group decision-maker variability.

HR001122S0031 IN THE MOMENT 15

Figure 5 - ITM Decision-maker characterization framework

TA1 will develop scenarios and probes for each decision-making domain designed to place a decision-maker in challenging situations. While it is the TA3 evaluation team’s responsibility to provide the scenario environment(s), TA1 must create scenarios and probes in the scenario environment that have psychological fidelity12,13,14 to real-world decisions to support the development of reference distributions for decision-maker attributes. TA1 teams are expected to execute development scenarios with decision-makers to provide multiple exemplar reference distributions to TA2 for algorithm development and tuning.

Examples of how to create psychological fidelity include creating time-pressure, controlling situational knowledge, and limiting potential choices. Scenarios and probes should be designed such that decisions in response to the probes reveal information about key decision-maker attributes. The types of attributes of interest must be identified in advance of scenario design and should be motivated by the TA1 team’s decision-making theory and based on models of human decision-making in difficult domains. Probes will primarily be forced-choice style questions, as constraining the possible choices will likely be important to extracting information about underlying decision-maker attributes. Probes will be presented to the members of the reference pool of trusted decision-makers in the scenario environment.

TA1 proposals should clearly describe what aspects of psychological fidelity are necessary in the scenario environment to effectively support their decision-making theory and planned scenarios.

12 Kozlowski, Steve WJ, and Richard P. DeShon. "A psychological fidelity approach to simulation-based training:

Theory, research and principles." Scaled worlds: Development, validation, and applications (2004): 75-99.

13 Quick, Jacob A. "Simulation training in trauma." Missouri medicine 115.5 (2018): 447.

14 Chiniara, Gilles, et al. "Moving beyond fidelity." Clinical Simulation. Academic Press, 2019. 539-554.

HR001122S0031 IN THE MOMENT 16

Each member of the reference pool is expected to make decisions independently of the other members of the reference pool.

TA1 should address the following additional challenges when designing scenarios and probes.

Algorithmic decision-makers will be presented the same set of scenarios and probes.

Responses by algorithmic decision-makers will also need to be captured in the computational representation and represented as decision-maker descriptors.

The approach needs to be sample-efficient (i.e., no more than hundreds of questions total across all scenarios in a domain).

Scenarios and probes will support the TA2 tasks of aligning to a group of decision-makers in Phase 1 and to a specific decision-maker in Phase 2.

Scenarios and probes should be constructed using the domain knowledge documents identified by the TA3 evaluation team and not require domain knowledge outside those documents.

Strong proposals will develop a scenario design process that results in scenarios and probes that some day could be used by service members in the field to fine tune decision-makers as the final step in accepting algorithmic decision-making systems for operational use.

Decision-maker responses to probes should be captured in a computational representation, shown in Figure 5 as decision-maker descriptors. TA1’s decision-maker descriptors should represent the presence and strength of key decision-maker attributes and enable computational analysis, such as computing the distribution of decision-maker attributes in the attribute space. This distribution will form the basis for quantifying whether an algorithm exhibits key decision-maker attributes that are similar to the reference pool of trusted decision-makers. Notionally, this distribution is expected to behave like a probability distribution. In the example in Figure 5, a notional two-dimensional attribute space is shown, with one axis defined by risk tolerant vs. risk avoiding behaviors and a second axis defined by maximizing vs. satisficing behaviors.

Proposers should define and justify their own attribute space based on their chosen theory of decision-making and should not limit themselves to the example distribution in the BAA.

Meaningful attribute spaces will likely have more dimensions than the two illustrated in the example. The decision-maker attribute space (or potentially multiple spaces) should capture key attributes of decision-makers that impact outcome preferences and decision-making process. An attribute space should have human-understandable attributes, such as sensitivity to risk or tendency to optimize on expected outcomes, that can guide selection of situation-appropriate decision-makers based on their key attributes. The definition of the attribute space must support the computation of a distribution over key attributes. Strong proposals will consider the impact of situational information, domain knowledge, and other contextual elements on decision-maker attributes and how that may affect decision-maker preferences. Proposers are free to include other elements in the reference distribution than those listed here, but should clearly describe why those elements are important to the decision-making process.

While the notional illustration of TA1 implies using feature vectors for decision-maker descriptors and a probability distribution for the distribution over the attribute space, proposers may choose other representations. If alternative representations are proposed, proposers should explain how their representation will support the development of TA2 algorithmic decision-makers that can be brought into alignment with selected decision-makers. The TA2 algorithmic

HR001122S0031 IN THE MOMENT 17

decision-makers will likely be designed to expect a probability distribution over decision-maker attributes, so TA1 proposals that depart from that representation should clearly describe what operations their representation will support and how those are similar to operations that are available on a probability distribution.

TA1 must develop a quantitative alignment score based on a decision-maker descriptor from a decision-maker under test and the reference distribution over the decision-maker attribute space.

The alignment score indicates how closely a decision-maker exhibits the attributes from a trusted pool of humans or from a single reference human decision-maker. The alignment score should be designed such that it informs perceived risk calculations and is indicative of end-user trust. In the notional example shown in Figure 5, algorithmic descriptors that map to the tails of the reference distribution indicate poor alignment between the algorithm and the trusted humans, i.e., the algorithm is not exhibiting the same sort of decision-maker attributes as the trusted humans.

Algorithmic descriptors that map to a high-density area of the reference distribution indicate good alignment between the algorithm and the trusted humans, i.e., the algorithm is exhibiting the same sort of decision-maker attributes as the trusted humans. The quantitative alignment score is notionally a continuous score between 0 (no alignment) and 1 (high alignment), but proposers are free to offer frameworks that provide alternate ranges. Proposers should specify how their alignment score will handle multiple modes in the reference distribution and how the alignment score will be designed to be correlated with measures of human trust. If proposers opt for multiple attribute spaces, such as organized by topics like domain knowledge or core values, they should describe the utility of the separate attribute spaces and how they will produce a summary alignment score across all the attribute spaces.

TA1 performers must collaborate with other TA1 and TA2 performers in the development of a program-wide standard Application Programming Interface (API) for decision-makers under test and an API (jointly referred to as the alignment and characterization API) for the reference distribution for trusted decision-makers. TA1 proposers should specify key functionality that they expect to be available in the reference distribution API. Each TA1 will be responsible for implementing the API for their reference distribution.

TA1 must participate in program evaluations. TA3 will evaluate TA1’s ability to characterize decision-makers and to generate meaningful alignment scores for an algorithmic decision-maker, such as those produced by TA2. TA1 will provide their implementation of the characterization framework to TA2 to enable algorithms that use the framework during development or evaluation. For instance, TA2 may interactively explore a set of probes to understand computationally how those probes map into the attribute space and how that impacts the alignment score. Given the interaction between TA1 and TA2 efforts, TA1’s responsibilities include collaboration with TA2 to provide their characterization framework and collaboration with TA3 in support of the evaluation. In support of evaluations, TA1 performers will also collaborate with other TA1, TA2, and the TA3 performers on computational representations for the scenarios, probes, and scenario context.

Strong TA1 proposals will:

Describe a decision-making theory for difficult decisions that is foundational for their characterization approach and framework and that applies across multiple DoD domains beyond the program specified domains.

HR001122S0031 IN THE MOMENT 18

Define a space of key decision-maker attributes that is informative of decision-maker behavior.

Consider the impact of situational information, domain knowledge, and other contextual elements on decision-maker attributes and how that may affect decision-maker preferences.

Create difficult decision-making scenarios with multiple probes that elicit decision-maker responses that provide insight into their decision-maker attributes.

Develop a scenario design process that results in scenarios and probes that someday could be used by service members in the field to fine tune decision-makers

Identify human decision-makers with relevant expertise in the ITM domains that will be used for development of the characterization framework.

Represent decision-maker responses to scenarios and probes in a computational framework.

Compute a reference distribution over decision-maker attributes for a small pool of trusted humans.

Compare a decision-maker under test to a reference distribution of trusted humans to compute a quantified alignment score between the test decision-maker and the reference distribution.

Provide for compute needs in support of framework development, internal testing, and program evaluations.

Have access to an Institutional Review Board (IRB) and experience with the process of acquiring and maintaining Government approval to conduct HSR. Significant Human Subjects Research (HSR) elements are anticipated in relation to decision-making theory refinement and computational framework development.

Technical Area 2: Human-aligned algorithmic decision-makers TA2 will develop human-aligned algorithms that leverage the TA1 computational characterization process and the quantitative alignment score (see blue panel in Figure 6). The human-aligned algorithms should be able to balance situational information with a preference for the key decision-maker attributes identified by TA1 and the reference distribution across the attribute space.

In the notional approach shown in Figure 6, a mathematical regularization approach is used to balance situational information with a preference for certain decision-maker attributes. Other mathematical formulations are encouraged, and proposers should NOT limit themselves to regularization approaches.

TA1 will provide their implementation of the characterization framework to TA2 to enable algorithms that use the framework during development or evaluation. For instance, TA2 may interactively explore a set of probes to understand computationally how those probes map into the attribute space and how that impacts the alignment score. TA2 algorithms will need to generalize to novel scenarios, probes, and reference distributions, such as those provided by multiple TA1 teams or those developed by TA3 to support evaluation. As a result, it will be critical for TA2 to develop approaches that can represent multiple types of key attributes and that

HR001122S0031 IN THE MOMENT 19

can build an understanding of how those key attributes relate to the structure of the scenarios and probes.

Figure 6 - ITM Alignment approach for an algorithmic decision-maker

To support the program goal of demonstrating human-aligned decision-makers for the program domains, TA2 algorithms must demonstrate the ability to align to the largest cluster within the reference distribution in Phase 1. In Phase 2, TA2 algorithms must demonstrate the ability to align with the decision-maker attributes from a single, trusted human. Alignment in this context means that the algorithm exhibits the same decision-maker attributes as the trusted humans and exhibits them in the same contexts, as decision-maker attributes may change across scenarios.

TA2 performers should expect that they will have to demonstrate alignment with multiple attribute spaces from different TA1 performers.

For the Phase 2 goal of fine-tuning an algorithm to a single trusted human decision-maker, it is assumed that the single decision-maker is similar to, or a member of, the trusted group of humans. As a result, the fine-tuning process may leverage the group alignment process as a good initialization. The expectation for fine-tuning is that there is less variability in a single human than there is across a group of humans, so the fine-tuning process may be more difficult than tuning to fit within a group’s attribute distribution. In future operational applications of ITM, the fine-tuning process would occur prior to operational deployment and involve a question-answer process between the algorithm and the human. For ITM, it is envisioned that TA2 will be provided with information on a particular human decision-maker by TA1 or TA3 in the form of the decision-maker attributes and their reference distribution. If needed for development, TA1 would have the responsibility of creating and providing scenarios and probes for the target individual to support the adaption of TA2 algorithms to the decision-making attributes of a particular human.

TA2 algorithms will answer forced-choice questions (probes) in the scenarios designed by TA1.

These questions will be the same questions used to elicit information from the trusted humans.

HR001122S0031 IN THE MOMENT 20

The responses will be used by TA1 when generating an alignment score with respect to a reference distribution. Strong TA2 proposals will describe how their systems will extend to scenarios and probes that allow open-ended answers by the end of Phase 2 and will describe how they will implement the alignment process envisioned for both Phase 1 and Phase 2.

TA2 performers will collaborate with TA1 performers and other TA2 performers to develop a program-wide standard interface and API (jointly referred to as the alignment and characterization API) for the reference distribution for trusted decision-makers. TA2 proposers should specify the key functionality that their algorithms will require in the reference distribution API. TA1 will be responsible for implementing the API for their reference distribution. TA2 teams will be responsible for using the API as part of their algorithm development to facilitate incorporating information from different reference distribution designs as a part of the alignment process.

TA2 algorithms will need to incorporate domain knowledge from a set of natural language documents provided by the TA3 evaluation team. Domain knowledge documents will define the scope of information needed to answer the probe questions. All sources should be in standard document formats, such as MS Word or Excel or Adobe PDF, that enable computation. TA1, TA2, and TA3 performers will collaborate to convert documents into a shared computable representation that can be ingested by the decision-making algorithms. If a unique representation is needed, TA2 will have primary responsibility for the conversion of domain knowledge into a format that supports their algorithm design. TA2 algorithms will be restricted to ITM-provided domain knowledge when responding to scenarios and probes.

TA2 performers must integrate the algorithmic decision-maker with the scenario environment provided by the evaluation team (TA3) in order to support execution of the scenarios and probes.

TA2 teams will collaborate with the TA1 performers and other TA2 performers on…

This is the start of the file's text. The full file is on GovTribe.

File details come from the government source that posted it. Updated .