HR001117S0031_Ground_Truth_BAA.pdf
PDF 501 KB Posted
- Attached to
- Ground Truth (GT) Federal contract opportunity
- Solicitation number
- HR001117S0031
View the file
Other files for this federal contract opportunity
| File | Type | Posted |
|---|---|---|
| HR001117S0031_Attachment_1_Abstract_Summary_Slide_Template.pptx | PPTX presentation | |
| HR001117S0031_Attachment_3_Proposal_Slide_Templates.pptx | PPTX presentation | |
| HR001117S0031_Attachment_4_Proposal_Template_Technical_&_Management_Volume.docx | DOCX document | |
| HR001117S0031_Attachment_5_Proposal_Template_Cost_Volume.docx | DOCX document | |
| HR001117S0031_Attachment_6_Proposal_Template_Administrative_&_National_Policy_Requirements.docx | DOCX document | |
| HR001117S0031_Attachment_2_Abstract_Template.docx | DOCX document |
On GovTribe
Work with this file on GovTribe
- Download the original file
- Contacts named in this file
- Similar government files
- Ask GovTribe AI about this file
Text version
HR001117S0031 GROUND TRUTH 1
Broad Agency Announcement
Ground Truth (GT) Defense Sciences Office
HR001117S0031
April 28, 2017
HR001117S0031 GROUND TRUTH 2
Table of Contents I. Funding Opportunity Description
A. Introduction B. Background C. Program Description/Scope D. Program Structure E. Technical Area Descriptions F. Schedule/Milestones G. Deliverables H. Other Program Objectives and Considerations
II. Award Information A. General Award Information B. Fundamental Research
III. Eligibility Information A. Eligible Applicants B. Organizational Conflicts of Interest C. Cost Sharing/Matching D. Other Eligibility Requirements
IV. Application and Submission Information A. Address to Request Application Package B. Content and Form of Application Submission C. Submission Dates and Times D. Funding Restrictions E. Other Submission Requirements
V. Application Review Information A. Evaluation Criteria B. Review and Selection Process C. Federal Awardee Performance and Integrity Information (FAPIIS)
VI. Award Administration Information A. Selection Notices B. Administrative and National Policy Requirements C. Reporting
VII. Agency Contacts VIII. Other Information
A. Frequently Asked Questions (FAQs) B. Collaborative Efforts/Teaming C. Proposers Day
ATTACHMENT 1: ABSTRACT SUMMARY SLIDE TEMPLATE
ATTACHMENT 2: ABSTRACT TEMPLATE
ATTACHMENT 3: PROPOSAL SUMMARY SLIDES TEMPLATE
ATTACHMENT 4: PROPOSAL TEMPLATE – TECHNICAL & MANAGEMENT VOLUME
ATTACHMENT 5: PROPOSAL TEMPLATE – COST VOLUME
ATTACHMENT 6: PROPOSAL TEMPLATE – ADMINISTRATIVE & NATIONAL POLICY REQUIREMENTS
HR001117S0031 GROUND TRUTH 3
PART I: OVERVIEW INFORMATION
• Federal Agency Name: Defense Advanced Research Projects Agency (DARPA), Defense Sciences Office (DSO)
• Funding Opportunity Title: Ground Truth (GT)
• Announcement Type: Initial Announcement
• Funding Opportunity Number: HR001117S0031
• Catalog of Federal Domestic Assistance (CFDA) Number(s): 12.910 Research and Technology Development
• Dates (All times listed herein are Eastern Time.)
o Posting Date: April 28, 2017 o Abstract Due Date: May 15, 2017, 4:00 p.m.
o FAQ Submission Deadline: June 22, 2017, 4:00 p.m. See Section VIII.A.
o Full Proposal Due Date: June 29, 2017, 4:00 p.m.
• Anticipated Individual Awards: DARPA anticipates multiple awards under both Technical Areas (TAs)
• Types of Instruments that May be Awarded: Procurement contracts, grants, cooperative agreements or other transactions.
• Agency contacts o Technical POC: Dr. Adam Russell, Program Manager, DARPA/DSO o BAA Email: GroundTruth@darpa.mil o BAA Mailing Address:
DARPA/DSO
ATTN: HR001117S0031
675 North Randolph Street Arlington, VA 22203-2114 o DARPA/DSO Opportunities Website: http://www.darpa.mil/work-with-us/opportunities
• Teaming Information: See Section VIII.B for information on teaming opportunities.
• Frequently Asked Questions (FAQ): FAQs for this solicitation may be viewed on the
DARPA/DSO Opportunities Website. See Section VIII.A for further information.
mailto:GroundTruth@darpa.mil http://www.darpa.mil/work-with-us/opportunities?tFilter=&oFilter=2&sort=name http://www.darpa.mil/work-with-us/opportunities?tFilter=&oFilter=2&sort=name
HR001117S0031 GROUND TRUTH 4
PART II: FULL TEXT OF ANNOUNCEMENT
I. Funding Opportunity Description
This Broad Agency Announcement (BAA) constitutes a public notice of a competitive funding opportunity as described in Federal Acquisition Regulation (FAR) 6.102(d)(2) and 35.016 as well as 2 CFR § 200.203. Any resultant negotiations and/or awards will follow all laws and regulations applicable to the specific award instrument(s) available under this BAA, e.g., FAR
15.4 for procurement contracts.
A. Introduction The Defense Sciences Office at the Defense Advanced Research Projects Agency (DARPA) is soliciting innovative research proposals in the area of new simulation capabilities to test the accuracy and robustness of causal modeling methods for understanding human social systems and behaviors. Proposed research should investigate innovative approaches that enable revolutionary advances in social science modeling, simulation, and causal inference.
Specifically excluded is research that primarily results in evolutionary improvements to the existing state of practice.
In particular, DARPA seeks to create artificial but socially plausible simulations that have known causal ground truth to validate the accuracy and robustness of social science modeling methods.
Ground truth simulations should allow for a wide range of qualitative, quantitative, and mixed social science modeling methods, and should provide different kinds of complexity to test causal modeling methods across a range of simulated behaviors and systems.
Using a series of staged tests, DARPA anticipates that these simulations will help quantify the capabilities and theoretical limitations of different modeling methods for explaining and predicting causal processes in complex social systems. Additionally, these simulations will provide opportunities to evaluate new modeling methods, or combinations of methods, to advance the rigor of causal inference and modeling in the pursuit of “solution-oriented” social sciences.
B. Background Military planners and decision-makers often rely upon the social sciences to help them understand and forecast a variety of scenarios that involve complex human social systems and behaviors. In particular, decision-makers often seek to identify, characterize, and model causal processes at different scales and for different social systems to help explain or predict certain patterns of behavior for a wide range of missions, including stability operations, humanitarian assistance, and counter-insurgency. However, human social systems and behaviors present enduring challenges for making “strong inference”1 about causal processes.
One of social scientists’ biggest challenges is often the lack of objective knowledge of the actual causes of observed behaviors (“ground truth”) in the real world. Conducting the experimental work necessary for understanding causality in social behaviors and systems is often impractical
1 Platt, JR. “Strong Inference.” Science 16 Oct 1964: Vol. 146, Issue 3642, pp. 347-353
HR001117S0031 GROUND TRUTH 5
or unethical, while observational (including “big data”) modeling approaches routinely use correlations to make conclusions about causality. These conclusions are often suspect due to inaccuracies and/or incompleteness inherent in social data. Hence the lack of causal ground truth limits capabilities to rigorously evaluate the explanatory and predictive accuracy of different modeling methods, particularly for causal processes at different scales - even as the need for such evaluation is growing in importance.2 Consequently, decision-makers cannot be confident in how accurate social science modeling methods are for making strong inference claims about causal processes in social systems and behaviors, or even if those are the correct modeling methods to use.
This challenge is exacerbated by the fact that human social systems often display emergent behaviors. These behaviors arise from dynamic, adaptive, longitudinal, multiscale interactions of different agents across different social structures with different social processes – all of which resist easy abstraction or simplification.
Enabled by increasing computational power and decreasing computational costs, social scientists are frequently incorporating simulation as a modeling method. However, these approaches generally suffer from the same validation challenges as other methods. For example, modelers often assume that if a simulation generates outcomes similar to observed behaviors, they can take this as evidence that their simulations have accurately captured candidate real-world causal mechanisms. Yet different simulations may often result in seemingly similar social phenomena.
Further, in the absence of causal ground truth, simulation validation approaches often end up being highly flexible, subjective, and ad hoc. As currently used, simulations cannot escape the same validation limitations that other modeling methods face.
C. Program Description/Scope DARPA posits that appropriately complex simulations might offer a high-risk, high-payoff opportunity for making significant advances beyond these current limitations in social science modeling. By using these simulations as social science modeling test beds, DARPA hypothesizes that there will be new opportunities to significantly enhance capabilities for evaluating the accuracy of different causal modeling methods.
The Ground Truth (GT) program is designed to test this hypothesis by creating and using artificial but socially plausible simulations with causal ground truth to quantify the explanatory and predictive performance of a range of social science modeling methods. Because causality is encoded into the simulations, DARPA and the teams creating simulations will have known causal ground truth. Modeling teams will then attempt to discover and predict causality in the simulations using their various methods. Based on the modeling teams’ success in meeting various metrics over a series of increasingly complex challenges, GT will afford unprecedented validation of quantitative, qualitative, and mixed research methods’ abilities to draw strong inference about causal mechanisms of different kinds of social behaviors under different conditions of complexity.
2 E.g., Hofman, JM, Sharma, A, Watts, DJ. “Prediction and explanation in social systems.” Science 03 FEB 2017:
486-488
HR001117S0031 GROUND TRUTH 6
Ground Truth Program Vision
Figure 1: Ground Truth Vision
As outlined in Figure 1 above, Ground Truth envisions a programmatic workflow that will result in currently unattainable knowledge about the explanatory and predictive accuracy–and theoretical limitations–of different social science modeling methods for different kinds of social complexity.
GT seeks to fund a number of teams to develop social simulations with causal ground truth that may each give rise to socially plausible but distinctive “alternate” worlds. These Ground Truth Simulators will use different kinds and combinations of first principles to form the causal rules and mechanisms that give rise to the observed simulated behaviors and systems in their particular alternate world. As the goal is not to reflect “real-world” first principles of social behaviors (since these are largely unknown), different Simulators are likely to use different principles to achieve the TA1 goals described in Section I.E, below. GT envisions three different phases of simulations, with TA1 teams creating or evolving greater systemic and behavioral complexity in their respective alternate worlds at each phase (so that the cycle starting with Ground Truth Simulators outlined in Figure 1 will occur three times during the Program, with increasing levels of simulated social complexity). GT envisions that this complexity may be achieved in principled but variable ways by different teams, such as increasing the number and kinds of dynamic interactions in a simulation, the complexity of agents or social structures, or increasing sources and levels of uncertainty in data and behaviors in that world.
At each phase, Ground Truth simulations are developed, deployed, and run in ways that enable modeling teams (see TA2 in Section I.E, below) to conduct “research” on, and in, those alternate worlds. GT foresees funding a number of researcher proposals with innovative approaches to rapid “solution-oriented” teaming, who then seek to discover the causal processes in these simulations. Researchers may observe behaviors in these alternate worlds, perhaps via datasets using web-based applications or a simulation observatory, as well as interact with them in necessarily limited but creative ways to use a wide range of social science research methods. A TA2 team, for example, might collect or be given various kinds and amounts of “observational”
HR001117S0031 GROUND TRUTH 7
data from different outputs of a certain simulation to enable regression-based statistical inference and machine-learning techniques as well as multi-level modelling and time-series analyses. At the same time, that alternate world might also allow TA2 researchers to interact with certain parts of the simulation – perhaps querying groups of agents on their state or intentions, or conducting limited experimental or qualitative work in a particular section of the simulation.
This interaction might let that team apply mixed methods network analyses, content or discourse analysis, grounded theory, or limited participant-observation. Using a combination of initial theorizing and modeling, that TA2 team then seeks to reverse-engineer the simulation’s first principles and hence explain that world’s causal rules and mechanisms.
TA2 teams will also test their modeling methods for predicting future states of the alternate worlds. GT envisions TA1 teams announcing an impending shock to their simulation, for example, by adding new agents to the simulation, removing certain resources, or simulating some other systemic disruption. TA2 teams will model the impact of this disruption on the alternate world in advance and then compare their predictions to the revealed actual outcomes in the simulations.
Finally, TA2 teams will seek to determine the limitations of their causal modeling of an alternate world by “prescribing” specific interventions to effect specific simulation states. GT envisions a TA2 team using their research methods to conclude, for example, that to decrease levels of resource inequality observed in a particular alternate world, their counter-intuitive prescription may be to increase the number of agents in that simulation. TA1 teams then will simulate these prescriptions as far as possible, and TA2 will have their prescriptions scored against the actual revealed simulated results.
If successful, the GT program will result in deliverables that will include artificial but socially plausible simulations with tunable complexity and causal ground truth against which to calibrate the accuracy and theoretical limitations of current and future social science modeling efforts.
Successful simulation teams will identify first principles that they then encode as different causal processes in simulations, which – when engineered and run in appropriate platforms – lead to emergent complex social behaviors.
These simulations will then provide capabilities to allow other researchers to test the accuracy of their , modeling methods for inferring causal processes in the simulations, using their models to predict future simulation states, and prescribing simulation parameters to guide future states. By having simulations that can serve as test-beds for social science modeling methods, GT will provide DARPA and the larger social science research communities with Quantitative understanding of why, when, and to what degree different causal modeling methods succeed or fail under conditions of varying social complexity.
If successful, GT will enable the Department of Defense (DoD) and the social sciences to better evaluate which causal modeling methods are most – and least – promising for a wide range of national security missions that involve complex social systems and behaviors.
HR001117S0031 GROUND TRUTH 8
D. Program Structure Ground Truth is a 30-month program comprising three phases with durations of 18 months, 6 months and 6 months, respectively. Phase I consists of an Initial Development period and the first of the three Challenges periods; if simulations and methods are sufficiently mature, Phases II and III will consist of additional Challenges. The Challenges are anticipated to increase in both complexity and difficulty as the program continues.
The GT program will be divided into two Technical Areas (TAs) with an independent Test and Evaluation (T&E) team providing oversight. The two TAs are:
• TA1: Simulations
• TA2: Methods
DARPA is soliciting proposals to TA1 or TA2, but is not soliciting proposals for participation on the T&E Team. Proposals to either TA must address the full program timeline. While TA1 simulators will know causal ground truth in the simulations during the Challenges, TA2 researchers will only be able to use their best modeling methods to infer causality.
To avoid a conflict of interest, no person or organization may be a performer for both TA1 and TA2, whether as a prime or as a sub-contractor.
E. Technical Area Descriptions
TA1: Simulations
The goal of TA1 is to leverage and advance complex social simulation capabilities to provide “minimally-viable” 3 test beds for a wide range of social science modeling methods. DARPA anticipates that TA1 simulations will have to build upon emerging capabilities that allow for the simulation of many heterogeneous agents with evolving objectives. The potential use of distributed and cloud computing and GPUs may be required for simulating agents and groups that can increasingly interact over structured but dynamic scales and networks and that may
3 “Minimally-viable” means that TA1 simulations will not necessarily require or include intricate or expensive graphics or interfaces to achieve Ground Truth goals. DARPA expects that successful TA1 proposals will focus primarily on making credible arguments for creating simulation capabilities discussed in this BAA. Accordingly, requests for resources for simulations that provide, e.g., graphical interfaces or avatars over and above these minimal capabilities, should be strongly justified.
HR001117S0031 GROUND TRUTH 9
exhibit purposive, adaptive, biased behaviors, and social learning – often leading to counter-intuitive behaviors at different levels of the simulation.4567
Additionally, DARPA anticipates that simulations may incorporate complex interactions among agents of systems that lead to learning, communication, group formation, and different responses to different perceived conditions such as resource scarcity or threats. Such interactivity might be instantiated via system dynamics simulations, agent-based models, or combinations thereof, but should presumably provide richness in time, space, and behavioral domains to allow for complex interactions among agents or subsystems.
Proposals: TA1 proposals should include clear, credible, and (where appropriate) quantitative descriptions that include (but need not be limited to) the following:
• Simulation type(s), platforms, requirements, and software proposed, including means of encoding and reporting first principles that form the causal rules and mechanisms that will lead to:
o Anticipated observable, socially plausible behavior;
o Types and volume of simulation output to be made available (e.g. visualizations, state descriptions, number of simulation runs, other datasets);
o Types and level of TA2-simulation interaction(s) to be instantiated (e.g., agent queries from menu, free-form, collective queries, limited experimentation).
• Anticipated/recommended data output formats;
• Principled approaches and mechanisms for increasing complexity within a simulation or across simulations, and whether increases are anticipated to be continuous or discrete;
• Nominated complexity measures for quantifying, comparing simulation complexity;
• Proposed mechanisms, level of control, and workflows for providing Predict and
Prescribe Tests for TA2 teams;
• Nominated simulation and scoring metrics for TA2 teams for Explain, Predict, Prescribe
Tests;
• Proposed methods for making some or all of the simulations and outputs available to a wider research community during open challenges;
• Identification of specific risks to the proposed simulation approaches and credible mitigation plans, including preventing and/or identifying efforts to “game” the simulations by TA2 or wider research community during open challenges;
• Additional information necessary to understand and evaluate the innovation of the approach(es) being proposed.
4 E.g., Lysenko, Mikola and D'Souza, Roshan M. (2008). 'A Framework for Megascale Agent Based Model Simulations on Graphics Processing Units'. Journal of Artificial Societies and Social Simulation 11(4)10.
GIScience. 2016.
5 Jin, Xiongbing, et al. "MIRACLE: A prototype cloud-based reproducible data analysis and visualization platform for outputs of agent-based models."
6 Ozik, Jonathan, et al. "From desktop to large-scale model exploration with Swift/T." Proceedings of the 2016 Winter Simulation Conference. IEEE Press, 2016.
7 Taylor, Simon JE, et al. "A tutorial on cloud computing for agent-based modeling & simulation with repast."
Proceedings of the 2014 Winter Simulation Conference. IEEE Press, 2014; https://repast.github.io/repast_hpc.html https://repast.github.io/repast_hpc.html
HR001117S0031 GROUND TRUTH 10
Performance Metrics: TA1 performers may elect to develop a single simulation approach that accommodates the varying complexity required across the three Challenges (see below) or develop multiple simulations that collectively provide the required range of complexity. TA1 performers will be assessed according to the following simulation capabilities, listed in order of importance:
• Verifiable ground truth – Ability to return known causal ground truth at multiple scales, over time, to include agent and system states, dynamics, and properties, in order to quantitatively test TA2 modeling methods’ accuracy and robustness for causal inference;
• Simulation Accessibility for researchers – Capabilities to provide simulation interfaces and simulated datasets that can accommodate and test a reasonably wide range of social science modeling methods including highly qualitative (interviews, surveys, etc.), highly quantitative (regression, time-series analyses, etc.), as well as combinations thereof (mixed methods) 8; The simulation should be accessible to at least 50% of TA2 methods at each Challenge;
• Simulation Flexibility – The degree to which the simulation complexity can be modified, and in what ways, e.g., by increasing dimensionality, dynamism, interactions, different sources of uncertainty; quantified by changes in e.g., entropy, agent graph connectivity, or parameter uncertainty and variability. The complexity of systems and agent/structure behaviors should be managed and manipulated in a principled manner. Simulations should also be able to accommodate perturbations for conducting the Predict/Prescribe tests described in Section I.D above. The simulation should allow for changes in >30% of rules, variables, agent interactions, sources of uncertainty;
• Social plausibility – Ability to grow simulations that are internally consistent, and do not require external interventions by TA1 teams to keep simulations running or to prevent simulation states from leading to total randomness or complete homogeneity.
Nominate Simulation Metrics: Proposers9 should nominate metrics for evaluating their simulations in terms of these requirements and include how their approaches will address each of the metrics within their proposal. In coordination with the T&E effort, TA1 teams should aim to quantify and compare simulation accessibility, flexibility, and plausibility using existing and novel metrics tailored to the space.
Anticipated/recommended data output formats: In order to facilitate efficient interaction among the GT performers, TA1 teams will be required to ensure that the output of their simulations conform to a common data model. Details of the data model will depend on the types of simulations proposed by selected TA1 teams and will be determined via collaboration between the TA1 and T&E teams. DARPA anticipates that TA1 teams will use a common format such as YAML or XML and require such information as simulation initial conditions, causal
8 For one example typology of different social science research methods, see http://eprints.ncrm.ac.uk/115/1/NCRMResearchMethodsTypology.pdf 9 As used throughout this BAA, “proposer” refers to the lead organization on a submission to this BAA. The proposer is responsible for ensuring that all information required by a BAA--from all team members--is submitted in accordance with the BAA. “Awardee” refers to anyone who might receive a prime award from the Government, including recipients of procurement contracts, grants, cooperative agreements, or Other Transactions.
“Subawardee” refers to anyone who might receive a subaward from a prime awardee (e.g., subawardee, consultant, etc.).
http://eprints.ncrm.ac.uk/115/1/NCRMResearchMethodsTypology.pdf
HR001117S0031 GROUND TRUTH 11
rules and processes, and complete state descriptions for each time step. TA1 proposals may nominate specific common data model formats.
Nominate Complexity Metrics: TA1 proposals should nominate candidate metrics for quantifying their simulation complexity, in part, to compare the complexity of their early and later Challenge simulations, as well as to assist with comparing simulation complexity across TA1 performers.
Tests: Teams should describe proposed mechanisms, level of control, and workflows for providing Explain, Predict and Prescribe Tests for TA2 teams. Each Challenge period will see three tests provided to TA2 teams.
• Test 1: “Explain” – TA1 teams should be able to provide TA2 teams with reasonable system observations, datasets, and abilities for limited interaction with various components of the simulation.
• Test 2: “Predict” – TA1 teams will define simulated “impactful” event relevant to the specific simulation platform and parameters. Events will be defined in coordination with DARPA and T&E and will be announced in advance to TA2 teams for their predictive modeling.
• Test 3: “Prescribe” – TA2 teams will work closely with T&E and TA1 teams to formalize their simulation prescriptions, and TA1 teams will then manipulate those parameters and report on the resulting simulation states.
Nominate TA2 Scoring Metrics: DARPA will assessTA2 teams’ modeling methods in terms of their accuracy, robustness, and efficiency across the various Challenge Tests (see section I.D for more information). TA1 proposals should nominate potential metrics that may be most appropriate given their specific simulation approach for evaluating TA2 teams.
Open Challenges: Teams should describe proposed methods for making some or all of the simulations and outputs available to a wider research community during open challenges.
Coinciding with the GT test and evaluation periods, DARPA currently expects to make the TA1 simulations open for a wider research community. These Open Challenges will give DARPA the opportunity to assess any solutions submitted from researchers outside of the GT Program.
Given the ambitious nature of the GT schedule and technical goals, DARPA does not expect TA1 teams to focus on non-GT solution providers. However, TA1 proposals should provide credible mechanisms and plans for easily making simulations open to non-GT solution providers.
Reasonable approaches that increase the likelihood of enabling DARPA, T&E, and TA1 teams to be responsive to non-GT solution providers - without increasing schedule or technical risk for GT goals and performers – will have a higher likelihood of receiving funding.
Human Subject Research Excluded: GT seeks to test social science modeling methods using TA1 simulations with known causal ground truth. Since including actual human participants as sources of data for the simulations would necessarily reduce known ground truth by introducing causal uncertainty that cannot be fully mitigated, DARPA anticipates that TA1 simulation approaches will not involve human subject research (HSR). Proposals that seek to include HSR
HR001117S0031 GROUND TRUTH 12
should clearly identify where and when such HSR might be necessary, why non-HSR alternatives would be insufficient, and strongly justify the proposed inclusion of HSR.
Publication of Research: Note that while DARPA anticipates that all research conducted for GT will be fundamental, unclassified research and therefore encourages performers to publish and/or distribute deliverables and results, given GT goals, TA1 teams may be expected to maintain some control during the program period over source code, platforms, generative principles, rules, etc.
Out of Scope: Simulations that are already widely available may be easily understood; therefore, while TA1 simulations should consider a balance of availability, speed, cost, and utility, easily acquired simulations may not be appropriate. Given GT goals, DARPA is also not looking to fund investments in standalone computational resources or sandboxes, so proposals seeking to develop large data storage facilities are also out of scope.
TA2: Methods
Successful TA2 teams will accomplish two major outcomes. First, they will provide solutions to the Explain/Predict/Prescribe Tests in TA1simulations by utilizing a wide range of social science modeling methods. In so doing, they will quantify the accuracy and robustness of those methods for causal inference and prediction under conditions of different kinds of social complexity.
TA2 teams will therefore provide DARPA with a sophisticated understanding of when, why, how, and to what degree different modeling methods succeed or fail under these conditions.
Second, successful TA2 teams will demonstrate innovative approaches to forming agile, multi-disciplinary, “solution-oriented” 10 modeling teams to address comprehensively and efficiently the various GT Tests they will face during the GT Challenge phases. Note that the GT schedule means TA2 teams may have only days or weeks to identify, create, and deploy modeling teams to develop solutions for specific Tests.
TA2 Proposals: TA2 proposals should include clear, credible, and (where appropriate) quantitative descriptions that include (but need not be limited to) the following:
• Approaches to testing modeling methods across the Explain, Predict, and Prescribe Tests;
• Proposed “solution-oriented” agile management and contracting plan for quickly forming teams and executing within the program timeline;
• Credible knowledge of a wide range of potentially GT-relevant modeling methods and expertise that teams could readily use;
• Hardware requirements for anticipated modeling methods;
• Data assumptions (e.g. scale, fidelity, frequency) required for anticipated methods with regard to TA1 simulation output;
• Examples of previous work on causal inference in social complexity;
• Nominated metrics and/or additional evaluation criteria for scoring TA2 methods’ accuracy, robustness, and efficiency;
10 E.g., Watts, D. (2017) “Should social science be more solution-oriented? Nature Human Behaviour 1 -doi:10.1038/s41562-016-0015
HR001117S0031 GROUND TRUTH 13
• Nominated measures for quantifying and comparing simulation complexity;
• Appropriate approaches and datasets for possible Initial Development modeling method demonstrations;
• Identification of specific risks to the proposed technical and management approaches and credible mitigation plans;
• Additional information necessary to understand and evaluate the innovation of the approach(es) being proposed.
Performance Metrics: TA2 will respond to the Challenge Tests described above in Section I.D, namely, providing their best solutions to Explain, Predict, and Prescribe various simulation states and behaviors. In collaboration with the T&E team, DARPA will assess TA2 performance according to the following criteria, listed below in order of importance:
• Accuracy: Teams recover at least 50% of causal processes in simulations for first explanatory Test; achieve statistical significance for predict/prescribe Tests
• Robustness: Teams ranked by average accuracy across simulations, capabilities to maintain accuracy under parameter variation and system stochasticity
• Efficiency: Teams ranked by computational efficiency, hardware requirements for methods
Nominate Potential Metrics: While the TA1 teams, T&E, and DARPA will know the principles and rules used to generate the behaviors seen in any given TA1 simulation, it is an open question whether recovering those known principles and rules can offer TA2 a complete description of complex interaction dynamics observed in a simulation. This may be particularly true for simulations that demonstrate dynamic and/or multilevel emergence of potentially new or unanticipated behaviors or properties. TA2 proposals therefore may propose additional metrics to help DARPA quantify the performance of TA2 methods. TA2 teams may also nominate potential measures for quantifying and comparing simulation complexity.
Solution-Oriented Agile Management Plan: Different TA1 simulations may require different modeling methods across the different Challenges and Tests, which could pose significant challenges to current approaches for supporting and conducting social science research. DARPA expects that successful TA2 teams will demonstrate “solution-oriented” approaches for quickly identifying, incorporating, and deploying a potentially wide range of social science modeling methods and expertise for any given Test. This focus on solutions may also mean that TA2 teams adopt and/or develop new modeling methods and approaches, targeted to specific complexity conditions and leveraging interactions across disciplines and TA2 teams. TA2 proposals should therefore identify how proposed approaches to GT agile management will provide new capabilities for solution-oriented social science modeling.
As TA1 simulations will not be known prior to Kickoff, TA2 proposals will have to estimate their approaches to providing teams and solutions for the Tests. For the purposes of their proposals, TA2 teams may assume at least 3 different TA1 simulations for each Challenge period, with datasets that may reach into gigabyte-scales. DARPA anticipates that these simulations and their data will be in formats that will be useful for TA2 teams, but TA2 teams
HR001117S0031 GROUND TRUTH 14
should not assume a specific Data Format. There may be additional data developed or collected by TA2 teams through their different research and modeling approaches, which DARPA cannot anticipate. To assist in evaluating proposals, therefore, DARPA anticipates that TA2 proposals will provide and justify Not To Exceed (NTE) costs that they believe to be realistic for each Phase, Challenge period, and Test. These costs should include potential sub-contracting costs, as well as anticipated incurred indirect costs.
DARPA anticipates that successful agile, solution-oriented TA2 teaming approaches will propose mechanisms for rapidly identifying and contracting with teams and expertise in response to TA1 simulation Tests. If selected for award, a TA2 team should expect to put in place a contracting mechanism for efficiently preparing a detailed cost estimate of the total amount required to develop and submit Test solutions, and obtaining the consent of the Contracting Officer for the placement of subcontracts. Successful TA2 teams will propose credible approaches to quickly providing detailed requests to the Contracting Officer, which DARPA anticipates will include:
(1) A description of the supplies or services to be subcontracted;
(2) Identification of the type of subcontract to be used;
(3) Identification of the proposed subcontractor;
(4) The subcontractor's current, complete, and accurate cost or pricing data;
(5) The Subcontractor's Disclosure Statement or Certificate relating to Cost Accounting
Standards for cost-reimbursement subcontracts;
(6) A negotiation memorandum (if available) reflecting:
(a) The total amount of the subcontract and the principal cost elements of the subcontract price negotiations;
(b) The most significant considerations controlling establishment of initial or revised prices;
(c) The reason certified cost or pricing data were or were not required;
(d) The reasons for any significant difference between the Contractor's price objective and the price negotiated.
While specific modeling methods for GT Tests will be unknown during the BAA open period, TA2 proposals should nonetheless demonstrate credible knowledge of potential solution spaces by identifying candidate teammates, their relevant expertise, and a wide range of potentially relevant modeling methods. Depending on the specific simulation, DARPA anticipates that successful TA2 teams may provide solutions that leverage combinations of a wide range of quantitative and qualitative modeling methods, including potentially novel combinations thereof.
Accordingly, TA2 proposals may also wish to recommend (or advise against) possible TA1 simulation output and data formats (see TA1 “Proposals” bullets, above). In this regard, TA2 proposers may wish to provide examples of previous modeling method work involving causal inference in complex social systems and behaviors. Further, to assist TA1 and T&E teams in designing simulation capabilities and scoring methods, TA2 performers selected for award may wish to propose limited early demonstrations or testing of modeling methods following GT kickoff. If so, TA2 proposals should identify the GT-relevant datasets and/or simulations they may use and provide credible evidence that they are able to use them for these purposes. Note that given the ambitious GT timelines, TA2 proposals should seriously consider the potential
HR001117S0031 GROUND TRUTH 15
schedule impacts and credibility of any work that may require Human Subject Research approval.
Challenges:
TA1 teams will develop principled approaches and mechanisms for increasing complexity within a simulation or across simulations, whether increases are anticipated to be continuous or discrete.
Teams will develop and deploy their simulations with sufficiently interactive agents, behaviors, and systems to support the evaluation of diverse research methods, models, and tools across all phases of the program. Simulations will target socially plausible systems of varying complexity, corresponding to Challenges 1, 2, and 3 of the program:
• Challenge 1: Simple systems. These may be simulations with relatively few variables, little dynamism, low uncertainty, and are potentially amenable to near-complete mathematical description.
• Challenge 2: Disorganized complex systems. These are simulations with many variables, increasing dynamism, increasing uncertainty, and are potentially amenable to probabilistic and statistical methods.
• Challenge 3: Organized complex systems. These are simulations that may include many interacting variables, high uncertainty, reflexivity, nonlinearity, multi-scale interactions, bifurcations and phase changes, adaptive behavior, goal-driven and/or gaming and deceptive behavior, heterogeneity of subcomponents, and emergent properties.
Tests:
In each Challenge, TA2 teams will address three Tests (note that TA2 teams may use different modeling methods for each Test):
• Test 1- Explain: What is causing the observed behaviors? TA2 teams will use modeling methods to determine causal processes generating the observed states and behaviors in TA1 simulations.
• Test 2 - Predict: What behaviors will come next? Based on a pre-announced “impactful event” that TA1 teams will instantiate in their simulation(s), TA2 teams will use their modeling methods to predict future states of the simulation at multiple time points and scales.
• Test 3 – Prescribe: Which parameters lead to different system states? TA2 teams will use modeling methods to recommend specific ways to influence a given simulation towards a desired state (universal cooperation, reduction of resource hoarding, etc.), and TA1 will instantiate those recommendations in their simulations to evaluate TA2 teams’ prescriptive accuracy.
Following the establishment of TA2 teams, there will be some time where the simulations and simulated datasets are only available to those teams while they develop and test their solutions for evaluation against simulation ground truth. However, after this period, DARPA intends to make the simulations and simulation data publicly available to allow for open responses to each of the Tests – inviting a wider community to participate and explore their modeling capabilities, as well as to serve as a further baseline against which to compare TA2 results.
HR001117S0031 GROUND TRUTH 16
Throughout the program, TA1 and TA2 performers will interact with the T&E team, and TA1 and TA2 teams should anticipate these interactions in their proposed costs, schedules, and deliverables. The T&E team will comprise subject matter experts from Government, Federally Funded Research and Development Centers (FFRDCs), and/or academia and domain experts.
T&E will score the TA2 performance in inferring causality and predicting/prescribing system states in the increasingly complex Challenges. The T&E team will also assess and test the TA1 teams’ simulations for suitability to the Challenge according to negotiated metrics. T&E will work with TA1 and TA2 teams to develop a Common Data Format to ensure interoperability of simulations and modeling methods; will lead development of simulation complexity metrics;
help shape requirements for minimal size, span, or scale of TA1 simulations; verify TA2 accessibility requirements for simulations; assist with Open Challenge development and deployment; and will work closely with the TA1 and TA2 teams to develop preliminary candidate measures of simulation and real-world computational equivalence.
F. Schedule/Milestones Proposals to either TA must address the full program timeline. Proposers should provide a technical and programmatic strategy that conforms to the entire program schedule and presents an aggressive plan to fully address all program goals, metrics, milestones and deliverables. A target start date of December 2017 may be assumed for planning purposes.
All GT performers should expect to attend a kickoff meeting in the Washington, D.C. area.
DARPA expects all performers to attend Principal Investigator (PI) Meetings every 6 months, as shown in Figure 2. The purposes of the PI Meetings are to (i) provide the Program Manager and other GT performers with updates on progress towards milestones and goals; (ii) summarize outstanding technical challenges; (iii) support test and evaluation; and (iv) provide Government and potential transition partners with opportunities to provide input, comments, and suggestions for the GT program and its performers. For budgeting purposes, proposers should assume a two-day kickoff meeting, while PI meetings will require three days and will alternate between Washington, D.C., and a west coast location. Regular teleconference meetings will be scheduled with the Government team for progress reporting as well as problems identification and mitigation. Proposers should also anticipate at least one site visit every 6 months by the DARPA Program Manager during which they will have the opportunity to demonstrate progress towards agreed-upon milestones. Additional anticipated programmatic events are included in Tables 1 and 2, below.
HR001117S0031 GROUND TRUTH 17
Figure 2: Ground Truth Schedule and Milestones
Table 1: Technical Goals by Phase
Phase I Phase II Phase III
Technical Area
Initial Development
(month 9)
Challenge 1 (month 18)
Challenge 2 (month 24)
Challenge 3 (month 30)
TA1
Simulations
• Develop simulations of simple systems
• Demonstrate simulation accessibility to >50% of TA2 methods
• Conform to common data model
• Demonstrate simulations of simple systems
• Run Explain, Predict, Prescribe Tests
• Develop simulations of disorganized complex systems
• Demonstrate simulations of disorganized complex systems
• Run Explain, Predict, Prescribe Tests
• Develop simulations of organized
• Demonstrate simulations of organized complex systems
• Run Explain, Predict, Prescribe Tests
TA2
Methods
• Conform to common data model
• Demonstrate management approach for
• Form and bid teams, methods
• Determine solutions
• Establish accuracy, robustness, and
• Form and bid teams, methods
• Determine solutions
• Establish accuracy, robustness, and
• Form and bid teams, methods
• Determine solutions
• Establish accuracy, robustness, and efficiency of each method for each Test
HR001117S0031 GROUND TRUTH 18
rapid team formation
• Early demonstrations of methods as justified efficiency of each method for each Test on TA1 simple systems simulations efficiency of each method for each Test on TA1 disorganized complex systems simulations on TA1 organized simulations
T&E
• Evaluate accessibility of TA1 simulations to TA2 methods
• Develop common data model with TA1 and TA2 collaboration
• Establish metrics for evaluating TA1 and TA2 performers
• Assess complexity, flexibility, plausibility of TA1 simulation
• Evaluate performance of TA2 methods on TA1 simulations
• Recommend appropriate increase in simulation complexity
• Assess complexity, flexibility, plausibility of TA1 simulation
• Evaluate performance of TA2 methods on TA1 simulations
• Recommend appropriate increase in simulation complexity
• Assess complexity, flexibility, plausibility of TA1 simulation
• Evaluate performance of TA2 methods on TA1 simulations
• Establish metrics for computational equivalence of simulations and real-world behavior
Table 2: Program events by month
Months after
Award Event Description
Initial Development
1 Program Kickoff
• TA1 and TA2 teams present technical approach and work plan
• T&E team provides test and evaluation plan, candidate metrics 6 PI meetings • All teams: review technical progress
Initial Common
Data Model Complete
• T&E presents initial common data model
• TA1 and TA2 begin integration with data model
6 Establish metrics • T&E establishes metrics for TA1 and TA2 performers
8 Simulation Assessment
• T&E evaluates TA1 simulations for accessibility to TA2 methods, appropriate complexity
Data Model Conformity Assessment
• T&E evaluates TA1 and TA2 conformity to data model
Challenge 1
9 Explain Test
• TA1 makes simulation and data available to TA2
• TA2 determines appropriate methods, gathers appropriate team members; begins analysis
HR001117S0031 GROUND TRUTH 19
after
Award Event Description
11 Open Simulations • TA1 teams make simulations open for wider research community to propose Explain solutions
12 Explain Evaluation • T&E evaluates TA2 results for accuracy, efficiency 12 PI meeting • All teams: review technical progress
12 Predict Test
• T&E and TA1 determine simulation parameter changes, provide to TA2
• TA2 determines appropriate methods, gathers appropriate team members; begins analysis
14 Open Simulations • TA1 teams make simulations open for wider research community to propose Predict solutions
15 Predict Evaluation
• T&E evaluates TA2 results for accuracy, robustness, computational performance; evaluates TA1 simulation flexibility for Challenge 2
15 Prescribe test
• T&E and TA1 determine desired simulation behavior/state;
provides to TA2
• TA2 determines appropriate methods, gathers appropriate team members; begins analysis
17 Open Simulations • TA1 teams make simulations open for wider research community to propose Prescribe solutions
18 Prescribe Evaluation
• T&E evaluates TA2 results for accuracy, robustness, computational performance; evaluates TA1 simulation accessibility for Challenge 2
18 PI meeting • All teams: review technical progress Challenge 2
18 Explain Test
• TA1 provides simulation data to TA2
• TA2 determines appropriate methods, gathers appropriate team members; begins analysis
19 Open Simulations • TA1 teams make simulations open for wider research community to propose Explain solutions
20 Evaluation • T&E evaluates TA2 results for accuracy, computational performance
20 Predict Test
• T&E and TA1 determine simulation parameter changes, provide to TA2
• TA2 determines appropriate methods, gathers appropriate team members; begins analysis
21 Open Simulations • TA1 teams make simulations open for wider research community to propose Predict solutions
22 Evaluation
• T&E evaluates TA2 results for accuracy, robustness, computational performance; evaluates TA1 simulation flexibility for Challenge 3
HR001117S0031 GROUND TRUTH 20
after
Award Event Description
22 Prescribe test
• T&E and TA1 determine desired simulation behavior/state;
provides to TA2
• TA2 determines appropriate methods, gathers appropriate team members; begins analysis
23 Open Simulations • TA1 teams make simulations open for wider research community to propose Predict solutions
24 Evaluation
• T&E evaluates TA2 results for accuracy, robustness, computational performance; evaluates TA1 simulation accessibility for Challenge 3, social plausibility, real-world computational equivalence
24 PI meetings • All teams: review technical progress Challenge 3
24 Explain Test
• TA1 provides simulation data to TA2
• TA2 determines appropriate methods, gathers appropriate team members; begins analysis
25 Open Simulations • TA1 teams make simulations open for wider research community to propose Explain solutions
26 Evaluation • T&E evaluates TA2 results for accuracy, computational performance
26 Predict Test
• T&E and TA1 determine simulation parameter changes, provide to TA2
• TA2 determines appropriate methods, gathers appropriate team members; begins analysis
27 Open Simulations • TA1 teams make simulations open for wider research community to propose Predict solutions
28 Evaluation • T&E evaluates TA2 results for accuracy, robustness, computational performance
28 Prescribe test
• T&E and TA1 determine desired simulation behavior/state;
provides to TA2
• TA2 determines appropriate methods, gathers appropriate team members; begins analysis
29 Open Simulations • TA1 teams make simulations open for wider research community to propose Prescribe solutions
30 Evaluation
• T&E evaluates TA2 results for accuracy, robustness, computational performance; evaluates TA1 simulation social plausibility, real-world computational equivalence
PI Meeting: Review and Program Closeout
• Performers (including T&E) present review of experimental results and process capture to DARPA and transition partners
HR001117S0031 GROUND TRUTH 21
G. Deliverables DARPA expects performers to provide at a minimum the following deliverables:
• Comprehensive quarterly…
This is the start of the file's text. The full file is on GovTribe.
File details come from the government source that posted it. Updated .