AMENDMENT 4 FA8750-19-S-7014 second repub.docx
DOCX document 65 KB Posted
- Attached to
- Operationalizing Machine Learning for Command and Control (OpML C2) Federal contract opportunity
- Solicitation number
- FA8750-19-S-7014
About this file
This Broad Agency Announcement (BAA) from the Department of the Air Force Materiel Command Research Laboratory solicits white papers and proposals to support two tasks related to operationalizing machine learning for command and control (OpML C2).
Task 1 seeks prototype applications of AI/ML to support operational C2 processes, focusing on planning, decision making, and execution management. Offerors select an operational thread and develop a prototype application. Multiple awards are anticipated. Task 2 focuses on applying reinforcement learning to operational planning, requiring the design of an evaluation environment to train planning algorithms with simulated feedback. Only one award is anticipated for developing both a planning prototype and associated evaluation environment.
White papers are due by specified dates in Fiscal Years 2020 through 2022. Proposals will be invited selectively based on white paper evaluations. The total estimated funding is $24.9 million over multiple awards from 1 to 24 months in duration ranging from $300,000 to $1 million. Procurement contracts, grants, cooperative agreements or other transactions are possible award types.
View the file
Other files for this federal contract opportunity
| File | Type | Posted |
|---|---|---|
| AMENDMENT 3 FA8750-19-S-7014 first repub.docx | DOCX document | |
| 19-14_Full_Text_Announcement.docx | DOCX document | |
| Operational_Threads_Attachment.doc | DOC document |
On GovTribe
Work with this file on GovTribe
- Download the original file
- Contacts named in this file
- Similar government files
- Ask GovTribe AI about this file
Text version
AMENDMENT 4 to BAA FA8750-19-S-7014
The purpose of this modification is to republish the original announcement, incorporating any previous amendments, pursuant to FAR 35.016(c).
This republishing also includes the following changes:
1. Part I, Overview Information:
0. No changes made;
1. Part II, Full Text Announcement:
1. Section IV.3.a, updates the NISOPM reference to reflect it being codified in the CFR;
1. Section IV.4, adds paragraph g;
1. Section V., adds paragraph 4 regarding adequate price competition;
1. Section VI.1, updates the Proposal Formatting language;
1. Section VI.7, updates the applicable provisions;
1. Section VII: updated the OMBUDSMAN and clause date.
No other changes have been made.
NAICS CODE: 541715
FEDERAL AGENCY NAME: Department of the Air Force, Air Force Materiel Command, AFRL - Rome Research Site, AFRL/Information Directorate, 26 Electronic Parkway, Rome, NY, 13441-4514
BAA ANNOUNCEMENT TYPE: Modification
BROAD AGENCY ANNOUNCEMENT (BAA) TITLE: Operationalizing Machine Learning for Command and Control (OpML C2)
BAA NUMBER: FA8750-19-S-7014
PART I – OVERVIEW INFORMATION
This announcement is for an Open, 2-Step BAA which is open and effective until 30 Sep 2022. Only white papers will be accepted as initial submissions; formal proposals will be accepted by invitation only. While white papers will be considered if received prior to 4 PM Eastern Standard Time (EST) on 30 Sep 2022, the following submission dates are suggested to best align with projected funding:
For Task 1 ONLY:
FY20 by 22 Aug 2019 FY21 by 31 Mar 2020 FY22 by 31 Mar 2021
For Task 2 ONLY:
The reliance on reinforcement learning (RL) algorithms for Task 2 introduces additional challenges covered in detail under the respective task description and it is anticipated that only one award will be made under this Task. Therefore, white papers for Task 2 ONLY are due on 22 AUG 2019 and if requested, proposals are due on 25 OCT 2019.
PLEASE NOTE: Task 2 was closed on 25 OCT 2019. No further white papers are being accepted for Task 2.
Offerors should monitor the Contract Opportunities on the Beta SAM website at https://beta.SAM.gov in the event this announcement is amended.
The OpML Program Team anticipates hosting an Industry Day to provide interested offerors an opportunity to learn more about our program and AFRL/RI’s multi-domain C2 (MDC2) activities on 2 AUG 2019. Details regarding time, place, and visitor access will be published in FedBizOps shortly following the announcement publication proper.
CONCISE SUMMARY OF TECHNOLOGY REQUIREMENT: AFRL/RI is seeking to identify, develop, and evaluate novel applications of Artificial Intelligence (AI) and Machine Learning (ML) to support operational aspects of Command and Control (C2). The OpML Program Team’s initial analyses considered wide-ranging applications of ML in the air combat and air mobility operations centers, as well as interesting applications for air battle management C2. The analyses have narrowed down the space of paired operational applications and their potential ML solutions to a manageable set of specifications without precluding the proposal of other unique and novel approaches. The sought-out solutions primarily focus on the problems of planning, operational and tactical level decision making, and operational execution management. The development of prototype applications, respective use-cases, workflows, and data requirements are critical to establish the viability and usefulness of the assessed candidate applications for assessment in an operational setting.
BAA ESTIMATED FUNDING: Total funding for this BAA is approximately $24.9M. Individual awards will not normally exceed 24 months with dollar amounts normally ranging from $300,000 to $1,000,000. There is also the potential to make awards up to any dollar value as long as the value does not exceed the available BAA ceiling amount.
ANTICIPATED INDIVIDUAL AWARDS: Multiple Awards are anticipated.
TYPE OF INSTRUMENTS THAT MAY BE AWARDED: Procurement contracts, grants, cooperative agreements or other transactions (OT) depending upon the nature of the work proposed. In the event that an Other Transaction for Prototype agreement is awarded as a result of this competitive BAA, and the prototype project is successfully completed, there is the potential for a prototype project to transition to award of a follow-on production contract or transaction. The Other Transaction for Prototype agreement itself will also contain a similar notice of a potential follow-on production contract or agreement.
AGENCY CONTACT INFORMATION: All white paper submissions and any questions of a technical nature shall be directed to the cognizant Technical Point of Contact (TPOC) as specified below (unless otherwise specified in the technical area) (email requests are preferred):
BAA MANAGER:
Mr. Gennady Staskevich
AFRL/RISC
525 Brooks Rd Rome, NY 13441-4505 Telephone: (315)330-4889 Email: Gennady.Staskevich@us.af.mil
Questions of a contractual/business nature shall be directed to the cognizant contracting officer, as specified below (email requests are preferred):
Ms. Amber Buckley Telephone (315) 330-3605 Email: amber.buckley@us.af.mil
Emails must reference the solicitation (BAA) number and title of the acquisition.
Pre-Proposal Communication between Prospective Offerors and Government Representatives: Dialogue between prospective offerors and Government representatives is encouraged. Technical and contracting questions can be resolved in writing or through open discussions. Discussions with any of the points of contact shall not constitute a commitment by the Government to subsequently fund or award any proposed effort. Only Contracting Officers are legally authorized to commit the Government.
Offerors are cautioned that evaluation ratings may be lowered and/or proposal rejected if proposal preparation (Proposal format, content, etc.) and/or submittal instructions are not followed.
PART II – FULL TEXT ANNOUNCEMENT
BROAD AGENCY ANNOUNCEMENT (BAA) TITLE: Operationalizing Machine Learning for Command and Control (OpML C2)
BAA NUMBER: BAA FA8750-19-S-7014
CATALOG OF FEDERAL DOMESTIC ASSISTANCE (CFDA) Number: 12.800
I. TECHNOLOGY REQUIREMENTS:
PROGRAM GOAL:
For the better part of a decade, there has been a massive resurgence of artificial intelligence (AI) and machine learning (ML) research and applications that presents both a tremendous opportunity for defense operations as well as a critical risk to the same. The application of such techniques at the operational command and control level is crucial for updating the usual modus operandi of mostly hand-crafted plans and effects (labor intensive and difficult to scale). The design of autonomous intelligent decision-making agents is particularly important in light of increasingly intelligent, autonomous, and capable adversaries. To overlook this modernization of Air Force operations is to introduce a critical vulnerability to the future warfighter in the form of increased uncertainty and predictability as well as diminished tempo and scale of the data-to-decision process.
While fundamental limitations of the human operator must be complemented via AI/ML, the human element will nonetheless remain involved. Indeed, future joint force operations will rely on the interaction between human subject matter experts (SMEs) and artificially intelligent agents as they cooperate to achieve mission-focused objectives in the presence of adversaries and the absence of data. In the meantime, heavy reliance on that exact SME input will help guide the development of an initial collection of intelligent agents that focuses human-machine interaction on several well-known operational challenges. The goal of this program then, is to identify, apply, and evaluate AI/ML techniques as they relate to well-defined operational C2 problems, such as, but not limited to planning, execution monitoring and battle management command and control (BMC2). While private industry and academia provide us with an abundance of useful tools and techniques in a broader context, these operational problems and the challenges therein are unique to the Air Force. As such, Offerors are expected to apply and adapt existing state-of-the-art AI/ML technologies in order to modernize operations while mitigating defense-specific challenges.
PROGRAM OVERVIEW:
The OpML program will achieve its goal of re-envisioning Air Force C2 operations through the use of AI/ML methods. Modern ML architectures have been revolutionized by the development of deep neural networks and the availability of massive amounts of data. However, while deep learning has yielded tremendous improvements in classification and decision-making tasks, its application is often limited to very restricted environments with convenient characteristics, such as Atari games and toy problems. A realistic operational environment will not likely share such convenient characteristics that come in the form of frequent reward signals and ease of data generation at speed and scale. Furthermore, applications in the defense sector are often precluded by a scarcity of labeled data due to operational restrictions as well as the manpower required to label existing data. This makes the problem of data and sample efficiency more pressing than in traditional industrial and academic settings. As such, Offerors are expected to address this issue of handling data efficiently. In particular, one challenge that arises in data-efficient learning is that of how to exploit the structure of data for a particular problem. In the case of unstructured data, such as image sequences, there is often a small set of salient features which largely define a given data point and which lend themselves well to be compared to other data points. This is known as a low-dimensional smooth manifold representation of the data and it is what allows deep learning techniques such as convolutional neural networks to drastically reduce the dimensionality of the input space, thereby leading to impressive results in tasks such as object recognition in videos and control in video games. Unfortunately, the assumption of such a low-dimensional representation seems to break down when dealing with structured data, such as an air tasking order (ATO) or course of action (COA), and methods for deriving efficient representations of such data are domain-dependent and not well-understood. Offerors should clearly identify C2 data requirements and sources or how they propose to overcome the C2 data challenges.
A critical component of this program is the application of RL techniques to operational-level planning, where the foregoing challenges are further exacerbated by the need to simulate decision-making policies of agents via an environment which captures the dynamics of agent interactions. For example, in the operational C2 problem space, a simulation environment which generates, processes, simulates, and evaluates ATOs and COAs with execution time in the order of seconds would be highly desirable. Ideally, such a simulation environment would also facilitate the integration of other AI/ML efforts into a cohesive whole. If the development of a high-fidelity operational environment that simulates and generates data at speed and scale is prohibitive, then the Offerors may propose a simpler low-fidelity environment to train on. However, the RL methods proposed for said environment should have an expected valuation of performance on high-fidelity environments based on evaluation metrics defined by the Offerors. One promising approach to accomplishing this is to use methods from transfer learning and domain adaptation to bridge the gap between fast low-fidelity environments and slow high-fidelity ones.
TECHNOLOGY REQUIREMENTS:
This program will consist of two primary tasks. Both tasks will concentrate on prototype development and evaluation for a single domain; however, there is a distinction between the two. The objective of Task 1 is focused on rapid prototype development for AI/ML applications supporting the operational C2 tools and processes. Due to the criticality of this task, several efforts will be spun up as quickly as possible to address as many operational challenges as resources permit, with each Offeror responsible for the evaluation of their solutions. Task 1 will also contain an optional task to extend the original workflow to multi-domain use-case(s) with a scaled-down number of solutions. Task 2 is specifically focused on applying RL to air combat and mobility operational-level planning processes, such as the development of ATOs, COAs, Airlift Schedules, Master Air Attack Plans (MAAP), and complex mission packages. In addition to algorithmic prototype development, the objective of Task 2 also includes the design of a necessary supporting simulation and evaluation environment for training and execution of AI/ML algorithms for C2 planning. The intent of this supporting development and evaluation environment is to emulate the environment and to provide feedback to the ML algorithm being trained at speed (time to evaluate plan), scale (size of the campaign), and fidelity (level of detail). Both tasks are subject to the challenges of data representation mentioned in the previous section. However, the reliance on RL algorithms for Task 2 introduces additional challenges alluded to earlier and covered in detail under the respective task description and it is anticipated that only one award will be made under this Task.
To guide the formulation of these tasks and provide scope for proposed approaches, AFRL has conducted an analysis of operational challenges that could benefit most from wide-ranging ML applications. This analysis ultimately focused in on the activities occurring in the air combat and mobility operations centers at the operational level, as well as air BMC2 at a lower level of command echelon to provide a natural progression to the higher echelons. The potential high-impact areas for AI/ML enhancements are listed below in Table 1 as eight operational challenges that are denoted in the shaded boxes. Each of the eight operational threads are further defined in accompanying Appendix document. Performers may also propose their own operational thread and candidate AI/ML solutions that are outside the shaded areas, but must also provide an accompanying description equivalent to those eight defined in the accompanying Appendix document.
Focus Areas
Operational Areas
Planning (Longer time horizon) Execution (Shorter time horizon)
| Air Mobility Operations (Air C2) |
| Ground MOG Planning/Scheduling |
| Flight Monitor Load Balancing |
| Air Battle Management C2 (BMC2) |
| Operations Execution Management |
| ML Applications for Tactical Chat |
| Air Combat Operations (Air C2) |
| MAAP-ATO-Package Planning |
| Dynamic Targeting/Retasking |
| AOC Productivity (Air C2) |
| Human-Machine Logistics Teaming |
| Operator Workload Balancing |
Table 1 – Operational Threads of Interest for AI/ML Application
Note that, for the purposes of this BAA, there is no one operational area or function that is more important than another. Novel approaches for any one of the eight operational threads and strong technical content are encouraged. AFRL chose to focus on a few mission threads in the Air domain among the vast spectrum of complex operational C2 processes to provide scope of potential areas and their solutions.
Task 1: AI/ML Prototype Development and Experimentation Task 1 performers must select an operational thread and address the technical problem(s) therein via the use of AI/ML technology. It is worth noting the difference between AI and ML for this effort. AFRL does not rule out potential AI solutions under this effort, but there is an emphasis on ML because it is a data-driven subset of AI with a narrower focus. This effort will focus on the predictive nature of ML and not necessarily on the prescriptive nature typically found in AI-based decision-making. The intent is to use Task 1 to develop prototypes that would demonstrate a specific ML-augmented capability supporting operational C2 processes.
The use of ML must leverage a repository of training data in order to improve the predictive assessment of future data points. However, such data must first be labeled or evaluated in order to be useful for training. One obvious way to resolve this is to simply acquire more labels by utilizing SME hours to assess an existing repository of data points. While this may be necessary, it is by no means sufficient due to the massive number of labels required by state-of-the-art data-hungry ML algorithms. This may be mitigated via the use of active learning methods that intelligently query the SME for labels in order to minimize the amount of labeled data required. Another promising approach is the use of semi-supervised learning to exploit the remaining unlabeled data for the creation of pseudo-labels. In operational domains where the availability of labeled data is restricted, performers will be expected to address this challenge of data-efficient learning or identify sufficient data sources to be leveraged to support the proposed approach. Both approaches are of equal interest.
Task 1 is also focused on rapid prototype development. The proposed work and scope should target roughly a 9-month prototype development (Phase 1a), followed by a 6-month period for algorithm evaluation and concept refinement (Phase 1b). The total time for Phases 1a and 1b should not exceed 15 months. Task 1 Offerors will also need to define the evaluation criteria and the necessary environment to evaluate the developed capabilities. The technology readiness level[footnoteRef:1] of the software is estimated to be about 4 at the end of Task 1. Multiple performers are anticipated for Phases 1a and 1b. [1: https://www.army.mil/e2/c/downloads/404585.pdf]
Task 1 prototype development will also contain an executable option (phase 1c) to build upon the prototype developed in Phases 1a and 1b and extend it from a single domain to multiple domains. During Phase 1b, performers are expected to identify other domain data-insertion points and submit an updated plan for the multi-domain use-case (Step 6 of the Task 1 technical tasks). The updated multi-domain plan should be submitted no later than the end of the 5th month of Phase 1b. The period of performance, scope and effort for Phase 1c should not exceed 12 months. The decision of which subset of solutions will proceed with an executable option will be made based on the use-case quality, cost, and anticipated level-of-effort proposed. The total period of performance for all phases of Task 1 will be limited to 27 months.
Task 1 Technical Tasks Task 1 performers shall:
1. (Phase 1a) Select an operational thread to apply ML to and identify specific functional objectives for the ML-augmented application; propose evaluation criteria for the candidate application.
1. (Phase 1a) Identify subject matter experts (SMEs) with relevant experience and demonstrate how their interaction will guide the evaluation process; develop a use-case that showcases the specific capability. Given the operational focus of the program, SME participation is required.
1. (Phase 1a) With SME input, identify and develop input training data sets in order to develop an initial model to initiate the evaluation process.
1. (Phase 1a) With SME input, identify and develop evaluation criteria reflective of the program schedule. Describe how SME interaction will guide the evaluation process.
1. (Phase 1a) Develop a functional prototype based on the selected application, data requirements and guidance from the SME; present the algorithm with new datasets and make predictive assessments; refine the models with SME guidance to assist in further training of the algorithms. During Phases 1a and 1b, Task 1 performers are required to evaluate the quality of their outputs and predictions in their own stand-alone environments. Document this environment and ensure it has the sufficient fidelity to capture meaningful feedback for model training.
1. (Phase 1a) Leverage the ecosystem and capabilities provided by AFRL’s Streamlined ML (SML) framework if appropriate. SML is an open, government-owned platform that enables rapid, low-cost deployment of state-of-the-art ML capabilities to AF and DOD data and applications via the Model Integration Software Toolkit (MISTK) Software Development Kit (SDK). This open platform enables new algorithms to be developed, integrated by industry and government, and made available to the larger DOD community. The Streamlined ML/MISTK SDK ecosystem is a flexible architecture focused on fostering the sharing, rapid development, and evaluation of ML data, algorithms, models, and workflows across the Air Force and DOD community. It is available (https://mistkml.github.io/index.html) for use by DOD government and industry partners. The beta version of the framework has been released and supports deployment at both local and cloud environments at various classification levels.
1. (Phase 1b) Identify potential multi-domain use-cases/workflows for your approach. Halfway through the effort, review existing workflows and use-cases, and identify potential data (or process) insertion points to or from another domain(s). This step does not involve any additional development work or any changes to be made to the developed prototype application and is due at the end of the 5th month of Phase 1b.
1. (Phase 1c) Transition the developed prototype to the multi-domain setting via the Flyleaf platform; develop data insertion points and micro-services as identified in the multi-domain workflow (Phase 1b); build upon and/or extend the prototype application developed under Phases 1a and 1b to now include the multi-domain data and services. The developed prototype at the end of the effort will be run and evaluated using the Flyleaf environment at AFRL and using AFRL hardware. The Flyleaf program provides a multi-domain experimentation environment to enable the development of operations at speed and scale. Key to the Flyleaf program is the focus on capturing business processes of operational threads, which enables the dynamic composition of MDC2 workflows via aggregation of micro-services. Flyleaf allows for modeling and rapid development of multi-domain capabilities within the operational and battle management echelons and provides a transition path for candidate multi-domain C2 (MDC2) technologies to the operational DevOps like ShadowOC. Insertion into the Flyleaf environment will require wrapping developed applications using appropriate web-based endpoints, and developing appropriate messaging and support for Universal C2 Interface (UCI) to interact with Flyleaf scenarios and modeling and simulation. As Flyleaf functionality will evolve at the pace of this effort, generic “wrapping” strategies should be described in Task 1 proposals, with more detail required at the end of the 5th month of Phase 1b in conjunction with the use-cases.
PLEASE NOTE: Task 2 is now closed and no further white papers will be accepted. The technical description below is included for reference purposes only.
Task 2: RL Prototype and Evaluation Environment Development The objective of Task 2 is to: 1) develop a prototype application supporting operational-level planning (Phase 2a); 2) design a prototype development and evaluation environment for training and execution of RL algorithm(s) (Phase 2b) and 3) connect the two components that take the actions generated from the planning algorithms and subject them to simulated events from the evaluation environment to drive algorithm refinement (Phase 2c).
Operational-level planning is a major component of operational C2. It is a continuous, iterative, and highly structured process for force management to meet commander’s requirements across a range of military options. Existing tools and processes supporting operational planning are labor intensive and difficult to scale. As an example, existing automated operational planning tools produce poor-quality ATOs, which, if executed, would result in a poor allocation of resources, significant loss of assets, and unnecessary exposure to the enemy. Furthermore, they may not even address the mission objectives. The OpML Program Team believes that operational planning can benefit from recent advancements in ML, more specifically from the area of RL, as both share a cyclic process that incorporates the necessary feedback from the environment and optimization across multiple (and often conflicting) objectives to derive near-optimal plan solutions.
The Task 2 Offeror can select from one of the air combat or mobility planning operational threads shown in Table 1 and identify the infrastructure environment required to train and evaluate the model to address the selected operational problem. Alternatively, performers may propose their own operational thread and provide a justification for their choice. This task is focused on developing a scalable environment for the development and evaluation of RL-based planning capabilities. Proposed RL algorithms supporting operational-level planning processes will need to incorporate feedback from the environment at speed, scale, and fidelity. Designing a realistic environmental feedback is necessary for Task 2 and should be accomplished via simulations. The OpML Program Team anticipates that this will result in a significant level of effort. A conceptual approach will likely require: 1) selecting a specific planning function, 2) identifying its evaluation criteria and data requirements, 3) identifying the underlying architecture and approach, 4) identifying existing simulation capabilities and bottlenecks supporting the existing process, 5) leverage, extend, and modify existing simulation capabilities, or develop new ones, to support the RL-augmented planning prototype capability. The technology readiness level of the software is estimated to be about 4 at the end of Task 2.
Operational planning is a complex process, and it has the potential to benefit from, and build upon the ideas from the Reinforcement Learning (RL) domain. Recent developments in RL, have been extended to complex planning applications. This often involves an agent using a rapid, and a computationally efficient iterative process (a deficiency of current C2 processes) to develop a policy or plan that is executed in a simulated environment to determine best outcomes. Feedback from the environment results in a change of state and an associated reward or penalty for decisions made in the original plan. The goal is to have both realistic policies and environmental feedback at the right level of fidelity to automatically drive training at the speed and scale that matches the complexity of the environment.
The effective use of RL requires an efficient data representation, a simulator which models the agent-environment dynamics, and the capacity to generate training data at speed and scale. Each of these major components faces many challenges and opportunities, including the following:
· Data Representation: Deep RL solutions are predicated on an efficient abstract representation of the underlying data. In the case of image sequences, there is a low-dimensional smooth manifold representation of the data which is implicitly used by convolutional neural networks to drastically reduce the dimensionality of the input space, thereby leading to greater efficiency. Such representations are not as well understood in other problem domains, such as those that arise in Air Force operations that deal with highly structured data. In particular, encoding the potentially high-dimensional and highly structured state of a C2 node into a low-dimensional data representation that is amenable to AI/ML methods is a challenge. Recent breakthroughs in representation learning may present an opportunity in the form of sequence-to-sequence transformers, which have been instrumental in advancing the field of natural language processing and the design of state-of-the-art autonomous agents with superhuman performance for control tasks, such as DeepMind’s AlphaStar and OpenAI’s Five.
· Agent-Environment Dynamics Model: Once the state of an agent and its environment are encoded, the transition dynamics which define the conditions under which a state transition occurs must be defined. Existing simulators and SME knowledge of the operational domain may be sufficient for defining these dynamics. We anticipate that the more pressing challenge will stem from the definition of the reward signal used by the RL algorithm to determine what constitutes a good action or sequence of actions. Whereas traditional RL approaches typically assume that a simple reward signal is given and sufficient to perform the task of interest, such a reward signal is not obvious in operational domains. The use of mission success as a binary reward in these domains is likely to suffer from the sparse rewards problem that arises when an agent must take many actions in the environment before observing a positive or negative reinforcement of its behavior. Thankfully, the field of reward shaping for mitigating this problem has seen significant advancement in recent years through the use of curiosity-driven exploration, imitation learning, and interactive learning, all of which can be used to provide an artificial reward signal for much faster convergence of RL solutions. Hierarchical RL and long short-term memories have also been used successfully to reason about sparse rewards by decomposing the mission objective into sub-tasks and utilizing neural networks with memory units to “remember” actions taken over a long time horizon.
· Data Generation: Existing simulation tools for several operational domains, such as those pertaining to ATO generation and execution, pose a bottleneck on the volume of data generation. Even massively parallel simulations using such tools may not generate enough data for RL applications since the execution time of each ATO can be in the order of hours and state-of-the-art RL algorithms based on experience replay require tens of millions of simulation instances. The poor execution time of these simulators largely stems from the high fidelity of the simulations. Thus, for the sake of generating training data from simulations at speed and scale, a lower-fidelity simulation environment may be proposed. The use of transfer learning and domain adaptation can then be applied to transfer the RL policies learned to an existing simulator or some other higher-fidelity environment. Data augmentation approaches, such as the recently proposed AutoAugment technology, may also be used to complement this. Other promising approaches include the use of DeepFakes and adversarial perturbations for data augmentation from existing data points.
For the operational domains AFRL has selected in this effort, realistic environmental feedback is likely to be leveraged using existing modeling and simulation (M&S) capabilities. Further, based on existing fidelities of these M&S solutions, the Task 2 Offeror will likely need to extend these capabilities. Known M&S capabilities are very powerful for training or exercise development, but tend to operate only a few times the normal speed of operations. To adequately train the algorithms, a more suitable simulation capability may need to be developed to compensate for lack of speed (parallelization, increasing simulation speed at the expense of level-of-detail, identifying a surrogate function with a much lower complexity, etc.). Running existing simulations in this case for the entire duration of this program may not yield enough data to refine the algorithms. Incorporating environmental feedback via simulation will typically involve a substantial amount of effort to develop an environment with sufficient fidelity, speed, and scale in which to train the algorithms. The risks for the Task 2 Offeror stem from the relative complexity of operational planning domain and limitation of existing simulation capabilities. Care must be taken to scope this effort in the context of program goals.
Task 2 Technical Tasks The Task 2 performer shall:
1. (Phase 2a) Select a planning operational thread.
1. (Phase 2a) Identify SMEs with relevant experience and how their interaction will guide the evaluation process. Identify how the chosen operational thread could be solved by a RL approach and identify specific functional objectives for the RL algorithm.
1. (Phase 2a) Identify the data (environment feedback) needs for your approach. Identify existing simulation capabilities, their limitations, and data requirements relevant to your approach.
1. (Phase 2a) Describe your approach and how it will address the identified limitation(s) enabling you to train your model. Develop input training data sets in order to refine your model.
1. (Phase 2a) With SME input, identify and develop evaluation criteria reflective of the program schedule. Describe how SME interaction will guide the evaluation process.
1. (Phase 2a) Develop the RL algorithm based on selected application, guidance from SME, and its data requirements.
1. (Phase 2b) Work with AFRL to identify existing M&S capabilities that could be leveraged or enhanced to meet Task 2 objectives.
1. (Phase 2b) Identify existing models and data sources that could be leveraged or modified to best evaluate planning algorithm performance.
1. (Phase 2b) Develop a simulated environment with sufficient fidelity that will provide the necessary feedback for policy learning / model training.
1. (Phase 2c) Integrate the planning agent with the evaluation environment such that a feedback loop is formed and plan actions are subject to simulated events that produce state changes and potential reward artifacts to improve algorithm performance.
1. (Phase 2c) Conduct an analysis of the feedback loop and determine the best trade-offs of simulation speed, scale and fidelity to refine the planning algorithms. Modify the evaluation environment based on the trade-off analysis and document the new observations.
1. (Phase 2c) If the evaluation environment is plagued by speed deficiencies as the overriding factor, as an example, experiment with virtualization or similar techniques that run multiple simulations at a given time to compensate. Experiment with scale and fidelity factors if necessary.
1. (Phase 2c) Evaluate the feedback loop and refine state/reward as necessary.
1. (Phase 2c) If time and resources permit, identify potential use-cases for multi-domain adaptation for the candidate approach. This does not involve any additional development work, and no additional changes to the developed prototype application are required.
OpML Program Phasing The program schedule describes how the two Tasks will be executed and synchronized over the course of three phases. In general, the core reviews each Offeror will conduct are, by design, aligned with one another to cut down on travel and maximize team availability. These reviews will likely be conducted consecutively at AFRL during a one- to two-week period to maximize input from all developers to each of the designs. Also, each task will end in a series of capstone demonstrations that culminates all the activity that occurred during the effort.
Figure 1 - Program Schedule
Phases 1a and 2a: Algorithm and Prototype Development Phases 1a and 2a are conducted within the first nine months of the program to ensure enough initial resources are dedicated to developing critical algorithms as quickly as possible. Performers will conduct design reviews that will be used to identify the data, algorithmic design, and evaluation criteria and schedule. An internal evaluation will be conducted by each of the developers at their locations at the end of this Phase. This review should include all local SMEs and developers so there is a consistent view of “system” behavior in its current state and where it needs to be feature-wise by the end of the Phase. There is a single evaluation event for each developer during this phase, but keep in mind that the start of Phases 1b and 2c, respectively, specifies a continuous evaluation and feedback process.
The Task 2 Offeror’s second design review should concentrate on the design, development, and integration of existing and new M&S capabilities that will be used to drive the RL feedback loop. This review should identify which existing capabilities will be leveraged, modified (if need be) and/or integrated with new capabilities. New development could include the use of virtualized environments to overcome current M&S deficiencies that typically require hours, or at best minutes, to execute scenarios that would likely require several orders of magnitude better performance to adequately train policies.
Phase 1b: Algorithm Evaluation Phase 1b is best summarized as a continuous evaluation process that includes algorithmic refinement to address shortfalls in the code. Phase 1a had specific evaluations so as not to impede core prototype development, but Phase 1b will require nearly continuous evaluation. This DevOps-like process will prepare all the developers for Phase 1c experimentation. All of the performers have the same tempo of two design reviews with a single comprehensive evaluation to end the phase. This evaluation will be used to determine the readiness of each of the implementations for inclusion in the MDO experimentation based on the defined metrics.
Whenever possible, Task 1 performers will leverage the SML framework and community expertise but still evaluate their software at their own facilities for the duration of the phase. Task 1 performers should develop a plan to make the software evaluation available to all OpML performers, to include the incorporation of new datasets into the evaluation process.
Phase 2b: Evaluation Environment Development Task 2 performers have the additional responsibility to not only establish an in-house evaluation environment, but also explore the stand-up of a mirror simulation environment at AFRL. This is critical because AFRL may be the only logical choice to host the necessary simulation capabilities (there could be several) based on the Task 2 approach and data classification requirements. At the AFRL evaluation environment, the same requirements hold to incorporate new datasets and the ability to modify the scenario threads using the legacy or newly developed simulation functionality.
Phase 1c: Multi-Domain Operations (MDO) Experimentation Phase 1c is similar to the previous phase except now development will be under the guise of Flyleaf processes with a focus on multi-domain operations of software produced for a single domain. It is not expected that OpML algorithms will undergo massive changes during this phase, but refinements that will begin to address MDC2 challenges by identifying technical gaps as the developers work the seams between domains and integration with Flyleaf scenarios and interaction with other C2 tools. All formal evaluations will be coordinated in the Flyleaf integration environment, but Flyleaf provides an ability for the developers to gain access to the environment from their facilities.
Phase 2c: Algorithm Evaluation This phase is also a continuous evaluation process that includes algorithmic and environmental refinement to maximize the RL feedback loop to achieve best program results. This phase requires a balance of algorithm versus environment enhancements driven by the trade-off analyses and experimental observations. In addition, careful resource consideration and planning may be required if the optimal evaluation environment is located at AFRL. Potential offerors should provide a resource plan and mitigation strategy in the event that an iterative evaluation process is not directly under their configuration control.
Program Metrics Because there are multiple approaches being evaluated under this effort, there are no uniform set of metrics that can be used as criteria to gauge how well the ML algorithms are performing. Therefore, a separate set of metrics will be used for Task 1 Offerors versus the Task 2 Offeror. In addition, for the metrics that do appear in Table below, they are not considered absolute. Rather, they should be a set of candidate metrics that should evolve and be addressed at each of the design reviews throughout all phases of the program. In this manner, the Government will have final oversight of the metrics to be used to evaluate the effectiveness of the proposed solutions in achieving the stated program tasks and objectives. It is important for all Task 1 Offerors to address the fact that the metrics will be coupled tightly to the MDC2 scenarios and individual mission threads that comprise the evaluation and experimentation events. Finally, in the detailed Operational Threads section, some of the Algorithm Performance and/or Evaluation Fidelity metrics may be expanded upon in the context of key performance parameters (KPP).
| Metric (KPP) |
| Prototype Dev Goal |
| Algorithm Evaluation Goal |
| End of Program Goal |
| Number of Air Threads |
| 5 |
| 15 |
| 15 |
| Dataset Fidelity |
| LOW |
(SME verified)
MED
(SME verified)
MED
(SME verified)
| Algorithm Performance |
| 2x speed up over manual methods |
| 5x speed up over manual methods |
| 5x speed up over manual methods |
| Evaluation Fidelity (Task 1) |
| LOW |
(SME verified)
MED
(SME verified)
MED
(SME verified)
| Simulation Performance (Task 2) |
| 5x speed up over existing M&S tools |
| 10x speed up over existing M&S tools |
| 10x speed up over existing M&S tools |
| Number of Domains (Task 1) |
| 1 |
| 1 |
| 2-3 |
Table 2 - OpML Program Metrics
See the BAA Attachment for details on the Eight Operation Threads.
The technical point of contacts (TPOC) for both of these focus areas are:
OMLC2 TPOCs:
Mr. Gennady Staskevich
AFRL/RISC
525 Brooks Rd Rome, NY 13441-4505 Telephone: (315)330-4889 Email: gennady.staskevich@us.af.mil
Mr. Carlos Merlos
AFRL/RISC
525 Brooks Road Rd Rome, NY 13441-4505 Telephone: (315)330-4128 Email: carlos.merlos@us.af.mil
Alternate TPOC:
Dr. Alvaro Velasquez
AFRL/RISC
525 Brooks Road Rome, NY 13441-4505 Telephone: (315) 330-2287 Email: alvaro.velasquez.1@us.af.mil
IMPORTANT NOTES REGARDING:
FUNDAMENTAL RESEARCH. It is DoD policy that the publication of products of fundamental research will remain unrestricted to the maximum extent possible. National Security Decision Directive (NSDD) 189 defines fundamental research as follows:
‘Fundamental research’ means basic and applied research in science and engineering, the results of which ordinarily are published and shared broadly within the scientific community, as distinguished from proprietary research and from industrial development, design, production, and product utilization, the results of which ordinarily are restricted for proprietary or national security reasons.
As of the date of publication of this BAA, the Government cannot identify whether work proposed under this BAA may be considered fundamental research and may award both fundamental and non-fundamental research. Proposers should indicate in their proposal whether they believe the scope of the research included in their proposal is fundamental or not. While proposers should clearly explain the intended results of their research, the Government shall have sole discretion to select award instrument type and to negotiate all instrument terms and conditions with selectees. Appropriate clauses will be included in resultant awards for non-fundamental research to prescribe publication requirements and other restrictions, as appropriate.
For certain research projects, it may be possible that although the research being performed by the awardee is restricted research, a sub-awardee may be conducting fundamental research. In those cases, it is the awardee’s responsibility to explain in their proposal why its sub-awardee’s effort is fundamental research.
CLOUD COMPUTING. In accordance with DFARS Clause 252.239-7010, if the development proposed requires storage of Government, or Government-related data on the cloud, offerors need to ensure that the cloud service provider proposed has been granted Provisional Authorization by the Defense Information Systems Agency (DISA) at the level appropriate to the requirement.
II. AWARD INFORMATION:
1. FUNDING: Total funding for this BAA is approximately $24,900,000. The anticipated funding to be obligated under this BAA is broken out by fiscal year as follows:
FY20 - $10,000,000
FY21 - $10,000,000
FY22 - $4,900,000
1. Individual awards will not normally exceed 24 months with dollar values normally ranging from $300,000 to $1,000,000. There is also the potential to make awards up to any dollar value as long as the value does not exceed the available BAA ceiling amount of $24,900,000.
1. The Government reserves the right to select all, part, or none of the proposals received, subject to the availability of funds. All potential Offerors should be aware that due to unanticipated budget fluctuations, funding in any or all areas may change with little or no notice.
2. FORM. Awards of efforts as a result of this announcement will be in the form of contracts, grants, cooperative agreements or other transactions depending upon the nature of the work proposed.
3. BAA TYPE: This is a two-step open broad agency announcement. This announcement constitutes the only solicitation.
As STEP ONE – The Government is only soliciting white papers at this time. DO NOT SUBMIT A FORMAL PROPOSAL. Those white papers found to be consistent with the intent of this BAA may be invited to submit a technical and cost proposal. See Section VI of this announcement for further details regarding the proposal.
III. ELIGIBILITY INFORMATION:
1. ELIGIBILITY: All qualified offerors who meet the requirements of this BAA may apply.
2. FOREIGN PARTICIPATION/ACCESS:
1. This BAA is closed to foreign participation. This includes both foreign ownership and foreign nationals as employees or subcontractors.
1. Exceptions.
1. Fundamental Research. If the work to be performed is unclassified, fundamental research, this must be clearly identified in the white paper and/or proposal. See Part II, Section I for more details regarding Fundamental Research. Offerors should still identify any performance by foreign nationals at any level (prime contractor or subcontractor) in their proposals. Please specify the nationals’ country of origin, the type of visa or work permit under which they are performing and an explanation of their anticipated level of involvement. You may be asked to provide additional information during negotiations in order to verify the foreign citizen’s eligibility to participate on any contract or assistance agreement issued as a result of this announcement
1. Foreign Ownership, Control or Influence (FOCI) companies who have mitigation plans/paperwork in place. Proof of approved mitigation documentation must be provided to the contracting office focal point, Amber Buckley, Contracting Officer, telephone (315) 330-3605, or e-mail Amber.Buckley@us.af.mil prior to submitting a white paper and/or a proposal. For information on FOCI mitigation, contact the contact the Defense Counterintelligence and Security Agency (DCSA). Additional details can be found at: https://www.dcsa.mil/mc/ctp/foci/
1. Foreign Nationals as Employees or Subcontractors. Applicable to any effort not considered Fundamental Research. Offerors are responsible for ensuring that all employees and/or subcontractors who will work on a resulting contract are eligible to do so. Any employee who is not a U.S. citizen or a permanent resident will be restricted from working on any resultant contract unless prior approval of the Department of State or the Department of Commerce is obtained via a technical assistance agreement or an export license. Violations of these regulations can result in criminal or civil penalties.
1. Information Regarding Non-US Citizens Assigned to this Project
0. Contractor employees requiring access to USAF bases, AFRL facilities, and/or access to U.S. Government Information Technology (IT) networks in connection with the work on contracts, assistance instruments or other transactions awarded under this BAA must be U.S. citizens. For the purpose of base and network access, possession of a permanent resident card ("Green Card") does not equate to U.S. citizenship. This requirement does not apply to foreign nationals approved by the U.S. Department of Defense or U.S. State Department under international personnel exchange agreements with foreign governments. It also does not apply to dual citizens who possess US citizenship, to include Naturalized citizens. Any waivers to this requirement must be granted in writing by the Contracting Officer prior to providing access. Specific format for waiver request will be provided upon request to the Contracting Officer. The above requirements are in addition to any other contract requirements related to obtaining a Common Access Card (CAC).
0. For the purposes of Paragraph 1, it an IT network/system does not require AFRL to endorse a contractor's application to said network/system in order to gain access, the organization operating the IT network/system is responsible for controlling access to its system. If an IT network/system requires a U.S. Government sponsor to endorse the application in order for access to the IT network/system, AFRL will only endorse the following types of applications, consistent with the requirements above:
1. Contractor employees who are U.S. citizens performing work under contracts, assistance instruments or other transactions awarded under this BAA.
1. Contractor employees who are non-U.S. citizens and who have been granted a waiver.
Any additional access restrictions established by the IT network/system owner apply.
3. FEDERALLY FUNDED RESEARCH AND DEVELOPMENT CENTERS AND GOVERNMENT ENTITIES: Federally Funded Research and Development Centers (FFRDCs) and Government entities (e.g., Government/National laboratories, military educational institutions, etc.) are subject to applicable direct competition limitations and cannot propose…
This is the start of the file's text. The full file is on GovTribe.
File details come from the government source that posted it. Updated .