HR001121S0040-Amendment-01.pdf

PDF 1 MB Posted

Attached to
Hardening Development Toolchains Against Emergent Execution Engines (HARDEN) Federal contract opportunity
Solicitation number
HR001121S0040
Issued by
Defense Advanced Research Projects Agency

About this file

This Broad Agency Announcement describes a research program called Hardening Development Toolchains Against Emergent Execution Engines (HARDEN). The Defense Advanced Research Projects Agency is soliciting proposals to develop tools that can anticipate, isolate and mitigate unintended programmable behaviors in software throughout the development lifecycle. The goal is to prevent adversaries from exploiting unexpected program execution in integrated systems.

The program involves four technical areas: tooling for developers, modeling emergent behaviors, an offensive security perspective, and integrating/evaluating technologies. Multiple awards are expected for the first two areas over three phases from 2022-2026. Proposals are due November 4, 2021 and will be evaluated based on technical and management factors. The notice provides details on program structure, evaluation metrics, intellectual property terms and eligibility requirements.

View the file

Other files for this federal contract opportunity

Other files attached to Hardening Development Toolchains Against Emergent Execution Engines (HARDEN), newest first.
File Type Posted
Proposal_Summary_Slide_-_final.pptx PPTX presentation
HR001121S0040.pdf PDF

On GovTribe

Work with this file on GovTribe

  • Download the original file
  • Contacts named in this file
  • Similar government files
  • Ask GovTribe AI about this file

Text version

Broad Agency Announcement

Hardening Development Toolchains Against Emergent

Execution Engines (HARDEN)

INFORMATION INNOVATION OFFICE

HR001121S0040

Amendment 1

10/5/2021

Summary of Amendment 1 Changes:

The purpose of this amendment is to include new Countering Foreign Influence Program (CFIP) guidance. The following sections have been amended and changes are highlighted in yellow throughout the document:

1. PART II. FULL TEXT OF ANNOUNCEMENT; II. Award Information; B.

Fundamental Research

2. PART II. FULL TEXT OF ANNOUNCEMENT; IV. Application and Submission Information; B. Content and Form of Application Submission

3. PART II. FULL TEXT OF ANNOUNCEMENT; V. Application Review Information; B. Review of Proposals

TABLE OF CONTENTS

PART I: OVERVIEW INFORMATION

PART II: FULL TEXT OF ANNOUNCEMENT

I. Funding Opportunity Description II. Award Information

A. General Award Information B. Fundamental Research

III. Eligibility Information A. Eligible Applicants B. Organizational Conflicts of Interest C. Cost Sharing/Matching

IV. Application and Submission Information A. Address to Request Application Package B. Content and Form of Application Submission

V. Application Review Information A. Evaluation Criteria B. Review of Proposals

VI. Award Administration Information A. Selection Notices and Notifications B. Administrative and National Policy Requirements C. Reporting D. Electronic Systems E. DARPA Embedded Entrepreneur Initiative (EEI)

VII. Agency Contacts VIII. Other Information

IX. APPENDIX 1 – PROPOSAL SUMMARY SLIDE

PART I: OVERVIEW INFORMATION

Federal Agency Name – Defense Advanced Research Projects Agency (DARPA), Information Innovation Office (I2O) Funding Opportunity Title – Hardening Development Toolchains Against Emergent

Execution Engines (HARDEN) Announcement Type – Initial announcement Funding Opportunity Number – HR001121S0040 Catalog of Federal Domestic Assistance Numbers (CFDA) – 12.910 Research and

Technology Development Dates o Posting Date: September 20, 2021 o Proposers Day: September 30, 2021 o Questions Due: November 1, 2021, 12:00 noon, Eastern Time o Proposal Due Date: November 4, 2021, 12:00 noon, Eastern Time o Solicitation Closing Date: March 21, 2022, 5:00 pm, Eastern Time

Program Overview – The HARDEN program will explore novel approaches that use formal verification methods and Artificial Intelligence (AI)-aided program models, analyses, and logics to develop practical tools to anticipate, isolate, and mitigate emergent execution engines throughout the entire software development lifecycle in order to disrupt the patterns of robust, reliable, and composable exploit primitives that empower attackers.

Anticipated Individual Awards – There are multiple technical areas for this solicitation.

Multiple awards are anticipated in Technical Area 1 and Technical Area 2, and a single award is anticipated in Technical Area 3 and Technical Area 4.

Types of Instruments that May be Awarded – Procurement Contracts, Cooperative Agreements, or Other Transactions (OT)

Agency Contacts o Points of Contact

The BAA Coordinator for this effort can be reached at:

Email: HARDEN@darpa.mil.

DARPA/I2O

ATTN: HR001121S0040

675 North Randolph Street Arlington, VA 22203-2114

PART II: FULL TEXT OF ANNOUNCEMENT

I. Funding Opportunity Description

This publication constitutes a Broad Agency Announcement (BAA) as contemplated in Federal Acquisition Regulation (FAR) 6.102(d)(2) and 35.016 and 2 C.F.R. § 200.203. Any resultant award negotiations will follow all pertinent laws and regulations, and any negotiations and/or awards for procurement contracts will use procedures under FAR 15.4, Contract Pricing, as specified in the BAA.

The Defense Advanced Research Projects Agency (DARPA) is soliciting innovative proposals in the following areas of interest: tools to anticipate, isolate, and mitigate adversarially programmable emergent behaviors in integrated software systems, and tools to protect intended software abstractions from adversarial reuse. Proposed research should investigate innovative approaches that enable revolutionary advances in theory, tools, devices, or systems. Specifically excluded is research that primarily results in evolutionary improvements to the existing state of practice.

A. Program Overview

Introduction

The Department of Defense (DoD) has a critical need to deny cyber attackers the capability to execute unintended, yet robust and often unobservable computations on DoD systems and critical infrastructure systems. The Hardening Development Toolchains Against Emergent Execution Engines (HARDEN) program will explore novel theories and approaches, and develop practical tools to anticipate, isolate, and mitigate emergent behaviors in computing systems throughout the entire software development lifecycle (SDLC). HARDEN will radically improve security outcomes in software for integrated systems by creating novel tools, metadata, and instrumentation for emergent computation, and it will efficiently mitigate exploitation of software abstractions and protect intended abstractions from adversarial reuse. HARDEN will integrate those capabilities into the standard processes of the SDLC.

Empirically, modern exploitation methods rely on long chains of emergent behaviors of the target’s unprotected computational abstractions, where attackers leverage one combination of abstractions to create an ephemeral state in which the next set of unprotected abstractions is exposed, until the goals of exploitation are achieved. Counterintuitively, instead of being brittle and easily disrupted, these chains are robust and portable between implementations independently created by different vendors. This phenomenon is colloquially described as “weird machines”—well-defined, robust, and abstractable engines of emergent execution (EE) and adversarial programmability—already pre-existing within the target and being merely unlocked for an attacker’s use through coding flaws.

Removal of individual code flaws by initial fixes and mitigations tends to be ineffective against methodical exploit programming because such fixes typically fail to disrupt the underlying emergent execution engine or “weird machine,” which remains accessible to the attackers through other flaws.

The HARDEN program will use formal verification methods and Artificial Intelligence (AI)-aided program models, analyses, and logics to develop practical tools to prevent exploitation of emergent execution engines by disrupting the patterns of robust, reliable exploits used by attackers.

HARDEN tools will facilitate analyses of integrated systems by leveraging modeling and analysis of multi-layered software abstractions, their interactions, and emergent properties.

HARDEN tools will analyze the extent of protection of each layer of abstraction and the semantic anchorings in lower layers to reason about its propensity for adversarial programmability and EE. Based on these analyses, the tools will alert system designers and developers about designs and implementations likely to result in adversarially programmable emergent behaviors. HARDEN tools will also suggest semantically equivalent transformations of implementations to mitigate composability of emergent behaviors and to disrupt exploit programming. Additionally, HARDEN tools will help validate both the design and implementation of integrated systems and inform architectural security standards for systems-of-systems.

To accomplish these goals, the HARDEN program will leverage the insight that composability of emergent behaviors and unprotected abstractions yield key advantages for the attacker.

Composability allows the attacker to create resilient programs out of sequences of emergent behaviors, chaining exploit primitives even where security mitigations reduce the effects of any single behavior or flaw. HARDEN’s insight is that unintended composability of the systems’ own abstractions and emergent behaviors is what enables attackers to robustly and effectively drive the system through long series of unintended illegal states without crashing or manifesting other observable signs of misbehavior.

The program seeks breakthrough approaches to the following technical challenges, including but not limited to:

Overcoming state explosion of typical models of software behavior;

Making annotation of expected behavior and predictions of emergent behavior accessible to typical software developers;

Developing efficient means of communicating about EE with software architects and developers;

Anticipating and preventing potential EE within common developer workflows and tools;

Creating models of EE that capture designed-in EE and abstract away irrelevant parts of the implementation;

Modeling interfaces and Application Programming Interfaces (APIs) at several layers of abstraction, together with the interactions between these layers; and, Developing effective tiered representations of abstractions to reason about EE and formats, and to efficiently store and retrieve these representations alongside software deliverables (e.g., by extending symbolic debugging data formats).

The HARDEN program will focus on validating approaches by applying broad theories and generic tools to concrete technological use cases of integrated software systems described in this BAA, with the overall goal of producing comprehensive security improvements in these systems as a whole.

Background

Today’s software development pipelines and testing methodologies do not typically include tools for reasoning about adversarial reuse of code that was correct for its original purpose. This leads to unwitting creation of stable, reliable patterns of emergent behaviors within integrated software systems that lend themselves to adversarial programming by attackers. Attackers have demonstrated an increasing ability to compose emergent behaviors of target systems into unintended and “weird,” yet effective and robust, exploit programming models and execution engines. Some examples are: a) modern web browser exploits co-opt functionality of the browser’s sophisticated memory management algorithms and just-in-time compilation of web scripts; b) the Spectre family of exploits adversarially reuse the Central Processing Unit’s (CPU's) microarchitecture and transactional memory mechanisms; and c) modern bootkits leverage elements of the trusted computing system management modes. In each case, attackers try to program an already present unintended engine of emergent behaviors with a sequence of suitable macro- or micro-events.

Since today’s developer tools focus only on intended execution paths and limited deviations from these paths, the developers remain unaware of designed-in emergent execution behaviors and adversarial reprogramming modes of their products for years after their release. Even when advanced exploitation models based on these behaviors are made public, creating effective mitigations takes years, as these mitigations need to consider widely used designed-in features rather than random developer errors.

Today, the software industry uses two approaches to experimentally characterize emergent behaviors in finished products: fuzz-testing (a.k.a. fuzzing) and Chaos Engineering. Fuzzing subjects the system-under-test to randomly generated malformed inputs and records violations of intended behaviors (such as crashes or out-of-bound memory accesses), while selecting or mutating inputs to maximize observed program coverage. Security researchers then manually analyze elicited violations for composability and judge whether they enable general patterns of programmability by the adversary. Fuzzing operates on the compiled binary form of the software product—in essence, the lowest form of computing abstraction above hardware, which is forgetful of most higher-level abstractions. Google and Microsoft’s recently open-sourced fuzzing infrastructures enable testing of individual code units rather than finished products, as they both recognize the effectiveness of fuzz-testing for finding code flaws.

Chaos Engineering operates at the highest form of architectural abstractions in distributed systems, such as services, nodes, or network functions, subjecting their instances to simulated random disconnections or excessive latency. Chaos Engineering eschews simulating data corruption (such as fuzzing), as it focuses on availability at scale regardless of the root causes of individual node failures.

Additionally, the industry is starting to adopt formal methods to ensure that the code behaves according to its specifications. However, today’s formal methods do not characterize behaviors of code that are outside the specifications, nor do they reveal behavior in the presence of violations of specification assumptions. Without the ability to represent or discover intermediate and implicit abstractions involved in implementations, today’s formal methods cannot help explore their composition properties, the resulting EE modes, and the adversarial programming models of their unintended reuse by attackers.

HARDEN will offer efficient mitigation at the early SDLC stages and protection of intended software abstractions from adversarial reuse. HARDEN will instrument the development toolchain to provide reasoning about emergent behaviors at all available layers of abstraction, from the compiled binary code through the compiler abstractions and intermediate representations, to the highest levels of architectural abstraction. It will develop metadata representations, logics, symbolic and binary instrumentation, as well as developer-focused tooling integrated with the standard build chains and integrated development environments (IDEs) to warn the designer and the developer about potential emergent behaviors at their inception point.

Insufficiency of Current Approaches

Neither fuzzing nor Chaos Engineering provide ways to reason about emergent behaviors at their inception. Both methods are limited to the lowest and the highest levels of architectural abstraction, and fail to exploit architectural knowledge of the intermediate abstraction layers in integrated systems. Neither method explores the composability of emergent behaviors.

These limitations are crucial. Prior studies have demonstrated that reliable adversarial reuse of code is enabled in multi-layered systems by particular implementations of higher-level abstractions via intermediate abstractions. Exploits leverage unintended, emergent, but fairly general and resilient abstractions impressed onto lower system layers by systematic design choices inherent in the target itself or in the development toolchain. Abuse of these impressions, known as “leaky abstractions,” is what lends exploitation methodologies their resilience and portability between platforms, despite the many low-level differences between these platforms.

What makes exploitation methodologies teachable and their mitigation hard is that disrupting the unintended abstractions that an exploit relies upon must be done without disrupting the intended design.

For example, the Return-Oriented Programming (ROP) exploit relies on the design of the stack activation frames (an intermediate compiler abstraction). Other exploits such as Jump Oriented Programming (JOP), Signal Return Oriented Programming (SROP), and Counterfeit Object Oriented Programming (COOP) rely on indirect control flow abstractions created by compilers or systems’ libraries. Heap memory manipulation techniques, “heap grooming” or “heap Feng-shui,” use abstractions of the heap metadata originating in the design, metadata operations, and memory management algorithms. This is why these techniques are effective with small variations across different instruction-set architectures and operating systems. Similarly, exploits leveraging CPU microarchitectures and chains-of-trust rely on intermediate abstractions of these designs and therefore persist across architectures. This way, the exploits can be modeled without regard to the details of specific target microarchitectures or chipset implementations.

Empirically, even the best-of-breed fuzzing methods fail to uncover flaws in software layers not immediately adjacent to the interface through which attackers inject their crafted inputs. This happens because the fuzzer’s guiding algorithm must essentially re-discover intended code paths, data structures, and other interfaces at great cost, and with little to no knowledge of the underlying abstractions. For that same reason, even when fuzzing triggers emergent behaviors in higher layers, it cannot reason about its composition.

Without the ability to model, represent, and discover intermediate abstractions inherent in designs and implementations, today’s methods, including formal verification methods, cannot explore their composition properties, the resulting EE modes, and the adversarial programming models of the unintended reuse by attackers.

Program Scope

HARDEN tools will facilitate modeling and analysis of integrated systems with technological stacks including, but not limited to, the following:

• Instrumentation of the development toolchain for reasoning about EE behaviors at all available layers of abstraction;

• Capabilities for effective searching and automated reasoning about EE behaviors for a wide variety of higher-layer abstractions, generalizing recent methods for reasoning about unintended behaviors without complete knowledge of implementations; and

• Prevention of composability of EE behaviors underlying robust exploit chain construction.

HARDEN tools for integrated systems will produce assurance evidence and trustworthiness outcomes superior to those of the current approaches of fuzz-testing and Chaos Engineering.

Although HARDEN seeks to create broad theories and generic tools, the program will focus on validating its approaches by applying them to concrete technological use cases of integrated software systems described further under “Technological Use Cases” within Section I.B.

B. Program Structure

The program will produce theories, technologies, tools, and formal methodologies leading to experimental prototype(s) that provide capabilities for the mitigation of emergent behaviors throughout the software lifecycle in order to improve security outcomes in software for complex integrated systems. It is expected that these prototypes will provide a starting point for technology transition and demonstrate that chained exploits can be impeded by disrupting them at all levels of abstraction in mission-critical software.

The HARDEN program is a 48-month program organized into three phases: Phases 1 and 2 will each be 18-months, followed by a 12-month Phase 3. Each of the Phases’ Metrics are described in Table 1 under Section I.C.

The program is divided into four Technical Areas (TAs) to support program goals:

TA1: Tooling for developers TA2: Modeling of emergent behaviors TA3: Voice of the offense TA4: Integration and systems engineering evaluation

Figure 1: HARDEN TAs 1-4 with notional subtasks and challenges

DARPA anticipates funding multiple technical approaches and performers across the HARDEN technical areas. Beyond Phase 1, subsequent phases will be considered options, and may or may not be exercised at the sole discretion of the Government. Funding of options will be based on demonstrated technical progress towards the goals of the HARDEN program and availability of funding.

Within the program phases, proposers are encouraged to identify a compact viable core subset of their proposed technologies and then associate them to proposal options that increase the practical coverage of the technological use cases discussed in “Technological Use Cases” and “Exemplary Evaluation and Transition Use Cases” within this section. These optional add-ons may or may not be exercised at the sole discretion of the Government.

TA1 and TA2 performers should be prepared to work closely with each other in order to support the integration of the TA1 tools for effective checking of EE models developed by TA2, and for TA1 tools that create effective mitigations for EE anticipated by the TA2 models. In addition, TA2 and TA3 performers should be prepared to work closely with each other, in order to ensure that TA2’s models reflect TA3’s insights of edge-of-the-art exploitation and enhance these insights.

To facilitate the open exchange of information, performers will have Associate Contractor Agreement (ACA) language included in their award, which is described further in Section VIII.

The TA4 performers will be responsible for executing the HARDEN ACA. While TA4 performers will be a party to the ACA, it is expected that TA4 outputs will be largely independent of TA1, TA2, and TA3 work, although robust interaction is expected.

Each proposal may address any one TA, or a combination of TA1 and TA2. Proposals covering a combination of TA1 and TA2 must make the respective efforts and costs proposed on the different TAs clearly separable to enable partial awards, and should explain their rationale for combining these two TAs, and their collaboration plans with the other TAs. Significant cost reductions for the combined TA1 and TA2 effort will be expected through synergies of the proposed approaches.

Proposers may submit multiple proposals. The Government reserves the right to decide which, if any, are selected for award. A proposer submitting combined TA1 and TA2 proposals may be selected to perform on one, or both, of these TAs. A proposer submitting proposals to TA3 and some other TA(s), if selected to perform on TA3, cannot be selected to perform on any other TAs, whether as a prime, subcontractor, or any other capacity from an organizational to individual level, to protect the integrity of the program evaluation. Similarly, a proposer submitting proposals to TA4 and some other TA(s), if selected to perform on TA4, cannot be selected to perform on any other TAs, whether as a prime, subcontractor, or any other capacity from an organizational to individual level.

DARPA encourages proposers to consider the investigation and creation of open-source and free software approaches. DARPA strongly encourages that proposals provide an overall open-source HARDEN framework that will result in open, modular tool architectures. Restricting technology transition of a proposed HARDEN technology may be considered a weakness of the proposal and DARPA believes that open-source solutions are critical to support program transitions.

The Government will assess performer progress against target goals set for each phase using a progression of technical use case challenges as outlined below. In addition, an advisory panel composed of participants from Government partners may participate in the meetings and informal challenges to provide feedback to the DARPA Program Manager.

Technological Use Cases

HARDEN’s TAs will address the following technological use cases of integrated software systems:

(1) Unified Extended Firmware Interface (UEFI), chain-of-trust, and trusted boot technologies that govern trusted boot processes and integrity of a modern computing system; and

(2) Integration technologies for securely connecting a tablet User Interface (UI) system with a trusted computer, such as a mission computer.

Performers will need to support both use cases for all tasks except TA2. TA2 proposers may choose to address only one, or both, of the use cases. If the choice is to address both use cases in the same proposal—for example, due to identified technological synergies and availability of platform expertise for both use cases—the proposers should make the tasking and the costs for each case clearly separable, so that the Government may select only one use case for a partial award, as explained later in this BAA.

In both of these integrated software system use cases, persistent vulnerabilities are known to exist. Patterns of exploitation and incomplete mitigations of these vulnerabilities suggest a wealth of uncontrolled emergent behaviors and abstraction leaks creating EE engines that can be successfully exploited by attackers across implementations of different provenance and by different vendors. HARDEN theories, models, methods, and tools will radically improve trustworthiness of these use cases by acting across multiple layers of its design and implementation.

Strong proposals to all TAs should discuss proposed approaches in terms of concrete exemplar systems matching the above use cases, for which the source code and the build processes are available for the majority of the system. Strong proposals should also discuss approaches for dealing with opaque system components in both software and hardware, such as methods for validating available specifications against the actual hardware, combining automated inference and automated interface exploration and interrogation.

More information about these use cases is provided below under “Exemplary Evaluation and Transition Use Cases” within this section.

TA1 – Tooling for developers

TA1 performers will develop and combine novel approaches for scalable reasoning about behaviors of computing systems’ units, layers, components, and subsystems, to support multi-level modeling of system state evolution and emergent behaviors of unprotected abstractions. To support such reasoning, TA1 will develop novel theories, models, metadata, instrumentation, and tools.

TA1 tools will receive use case-specific models developed by TA2 throughout the program and will support reasoning about these models at the scales and granularities necessary to harden the software technologies essential to the trustworthiness of integrated system use cases described under “Technological Use Cases” within Section I.B. TA1 is expected to inform TA2’s multi-level modeling of the integrated system use cases by providing feedback on which kinds of models can be concretely reasoned about.

TA1 will integrate these approaches and tools with popular development environments to produce effective and intelligible warnings of EE for system designers and developers, and will help them protect intended abstractions and create effective mitigations of EE across layers.

These tools will be evaluated through use by the TA4 performer.

In addition, TA1 will provide its tools to TA3, who will use them in white-box testing of the exemplar systems.

Strong proposals should present a cohesive theory of EE that:

(1) Enables automated reasoning about EE that is algorithmically efficient and implementation agnostic;

(2) Consistent with insights from the latest edge-of-the-art exploitation experience;

(3) Capable of being developed and applied incrementally to improve trustworthiness of the use cases; and

(4) Enables tractable reasoning at the scales of the integrated software use cases.

Quantitative arguments that support the rationale for anticipated success of the proposed methods would strengthen proposals.

TA1 performers may employ any methods including, but not limited to, AI methods for searching large state spaces of EE models; hybrid, compiler-assisted analyses that leverage higher-level abstractions available at compilation; parametrized unit harnesses for testing expected behavior and automatic exploration of EE; interactive counterexample-guided methods;

or any hybrid method. TA1 may complement static methods of analysis and EE disruption with dynamic methods, so long as the static and dynamic methods provide strong and cohesive assurance guarantees.

Developers working with the TA1 integrated environment will receive timely, interactive, and intelligible feedback on the propensity of their designs and implementations to create EE behaviors and will receive interactive guidance from the automated tools on how to mitigate it.

The TA1 integrated environment will take advantage of the existing specification, if any, interface control documents, if any, source code and related metadata, build chain, unit tests, and other information available for the use case code base, subject to the caveats under “Technological Use Cases” within Section I.B.

A TA1 capability will require research breakthroughs in overcoming state explosion of typical models of software behavior, making annotation of expected behavior and predictions of emergent behavior accessible to regular software developers, developing efficient means of communicating EE to developers, and integrating the ability to anticipate EE with common developer workflows and tools.

Strong TA1 proposals should discuss how the proposed effort will progress from basic models of the exemplar systems to increasing coverage and assurance of these systems. The Government would prefer that this progression is discussed in the context of concrete open-source software (or open software/hardware combination) that is representative of industry deployments and is relevant to the core technological use cases’ trustworthiness. The discussion should clearly and quantitatively identify the challenges and obstacles on achieving superior security outcomes and present a compelling rationale for why the proposed approach will be successful at the scales and granularities necessary for large-scale commercial or open-source software development.

Proposals should present additional metrics and milestones for evaluating the progress of the envisioned approaches in the context of the discussed use case, in line with the general program metrics described below in Table 1.

Strong TA1 proposals will present a review of the existing approaches, techniques, and challenges, emphasizing industry experience, and applications to large integrated systems supported by appropriate literature citations.

Strong proposals will offer metrics and benchmarks for evaluating the success of the newly developed technologies in comparison to existing approaches in open reproducible settings.

Intellectual property rights asserted by proposers are strongly encouraged to be aligned with open-source regimes. See Section IV.B.2.i for more details on Intellectual Property.

TA2 – Modeling of emergent behaviors

TA2 performers will focus on creating the capability for platform experts to model EE behavior in the use cases described under “Technological Use Cases” within Section I.B. TA2 performers will develop modeling methods, languages, and tools for characterizing emergent execution at different abstraction levels; produce and represent tiered and scalable models suitable for fast TA1 feedback; and will create models of EE for all relevant interfaces in the integrated system use cases.

In particular, TA2 performers will formulate approaches for modeling emergent behaviors across multiple layers of abstraction in integrated systems, and will use their subject matter expertise with the use case platforms and technologies to develop use-case specific models of emergent behaviors.

TA2’s models will inform TA1’s instrumentation of the development toolchain for reasoning about emergent behaviors at all available layers of abstraction, from the compiled binary code through the compiler abstractions and intermediate representations, to the highest levels of architectural abstraction. TA2’s use-case specific models will be ingested and reasoned by TA1 tools. TA2 performers will receive feedback from TA1 regarding the models’ amenability to reasoning at scale, and will iterate on the design and content of these models to support scaling.

Strong proposals should present cohesive modeling approaches to produce models that:

(1) Capture relevant descriptions of emergent behaviors empirically known to be of importance to the use cases' trustworthiness;

(2) Can be effectively reconciled with actual software and hardware behaviors;

(3) Are suitable for automated reasoning to anticipate emergent behaviors at the use case software scale envisioned by the BAA, so that the resulting solver computations needed to process the model should be within reach, possibly assuming some algorithmic breakthroughs;

(4) Can be formulated incrementally for the integrated system use cases; and

(5) Reflect expert-level knowledge of the use cases platforms and exploitation thereof.

Quantitative arguments to support the rationale for anticipated success of the proposed methods would strengthen proposals.

Strong TA2 proposals should address automation for deriving TA2 models from source code, build systems, or compiled binary code, and should seek to reduce the amount of platform subject matter expertise and labor needed to produce such models. TA2 modeling may take advantage of the existing specification (if any), interface control documents (if any), source code and related metadata, build chain, unit tests, and other information available for the use case code base, subject to the caveats under “Technological Use Cases” within this section.

Successful TA2 modeling approaches and tools are expected to be effective without complete knowledge of the implementations of underlying abstraction layers where such knowledge is not needed for anticipating adversarial programmability. Strong TA2 proposals should address the development of models for anticipating composability of emergent behaviors and for the reliable chaining of exploit primitives, even where the effects of any single behavior or flaw are reduced by current security mitigations.

Strong TA2 proposals will have detailed plans for supporting TA1’s automated techniques for identifying implementations that are likely to result in composable emergent behaviors and for suggesting semantically equivalent implementation transformations that mitigate emergent composability and disrupt exploit programming. TA2 performers will be expected to work closely with TA3 and should provide a plan for interactions that leverage TA3’s insights about state-of-the-art exploitation.

Strong proposals should discuss their modeling approaches and tools within the context of concrete open-source software (or an open software/hardware combination) that is representative of industry deployment and is relevant to the core technological use cases’ trustworthiness. The discussion should clearly and quantitatively identify the challenges and obstacles to achieving superior security outcomes and present a compelling rationale for why the proposed approach will be successful at the scales and granularities necessary for large-scale commercial or open-source software development. Proposals should present metrics and milestones for evaluating the progress of the proposed approaches in the context of the discussed use cases, in line with the general program metrics described below in Table 1.

Strong TA2 proposals should present a review of the existing approaches, techniques, and challenges in academia and industry, supported by appropriate literature citations.

The program will emphasize creating and leveraging open-source technology and open-source architectures. Strong proposals are encouraged to offer metrics and benchmarks for evaluating the success of existing and newly developed technologies, in open reproducible settings.

Intellectual property rights asserted by proposers are strongly encouraged to be aligned with open-source regimes to support broad transition. See Section IV.B.2.i for more details on Intellectual Property.

TA3 – Voice of the offense

The TA3 performer will focus on the generalization of edge-of-the-art exploitation patterns in close coordination with TA2 performers to help model EE and exploitation. The TA3 performer will identify and describe integrated system understanding and exploration techniques used in public exploitation of complex targets to locate composable primitives for multi-step exploitation methods.

The TA3 performer is expected to have or be able to procure strong subject-matter expertise in the technological stacks relevant to the use cases discussed in “Technological Use Cases” within Section I.B. Although the selection of specific technological exemplars of use case and transition platforms will be done by the TA4 performer, the TA3 performer will advise on the selection to ensure that these exemplars can be effectively hardened to prevent exploitation.

DARPA encourages TA3 proposers to address, at a minimum, the following topics:

(1) Explain the methodology for evaluating and selecting different architectural design choices for the HARDEN use cases in collaboration with the TA4 performer;

(2) Describe how the exemplar architectures will address potential partial deployment of

HARDEN technologies across integrated systems (e.g., only some systems incorporate HARDEN functionality);

(3) Explore various options relating to the security of the HARDEN architecture itself, especially with respect to resistance to tampering by compromised devices and leakage via the TA1/TA2 mechanisms; and

(4) Integrate a system containing devices that represent a variety of platforms, operating systems, and application environments. Proposers should discuss the coordination of different TA1/TA2 mechanisms that operate across the various layers in the software stack and across different devices in the integrated system.

TA3 proposals should provide additional metrics to show increased efficiency of security analyses of targeted use case systems with TA1 tools and TA2 models over representative state-of-the-art red teaming methods. Strong TA3 proposals should also describe methods for developing tactics, techniques, and procedures capable of demonstrating specific weaknesses in the anticipated classes of TA1 and TA2 performers’ hardening and assurance technologies, and integrated system software security analyses and assurance evidence.

The TA3 performer will assist the Government team in the development of evaluations for the technological capabilities developed by the TA1, TA2, and TA4 performers, and will provide feedback to the TA1, TA2, and TA4 performers. The TA3 performer will be responsible for defining and executing a pragmatic security testing approach that enables the incremental development, demonstration, evaluation, and eventual DoD transition of HARDEN capabilities.

Strong TA3 proposals will provide open-source activity approaches that represent diverse, DoD-relevant use cases that support translating exploit development intuition and tradecraft into the formal modeling approaches of TA2.

The program will emphasize creating and leveraging open-source technology and open-source architectures. Strong proposals should offer metrics and benchmarks for evaluating the success of the newly developed technologies in comparison to existing approaches in open, reproducible settings. Intellectual property rights asserted by proposers are strongly encouraged to be aligned with open-source regimes. See Section IV.B.2.i for more details on Intellectual Property.

TA4 – Integration and systems engineering evaluation

The TA4 performer will provide system integration and evaluation, applying tools developed by TA1 performers, and models developed by TA2 performers, to demonstrate reliable and effective hardening of a sensor system based on UEFI chain-of-trust (“trusted sensor”) and of pilots’ tablets and trusted mission system integration layer (“Cockpit tablet/User Interface (UI)”).

TA4 proposals should initiate the application of concepts and techniques to critical system elements and high-assurance integrated military software systems with the goal of demonstrating the HARDEN capability of mitigating complex scenarios of exploitation, leveraging EE engines starting at early SDLC stages.

The TA4 performer will produce the testbed for demonstrating the HARDEN technological capabilities developed by TA1, TA2, and TA3 performers, and will evaluate these capabilities against the program metrics in coordination with the TA3 performer. The evaluation will include both security testing of the HARDEN technologies, and rigorous testing of the functional enhancements produced with these technologies via a series of evaluation/challenge exercises of increasing complexity and difficulty. The challenges should be suitable for fundamental research, shared without limitations with the TA1, TA2, and TA3 fundamental research teams, and should not contain Controlled Technical Information (CTI) or Controlled Unclassified Information (CUI).

TA4 proposals should also identify platforms and use cases for DoD transition, and apply tools and methodologies developed by the TA1 and TA2 performers to these use cases.

Strong TA4 proposals are encouraged to provide open-source approaches that support diverse technological use cases of software and hardware platforms relevant to the DoD.

The TA4 performer will provide the transition use cases and work with the DoD Services (Army, Navy, Air Force, and/or Marine Corps) to establish a HARDEN technology transition plan. The TA4 performer will also work with transition partner(s) for any transition accreditations or certifications required for transition of the resulting HARDEN capability to the services. The Government will require the TA4 performer to include personnel cleared, at a minimum, for SECRET level work with transition partner(s).

Evaluation Testbed Development

The series of evaluation/challenge exercises will progress in scale and complexity as detailed in Table 1, increasing the technical complexity of the challenges and the assurance guarantees that TA1 and TA2 technologies must meet.

In each evaluation/challenge exercise, TA1, TA2, and TA3 performers will receive source-level software/firmware or equivalent high-level representations of the code that will be their target to reason about EE and protect intended software abstractions from adversarial reuse. The TA1/TA2 performer evaluation code will be tested for desired functionality, robustness, and security.

Specific HARDEN end-of-program goals will be used during the Phase 3 evaluation/challenge assessment exercises. Strong proposals should present detailed plans for organizing the integrated hackathon demonstrations and evaluation/challenge exercises for the TA1, TA2, and TA3 performers using the TA4 testbed, as well as plans for allowing the performers suitable access to the testbed to prepare for these events.

Challenges should be drawn from the technologies relevant to the use cases discussed in “Technological Use Cases” within Section I.B, and work towards improved security outcomes in these integrated system use cases. The challenges should be suitable for fundamental research, should be shared without limitations with the TA1, TA2, and TA4 performers, and should not contain CTI or CUI.

Exemplary Evaluation and Transition Use Cases

The following discussion of an exemplary evaluation use case from the Air Force and Navy are provided below. Proposers are encouraged to enhance it with other relevant challenges, systems, and domains as needed to demonstrate safe and effective composition with the code base capabilities.

HARDEN will immediately address two multi-layer technologies important to the DoD:

(1) The root-of-trust and supply chain trust management, such as the UEFI architecture and the DoD-specific Sensor Open System Architecture (SOSA); and

(2) An architectural basis for a warfighter UI, such as pilots’ tablets interfacing with the plane’s mission computers.

The UEFI architecture was broadly adopted by industry to replace the legacy Basic Input/Output System (BIOS) and chipset firmware and provide a trustworthy architectural basis for root-of-trust and supply chain trust management. However, the UEFI ecosystem harbors rich classes of EE behaviors and offers a complex attack surface that permeates layers across the technology stack—from boot processes, to standalone drivers, to protected regions of memory and the main processor, to network connectivity—and cascades across the supply chain. HARDEN analyses and tools will disrupt the composability of EE behaviors at all layers of abstraction of the UEFI architecture to mitigate state-of-the-art threats and anticipate future threats.

SOSA is an Air Force Life Cycle Management Center (AFLCMC) initiative with broad industry engagement. SOSA’s objective is to create standards for a variety of next generation DoD sensor systems. SOSA’s key area of interest is the modeling and validation of the startup process of a sensor system to ensure system integrity before the sensor becomes operational. Complementary to SOSA’s efforts, HARDEN will explore relevant firmware interfaces and events involved in a trusted computing system’s startup, will help formulate standards to ensure system trustworthiness, and will help create tools to validate compliance with these standards in SOSA systems.

The warfighter UI use case will explore secure integration of modern UI elements derived from Commercial Off-The-Shelf (COTS) technologies, such as pilots’ tablets, with aircraft’s mission systems and networks. Communications between pilots’ tablets and the aircraft mission computers present a large attack surface inside and outside the aircraft. The HARDEN program will create state-of-the-art integrated systems analysis capabilities responsive to the assurance goals of trustworthy pilots’ tablets.

If successful, HARDEN methodologies and tools will serve other types of DoD integrated systems, anticipating and pre-empting the root cause of exploitability in their design and implementation.

C. Program Phases and Metrics The HARDEN program is a 48-month program organized into three phases: Phase 1 is an 18-month open-source component-scale phase, Phase 2 is an optional 18-month open-source subsystem-scale phase, and Phase 3 is an optional 12-month phase focused on scaling the technology to a DoD-relevant integrated system. The HARDEN technical and management milestones are depicted in Figure 2 below. As shown, program evaluation exercises are planned for months 9, 16, 23, 29, 35, 41, and 47. The month 9 evaluation will occur at the TA4 performer’s facility and will be a TA1/TA2 integration hackathon demonstration. The integrator’s testbed is to be established by month 9. The second program evaluation exercise at month 16 will be used to determine whether or not the performers should continue into Phase 2.

Month 23 and 41 will also be TA1/TA2 integration hackathon demonstrations with the necessary testbed functionality improvements required for that phase to demonstrate program metrics in a realistic environment.

The capability milestones and metrics for the HARDEN program, shown in Table 1 below, are related to the capability to reason about EE and protect intended software abstractions from adversarial reuse. The specific metrics provided in the remainder of this section are indicative of the expected progress. Proposers should describe specific approaches that they will use for testing and evaluation purposes during each of the program phases, and propose additional quantitative metrics tailored to measuring progression of these approaches. Note that the platforms named in Table 1 for “Exemplary software complexity” are for gauging the approximate size and complexity of system under test, and not necessarily the actual use case exemplars.

Table 1: HARDEN Metrics The Government will assess individual performer efforts in terms of the viability of their technical approaches, the trend in the performance of their systems over time, and their overall progress toward HARDEN program objectives.

Schedule and Milestones

For each year of effort, there will be quarterly meetings with the Program Manager (PM), consisting of two site visits and two Principal Investigator (PI) meetings. During these meetings/reviews, the PM will assess progress towards the solution via performer briefings, technical discussions, demonstrations, and informal end-of-phase evaluation/challenge exercises based on the target goals of each phase.

PI meetings will focus on open technical exchange. Difficulties encountered and possible solutions will also be discussed. The goals of the PI meetings will be to: (1) review and share innovations/accomplishments of the HARDEN program; (2) review and discuss plans and options for technology demonstrations and prototypes and HARDEN evaluation/challenge exercises; (3) review and discuss results from meetings and events conducted prior to and after the tests and evaluation/challenge exercises; (4) demonstrate prototypes; and (5) plan for the next six-month period.

The Government will specify the locations for the technical interchanges and PI meetings.

Evaluation/challenge exercises will be held at the TA4 performer’s site. For budgeting purposes, assume the locations of the two PI meetings held each year will alternate between Washington, D.C., and San Diego, CA. For budgeting travel to the TA4 performer’s site, assume the location will be on the opposite coast from your location, or if regionally located in the Midwest, choose the more expensive coastal travel destination between San Diego, CA, or Washington, D.C. In addition to site visits, regular teleconference meetings are encouraged to enhance communications and collaborations, as required, among the performers. Should important issues arise between program reviews, the Government team will be available to support informal meetings. In-person meetings, evaluations, and site visits may be replaced with virtual ones if necessary.

Figure 2 below provides a tentative program schedule. Proposers should propose a detailed schedule that is consistent with the maturity of their approaches and the risk reduction required for their concepts and their program plan. These schedules will be synchronized across performers, as required, and monitored and revised as necessary throughout the HARDEN program’s period of performance. A start date of July 1, 2022, should be assumed for budgeting purposes.

Figure 2: HARDEN Tentative Program Schedule Deliverables

Performers are responsible for providing the following deliverables, as applicable:

• Slide Presentations – Annotated slide presentations will be submitted within two weeks after program kick-off meeting and after each review.

• Quarterly Coordination Reports – A quarterly technical…

This is the start of the file's text. The full file is on GovTribe.

File details come from the government source that posted it. Updated .