HR001118S0040.pdf
PDF 1 MB Posted
- Attached to
- Computers and Humans Exploring Software Security (CHESS) Federal contract opportunity
- Solicitation number
- HR001118S0040
About this file
Not Listed
View the file
Other files for this federal contract opportunity
| File | Type | Posted |
|---|---|---|
| CHESS_BAA_Attachment_Proposal_Summary_Chart_Template.pptx | PPTX presentation | |
| CHESS_BAA_proposal_LoE_table_template_-_SkillSets.xlsx | XLSX spreadsheet |
On GovTribe
Work with this file on GovTribe
- Download the original file
- Contacts named in this file
- Similar government files
- Ask GovTribe AI about this file
Text version
Broad Agency Announcement Computers and Humans Exploring Software Security (CHESS)
HR001118S0040
April 18, 2018
Defense Advanced Research Projects Agency Information Innovation Office 675 North Randolph Street Arlington, VA 22203-2114
HR001118S0040 COMPUTERS AND HUMANS EXPLORING SOFTWARE SECURITY (CHESS) 2
Table of Contents
I. Funding Opportunity Description A. Introduction B. Program Description C. Program Structure D. Technical Areas E. Evaluation F. Schedule and Milestones G. Deliverables to DARPA H. Intellectual Property I. Glossary
II. Award Information A. Awards B. Fundamental Research C. Disclosure of Information and Compliance with Safeguarding Covered Defense Information
Controls III. Eligibility Information
A. Eligible Applicants B. Organizational Conflicts of Interest C. Cost Sharing/Matching D. Other Eligibility Requirements
IV. Application and Submission Information A. Address to Request Application Package B. Content and Form of Application Submission C. Submission Dates and Times D. Funding Restrictions E. Other Submission Requirements
V. Application Review Information A. Evaluation Criteria B. Review and Selection Process
VI. Award Administration Information A. Selection Notices B. Administrative and National Policy Requirements C. Reporting
VII. Agency Contacts VIII. Other Information
A. Frequently Asked Questions (FAQs) B. Collaborative Efforts/Teaming C. Proposers Day D. Submission Checklist E. Associate Contractor Agreement (ACA)
HR001118S0040 COMPUTERS AND HUMANS EXPLORING SOFTWARE SECURITY (CHESS) 3
PART I: OVERVIEW INFORMATION
Federal Agency Name: Defense Advanced Research Projects Agency (DARPA), Information Innovation Office (I2O)
Funding Opportunity Title: Computers and Humans Exploring Software Security
(CHESS)
Announcement Type: Initial Announcement
Funding Opportunity Number: HR001118S0040
Catalog of Federal Domestic Assistance Numbers (CFDA): 12.910 Research and Technology Development
Dates o Proposers Day: April 19, 2018 o Posting Date: April 18, 2018 o Abstract Due Date: May 3, 2018, 12:00 noon (ET) o Proposal Due Date: June 15, 2018, 12:00 noon (ET) o BAA Closing Date: June 15, 2018, 12:00 noon (ET)
Anticipated Individual Awards: DARPA anticipates multiple awards for technical areas 1 and 2; and single awards for technical areas 3, 4 and 5.
Types of Instruments that May be Awarded: Procurement contracts or cooperative agreements
Agency Contacts o Technical POC: Mr. Dustin Fraze, Program Manager, DARPA/I2O o BAA Email: CHESS@darpa.mil o BAA Mailing Address:
DARPA/I2O
ATTN: HR001118S0040
675 North Randolph Street Arlington, VA 22203-2114 o I2O Solicitation Website: http://www.darpa.mil/work-with-us/opportunities http://www.darpa.mil/work-with-us/opportunities
HR001118S0040 COMPUTERS AND HUMANS EXPLORING SOFTWARE SECURITY (CHESS) 4
PART II: FULL TEXT OF ANNOUNCEMENT
I. Funding Opportunity Description
DARPA is soliciting innovative research proposals to develop techniques and systems that will substantially accelerate software vulnerability research (VR). The goal of the CHESS program is to develop computer-human systems to rapidly discover all classes of vulnerability in complex software. These novel approaches for the rapid detection of vulnerabilities will focus on identification of system information gaps that require human assistance, generation of representations of these gaps appropriate for human collaborators, capture and integration of human insights into the analysis process, and the synthesis of software patches based on this collaborative analysis.
Proposed research should investigate innovative approaches that enable revolutionary advances in science, devices, or systems. Specifically excluded is research that primarily results in evolutionary improvements to the existing state of practice.
This Broad Agency Announcement (BAA) is being issued, and any resultant selection will be made, using procedures under Federal Acquisition Regulation (FAR) 6.102(d)(2) and 35.016.
Any negotiations and/or awards will use procedures under FAR 15.4 (or 32 CFR § 200.203 for cooperative agreements). Proposals received as a result of this BAA shall be evaluated in accordance with evaluation criteria specified herein through a scientific review process.
DARPA BAAs are posted on the Federal Business Opportunities (FBO) website (https://www.fbo.gov/) and the Grants.gov website (http://www.grants.gov/).
The following information is for those wishing to respond to this BAA.
A. Introduction
The Department of Defense (DoD) maintains information systems that depend on Commercial off-the-shelf (COTS) software, Government off-the-shelf (GOTS) software, and Free and open-source (FOSS) software. Securing this diverse technology base requires highly skilled hackers who reason about the functionality of software and identify novel vulnerabilities. This process requires hundreds or thousands of hours of manual effort per discovered vulnerability and does not scale sufficiently to secure the continuously growing technology base.
Hackers use program analysis techniques and tools to identify and mitigate vulnerabilities, but this process requires considerable expertise, manual effort, and time. These techniques include dynamic analysis, static analysis, symbolic execution, constraint solving, data flow tracking, and fuzz testing. Automated program analysis capabilities can reason over only a few vulnerability classes without human involvement, such as memory corruption or integer overflow, but cannot address the majority of vulnerabilities. These unaddressed vulnerability types depend on subtle semantic and contextual information, which is beyond the grasp of modern automation. Scaling up existing approaches to address the size and complexity of modern software packages is not possible given the limited number of expert hackers in the world, much less the DoD.
https://www.fbo.gov/ http://www.grants.gov/
HR001118S0040 COMPUTERS AND HUMANS EXPLORING SOFTWARE SECURITY (CHESS) 5
The CHESS program will develop capabilities to discover and address vulnerabilities of all types in a scalable, timely, and consistent manner. DARPA believes that achieving the necessary scale and timelines in vulnerability discovery will require innovative combinations of automated program analysis techniques with support for advanced computer-human collaboration (CHC).
Due to the cost/scarcity of expert hackers, such capabilities must be able to collaborate with humans of varying skill levels, even those with no previous hacking experience or relevant domain knowledge.
B. Program Description
The CHESS program will research the effectiveness of enabling computers and humans to collaboratively reason over software artifacts (source code, compiled binaries, etc.) with the goal of finding 0-day vulnerabilities at a scale and speed appropriate for the complex software ecosystem upon which the U.S. Government, military, and economy depend.
Achieving these goals will require research breakthroughs in:
o Developing instrumentation to capture and analyze the process by which hackers reason over software artifacts to provide a basis for developing new forms of highly effective communication and information sharing between computers and humans;
o Creating techniques for addressing classes of vulnerability that are currently hampered by information gaps and require human insight and/or contextually sensitive reasoning;
o Generating representations of the information gaps for human collaborators of varying skill levels to reason over;
o Integrating human-generated insights into the vulnerability discovery process;
o Emitting a Proof of Vulnerability (PoV) to confirm existence of the 0-day vulnerability, and generating a non-disruptive, specific patch to neutralize the 0-day vulnerability; and o Synthesizing vulnerable Challenge Set (CS) corpora representative of large, real world, complex software packages.
HR001118S0040 COMPUTERS AND HUMANS EXPLORING SOFTWARE SECURITY (CHESS) 6
The following figure illustrates a high-level overview of the CHESS system:
Figure 1: CHESS System Overview
The CHESS program will involve human subjects research (HSR), and defines three (3) classes of human subjects as follows:
Human Subject Class Description
Expert Hacker A professional with 5+ years’ experience in software reverse engineering, program analysis or exploit development, should have proven experience finding vulnerabilities in operating systems and large software packages.
Novice Hacker
A professional with less than 2 years’ experience in software development, reverse engineering, program analysis and exploit development, should have 4 or fewer years formal education/background in computer science, computer engineering, software development or related disciplines.
Non-Hacker Any adult with basic computing skills and no formal background in computer science, computer engineering, software development or any related disciplines.
Table 1: Human Subject Classes
HR001118S0040 COMPUTERS AND HUMANS EXPLORING SOFTWARE SECURITY (CHESS) 7
All CHESS proposals should be applicable to at least one of the following source code languages or compiled binary targets:
Source Code Languages C/C++ Python
Javascript
Table 2: Source Code Languages
Compiled Binary Targets Platform Architecture
Linux x86-64 Windows x86-64
Table 3: Compiled Binary Targets
Specific platform versions should be representative of the systems currently deployed and of interest to the DoD. Proposals should consider approaches that apply to multiple languages/platforms or are largely language/platform agnostic.
The CHESS program will target the vulnerability classes shown in Table 4, which also provides the relevant parent Common Weakness Enumerations (CWEs). Each class subsumes the listed CWEs, including all child CWEs, per MITRE’s CWE List Version 3.0.1 Proposals need not address all the vulnerability classes or CWEs of interest. While reasoning over vulnerabilities in all systems of interest (Table 2, Table 3) is the primary CHESS goal, it is understood that certain CWEs are relevant only to a subset of these systems. CHESS will achieve maximal effective coverage of the vulnerability classes through collaboration between all performers.
Vulnerability Class Parent CWEs Data/Code Injection 74
Data Misuse 471, 501, 610, 628, 642, 662, 665, 673, 704, 706 Logic Errors 691, 697, 703, 758, 768
Authentication Issues 287 Input Validation 20, 138, 170, 172, 228, 463
Access Control Errors 269, 285, 282, 286, 923 Cryptographic Issues 324, 325, 326, 330, 347
Path Traversal 22, 41, 59 Resource Management
Errors 400, 404, 405, 665, 666
Information Disclosure 668 Memory Corruption 118
Arithmetic Errors 682 Table 4: Target Vulnerability Classes
1 https://cwe.mitre.org/data/index.html https://cwe.mitre.org/data/index.html
HR001118S0040 COMPUTERS AND HUMANS EXPLORING SOFTWARE SECURITY (CHESS) 8
Proposals may also address additional vulnerabilities that are not on the list, in which case there should be a justification describing the importance of the additional vulnerabilities and how the proposed techniques substantially improve the state of the art.
C. Program Structure
The CHESS program is divided into five (5) technical areas (TAs) that will be working in parallel, starting at program kickoff. The program will span 42 months divided into one 18-month phase and two 12-month phases. Each phase will focus on increasing the application complexity for which the CHESS system is able to effectively analyze. Based on consultation with performers, DARPA will decide which vulnerability classes (Table 4) are in scope during each phase of the CHESS program. (Figure 3, Figure 4)
Phase 1 (18 months) will emphasize initial development of the tools and techniques needed to capture human insight and communicate context to computers covering four (4) vulnerability classes. The scale of the target software in Phase 1 will be on the order of small software libraries (low complexity).
Phase 2 (12 months) will emphasize refining the insight capture and communication mechanisms while expanding coverage to eight (8) vulnerability classes. The scale of the target software in Phase 2 will be on the order of whole software packages (moderate complexity).
Phase 3 (12 months) will emphasize scaling techniques while expanding coverage to twelve (12) vulnerability classes. The scale of the target software in Phase 3 will be on the order of modern web browsers (high complexity).
In Phase 1, there will be one kickoff meeting, four hackathons, two demonstrations, and a final evaluation event. Phases 2 and 3 will each have two hackathons and one late-phase demonstration to identify and correct any weaknesses, which should provide ample time to address any shortcomings before each end-of-phase evaluation. Hackathons and evaluations will also serve as PI meetings. (Figure 3)
D. Technical Areas
CHESS will be structured with five (5) technical areas as shown in Figure 2:
TA1 - Human Collaboration TA2 - Vulnerability Discovery TA3 - Voice of the Offense TA4 - Control Team TA5 - Integration, Test, and Evaluation
HR001118S0040 COMPUTERS AND HUMANS EXPLORING SOFTWARE SECURITY (CHESS) 9
TA1 performers will focus on capturing and decomposing hacker workflow and all other human-computer interaction (HCI) aspects of the program. TA2 performers will develop technologies for discovery and patching of specified vulnerability classes in both source code and compiled binaries. These vulnerabilities will be synthetic, but representative of vulnerabilities facing the DoD and the U.S. Government. The TA3 performer will create challenges for the evaluations.
The TA4 performer will provide a baseline for measuring CHESS improvement over the current edge of the art by tackling the TA3 challenges with existing tools and techniques. The TA5 performer will manage evaluations, integration, and transition to government, and commercial partners.
It is anticipated that TA1 and TA5 may both involve human subjects research (HSR). TA1 may need to study the workflow of expert and novice hackers, as well as identify tasks appropriate for non-hackers. TA5 may need to recruit human subjects in each of the three (3) categories of expert, novice, and non-hackers for evaluations (Table 1). Any HSR efforts will require approval by an institutional review board (IRB). TA5 may leverage IRB-approved protocols, data anonymization strategies and other artifacts from TA1, but TA5 will still need separate IRB approval if TA5 recruits human test subjects for evaluation.
TA1 proposals involving HSR should include at least one draft HSR protocol, a plan for IRB approval, and an anonymization strategy for HSR-related data. TA5 proposals involving HSR should include a plan for transitioning and adapting IRB-approved protocols and data anonymization strategies from TA1 for evaluations, as well as a draft plan for IRB approval of TA5’s test subject recruitment strategy. TA1 and TA5 should consolidate all HSR work into separate sections of the Statement of Work (SOW) and cost proposal as options. Any HSR involving invasive medical procedures or implantation is out of scope.
Figure 2: CHESS Technical Areas
HR001118S0040 COMPUTERS AND HUMANS EXPLORING SOFTWARE SECURITY (CHESS) 10
The Government anticipates one (1) or more awards for TA1 and TA2, and single awards for TA3, TA4, and TA5. Each abstract and proposal submitted against this solicitation shall address only one (1) TA. Organizations may submit multiple abstract/proposals to any one TA, or they may propose to multiple TAs. A proposer submitting a proposal to TA1 and another to TA2 may be selected to perform on both TAs. However, TA3, TA4, and TA5 performers cannot perform on any other TA.
Proposers for TA1, TA2, TA3 or TA4 are not required to hold or obtain security clearances;
however, having cleared personnel will be viewed positively. It is preferable (but not required) that the Principal Investigator in each TA1, TA2, TA3, and TA4 proposal be cleared at the final Top Secret level. Academic and small company participation is explicitly encouraged, regardless of possession of a security clearance.
At the time of proposal submission, all proposers submitting proposals under TA5 must have some personnel with a final Top Secret clearance that are eligible for Sensitive Compartmented Information (SCI). It is preferable (but not required) that the Principal Investigator in TA5 proposals be cleared at the SCI level.
TA1 and TA2 performers may elect to build and integrate their tools with the rest of the CHESS system such that the performers themselves are not exposed to DoD transition systems. In this case, the research performed by a TA1 or TA2 performer could be considered fundamental research. See the CHESS Program Controlled Unclassified Information (CUI) guide for further information.2 Any proposal for work that requires access to specific information regarding a DoD system will not be considered fundamental research and will require security clearances and secure facilities commensurate with any relevant security classification guides. At a minimum, such research will be considered Controlled Technical Information (CTI) and subject to mandatory pre-publication review. (See also Section II.B).
There are several points of potential collaboration among TAs, and the Government expects that all performers producing software will interact closely with the Integrator/Evaluator (TA5). TA1 and TA2 will interact closely with each other, and with TA5, as TA1 and TA2 develop techniques for information gap identification, insight extraction, and communication to computers. All proposers should read the descriptions of all TAs, as well as the evaluation section (Section I.E), to ensure there is a full understanding of the program context, structure, and anticipated relationships required among performers. To facilitate the open exchange of information, performers will have Associate Contractor Agreement (ACA) language included in their award contract or agreement. TA5 will lead the development of the ACA for the program.
See Table 6 for further detail on the collaboration between TAs, and Section VIII.E for more information regarding an ACA.
2 http://www.darpa.mil/work-with-us/opportunities http://www.darpa.mil/work-with-us/opportunities
HR001118S0040 COMPUTERS AND HUMANS EXPLORING SOFTWARE SECURITY (CHESS) 11
Human Collaboration (TA1)
TA1 performers will research how expert hackers discover software vulnerabilities, and develop technologies to enable humans and machines to collaboratively reason over software artifacts (source code, compiled binaries, abstract syntax trees, and any other intermediate format involved in the build process) for the purpose of vulnerability discovery.
Humans have world knowledge and semantic/contextual understanding that is beyond the reach of automated program analysis alone. These information gaps inhibit machine understanding of many classes of software vulnerability. Properly communicated, human insights can fill these information gaps and enable expert hacker-level vulnerability analysis at machine speeds.
TA1 performers will collaborate with TA2 to identify information gaps impeding the vulnerability discovery process. Once identified by TA2, TA1 will generate representations of these gaps with the goal of highlighting them for human collaborators. As the humans reason over these representations, a context processor will capture human insights that may improve the vulnerability discovery process. These insights will be collected and aggregated by a context processor, and then delivered back to TA2.
TA1 performers will also be responsible for decomposing and identifying sub-tasks that are appropriate for expert hackers, novice hackers, and non-hacker human collaborators.
TA1 may involve HSR, and performers will ensure all human subject interactions are documented in an IRB-approved protocol. Proposals that will rely on HSR should segregate those tasks in the TA1 SOW so that the Government can allow them to start on non-HSR tasks prior to receiving IRB approval. This segregation should be reflected in the cost proposal with a cost option that includes all HSR-related tasks. All data derived from human subjects must be anonymized for sharing with other TAs and/or use in future research.
The technologies developed by TA1 performers may use a combination of active and passive techniques for capturing and interacting with human collaborators. These may include, but are not limited to, text-based, graphical user interface (GUI), and non-invasive biofeedback, etc. Any HSR involving invasive medical procedures or implantation is out of scope.
TA1 performers must produce software that is sufficiently mature to preclude lengthy bug finding on the part of TA5 during integration. Proposals should describe approaches for doing so.
TA1 proposals should address the following topics:
1. The vulnerability discovery process is long, complex, and places a heavy cognitive load on expert hackers. Proposers should discuss how TA1 solutions will convert insights from human observations into measurable, succinct, consistent characteristics of successful vulnerability discovery. Proposals should address how these characteristics will support reducing the effort required by expert hackers throughout the vulnerability discovery process.
HR001118S0040 COMPUTERS AND HUMANS EXPLORING SOFTWARE SECURITY (CHESS) 12
2. The methods of representing and communicating program analysis data to expert hackers have not changed significantly in 20+ years. Proposers should discuss the strengths and weaknesses of existing representations and provide improvements to these existing representations and/or new representations to substantially accelerate the vulnerability discovery process. Proposals should consider the applicability of state of the art human-computer interaction (HCI) technologies from other domains.
3. The community of expert hackers is small and tight-knit. TA1 proposers should discuss how they will gain sufficient access to enough of these individuals to ensure robust, accurate results.
4. Expert hacker time is expensive and scarce. Proposers should discuss how TA1 solutions will capture/make use of feedback from both expert and novice hackers as well as non-hackers. Proposals should describe approaches that enable novices and non-hackers to take on challenges that currently require expert hackers.
5. Proposers may need to study the workflow of expert and novice hackers as well as identify tasks appropriate for non-hackers. Proposals should discuss how they will address any relevant HSR issues. Proposals relying on HSR should describe the protocols to be followed, how they will obtain IRB approval and an anonymization strategy for any data collected on human subjects.
Vulnerability Discovery (TA2)
TA2 performers will develop techniques and systems to discover vulnerabilities matching the vulnerability classes of interest to CHESS (Table 4). The world knowledge and semantic/contextual information that humans can provide create higher-order models of software behavior that computers cannot currently discern on their own. Once captured, these insights may enable modern program analysis techniques to reason about vulnerabilities at scale and speed appropriate for modern software. Approaches primarily based on fuzzing are out of scope.
A key part of this process is identifying missing, but relevant, information to vulnerability analysis. TA2 will develop techniques to detect information gaps impeding the vulnerability analysis and leverage human-generated insights from TA1 to fill those gaps.
TA2 will collaborate with TA1 to define a common format for communicating identified information gaps to TA2. Both TAs must jointly produce a common data format to exchange information gap queries and insights from humans.
TA2 performers will generate a PoV and patch for each discovered vulnerability. A PoV is an executable that activates and proves the existence of a hidden vulnerability. The vulnerabilities in scope for the CHESS program are only those that enable a remote adversary with minimal or no privileges to compromise a target system. PoVs may span multiple CWEs of interest (Table 4).
Patches should be as specific as possible to the identified vulnerability or vulnerabilities while not interfering with the normal CS functionality. All approaches must be able to meet the program objectives for vulnerability class coverage and speed.
HR001118S0040 COMPUTERS AND HUMANS EXPLORING SOFTWARE SECURITY (CHESS) 13
TA2 will be provided documentation by TA3 for each vulnerability class and CWE early in each phase, along with an example CS corpus prior to the midpoint of each phase to allow TA2 sufficient time to prepare for evaluation and demonstration.
TA2 performers will produce software that is sufficiently mature to preclude lengthy bug finding on the part of TA5 during integration. Strong proposals will describe approaches for doing so.
TA2 performers may work on source, binary, or both. TA2 is therefore divided into two tasks, the first for working with source, the second with binaries. TA2 proposers should declare one of these tasks as part of the TA2 base proposal. If the proposed approach will address both, the task not declared as part of the base proposal must be proposed as a discrete option.
Approaches based on binary or source code differencing tools that discover TA3-inserted vulnerabilities are out of scope. DARPA will consider a proposal using such an approach as non-responsive to the BAA and will not be evaluated.
The Government prefers focused proposals on a single TA2 task rather than shallow proposals covering both TA2 tasks. Approaches that address both tasks should be structured as distinct and separable tasks in the SOW.
TA2 Task 1: Source-Assisted Vulnerability Discovery
TA2 Task 1 performers will be allowed access to the source code and all related build artifacts of each CS for analysis.
TA2 Task 1 proposals should address the following:
1. There is currently a wide variety of source code analysis tools available for general use.
Proposers should discuss the applicability or lack thereof of these tools and the underlying techniques. Proposals should identify and discuss semantic and/or contextual information gaps in proposed techniques that human collaborators can assist with and what level of expertise (expert hacker, novice hacker or non-hacker) is required to address each gap.
2. While TA2 Task 1 performers will have access to source code for challenges, proposers should discuss how they could make use of the software artifacts involved in the build process, up to and including the compiled binary. While use of Requests for Comments (RFC) and other public documentation are also in scope for this task, proposals should use the source and build artifacts as the primary source of knowledge, and leverage human-provided insights to address identified information gaps.
3. Many generic software protections exist that make exploitation of vulnerabilities more challenging, but do not fix the underlying vulnerabilities. Generated patches should address detected vulnerabilities completely and specifically without interfering with normal program behavior. Proposals should also minimize modifications to software
HR001118S0040 COMPUTERS AND HUMANS EXPLORING SOFTWARE SECURITY (CHESS) 14
behavior and testbed resource usage. Approaches that make exploitation more challenging but do not fix the underlying vulnerability are out of scope.
4. TA2 Task 1 proposers should discuss which source code language(s) of interest to CHESS their techniques will address. (Table 2) Proposals should consider techniques that are source code language agnostic.
TA2 Task 2: Binary Vulnerability Discovery
TA2 Task 2 performers will not be allowed access to source code for the CS corpus. Each CS will include a compiled binary.
TA2 Task 2 proposals should address the following topics:
1. There are a wide variety of binary analysis tools available for general use. Proposers should discuss the applicability or lack thereof of these tools and their underlying techniques. Proposals should identify semantic/contextual information gaps that proposed techniques will address, and discuss how non-hackers, novices, or experts could assist with any gaps beyond the proposed techniques.
2. TA2 Task 2 performers will not have access to source code for challenges but may be allowed access to debug symbols for some challenges. Use of RFCs and other public documentation are also in scope for this task. Proposers should discuss the potential value of these artifacts. Proposals should use the binary as the primary source of knowledge and leverage human-provided insights to address identified information gaps.
3. Many generic software protections exist that make exploitation of vulnerabilities more challenging but do not fix the underlying vulnerabilities. Generated patches should address detected vulnerabilities completely and specifically without interfering with normal program behavior. Proposals should also minimize modifications to software behavior and testbed resource usage. Approaches that make exploitation more challenging but do not fix the underlying vulnerability are out of scope.
4. TA2 Task 2 proposers should discuss which operating systems and platforms of interest to CHESS (Table 3) their techniques will address. Proposals should consider techniques that are operating system or platform agnostic.
Voice of the Offense (TA3)
The TA3 performer will research and produce Challenge Set (CS) corpora representative of the vulnerability classes of interest to CHESS (Table 4) to challenge the effectiveness of the vulnerability discovery and mitigation techniques developed by TA1 and TA2. TA3 will also produce PoV specifications describing the features and scope of each vulnerability class and CWE of interest to the CHESS program.
Each CS is comprised of the following components:
HR001118S0040 COMPUTERS AND HUMANS EXPLORING SOFTWARE SECURITY (CHESS) 15
Challenge Executable (CE) A CE can be either a network service that accepts remote network connections, or client meant to connect to arbitrary servers and perform processing on network-supplied data, and interact with remote hosts over network connections. CEs will be used as analysis challenges for the vulnerability discovery and mitigation techniques developed by TA1 and TA2. Each CE will be implemented performing tasks emulating real world software;
examples include but are not limited to file transfer, remote procedure call, remote login, peer-to-peer (P2P) networking. Each CE will contain at least one vulnerability hidden in the program and reachable via network input. Vulnerabilities should cover the classes and CWEs listed in Table 4 and the compilation targets listed in Table 3. Non-compiled source code languages of interest (Table 2) should have an equivalent portable executable, or containerized configuration, such that running the CE is straight forward.
(e.g., Javascript shell, frozen Python binary, Docker container, etc.)
Challenge Source Code (CSrc) The source code for the CEs, which will also be used as analysis challenges for the source-assisted vulnerability discovery and mitigation techniques developed by TA1 and TA2. Intentionally inserted vulnerabilities should be well commented and documented for TA5 review. Each CSrc should include an accompanying script for straight forward automated removal of these comments before delivery for an evaluation.
Reference Patched Binary (RPB) The patched CE delivered in each CS will function identically to the unpatched CE, but will not contain known, hidden vulnerabilities. This Reference Patched Binary should take realistic actions when encountering malicious/irregular inputs like those in PoVs or the RPoV. A realistic action might be continuing to process an input or releasing program resources and exiting. Any actions that require prior knowledge of the known, hidden vulnerabilities are out of scope.
Reference Proof of Vulnerability (RPoV) A Reference Proof of Vulnerability (RPoV) is an executable that activates and proves the existence of a hidden vulnerability in each CE. For CEs with multiple, hidden vulnerabilities, a separate RPoV must be delivered for each vulnerability.
Service Poller (SP) A Service Poller (SP) will implement a functionality test suite to detect whether the baseline function or performance of the corresponding CE has been impaired. A successfully patched CE will cause the corresponding SP to report no errors. Each SP should generate repeatable queries when provided with a random seed value and queries generated by different seed values should be highly diverse. Over a large number of randomly seeded network tests, SPs should be capable of exercising a majority of expressed CE code.
Each CS should be designed such that it does not function outside the CHESS testbed environment developed by TA5. Although TA3 proposers may propose to survey existing vulnerabilities found in the wild in order to inform TA3 efforts, searching for exploitable
HR001118S0040 COMPUTERS AND HUMANS EXPLORING SOFTWARE SECURITY (CHESS) 16
vulnerabilities in deployed systems is out of scope. Independent verification and validation (IV&V), red teaming, and penetration testing approaches are also out of scope.
Many of the vulnerabilities of interest to the CHESS program lack standardized, computer-generated oracles (e.g., signals, exceptions, interrupts). This is one of the challenges TA2 performers must overcome in TA2 development of novel techniques to detect these vulnerabilities, and as such, is not in scope for TA3. However, to ensure alignment between TA3 CSs and TA2 vulnerability detection techniques, TA3 will provide TA2 with a PoV specification for each of these CWEs. The PoV specification will inform TA2 on the features, constraints, and scoping of each CWE.
The TA3 performer will also provide TA2 performers with an example CS corpus covering vulnerability classes/CWEs in scope for each phase of the CHESS program.
TA3 proposals should address the following topics:
1. The target vulnerability classes (Table 4) cover a wide range of CWEs, some of which are less common than others. TA3 proposers should discuss how they will ensure effective coverage of all vulnerability classes of interest and provide appropriate justification if certain classes receive more attention than others.
2. The role of TA3 in measuring and pushing the progress of the CHESS system requires that the TA3 evaluation CS corpus is representative of vulnerabilities in realistically large, complex codebases. TA3 proposers should discuss how they will address this scaling issue across all phases of the CHESS program.
3. The TA3 performer may construct each CS from scratch, or may modify existing software and firmware as part of CS development. However, when a new CS contains code reused from existing open source programs or previous challenge programs, TA1 and TA2 teams may be able to find vulnerabilities simply by looking for differences between the new code and the old, instead of employing novel techniques and tools.
Proposals should describe a method for avoiding or mitigating this problem.
4. While introducing vulnerabilities into existing software is in scope, each TA3 CS must be constructed or modified so that they only run in the testbed environment developed by TA5. TA3 proposals should discuss approaches for limiting execution to the TA5 testbed.
5. To ensure sufficient realism, the evaluation CS corpus should contain some vulnerabilities that require leveraging the effects of multiple CWEs in concert. TA3 proposers should discuss how they will select the CWEs for these combined vulnerabilities, and develop the necessary CSs and PoV specifications.
6. TA3 proposers should discuss how they will ensure effective coverage of CS functionality and how the SP will identify any interference caused by TA2 patches.
7. TA3 proposers must provide an example CS corpus and documentation at least 6 months
HR001118S0040 COMPUTERS AND HUMANS EXPLORING SOFTWARE SECURITY (CHESS) 17
prior to each evaluation. (Figure 4) This corpus must to provide other TAs with working CS examples covering vulnerability classes in scope for each phase. These should illustrate baseline CE performance, how the CE behaves when successfully exploited via the RPoV and how the RPB behaves when presented with the RPoV or SP.
Control Team (TA4)
The TA4 performer will create an expert hacker performance baseline against TA3 evaluation CS corpus. The TA4 performer will not be allowed to use any of the research technologies that TA1 and TA2 produce. The TA4 performer will instead research and leverage the current edge of the art outside of the CHESS program, and decide which of these tools to use against the evaluation CS corpus and produce PoVs for each set. The results of the TA4 evaluations will serve as a baseline for measuring improvements in performance attributable to TA1 and TA2.
TA4 proposals should address the following topics:
1. Many tools and techniques exist for analyzing software for vulnerabilities. TA4 proposers should discuss how they will ensure maximal coverage of emerging tools and techniques.
Strong proposals will produce detailed and methodical analyses of the strengths and weaknesses of each technique and tool considered against each specific challenge. An initial Edge of the Art report on the performer’s research will be due to DARPA five (5) months after kickoff. Thereafter, the TA4 performer will be expected to provide an updated Edge of the Art report delivered 6 months prior to each evaluation and at the beginning of Phase 2 and Phase 3. (Figure 4)
2. The TA4 performer will perform against the CS corpus with and without source code.
TA4 proposers should discuss how they will approach analysis of the CS corpus in both cases.
3. The CWE categories in scope for this program encompass a wide array of vulnerabilities, some without real world examples. TA4 proposers should demonstrate a deep and broad understanding of the edge of the art in vulnerability research, program analysis and reverse engineering.
4. The TA4 performer will provide detailed feedback on their experience reasoning over the CS corpus after each evaluation. This report should be sufficiently detailed so as to capture all attempted analysis tasks, both successful and unsuccessful. This report will be shared with other performers after each evaluation to help characterize how expert hackers reason over and find software vulnerabilities.
HR001118S0040 COMPUTERS AND HUMANS EXPLORING SOFTWARE SECURITY (CHESS) 18
Integration, Test, and Evaluation (TA5)
The TA5 performer will handle integration of technology produced by TA1 and TA2, evaluation of the integrated CHESS system against the TA3 evaluation CS corpus, evaluation of TA4’s performance against the TA3 evaluation CS corpus and transition efforts to Government and commercial partners. TA5 will also be responsible for leading the development of the required Associate Contractor Agreement (ACA), in close collaboration with all other performers.
Evaluation:
The TA5 performer will be responsible for evaluating the performance of TA1 and TA2 systems using the CHESS program metrics described in Section I.E (Table 5) and integrating TA1 and TA2 systems into source-assisted and binary subsystems. TA5 proposals should discuss additional, objective metrics for determining the vulnerability discovery improvements of the overall CHESS system. These metrics should account for and scale representative software characteristics, including but not limited to the number of source lines of code, system calls, multiple threads, and multiple processes.
Human-Subjects Research (HSR):
TA5 may involve HSR during evaluations, as TA1 techniques may require human collaborators for evaluation. TA5 may need to recruit human subjects in each of the three (3) categories of expert, novice, and non-hackers (Table 1). TA5 may leverage IRB-approved protocols, data anonymization strategies and other artifacts from TA1, but will still need separate IRB approval for TA5 recruitment of human test subjects for evaluation. TA5 proposals involving HSR should include a plan for transitioning and adapting IRB-approved protocols and data anonymization strategies from TA1 for evaluations, as well as a draft plan for IRB approval of TA5 test subject recruitment strategy.
The technologies evaluated by TA5 performers may use a combination of active and passive techniques for capturing and interacting with human collaborators. These may include but are not limited to text-based, graphical user interface (GUI), non-invasive biofeedback, etc. HSR involving invasive medical procedures or implantation is out of scope and will not be part of the evaluations.
TA5 proposals should segregate HSR-related tasks in the TA5 SOW so that the Government can allow them to start on non-HSR tasks prior to receiving IRB approval. This segregation should be reflected in the cost proposal with a cost option that includes all HSR-related tasks. All data derived from human subjects must be anonymized for sharing with other TAs and/or use in future research.
HR001118S0040 COMPUTERS AND HUMANS EXPLORING SOFTWARE SECURITY (CHESS) 19
Integration Frameworks:
TA5 must propose a simple integration framework for each target platform (Table 3) that can be provided (with initial functionality) to CHESS performers within four (4) months after kickoff.
TA5 should augment and expand this framework for the duration of the effort to facilitate automated regression testing and evaluation. The TA5 performer will be responsible for coordinating the development of interface specifications and overall system design with TA1 and TA2. TA5 should expect containerized software from TA1 and TA2 that interfaces via a common data format.
The integration framework should be able to isolate source-assisted and binary analysis techniques from TA2 into separate workflows for evaluation. TA5 will manage the integration process with the assumption that TA1 and TA2 performers will produce software that is sufficiently mature to preclude lengthy bug finding on the part of TA5. The framework should be tested in an automated fashion after each code delivery from TA1 and TA2 delivery on the aforementioned simple surrogate system of TA5’s devising to prevent regressions.
Testbed Environment:
TA5 must propose a testbed environment for evaluation of the CS corpus developed by TA3.
This testbed should include instrumentation for evaluations, automated deployment of challenges sets, and testing of PoVs and patches from other performers. The testbed environment should also isolate the running CS from any production networks. TA5 will coordinate the development of common data formats and interfaces for the testbed with all other TAs. TA5 will provide specifications for the interfaces and formats to all other TAs at least six (6) months prior to each evaluation.
Over the three phases of the CHESS program, the TA5 performer will lead eight (8) hackathons, four (4) demonstrations, and three (3) evaluations as described in Figure 3. The TA5 performer will submit plans for these events as described in Section I.E to the Government team at least two months prior to each event.
TA5 proposals should address the following topics:
1. The tools and techniques produced by the research TAs should be integrated into a single framework for each target platform (Table 3) with an installation guide and a usage guide. TA5 proposers should discuss their approach to the development of the integrated CHESS framework, and how they will ensure interoperability with TA1 and TA2 software. The most current, stable version of the TA5 framework should be delivered to the Government at the end of each phase of the program. Proposals should deliver this integrated framework in a common installation package (.msi, .deb, etc.), Docker container or similar format.
HR001118S0040 COMPUTERS AND HUMANS EXPLORING SOFTWARE SECURITY (CHESS) 20
2. The TA5 performer must provide a testbed environment for each target platform (Table 3). These testbeds should provide appropriate instrumentation for evaluation of the TA3 CS corpus and compatibility with the integrated CHESS system. The testbeds should also provide sufficient isolation that any CS cannot influence or run on commercial or open source software platforms.
3. The TA5 performer must research the designated vulnerability classes (Table 4) and assess each TA1 and TA2 performer’s coverage of those classes, both in terms of TA1 and TA2’s current technical progress, and in terms of what TA1 and TA2 would likely cover if they ultimately met all of their technical goals. Relevant staff from the TA5 performer must accompany DARPA representatives on visits to the other TA1 and TA2 performer sites, and study TA1 and TA2 technical approaches.
4. To properly evaluate the technologies and techniques produced by the research TAs, human subjects will be required. TA5 proposers should discuss how they plan to provide human test subjects that fulfill the necessary criteria for the “expert hacker,” “novice hacker,” and “non-hacker” groups (Table 1). TA5 proposers should discuss how they will address any HSR issues. Proposals should include at least one draft experiment protocol, a plan for IRB approval, and an anonymization strategy for any data collected on human subjects.
5. To jumpstart research and development efforts and collaboration across all performers and TAs, a single kickoff meeting will be held at the onset of the CHESS program. The kickoff meeting will focus on open technical exchange, discussion of the research problems encompassed by the CHESS program, and how effective cross-discipline collaboration may address these research problems. TA5 proposers should discuss how they will facilitate these events, including the acquisition and provisioning of appropriate event facilities and resources. (See Section I.F for further detail).
6. To encourage innovative research and prevent duplicate effort, eight (8) hackathons involving participants from TA1, TA2, TA3, and TA5 will be held throughout the program. All TA1, TA2, TA3, and TA5 performers are expected to attend the entirety of all hackathons. TA5 proposers should discuss how they will facilitate these events, including the acquisition and provisioning of appropriate event facilities and resources.
(See Section I.F for further detail).
7. To engage and solicit feedback from potential transition partners, the TA5 performer will be responsible for facilitating four (4) demonstrations involving the integrated CHESS framework. To measure the progress of the CHESS system, the TA5 performer will organize three (3) evaluation events, which will occur at the end of each phase. These evaluation events will involve representatives from each performer as well as human subjects. TA5 proposers should discuss how they will facilitate these events, including the acquisition and provisioning of appropriate event facilities and resources. (See Section I.F for further detail).
HR001118S0040 COMPUTERS AND HUMANS EXPLORING SOFTWARE SECURITY (CHESS) 21
8. To encourage collaboration between CHESS performers, all performers producing software will regularly commit to a common shared, version-controlled source code repository administered by TA5. TA5 will ensure appropriate access control between performers that enable sharing when necessary and preclude inappropriate information exchange when deemed necessary by DARPA.
TA5 proposers should submit a base proposal assuming one performer in each of the first two technical areas (TAs 1-2), and two additional cost options. The additional costs should cover: the incremental cost of an additional performer in TA1, and the incremental cost of an additional performer for TA2. The Government will determine the total contract value to award during contract negotiations based on the selection of performers for awards and the nature of their proposed efforts. TA5 proposals should also include an option scoped for 1-2 full-time equivalents (FTEs) to begin performance at the end of the 42 month program period of performance and continue for an additional 18 months to assist in supporting CHESS transition.
E. Evaluation
The CHESS program Integrator/Evaluator (TA5) will provide engineering input to the Government team in the development of evaluations to provide feedback to the TA1 and TA2 performers. These evaluations will take place in an isolated testbed environment hosting the CS corpus developed by TA3 to characterize the capabilities TA1 and TA2 performers produce.
TA4 will analyze this same CS corpus to provide a baseline for expert hacker performance to measure the CHESS system against.
DARPA will assess individual performer effort in terms of the viability of their technical approaches, the trend in the performance of their systems over time, and their overall progress toward CHESS program objectives. Proposers are encouraged to provide additional metrics, as appropriate for their technical approach and methodology.
Phase Duration
Phase 1 18 months
Phase 2 12 months
Phase 3 12 months
Vulnerability Discovery Speed
As fast as control
10x faster than control
100x faster than control
Vulnerability Discovery Accuracy with Source Code 70% 85% 99%
Vulnerability Discovery Accuracy without Source
Code 50% 75% 99%
Software Complexity
Small Software Library (Low)
Whole Software Package
(Moderate)
Web Browser (High)
Table 5: CHESS System Metrics
HR001118S0040 COMPUTERS AND HUMANS EXPLORING SOFTWARE SECURITY (CHESS) 22
Vulnerability Discovery Speed will measure the vulnerabilities found in the evaluation CS corpus per unit of expert hacker collaborator time. TA4 will act as the control for this metric.
Novice and non-hacker time will not factor directly into this metric.
Vulnerability Discovery Accuracy will measure how many known vulnerabilities in the evaluation CS corpus are discovered versus total known…
This is the start of the file's text. The full file is on GovTribe.
File details come from the government source that posted it.