HR001117S0055-Amendment-01.pdf

PDF 1 MB Posted

Attached to
Page 3 Architectures Federal contract opportunity
Solicitation number
HR001117S0055
Issued by
Defense Advanced Research Projects Agency

About this file

Not Listed

View the file

Other files for this federal contract opportunity

Other files attached to Page 3 Architectures, newest first.
File Type Posted
HR001117S0055-Amendment-02.pdf PDF
HR001117S0055_Arch_Att2_Proposal_Summary_Chart.pptx PPTX presentation
HR001117S0055_Arch_Att1_Proposer_Checklist.pdf PDF
HR001117S0055.pdf PDF

On GovTribe

Work with this file on GovTribe

  • Download the original file
  • Contacts named in this file
  • Similar government files
  • Ask GovTribe AI about this file

Text version

HR001117S0055

Microsystems Technology Office Broad Agency Announcement

Electronics Resurgence Initiative: Page 3 Investments Architectures Thrust

HR001117S0055

September 13, 2017

Amendment No. 01 As Amended on September 19, 2017

Foreword

In his seminal 1965 paper, Gordon Moore, one of the pioneers of the ongoing microelectronics revolution, famously predicted a trajectory of progress in which the transistor count of integrated circuits would double every two years while the cost per transistor would decrease1. This projection became known as Moore’s Law. It set the electronics industry on a quest for continued scaling for more than 50 years and those who have mastered the technology have enjoyed the greatest commercial benefits and the greatest gains in defense capabilities. However, it is clear that the design work and fabrication now required to keep pace is becoming increasing difficult and expensive. The current trajectory of scaling has strained both the commercial and defense sectors, as much for economic as for technical reasons.

From a national security perspective, the dynamics that resulted from Moore’s observations and analysis have become increasingly complex. The DoD has ridden the relentless wave of technical progress in electronics to create exceptionally complex and high-performance systems.

However, the current cost of development is challenging the national security enterprise. The rate of development of novel and unique electronics, based on advances in fundamental science and engineering research, has dwindled within the DoD. The number of leading-edge manufacturing sites that are considered a part of the national security enterprise is diminishing, and the fundamental tie between national security and the health of the electronics industry is strained. The shift in focus of all major electronics entities has been towards large-volume global supply chains as a means to manage the dynamics and economics of scaling. But this has also made it more difficult for the DoD to leverage industry capabilities for the small-volume, high-performance technology that the defense sector needs. With the Electronics Resurgence Initiative (ERI), DARPA seeks to address these imbalances, ultimately working hand-in-hand with industry to embrace the coming inflection in Moore's Law. The goal of the ERI is to more constructively enmesh the technology needs and capabilities of the defense enterprise with the commercial and manufacturing realities of the electronics industry.

1 G. E. Moore, "Cramming more components onto integrated circuits, Reprinted from Electronics, volume 38, number 8, April 19, 1965, pp.114," in Proceedings of the IEEE, vol. 86, no. 1, pp. 82-85, Jan. 1998.

doi: 10.1109/JPROC.1998.658762 <https://doi.org/10.1109/JPROC.1998.658762>

Electronics Resurgence Initiative (ERI)

During these unique times, it is instructive to read all of Moore’s prescient paper. On page two, he laid out what became his famous projection for scaling transistor count. However, on page three, with an eye toward the times we now live in, he laid out the technical directions to explore when the conditions under which scaling will be the primary means for advancement are no longer met. A trio of simultaneously-released ERI BAAs—this one among them—parallel the research areas detailed on page three of Moore’s paper: materials and integration, architecture, and design. These new page-three-inspired investments, along with a series of related investments from the past year, comprise the overall Electronics Resurgence Initiative.

The “ERI Page 3 Investments” are the next steps in creating an electronics capability that will provide a foundational contribution to U.S. national security. They reflect a collaborative spirit that we hope will lead an electronics industry capable of meeting its own commercial needs and ambitions while simultaneously advancing national defense in the 2025 to 2030 time frame.

DARPA is eager to receive proposals from entities that can further the broader cause of the electronics industry while simultaneously embracing national security, based on the development and consistent availability of advanced, high-performance electronics technologies.

ERI Page 3 Investments Overview

Table of Contents

Foreword

PART I: OVERVIEW INFORMATION

PART II: FULL TEXT OF ANNOUNCEMENT

I. Funding Opportunity Description A. Page 3 Architectures Description B. SDH Program

1. Background

2. Program Structure

3. SDH Schedule and Milestones

4. TA-1: Reconfigurable Processors

5. TA-2: Dynamic HW/SW compilers for high-level languages

6. Government-furnished Property/Equipment/Information

7. Program Metrics and Targets

8. Intellectual Property C. DSSoC Program

1. Background

2. Program Description

3. Program Structure

4. Intelligent Scheduler

5. Software

6. Domain Representation – Ontologies

7. Medium Access Control (MAC) Layer

8. Hardware Integration

9. Program Schedule and Milestones

10. Program Metrics

II. Award Information A. General Award Information B. Fundamental Research

III. Eligibility Information A. Eligible Applicants B. Federally Funded Research and Development Centers (FFRDCs) and Government

Entities C. Organizational Conflicts of Interest D. Cost Sharing/Matching E. Other Eligibility Criteria

1. Collaborative Efforts F. Associate Contractor Agreement Clause

IV. Application and Submission Information A. Address to Request Application Package B. Content and Form of Application Submission

1. Abstract Format

2. Full Proposal Format

3. Proprietary Information

4. Security Information

a. Unclassified Submissions

b. Classified Submissions

5. Disclosure of Information and Compliance with Safeguarding Covered Defense

Information Controls

6. Human Research Subjects/Animal Use

7. Approved Cost Accounting System Documentation

8. Section 508 of the Rehabilitation Act (29 U.S.C. § 749d)/FAR 39.2

9. Grant Abstract

10. Small Business Subcontracting Plan

11. Intellectual Property

a. For Procurement Contracts

b. For All Non-Procurement Contracts

12. Patents

13. System for Award Management (SAM) and Universal Identifier Requirements

14. Funding Restrictions C. Submission Information

1. Submission Dates and Times

a. Abstract Due Date

b. Full Proposal Date

c. Frequently Asked Questions (FAQ)

2. Abstract Submission Information

3. Proposal Submission Information

a. For Proposers Requesting Grants or Cooperative Agreements:

b. For Proposers Requesting Contracts or Other Transaction Agreements

c. Classified Submission Information

V. Application Review Information A. Evaluation Criteria

1. Overall Scientific and Technical Merit

2. Potential Contribution and Relevance to the DARPA Mission of Supporting

National Security

3. Impact on the Overall Electronics Landscape

4. Cost Realism B. Review and Selection Process

1. Review Process

2. Handling of Source Selection Information

3. Federal Awardee Performance and Integrity Information (FAPIIS)

VI. Award Administration Information A. Selection Notices

1. Abstracts

2. Proposals B. Administrative and National Policy Requirements

1. Meeting and Travel Requirements

2. FAR and DFARS Clauses

3. Controlled Unclassified Information (CUI) on Non-DoD Information Systems

4. Representations and Certifications

5. Terms and Conditions (for grants and cooperative agreements only)

C. Reporting D. Electronic Systems

1. Wide Area Work Flow (WAWF)

2. i-Edison

3. Contract Execution Reporting Service (CERS)

VII. Agency Contacts VIII. Other Information

A. Proposers Day B. Protesting

ATTACHMENT 1: Cost Volume Proposer Checklist ATTACHMENT 2: Proposal Summary Slide Template

PART I: OVERVIEW INFORMATION

Federal Agency Name: Defense Advanced Research Projects Agency (DARPA), Microsystems Technology Office (MTO) Funding Opportunity Title: Electronics Resurgence Initiative: Page 3 Architectures Announcement Type: Initial Announcement Funding Opportunity Number: HR001117S0055 Catalog of Federal Domestic Assistance Numbers (CFDA): 12.910 Research and

Technology Development Dates: (All times listed herein are Eastern Time) o Posting Date: September 13, 2017 o DSSoC Proposers Day: September 18, 2017 o SDH Proposers Day: September 19, 2017 o Abstract Due Date: October 11, 2017 o FAQ Submission Deadline: November 20, 2017 o Proposal Due Date: December 4, 2017 o Estimated period of performance start: June 1, 2018

Concise description of the funding opportunity:

DARPA is soliciting innovative research proposals in the area of novel computing architectures.

The Page 3 Architectures thrust of the Electronics Resurgence Initiative (ERI) seeks to demonstrate heterogeneous computing systems that provide the performance advantages of specialized processors, while maintaining the programmability of general purpose processors.

The goal of the Software Defined Hardware (SDH) program is to build runtime-reconfigurable hardware and software that enables near ASIC performance without sacrificing programmability for data-intensive algorithms. SDH will create a hardware/software system that allows data-intensive algorithms to run at near ASIC efficiency without the cost, development time or single application limitations associated with ASIC development.

The overall goal of the Domain-specific System on Chip (DSSoC) program is to develop a heterogeneous SoC comprised of many cores that mix general-purpose processors, special-purpose processors, hardware accelerators, memory, and input/output (I/O). DSSoC seeks to enable rapid development of multi-application systems through a single programmable device.

Anticipated individual awards: Multiple awards are anticipated.

Anticipated funding type: 6.1, 6.2 (across both programs) Types of instruments that may be awarded: Procurement contract, grant, cooperative agreement or other transaction.

Agency contacts:

Dr. Wade Shen, Program Manager BAA Coordinator:

SDH@darpa.mil

DARPA/MTO

ATTN: SDH

675 North Randolph Street Arlington, VA 22203-2114

Dr. Tom Rondeau, Program Manager BAA Coordinator:

DSSoC@darpa.mil

DARPA/MTO

ATTN: DSSoC 675 North Randolph Street Arlington, VA 22203-2114

PART II: FULL TEXT OF ANNOUNCEMENT

I. Funding Opportunity Description

The Defense Advanced Research Projects Agency (DARPA) often selects its research efforts through the Broad Agency Announcement (BAA) process. This BAA is being issued, and any resultant selection will be made, using the procedures under Federal Acquisition Regulation (FAR) 6.102(d)(2) and 35.016 and 2 C.F.R. § 200.203. Any negotiations and/or awards will use procedures under FAR 15.4, Contract Pricing. Proposals received as a result of this BAA shall be evaluated in accordance with evaluation criteria specified herein through a scientific review process. Any negotiations and/or awards will use procedures under FAR 15.4, Contract Pricing, as specified in the BAA (including DoDGARS Part 22 for Grants and Cooperative Agreements, and Part 37 for Technology Investment Agreements). Proposals received as a result of this BAA shall be evaluated in accordance with evaluation criteria specified herein through a scientific review process.

DARPA BAAs are posted on the Federal Business Opportunities (FedBizOpps) website, http://www.fbo.gov/, and, as applicable, the Grants.gov website at http://www.grants.gov/. The following information is for those wishing to respond to the BAA.

DARPA is soliciting innovative research proposals in the areas of reconfigurable computing hardware architectures and compiler/programming approaches to provide adaptive, optimizable hardware/software solutions for data intensive applications. Proposed research should investigate innovative approaches that enable revolutionary advances in science, devices, or systems.

Specifically excluded is research that primarily results in evolutionary improvements to the existing state of practice.

A. Page 3 Architectures Description

The Page 3 Architectures thrust includes two main programs that will operate independently of each other:

Software Defined Hardware (SDH): Build runtime reconfigurable hardware and software that enables near ASIC performance without sacrificing programmability for data-intensive algorithms mailto:SDH@darpa.mil mailto:DSSoC@darpa.mil http://www.fbo.gov/ http://www.grants.gov/

Domain-Specific System on Chip (DSSoC): Enable rapid development of multi-application systems through a single programmable device

Proposers should not propose to more than one program in a single proposal. Further information about the programs can be found in Section B (SDH) and Section C (DSSoC).

B. SDH Program

1. Background

In modern warfare, decisions are driven by information. That information comes to us in the form of thousands of sensors providing information, surveillance, and reconnaissance (ISR) data;

logistics/supply-chain and personnel performance measurements; etc. Our ability to exploit this data to understand and predict the world around us is an asymmetric advantage for the DoD.

Utilization of this data relies on computational algorithms running at a huge scale. Today, developers are limited in their ability to run these algorithms efficiently because they generally have to trade the efficiency of their algorithms with the efficiency of available hardware architecture implementations. To gain the most efficiency possible, we design and fabricate application specific integrated circuits (ASICs)—customized hardware designed to maximize the runtime efficiency of a specific algorithm. ASICs typically cost hundreds of millions of dollars and take many years to develop. Once developed, they can perform exactly one class of computation because they were designed and specialized for specific tasks. Because these systems are so specifically tailored and costly, we cannot create ASICs for all but the highest priority algorithms. For problems that cannot afford this level of investment, we sacrifice compute efficiency by implementing solutions as software on general-purpose processors or field programmable gate arrays (FPGAs). Often, this results in application implementations that are thousands of times worse than optimal.

2. Program Structure

The goal of this program is to build runtime-reconfigurable hardware and software that enables near ASIC performance without sacrificing programmability for data-intensive algorithms. For the purposes of this BAA, data-intensive algorithms are machine learning and data science algorithms that process large volumes of data and are characterized by their usage of intense linear algebra, graph search operations and their associated data-transformation operators. SDH will create a hardware/software system that allows data-intensive algorithms to run at near ASIC efficiency without the cost, development time or single application limitations associated with ASICs. If successful, SDH will result in the ability to develop and run data-intensive, data-exploitation algorithms at very low cost, and, consequently, enable pervasive use of big-data solutions for a wide range of DoD applications including ISR, predictive logistics, decision support, etc.

Technical Approach:

The SDH program will create a malleable hardware/software architecture that, unlike ASICs, allows an application to defer hardware configuration to runtime. SDH will enable 1) the optimization of code and hardware dynamically when input data change, and 2) the reuse of hardware for new problems and new algorithms to solve existing problems. In order to accomplish these objectives, SDH targets very fast hardware reconfiguration speeds and dynamic compilation.

If successful, SDH will be able to take advantage of data-dependent optimizations that even today’s ASICs cannot exploit.

The program is organized into two technical areas (TAs), as follows:

Reconfigurable processor technology (TA-1): The goal in TA-1 is to build a processor that is reconfigurable at runtime. This means SDH processors will be faster to reconfigure than FPGAs. In order to achieve ASIC-like efficiencies, these processors will use a small fraction of the power required by current FPGA. These systems will allow for compute and internal memory reconfiguration and incorporate external memory subsystems that can be configured for different data access patterns.

Dynamic hardware/software compilers for high-level languages (TA-2): TA-2 performers will build programming languages and compilers that optimize software and hardware at runtime. TA-2 systems will do this by translating high-level programs into machine code and application-specific hardware configurations, such that it is easy for programmers to develop their applications. TA-2 systems will iteratively optimize code and hardware as the program is running to maximize performance dynamically based on the data being processed.

Figure 1 shows how these technical areas will interoperate.

Reconfigurable processors (TA1) Properties:

1. Reconfiguration times: 300 - 1,000 ns

2. Re-allocatable compute resources – i.e. ALUs for address computation or math

3. Re-allocatable memory resources – i.e. cache/register configuration to match data

4. Malleable external memory access – i.e. reconfigurable memory controller

Config1 Config2 Config3 Config4 ConfigN

High-level program

Dynamic HW/SW compilers for high-level languages (TA2)

Code1 Code2 CodeN

1. Generate optimal configuration based on static analysis code

2. Generates optimal code

3. Re-optimize machine code and processor configuration based on runtime data

Code3

Time

Figure 1: Overview of SDH concept of operations

3. SDH Schedule and Milestones

Figure 2 shows the anticipated program schedule for SDH. The program is divided into three phases (12 months, 18 months, and 18 months). During the first phase of the program, TA-1 and TA-2 teams will develop initial SDH hardware architecture models along with prototype languages/compilers. TA-1 performers will develop initial architectural design simulations for architectural evaluation and demonstration and use by the TA-2 performers. At the end of this phase, TA-1 and TA-2 systems will combine for a simulation demo.

During Phase 2 of the program, detailed hardware architecture prototype designs for TA-1 systems will be developed and compilers in TA-2 will be revised accordingly. A detailed TA-1 design simulation will be developed in Phase 2. Phase 3 will result in an integrated demonstration with the fabrication of prototype hardware architecture devices developed in TA-1 and compiler/software implementations developed in TA-2.

Teams may propose joint TA-1/TA-2 efforts or propose to each technical area separately. In the latter case, DARPA will pair teams based on technical affinities after source selection. The program is designed to create joint TA-1/TA-2 prototypes and demonstrators at the scale of current microprocessors and compiler tools. This means that TA-1 systems should be sized to run on real-world data problems (i.e. millions of rows of dense data, graphs of 10 billion+ edges, matrix completion problems involving 10 million+ sparse columns, matrix computations involving 100+ TFLOPs, stochastic search networks with perplexity greater than 200, etc.) and TA-2 tool suites usable by lay programmers.

TA1: Reconfigurable processor

TA2: SDH

languages/compilers

Phase I PDR Phase 1.5 CDR

Phase 1 (12 months) Phase 2 (18 months)

Compiler prototypes Data-driven JIT compilers on simulation platform

Initial language design V2 language design

Initial simulator and FPGA prototype

Initial architecture design Full architecture design

Annual programmability evaluation

Phase 3 (18 months)

Final FPGA Prototype + final simulator

Phase 2 production milestone

Phase 2 efficiency measurement

Production prototype

Final language specifications

Reference compiler implementations

Phase 3 Integrated Demo

Language/hardware Co-design

Figure 2: Program Schedule

4. TA-1: Reconfigurable Processors

The goal in TA-1 is to build a processing architecture that can be configured to optimize performance on a range of data-intensive problems.

These systems must be reconfigurable at runtime. As such, the program targets fast reconfiguration times (less than 1 µs by the end of the program). Proposers may propose either reconfiguration pipelining strategies or full reconfiguration solutions that address this target. The class of problems that SDH targets includes problems that are currently memory-bound (e.g. sparse graph/matrix operations) and those that are typically Arithmetic Logic Unit/Floating-Point Unit (ALU/FPU) bound (e.g. dense tensor math). This means that proposed systems must allow for different memory access patterns and different compute facilities that can be tailored at both compile-time and runtime by TA-2 compilers.

SDH hypothesizes that many ASIC designs could be realized by using a fabric of underlying building blocks of internal memories and ALUs that can be combined with reconfigurable interconnects. This fabric, when combined with reconfigurable and/or programmable memory controllers, could address many Machine Learning/Artificial Intelligence (ML/AI) and graph workloads as such a system would allow applications to specialize external memory access methods and on-chip compute/cache resource allocation. Figure 33 shows a conceptual approach for how SDH processors could be designed.

Programmable Memory Controller

Programmable Memory Controller

MALU

M MALU

ALU

ALU ALUM

Reconfigurable Interconnect

Figure 3: SDH Programmable Fabric

This example is essentially a Coarse-Grained Reconfigurable Array (CGRA) as described in research literature with programmable peripheral memory controllers. Solutions to SDH hardware systems are not limited to this conceptual design as long as they can credibly reach SDH’s stated performance targets as described in Section I.B.7.

TA-1 systems should facilitate performance measurement to enable TA-2 compilers/optimizers.

This could come in the form of low-overhead trace point and/or runtime instrumentation facilities.

DARPA expects that selected designs, after phase 2 of SDH, will be built as a demonstration prototype in silicon at the fabrication geometry of current state-of-the-art GPUs and microprocessors (and in both density and die size terms). To enable fabless and academic design teams to participate in TA-1, proposals can make use of DARPA/MTO fab shuttle runs (USG-furnished) as part of their Phase 3 efforts. Proposals that rely on these facilities should state this fact clearly and should be costed accordingly. To be selectable, proposers must provide a viable path to TA-1 architecture design fabrication. Proposals may incorporate novel external memory subsystems and bus architectures. In addition to demonstration silicon, TA-1 teams must provide running prototype boards including debugging access, code upload facilities, and large external memories (256GB minimum) to support testing.

DARPA expects that TA-1 and TA-2 teams will need to cooperatively design both TA-1 hardware and TA-2 languages. In order to help facilitate the design of languages for specific TA-1 solutions, the program requires that TA-1 performers allow for TA-2 teams that are not part of an explicit TA-1+TA-2 team to be sited with TA-1 designers. Further, to facilitate TA-2 development, TA-1 teams will develop and provide for TA-2 teams use of TA-1 architectural simulators. These simulators shall be of sufficient quality to allow TA-2 teams to develop and evaluate their hardware reconfiguration and optimization approaches.

TA-1 teams can make use of the supplied D3M/HIVE corpus (2,000+ programs and data developed under DARPA’s Data-driven Discovery of Models (D3M) and HIVE programs, to be made available to SDH performers after source selection) to discover common kernels/functionality, data access patterns and for the consideration of potential hardware reconfiguration architectural elements.

5. TA-2: Dynamic HW/SW compilers for high-level languages

To make reconfigurable hardware useful, we need compilers and programming languages that support joint optimization of hardware and software. Traditional compilers assume a fixed processor. They can generate code for that processor from input programs and they can iteratively refine the code that they produce by analysis of a runtime trace. This process allows for data-dependent optimization but is limited by the available processing hardware.

Because the SDH-envisioned hardware is highly malleable, existing programming languages may not be appropriate for TA-2 as most of these languages conflate the declarative compute objective of a program with procedural and optimization details of how that algorithm is computed. This results in code that is hard to optimize when the hardware architecture changes, even when that code is “high-level”. TA-2 proposers are encouraged to develop declarative languages and compilers that can select or synthesize optimal implementations of code while selecting optimal hardware configurations to support that implementation. Methods that use low-level synthesis techniques, dynamic auto-tuning and implementation selection, and code mining are strongly encouraged.

SDH compilers will co-optimize processor configuration and code using a similar iterative procedure. The problem is that this optimization space is now much larger than what traditional compilers have to contend with. In order to make SDH compilers practical, TA-2 performers may use, for example, stochastic optimization and reinforcement learning across programs to learn how to predict the best architectural configurations for a given problem and runtime data associated with that problem.

TA-2 teams should propose compilers and high-level programming languages to support 1) high programmer productivity for data-intensive algorithms and 2) optimal runtime performance given TA-1 systems. TA-2 proposals are not required to propose new languages (but they may). In these cases, proposers should clearly specify their programming language targets and any modifications they intend for existing languages.

TA-2 teams can make use of the supplied D3M/HIVE corpus (programs and data) for discovering and training stochastic optimizers offline. TA-2 teams may use this corpus to discover kernels and data access patterns and to inform the construction of hardware primitives.

TA-2 systems will be expected to perform runtime, data-dependent optimizations. Proposal may include runtime, compile-time, and offline optimization with a TA-1 processor in-the-loop. Note that compile-time optimizations and dynamic recompilation costs will be measured as part of performance evaluation. Proposals should state clearly what hardware tracing/instrumentation facilities they expect or require.

TA-2 teams should include supporting tools such as debuggers, profilers, etc. to support development for their proposed languages. As TA-2 systems will be partially evaluated on programmer efficiency, DARPA strongly encourages the development of programmer affordances as part of an overall TA-2 solution.

Proposals that define new programming languages should note that their programming efficiencies will be measured relative to Python+NumPy/SciPy. Proposed languages that are less efficient than this stack from a programming productivity standpoint will not be considered.

DARPA expects that TA-1 and TA-2 teams will need to cooperatively design both TA-1 hardware and TA-2 languages. In order to help facilitate the design of TA-1 systems, the program requires that TA-2 performers co-locate language design experts with TA-1 teams (specific TA-1 performers will be determined by the DARPA Program Manager after source selection). TA-2 teams should budget a staff researcher, post-doc, or graduate student for temporary assignment on location full-time during phase 1 (full time) and for 45 working days during phase 2.

6. Government-furnished Property/Equipment/Information

DARPA will provide a corpus of data-intensive programs and associated data inputs (from I2O’s D3M program and MTO’s HIVE program) upon kickoff of the program. The USG Test and Evaluation (T&E) team will provide high-level implementations of these programs (i.e. Python + NumPy/SciPy) and representative optimized example implementations for CPU and GPU as comparative baselines for evaluation of program targets as described below. The USG evaluation team will also develop FPGA-based solutions for a subset of these programs to measure SDH efficiencies in comparison with potential ASIC solutions. These will be made available to eventual performers throughout the length of the program.

7. Program Metrics and Targets

The program will measure compute efficiency (in terms of giga-operations per watt (GOPs/Watt) or tera-operations per watt (TOPs/Watt)) and programmability via annual user experiments. Table 1 shows the evaluation metrics to be used for SDH and the program targets for each phase.

vs. CPU vs. ASIC vs. ASIC (sparse math, search, graphs) Programmability

Phase 1 100-300x better within 50x ~1x within 10x

Phase 2 100-300x better within 10x 2x better within 3x

Phase 3 500-1000x better within 5x 8-10x better ~1x Table 1: Program metrics and target

By the end of the program, we expect SDH systems to be at efficiencies within 5x of ASICs. For a sub-class of highly data-dependent problems that will likely include data-dependent sparse math, search and graph programs, and other areas to be announced post-award, SDH compilers and hardware should enable up to 10x better than ASIC efficiencies (as shown in Table 1). In order to evaluate SDH against potential ASICs, the USG team will develop baseline solutions on FPGA proxies and extrapolate these solutions to ASIC power efficiencies. CPU baselines will be based on Intel E7-8894 v4.

Annual evaluations of programmability will be conducted as user studies by the USG, working with TA-2 performers. These evaluations will measure the ability of programmers to make use of TA-2 supplied programming languages and compiler tools to build SDH solutions to D3M/HIVE problems and specifically for a subset of evaluation problems. Program evaluations will measure the time and effort required to complete working solutions using DSH tools in comparison with time/effort required to complete solutions using the baseline Python + NumPy/SciPy stack.

Component-level evaluations will be conducted at each phase of the program. TA-1 performers will be required to show that their systems can be manually optimized for a wide range of D3M/HIVE problems. Similarly, TA-2 systems will be evaluated on their ability to automatically find optimal or near optimal configurations given high-level programs as input.

8. Intellectual Property

Performers will develop proof-of-concept designs and “at-scale” prototypes of SDH hardware/software during the course of the program. DARPA hopes that these technologies will be further developed after program completion into fully commercialized products so that the DoD can benefit from the commoditization of SDH solutions.

The program encourages creating and leveraging open source technology and architecture, when possible, to help facilitate adoption of SDH research. SDH encourages the use of non-viral open source licenses to enable commercial adoption.

That said, DARPA recognizes that its investment is only part of the overall cost to develop SDH solutions and open source licensing may not be the best solution for all proposers, especially for solutions using existing IP or for joint cost-sharing investments. As such, proposals may include an IP rights assertion with appropriate justification. Proposals will be gauged (in part) by the probability of successful technical execution and commercialization. See Section IV.11 for more details on intellectual property.

C. DSSoC Program

1. Background

The invention of the digital computer came about as a proof of computable numbers, and so numerical processing was solved by the Turing Machine that has led to the general purpose computer. These computers are good at solving multiple types of problems with a single machine.

However, within the scope of computable numbers are specific computational problems where purpose-built machines can solve them faster while using less energy. An example of this is the digital signal processor (DSP) that performs the multiply and accumulate operation important to signal processing, or the more recent development of the tensor processing unit (TPU) to compute dense matrix multiplies that are at the core of deep neural networks. These specialized machines solve a single problem but with optimized efficiency.

As a result, there exists a tension between the flexibility in general purpose processors and the efficiency of specialized processors. Domain-Specific System on Chip (DSSoC) intends to demonstrate that the tradeoff between flexibility and efficiency is not fundamental. The program will develop a method for determining the right amount and type of specialization while making a system as programmable and flexible as possible. It will take a vertical view of integrating today’s programming environment with a goal of full-stack integration. DSSoC will de-couple the programmer from the underlying hardware with enough abstraction but still be able to utilize the hardware optimally through intelligent scheduling. DSSoC specifically targets embedded systems where the domain of applications sits at the edge and near the sensor. Workloads consist of small chunks of data but often with a large number of algorithms required in the processing, meaning that high compute power and low latency at low power are required.

Themes of specialization and parallelism reoccur whenever industry runs into processing bottlenecks. The last time this discussion occurred in the early 2000s was when the general purpose processor hit the frequency wall. This led to multiple homogeneous cores and hyper-threading processor paradigms. During this period, an attempt to analyze processing specialization was to classify programs into a taxonomy. The first well-known attempt at this is from 20042 with a list of seven classes, or motifs, of programs in high performance computing. The Berkeley view of this in 20063 referred to these as the “Seven Dwarfs” and then went on to expand this list to thirteen.

Later, another taxonomy was developed for the seven dwarfs of symbolic computation.4 Taxonomies are useful but have limitations as they only provide a list of categories and not relationships between categories. Instead, DSSoC seeks to consider the concept of ontologies for the domains.

A domain addresses a set of problems that share common concepts and mathematical underpinnings that can be described by such ontologies. A domain is smaller than all computable

2 P. Colella, “Defining Software Requirements for Scientific Computing,” presentation, 2004.

3 K. Asanovic, et al. The landscape of parallel computing research: A view from Berkeley. Technical Report UCB/EECS-2006-183, EECS Department, University of California, Berkeley, 2006.

4 E. L. Kaltofen, “The ‘Seven Dwarfs’ of Symbolic Computation,” In: Langer U., Paule P. (eds) Numerical and Symbolic Scientific Computing. Texts & Monographs in Symbolic Computation (A Series of the Research Institute for Symbolic Computation, Johannes Kepler University, Linz, Austria). Springer, Vienna, 2012.

numbers, but much larger than a specific application. The concept of DSSoC is an informed set of general purpose, special purpose, and hardware accelerator coprocessors that is easily programed for applications within a domain. The concepts and methods developed in DSSoC should be extensible into any domain, but proposing teams should focus on a single domain. The nominal domain of the program will be software radio, which includes problems such as mobile communications, satellite communications, personal area networks, all types of radar, and applications in the electronic warfare space important to national security. Other domains may be proposed if accompanied with a list of exemplar applications that can be used to measure the performance improvements and programmability of the DSSoC.

Much of working with computers is resource management, and DSSoC is a program about addressing resources. The design of a heterogeneous chip means resource optimization of the processor elements on the chip. The programming environment means compilers to address static compute resources. And the runtime hardware resource optimization is accomplished through the intelligent scheduler.

2. Program Description

DSSoC aims to develop a heterogeneous SoC comprised of many cores that mix general purpose processors, special purpose processors, hardware accelerators, memory, and input/output (I/O).

These types of cores will be referred to as processor elements (PEs) throughout the program. The mix and number of PEs will be guided by the domain of interest for the given processor. In this way, a domain is larger than any one application. For many applications, the market cannot sustain the development time and cost to create an application-specific integrated circuit (ASIC). DSSoC focuses on a larger set of applications, identified as a domain, which can solve the problem space more efficiently than a general purpose processor while maintaining programmability within that domain.

Much of the work that goes into managing heterogeneous processor resources in current systems and SoCs is done through manual optimization. An expert engineer analyzes the program for optimization points, maps the computation and I/O between the processors, and then codes the implementation to make the program work on that set of processors. This trend has two big issues.

First, the resulting programs do not scale well to multi-application systems where sharing of resources cannot be pre-allocated. Second, specialization has always required direct knowledge of how to use that chip, and often, chip-specific information must be hard-coded into the software.

Changes to the platform or the attempt to port the code between platforms with different processors is often not possible without full reengineering efforts. DSSoC seeks to improve on this development style by embedding resource management into the hardware as an intelligent scheduler along with the programming tools that allow developers to focus on their applications and not on the underlying hardware.

The full concept of the DSSoC program is intelligent resource allocation at three levels. The first level is design-time resource management for deciding the type, number, and distribution of PEs.

The second level is compiler-time resource management to statically compile optimizations into each program. The third level is run-time optimization where the scheduler dynamically makes online updates to the use of the PEs to support multiple, simultaneous applications.

At design time, DSSoC manages the computing resources by appropriately scoping the selected domain. The challenge here is to define this scope in terms of the mathematical motifs that are common within the domain. This domain description should then map to the required set of heterogeneous processor elements. It should describe the important accelerators needed, the number of each PE, and how to distribute those PEs for optimal use of resources. Specialization can increase the efficiency of solving problems, but making and integrating specialized processors for all possible problems is intractable and infeasible. Instead, the domain will guide the balance between general and special purpose computing without the narrow focus on a single application.

Along with the selection and distribution of heterogeneous PEs is how data, instructions, and power are managed between and among PEs. In addition to designing an SoC, the program aims to develop new methods or techniques to best manage the ability to move data between PEs, allow programmers to seamlessly utilize the specialization, and enable the system to control power distribution to manage overall system power consumption. To support programming abstraction and the embedded, intelligent scheduler, DSSoC will provide an efficient yet standardized interconnect between PEs in both as a hardware interconnect and as a software programmable interface.

Given that a DSSoC is a collection of heterogeneous PEs, one of the program’s goals is to allow easy integration and programming of new PEs. Compiling code to a processor or configuring the parameters on an accelerator should be made available to any higher-level program and/or operating system. This means that a new processor may have a different instruction set architecture (ISA), but that the ISA must be supported in a standard compiler (e.g., a new target to LLVM’s Clang compiler5). Similarly, the interfaces for an application to use an accelerator should be standardized for easy integration with programs and operating systems. At a minimum, at the operating system level, a user should be able to pipe data to a device file (e.g., /dev/dssoc_accel0), and programming languages should have common structures for input/output calls such as ioctl6 for C/C++ or something similar to a SensorManager class7 supported under Android in Java. Part of this program will be to allow easy integration and programming of new PEs.

In addition to compilers, a full development ecosystem requires many other tools to support programming tasks. The list of tools important to this program includes, but is not limited to, debuggers, performance analysis, benchmarking and example applications, and libraries. As much as possible, these tools should be free and open source software (FOSS). FOSS implementations of the tools will reduce risk and burden to any one company or development effort by sharing the responsibilities while also providing a common framework to work on applications and ideas. The libraries in these cases should support the domain of interest, such as signal processing, computer vision, machine learning, etc. Examples to follow are the Vector Optimized Library of Kernels (VOLK)8 from the GNU Radio project9, OpenCV10 for computer vision, and FFTW11 for Fourier transforms.

5 https://llvm.org/ 6 http://man7.org/linux/man-pages/man2/ioctl.2.html 7 https://developer.android.com/reference/android/hardware/SensorManager.html 8 http://libvolk.org/ 9 https://www.gnuradio.org/ https://llvm.org/ http://man7.org/linux/man-pages/man2/ioctl.2.html https://developer.android.com/reference/android/hardware/SensorManager.html http://libvolk.org/ https://www.gnuradio.org/

DSSoC will be a 48-month program, divided into four 12-month phases. During Phase 0 of the program, the scheduler should be designed in software against a current heterogeneous SoC while work on designing the DSSoC chip, MAC, and other software support begins. In Phase 1, development work should go into porting the initial design of the DSSoC into an emulation on discrete hardware of a heterogeneous design with software cores representing the accelerators. At the end of Phase 1, the program will show the effectiveness of the scheduler managing the discrete hardware emulation of the DSSoC and provide results on the effectiveness of the software, development tools, scheduler, and ontology. In Phase 2, teams will complete the first spin of their DSSoC, integrating lessons learned from the discrete hardware emulation and with at least two simultaneously running domain applications running through the intelligent scheduler. In Phase 3, teams will complete a second spin of the DSSoC and demonstrate it with at least five simultaneously running domain applications. Throughout all phases, all aspects of the software tools and scheduler should be continuously advanced.

3. Program Structure

DARPA has identified five key development areas that need to be addressed in order to respond to the challenges described above.

All solutions must be developed across the areas of:

1. Intelligent scheduling to manage the set of domain resources in the context of specific applications,

2. Software tools to enable a development ecosystem that exercises the full capability of the highly programmable system,

3. Forming domain representations as ontologies,

4. Medium access control (MAC) to interconnect the PEs and to allow the data throughput, taking into consideration latency, power, and other domain constraints, and

5. Hardware integration of the right set of PEs on the MAC layer with the operating scheduler and software into a fabricated DSSoC.

These five development areas form the core of the program and look to address all of the problems mentioned above to produce a usable DSSoC. Proposals must be fully responsive to all areas of the program. DARPA expects DSSoC teams will need to be multidisciplinary to manage the full scope of the program as described here.

The goal of this program is to look at vertical integration of the full development stack to decouple the developer from the underlying system. DSSoC contains a single technical area with five separate components. Proposing teams must make sure they are able to support all five of these components, which are represented by the

10 http://opencv.org/ 11 http://www.fftw.org/

Figure 4. Vertical integration to enable decoupled application development http://opencv.org/ http://www.fftw.org/ different colored blocks in Figure 4. Red is the intelligent scheduling that manages the programs and DSSoC resources. Blue is the software and tools required to use the DSSoC and provide the necessary ecosystem. The purple is the domain ontology that describes the design and layout of the DSSoC. Orange is the MAC layer for physically moving data around the DSSoC. And gray is the DSSoC itself that represents the hardware integration task.

4. Intelligent Scheduler

In DSSoC, the traditional concept of an operating system scheduler is no longer sufficient. To optimize the use of the large set of heterogeneous PEs by multiple applications, a scheduler must monitor the current state of all of the PEs, observe the instructions of all application binaries, track inter- and intrachip data movement, determine the optimal, available PE given the system’s constraints, and direct the DSSoC to execute the algorithm and data on the optimal PE. The scheduler should sit in the hardware as an overlay processor that works in conjunction with the DSSoC hardware and operating system to interpret the execution of multiple concurrent applications and map the current application requirements to the appropriate PE. The term “overlay processor” indicates that this is a processor with a different function than another PE. It is a monitoring and control element to observe the state of the DSSoC and to make command and control decisions for how to handle power and data movement. It aims to provide dynamic, run-time optimization of the computing resources. The DSSoC program is seeking the implementation of such a scheduler both algorithmically as well as building an actual embedded processor to operate with the DSSoC.

From the scheduler’s perspective, selecting the optimal PE for an operation or portion of an operation will have to consider a number of issues, broadly scoped and not limited to computability and environmental. Computability looks to determine which PE best handles a particular algorithm with the constraint that the PE may be currently in use or the latency to move the data to the PE may be more costly than the speed that it can compute. The appropriateness of the PE to the algorithm should be balanced against the time or energy cost of moving the data, depending on distance between the PEs, amount of data to be computed, and if the optimal PE is already scheduled and operating on other data. The environmental issues include power and thermal conditions. The scheduler should handle power management to regulate power distribution and be able to turn PEs on and off to minimize power consumption. For example, the use of a given PE at a given time might depend on how fast a PE can be modulated on or off or to other scaled power states and whether it might cause thermal issues at specific localities. In those cases, a sub-optimal PE may be selected to reduce hotspots.

A concept important for proper resource management by the scheduler is device introspection.

Every PE should be able to express itself, know what is happening inside itself through internal performance counters (e.g., cache hits and misses, temperature, etc.), and allow external programs to query for this information. The scheduler should utilize this information to select the best PE for an operation given the current state of the system.

Device introspection is also important for code portability. Code portability has been a continuous problem for heterogeneous processing systems, even in systems composed of discrete processors and not just integrated SoCs. General purpose processors are architected around a logic core to manage the complexity of any problem at the expense of efficiency. But a program designed for one general purpose machine can be easily ported to any other. When using specialized processors, they often have size, resource, or other constraints on the type of computation or parameterization of the problem. GPUs are good examples of this problem. Often, algorithms are scoped to the size of the GPU, such as the number of blocks and threads, and therefore locked into GPUs of the same or greater dimensions. Code portability requires information about a PE to be discoverable at runtime in order to automatically and optimally use resources.

5. Software

All teams must incorporate software development into their proposal, which includes development tools such as compilers and debuggers, libraries of algorithms, and applications and examples within the proposed domain. Any new processor must be developed with these software components in order to make it usable by the rest of the development community. This requirement therefore demands that construction of these tools be focused on usability and upgradability to support new DSSoCs. The domain applications will be used to measure performance, and a DSSoC must be able to simultaneously run multiple applications.

The DSSoC program…

This is the start of the file's text. The full file is on GovTribe.

File details come from the government source that posted it. Updated .