HR001117S0056-Amendment-01.pdf

PDF 1 MB Posted

Attached to
Page 3 Materials and Integration Federal contract opportunity
Solicitation number
HR001117S0056
Issued by
Defense Advanced Research Projects Agency

About this file

Not Listed

View the file

Other files for this federal contract opportunity

Other files attached to Page 3 Materials and Integration, newest first.
File Type Posted
HR001117S0056_Materials_Att2_Proposal_Summary_Chart.pptx PPTX presentation
HR001117S0056_Materials_Att1_Proposer_Checklist.pdf PDF
HR001117S0056.pdf PDF

On GovTribe

Work with this file on GovTribe

  • Download the original file
  • Contacts named in this file
  • Similar government files
  • Ask GovTribe AI about this file

Text version

HR001117S0056

Microsystems Technology Office Broad Agency Announcement

Electronics Resurgence Initiative: Page 3 Investments Materials and Integration Thrust

HR001117S0056

September 15, 2017

Amendment No. 01 As Amended on September 19, 2017

Foreword

In his seminal 1965 paper, Gordon Moore, one of the pioneers of the ongoing microelectronics revolution, famously predicted a trajectory of progress in which the transistor count of integrated circuits would double every two years while the cost per transistor would decrease1. This projection became known as Moore’s Law. It set the electronics industry on a quest for continued scaling for more than 50 years and those who have mastered the technology have enjoyed the greatest commercial benefits and the greatest gains in defense capabilities. However, it is clear that the design work and fabrication now required to keep pace is becoming increasing difficult and expensive. The current trajectory of scaling has strained both the commercial and defense sectors, as much for economic as for technical reasons.

From a national security perspective, the dynamics that resulted from Moore’s observations and analysis have become increasingly complex. The DoD has ridden the relentless wave of technical progress in electronics to create exceptionally complex and high-performance systems.

However, the current cost of development is challenging the national security enterprise. The rate of development of novel and unique electronics, based on advances in fundamental science and engineering research, has dwindled within the DoD. The number of leading-edge manufacturing sites that are considered a part of the national security enterprise is diminishing, and the fundamental tie between national security and the health of the electronics industry is strained. The shift in focus of all major electronics entities has been towards large-volume global supply chains as a means to manage the dynamics and economics of scaling. But this has also made it more difficult for the DoD to leverage industry capabilities for the small-volume, high-performance technology that the defense sector needs. With the Electronics Resurgence Initiative (ERI), DARPA seeks to address these imbalances, ultimately working hand-in-hand with industry to embrace the coming inflection in Moore's Law. The goal of the ERI is to more constructively enmesh the technology needs and capabilities of the defense enterprise with the commercial and manufacturing realities of the electronics industry.

1 G. E. Moore, "Cramming more components onto integrated circuits, Reprinted from Electronics, volume 38, number 8, April 19, 1965, pp.114," in Proceedings of the IEEE, vol. 86, no. 1, pp. 82-85, Jan. 1998.

doi: 10.1109/JPROC.1998.658762 <https://doi.org/10.1109/JPROC.1998.658762>

Electronics Resurgence Initiative (ERI)

During these unique times, it is instructive to read all of Moore’s prescient paper. On page two, he laid out what became his famous projection for scaling transistor count. However, on page three, with an eye toward the times we now live in, he laid out the technical directions to explore when the conditions under which scaling will be the primary means for advancement are no longer met. A trio of simultaneously-released ERI BAAs—this one among them—parallel the research areas detailed on page three of Moore’s paper: materials and integration, architecture, and design. These new page-three-inspired investments, along with a series of related investments from the past year, comprise the overall Electronics Resurgence Initiative.

The “ERI Page 3 Investments” are the next steps in creating an electronics capability that will provide a foundational contribution to U.S. national security. They reflect a collaborative spirit that we hope will lead an electronics industry capable of meeting its own commercial needs and ambitions while simultaneously advancing national defense in the 2025 to 2030 time frame.

DARPA is eager to receive proposals from entities that can further the broader cause of the electronics industry while simultaneously embracing national security, based on the development and consistent availability of advanced, high-performance electronics technologies.

ERI Page 3 Investments Overview

Table of Contents

Foreword I. Funding Opportunity Description

A. Overall Program Description B. Three Dimensional Monolithic System-on-a-Chip (3DSoC) Program

1. Background

2. Program Description

3. Program Structure

4. TA-1: Develop a 3D monolithic fabrication process capable of high-yield fabrication of complex compute/memory systems

5. TA-2: Fabrication Development Test Chip SoC Design

6. TA-3: Development of 3DSoC Design EDA Tools

7. Program Schedule

C. Foundations Required for Novel Compute (FRANC) Program

1. Background

2. Program Description

3. Technical Areas

4. Program Structure

5. Schedule/Milestones

D. Deliverables II. Award Information

A. General Award Information B. Fundamental Research

III. Eligibility Information A. Eligible Applicants

1. Federally Funded Research and Development Centers (FFRDCs) and Government Entities

B. Organizational Conflicts of Interest C. Cost Sharing/Matching D. Other Eligibility Criteria

1. Collaborative Efforts IV. Application and Submission Information

A. Address to Request Application Package B. Content and Form of Application Submission

1. Full Proposal Format

2. Proprietary Information

3. Security Information

a. Unclassified Submissions

b. Classified Submissions

4. Disclosure of Information and Compliance with Safeguarding Covered Defense

Information Controls

5. Human Research Subjects/Animal Use

6. Approved Cost Accounting System Documentation

7. Section 508 of the Rehabilitation Act (29 U.S.C. § 749d)/FAR 39.2

8. Grant Abstract

9. Small Business Subcontracting Plan

10. Intellectual Property

a. For Procurement Contracts

b. For All Non-Procurement Contracts

11. Patents

12. System for Award Management (SAM) and Universal Identifier Requirements

13. Funding Restrictions

C. Submission Information

1. Submission Dates and Times

a. Full Proposal Date

b. Frequently Asked Questions (FAQ)

2. Proposal Submission Information

a. For Proposers Requesting Grants or Cooperative Agreements:

b. For Proposers Requesting Contracts or Other Transaction Agreements

c. Classified Submission Information

V. Application Review Information A. Evaluation Criteria

1. Overall Scientific and Technical Merit

2. Potential Contribution and Relevance to the DARPA Mission

3. Impact on the Overall Electronics Landscape

4. Cost Realism

B. Review and Selection Process

1. Review Process

2. Handling of Source Selection Information

3. Federal Awardee Performance and Integrity Information (FAPIIS)

VI. Award Administration Information A. Selection Notices

1. Proposals B. Administrative and National Policy Requirements

1. Meeting and Travel Requirements

2. FAR and DFARS Clauses

3. Controlled Unclassified Information (CUI) on Non-DoD Information Systems

4. Representations and Certifications

5. Terms and Conditions

C. Reporting D. Electronic Systems

1. Wide Area Work Flow (WAWF)

2. i-Edison

3. Contract Execution Reporting Service (CERS)

VII. Agency Contacts VIII. Other Information

A. Proposers Day B. Protesting

ATTACHMENT 1: Cost Volume Proposer Checklist ATTACHMENT 2: Proposal Summary Slide Template

PART I: OVERVIEW INFORMATION

Federal Agency Name: Defense Advanced Research Projects Agency (DARPA), Microsystems Technology Office (MTO)

Funding Opportunity Title: Electronics Resurgence Initiative: Page 3 Materials and Integration

Announcement Type: Initial Announcement Funding Opportunity Number: HR001117S0056 Catalog of Federal Domestic Assistance Numbers (CFDA): 12.910 Research and

Technology Development Dates: (All times listed herein are Eastern Time) o Posting Date: September 13, 2017 o 3DSoC Proposers Day: September 22, 2017 o FRANC Proposers Day: September 15, 2017 o FAQ Submission Deadline: October 23, 2017 at 1:00 PM o Proposal Due Date: November 6, 2017 at 1:00 PM o Estimated period of performance start: May 2018

Concise description of the funding opportunity:

The overall goal of the Three Dimensional Monolithic System-on-a-Chip (3DSoC) program is to develop 3D monolithic technology that will enable > 50X improvement in SOC digital performance at power. 3DSOC aims to drive research in process, design tools, and new compute architectures for future designs while utilizing U.S. fabrication capabilities.

The goal of the Foundations Required for Novel Compute (FRANC) program is to define the foundations required for assessing and establishing the proof of principle for beyond von Neumann compute architectures. FRANC will seek to demonstrate prototypes that quantify the benefits of such new computing architectures.

Anticipated individual awards: Multiple awards are anticipated.

Anticipated funding type: Both programs 6.1 and 6.2 Types of instruments that may be awarded: Dependent on program.

o The 3DSoC program will award procurement contract and other transactions only.

o The FRANC program may award procurement contract, grant, cooperative agreement, or other transaction.

Agency contact:

Dr. Linton Salmon, Program Manager BAA Coordinator:

3DSoC@darpa.mil

DARPA/MTO

ATTN: 3DSoC 675 North Randolph Street Arlington, VA 22203-2114

Dr. Dan Green, Program Manager BAA Coordinator:

FRANC@darpa.mil

DARPA/MTO

ATTN: FRANC

675 North Randolph Street Arlington, VA 22203-2114

PART II: FULL TEXT OF ANNOUNCEMENT

mailto:3DSoC@darpa.mil mailto:FRANC@darpa.mil

I. Funding Opportunity Description

The Defense Advanced Research Projects Agency (DARPA) often selects its research efforts through the Broad Agency Announcement (BAA) process. This BAA is being issued, and any resultant selection will be made, using the procedures under Federal Acquisition Regulation (FAR) 6.102(d)(2) and 35.016 and 2 C.F.R. § 200.203. Any negotiations and/or awards will use procedures under FAR 15.4, Contract Pricing, as specified in the BAA (including DoDGARS Part 22 for Grants and Cooperative Agreements, and Part 37 for Technology Investment Agreements). Proposals received as a result of this BAA shall be evaluated in accordance with evaluation criteria specified herein through a scientific review process.

DARPA BAAs are posted on the Federal Business Opportunities (FedBizOpps) website, http://www.fbo.gov/, and, as applicable, the Grants.gov website at http://www.grants.gov/. The following information is for those wishing to respond to the BAA.

A. Overall Program Description

The ERI Page 3 Materials and Integration thrust has two programs that will operate independent of each other:

Program 1, Three Dimensional Monolithic System-on-a-Chip (3DSoC): Develop 3D monolithic technology that will enable > 50X improvement in SoC digital performance at power.

Program 2, Foundations Required for Novel Compute (FRANC): Develop the foundations for assessing and establishing the proof of principle for beyond von Neumann compute topologies enabled by new materials and integration.

When proposing to multiple programs, proposers should not propose to more than one program in a single proposal. Further information about the programs can be found in Section B (3DSoC) and Section C (FRANC).

http://www.fbo.gov/ http://www.grants.gov/

B. Three Dimensional Monolithic System-on-a-Chip (3DSoC) Program

1. Background

Electronic system performance is increasingly dominated by the time and power required to access system memory. Figure 1 shows the relative contribution of logic delay and memory access delay.

As can be seen from in Figure 1, between 80% and 90% of the execution time is spent in memory access compared to the 10% to 20% of the execution time spent in computation. These numbers were obtained from a state-of-the-art machine learning accelerator and the results are even more dramatic when the computation is executed using a general-purpose processor. From these data, it is clear that any radical improvement in electronic system performance will require a radical reduction in memory access time and power. This limitation is often called the “memory bottleneck.”

Figure 1: Comparison of clock cycles spent in computation and in memory access for several Deep Neural Network algorithms as run on a Machine Learning accelerator executed in 7nm CMOS technology.

One powerful approach to address the memory bottleneck is to integrate the memory and logic in a monolithic 3D SoC stack. This approach dramatically increases the width of buses to memory through much finer-pitch interconnect while simultaneously decreasing the propagation delays through much shorter interconnect lines. This combination can result in memory access bandwidths as high as 40Tb/sec for a 3D SoC as compared with 400Gb/sec bandwidth in a state-of-the-art 2D memory system. Simulations indicate that this increase of 100X in memory access bandwidth can increase 3D SoC compute/memory system performance by several orders of magnitude when compared with 2D compute/memory systems.

In fact, simulations have shown that 3DSoC technology at less capable lithography nodes can still show > 50X improvement over the latest leading-edge 2D CMOS node. Figure 2 shows the simulation results comparing a 7nm 2D implementation of three Long Short-Term Memory (LSTM) machine-learning networks to a 7nm and a 90nm 3DSoC implementation of the same networks. The benefit is calculated as the ratio for each technology of the product of the execution time and energy required to execute the network. As can be seen from the table, the 3DSoC benefit ranges from 645X to 323X for a 7nm 3DSoC technology and from 35X to 75X for a 90nm 3DSoC technology. These results for LSTM machine learning networks are similar to those obtained for other memory intensive applications such as graph analytics.

Figure 2: From Subhasish Mitra of Stanford University

The goal of the 3DSoC program is to exploit the promise of monolithic 3D integration by developing manufacturable 3D process technology and enabling use of that technology to design and fabricate compute/memory systems that provide >50X performance at power than leading-edge CMOS technology nodes. These systems include both traditional microprocessor designs running memory intensive programs and new machine learning and graph analytic accelerators.

The end goal is to develop the process technology in a manufacturing environment, provide the representation of that technology that will enable design of revolutionary 3DSoC designs, develop the EDA design tools required to make those designs, and demonstrate the ability to fabricate and yield high-performing 3DSoCs.

2. Program Description

The goal of the 3DSoC program is to develop the 3D technology required to build logic, memory and I/O on a single die while improving the performance by >50X when compared with 2D 7nm technology. 3DSoC seeks to leverage current industrial and university research in monolithic 3D processes and propel research in the areas of 3DSoC design tools and novel architectures that can be utilized to build highly efficient computation systems.

3DSoC will develop the process, design tools and novel architectures to enable a revolution in the computational capability. The developed approaches will allow low-power, high-compute capability at the leading-edge while also allowing the integration of new high bandwidth compute/memory systems needed to implement advanced machine learning, AI, and communication algorithms. 3DSoC will also drive the development of 3DSoC technology that demonstrates >50X performance at power improvement in compute/memory systems.

Commercial grade fabrication process and design tools will be developed and delivered at the end of the program to be used for DoD and commercial products.

The 3DSoC program encourages proposals with clear, detailed plans to develop novel materials, process techniques, integrated process flows, and application of technology, e.g. digital logic, memory etc. 3DSoC design tools should be able to integrate with current commercial digital design flows to take advantage of the capabilities of 3DSoC technology. Some examples of the required design capabilities are 3D floorplanning, 3D place and route optimization based on design requirements, 3D parasitic extraction, 3D design exploration tools, and 3D memory compilers with the capability to use the high bandwidth 3D memory provided by 3DSoC technology. It is expected that 3DSoC technology will enable revolutionary capabilities for applications in processing, communications, data analytics, AI, mobile computing, and other critical power constrained applications. The end goal of the program is to develop the capability to manufacture 3DSoC technology in the United States.

Out of Scope Technical Areas

The 3DSoC BAA does not solicit research in the following areas:

1. Extensions of 2D CMOS technology

2. Packaging technologies that do not meet the interconnect bandwidth and density goals of the program

3. Sensor fusion

3. Program Structure

This BAA addresses the challenge to develop 3DSoC technology that enables > 50X performance at power improvement over traditional 2D 7nm CMOS technology. The intention is that the > 50X improvement will be driven by the inherent 3D nature of the technology, the low-parasitic and massive interconnectivity between 3D layers, and the reduction in memory access time/power that the technology will provide.

To achieve these goals, the program is divided into three phases.

Phase 1 (18 months): Proof of feasibility and initial set up of the technology Phase 2 (12 months): Increase in yield and implementation of the technology for select designs Phase 3 (12 months): High-yield manufacturing and implementation for multiple designs

At the conclusion of the program, the 3DSoC technology requested in this BAA should have the following characteristics:

1. Capability of > 50X the performance at power when compared with 7nm 2D CMOS technology.

2. Interconnect densities > 3K interconnects/mm (9M interconnects/mm2) between 3D layers.

3. Interconnect bandwidth > 50Tb/s between 3D layers.

4. Memory access energy < 2pJ/bit.

5. Inclusion of > 4GB of non-volatile memory in a monolithic SoC that has a 2D footprint of no more than 200mm2 and dissipates < 500mW of average operating power.

6. Provision for logic densities > 1M gates/mm2 (as measured in a 2D projection) across multiple logic layers in the 3D stack. Note that the process temperatures required for the 3D layers must be low enough not to compromise earlier logic or memory layers in the 3D stack.

7. 3D transistor performance approximately equal to or exceed that of standard 90nm CMOS transistors:

a. NFET: Ion=640µA/µm at Ioff=5nA/µm, ft=100GHz, Subthreshold Slope=90mV/dec.

b. PFET: Ion=280µA/µm at Ioff=5nA/µm, Subthreshold Slope=90mV/dec.

8. Projected fabrication cost comparable to 7nm 2D technology.

9. Capability to fabricate, at reasonable yield, 3DSoCs with > 4GB (Giga-Byte) of memory and > 50M logic gates.

10. Capability to be fabricated as an on-demand foundry service for multiple DoD and commercial customers using a U.S. fabrication facility.

Technical Areas

This 3DSoC BAA is soliciting proposals in three technical areas:

Technical Area 1 (TA-1) will develop the 3D monolithic manufacturing process that will be utilized to build DoD and commercial SOCs. The end goal of this technical area is to be prepared to offer the 3D monolithic process as a commercial foundry offering to external users, including defense contractors, for fabrication of relevant large-scale compute/memory systems.

Technical Area 2 (TA-2) will focus on developing a test chip worthy SoC that will facilitate process development, drive yield improvement, and exercise the design elements required for realization of the program goals for SoC performance improvement.

Technical Area 3 (TA-3) will develop the EDA design tools required to enable design of large-scale compute/memory systems that utilize 3D monolithic processes to achieve the performance at power goals of the program.

Proposers proposing to TA-1 or TA-2 must propose to both technical areas, TA-1 + TA-2.

This is to ensure that the design of the Development Evaluation Circuit test chip is strongly connected with the team developing the 3DSoC process. Proposers to TA-3 may propose to TA- 3 alone or to all three technical areas, TA-1 + TA-2 + TA-3.

4. TA-1: Develop a 3D monolithic fabrication process capable of high-yield fabrication of complex compute/memory systems

TA-1 proposers should propose process technology that will enable the goal of a 50X improvement in overall performance at power for complex compute/memory systems when compared with comparable systems implemented in leading-edge 2D CMOS technology, e.g. 7nm. Proposals should contain a clear achievable plan to accomplish the following tasks.

1. Develop the process technology from initial module development through final, high-yielding compute/memory system fabrication.

2. Drive yield improvement activities to achieve yield targets listed in Table 2.

3. Establish a technology enablement infrastructure that will enable design of compute/memory systems using the technology. The infrastructure will include delivery of a full Process Design Kit (PDK) that includes device and extraction models, design rules, and fabrication foundation IP (bit cells, etc.).

4. Demonstrate the ability to provide foundry services for the 3DSoC technology by providing regular Multi-Project Wafer (MPW) runs for 3DSoC technology users. Design groups will use the published PDK to create designs exhibiting the characteristics outlined in the program plan.

5. Accurately predict improvements in compute/memory systems achievable through use of the process technology.

Table 1. TA-1 Deliverables by Phase Task Phase 1 Phase 2 Phase 3

Process Technology

1. Process module description

2. Initial integrated process description

1. Demonstration of integrated process flow

2. Cost of fabrication model for the 3DSoC technology

1. Utilization of integrated, high-yielding process flow to successfully fabricate multiple 3DSoC designs

Yield Improvement

1. Yield improvement plan in place

2. Measurement test plan in place

1. Successful fabrication of Development Evaluation Chip (DEC)

2. Successful fabrication of initial 3DSoC demonstration circuits

1. Successful fabrication of multiple 3DSoC designs

Technology Enablement

1. Delivery of a V0.5 (Beta) PDK that is hardware-based and includes at a minimum;

a. Transistor and memory cell models

b. Parasitic extraction models

c. Design rules

1. Delivery of a V1.0 (Design Ready) PDK that is hardware-verified

2. Development of the enablement infrastructure required for initial 3DSoC demonstration circuit design

1. Refinement of PDK and other enablement infrastructure

Foundry Service 1. Prepare for MPW runs in Phase 2

1. Establish the infrastructure required for MPW participation

1. Refinement of MPW infrastructure

2. Schedule and fabricate 2 MPW runs

System Performance Evaluation

1. Provide hardware performance comparison based on 3DSoC simulations

1. Provide hardware performance comparison based on 3DSoC test chip (DEC) hardware

1. Provide hardware performance comparison based on 3DSoC demonstration hardware

Table 2. TA-1 Metrics by Phase

Process Technology

1. > 90% of all process modules demonstrated

2. Initial integrated fabrication lots complete

1. 100% of all process modules demonstrated

2. Full integrated process demonstrated

1. Full integrated process demonstrated on 3DSoC designs

Yield Improvement

1. > 60% yield on representative 3DSoC circuit building blocks

2. Sufficient yield on test structures to predict 30% yield at final DEC circuit complexity

1. > 30% yield on the Development Evaluation Circuit (DEC)

2. Successful fabrication of initial 3DSoC demonstration circuits

1. > 60% yield on the Development Evaluation Circuit (DEC)

2. > 30% yield on initial 3DSoC demonstration circuits

Technology Enablement

1. < 10% difference between measured test structure results and PDK prediction

1. < 5% difference between measured test structure results and PDK prediction

1. < 2% difference between measured test structure results and PDK prediction

Foundry Service 1. Full MPW infrastructure in place

1. < 6 month MPW turn-around time from design delivery to chip delivery

1. < 4 month MPW turn-around time from design delivery to chip delivery

System Performance Evaluation *

1. > 50X compute/memory performance at power improvement over 2D 7nm technology predicted through simulation

1. > 10X compute/memory performance at power improvement over 2D 7nm technology demonstrated using the DEC

1. > 50X compute/memory performance at power improvement over 2D 7nm technology demonstrated with 3DSoC demonstration circuits

* Performance at Power (PaP) can be calculated as performance at a given power level, power at a given performance level, or by calculating [power*execution time]-1.

As noted above, the 3DSoC technology demonstrated at the end of the program should also have the following characteristics:

1. Capability of > 50X the performance at power when compared with 7nm 2D CMOS technology.

2. Interconnect densities > 3K interconnects/mm (9M interconnects/mm2) between 3D layers.

3. Interconnect bandwidth > 50Tb/s between 3D layers.

4. Memory access energy < 2pJ/bit.

5. Inclusion of > 4GB of non-volatile memory in a monolithic SoC that has a 2D footprint of no more than 200mm2 and dissipates < 500mW of average operating power.

6. Provision for logic densities > 1M gates/mm2 (as measured in a 2D projection) across multiple logic layers in the 3D stack. Note that the process temperatures required for the 3D layers must be low enough not to compromise earlier logic or memory layers in the 3D stack.

7. 3D transistor performance approximately equal to or exceed that of standard 90nm CMOS transistors:

a. NFET: Ion=640µA/µm at Ioff=5nA/µm, ft=100GHz, Subthreshold Slope=90mV/dec

b. PFET: Ion=280µA/µm at Ioff=5nA/µm, Subthreshold Slope=90mV/dec

8. Projected fabrication cost comparable to 7nm 2D technology.

9. Capability to fabricate, at reasonable yield, 3DSoCs with > 4GB (Giga-Byte) of memory and > 50M logic gates.

10. Capability to be fabricated as an on-demand foundry service for multiple DoD and commercial customers using a U.S. fabrication facility.

Performance at Power (PaP) can be calculated as performance at a given power level, power at a given performance level, or by calculating [power*execution time]-1. PaP can be improved by both process improvements and design improvements. In TA-1, proposers will be responsible for demonstrating the PaP advantages due to process improvements and projecting, using data from the Design Evaluation Chip (DEC), the PaP advantages due to design improvements that are enabled by 3DSoC technology. These projections will be validated with full system designs, but the TA-1 proposers are responsible for the 50X PaP improvement metrics listed in Table 2.

Some possible areas of improvement through process enhancement are: new 3D transistor design technology, new memory design technology, and new high-bandwidth interconnect technology.

Possible areas of design improvements enabled by 3DSoC technology are: high-bandwidth connections between 3D logic and 3D memory, architectures built to accommodate new memories, use of non-volatile memories, efficient placement of blocks, power routing improvements, and reduced routing due to 3D architectures.

5. TA-2: Fabrication Development Test Chip SoC Design

The focus of TA-2 is to design and test a test chip SoC that will facilitate process development, drive yield improvement, and exercise the design elements required for realization of the program goals for SoC performance improvement. TA-2 is comprised of a single key task that will evolve across the first two phases of the program:

1. 3DSoC DEC Design: This task will be the design of two versions of the DEC. The first design will be a preliminary design based on the initial V0.1 PDK provided in TA-1 and designed to enable the yield improvement activities of Phases 1 and 2. The second design will be a respin of the DEC that utilizes V0.5 of the PDK provided in TA-1 and will be expanded to enable the more aggressive yield improvement activities of Phases 2 and 3.

Table 3. TA-2 Metrics and Deliverables by Phase Task/Metric Phase 1 Phase 2

DEC Design Metrics

1. > 100MB of memory on 3 or more 3DSoC layers

2. > 20M gates of logic on 3 or more 3DSoC layers

3. At least 1 microprocessor core of the complexity of a RISC-V or ARM Axx

4. Implementation of a complex logic function such as DSP or DNN

1. > 1GB of memory on 3 or more 3DSoC layers

2. > 50M gates of logic on 3 or more 3DSoC layers

3. At least 1 microprocessor core of the complexity of a RISC-V or ARM Axx

4. Implementation of a complex logic function such as DSP or DNN

DEC Design Task

1. Complete initial DEC design using the V0.1 (initial) PDK

1. Complete DEC design using the V0.5 (beta) PDK

6. TA-3: Development of 3DSoC Design EDA Tools

The focus of TA-3 is to develop the EDA tools that will be utilized to design demonstration SOCs that utilize the 3D monolithic process to explore and enable new digital architectures. The TA-3 effort should be proposed in a way that could support multiple TA-1/TA-2 teams. TA-3 is comprised of three key tasks that will evolve across the three phases of the program:

1. 3DSoC Logic EDA Tool Development: The goal of this task is to develop the EDA tools required to address the design, partitioning, and place/route challenges introduced by the intimate connection between multiple 3D layers of logic in 3DSoC technology.

2. 3DSoC Design Memory EDA Tool Development: The goal of this task is to develop the EDA tools required to design, compile, and connect large scale monolithic memories that span multiple layers of memory in the 3DSoC technology.

3. Monolithic 3DSoC Design EDA Tool Support: The goal of this task is to develop the EDA design flow required to exploit the revolutionary capabilities of 3DSoC technology to design and fabricate greatly more capable 3D compute/memory systems.

Table 4. TA-3 Deliverables by Phase

3DSoC Logic EDA Tool Development

1. Define an EDA flow that will place and route logic monolithically across 3DSoC layers

2. Distribute the EDA tools to the development team

1. Provide an EDA flow that can be used by multiple external 3DSoC design teams

1. Provide an EDA flow that is generally available for external use, specifically by DoD users

3DSoC Memory EDA Tool Development

1. Define an EDA flow that will compile, place, and connect memories with logic and across 3DSoC layers

2. Distribute the EDA tools to the development team

1. Provide an EDA flow that can be used by multiple external 3DSoC design teams

1. Provide an EDA flow that is generally available for external use, specifically by DoD users

Monolithic 3DSoC Design EDA Tool Development

1. Define an EDA flow that will design an entire 3DSoC that spans across the entire 3DSoC stack

2. Distribute the EDA tools to the development team

1. Provide an EDA flow that can be used by multiple external 3DSoC design teams

1. Provide an EDA flow that is generally available for external use, specifically by DoD users

Table 5. TA-3 Metrics by Phase

3DSoC Logic EDA Tool Development

EDA flow capable of designing the DEC chip

EDA flow capable of designing a 500M gate 3DSoC design that meets program metrics

EDA flow that is sufficiently robust to be used by multiple external users to design 3DSoCs that meet program metrics

3DSoC Memory EDA Tool Development

EDA flow capable of designing the DEC chip

EDA flow capable of designing 4GB of memory in a 3DSoC design that meets program metrics

EDA flow that is sufficiently robust to be used by multiple external users to design 3DSoCs that meet program metrics

Monolithic 3DSoC Design EDA Tool Development

EDA flow capable of designing the DEC chip

EDA flow capable of designing a full 3DSoC design with 4GB of memory and 500M logic gates that meets program metrics

EDA flow that is sufficiently robust to be used by multiple external users to design 3DSoCs that meet program metrics

7. Program Schedule

The 3DSoC program schedule comprises an 18-month base period (Phase 1) followed by two 12-month periods (Phases 2 and 3 respectively), for a total of 42 months, subject to the availability of funds.

Figure 3: 3DSoC Program Schedule

Phase 1: Base Period (18 months)

TA-1: The initial steps of developing the process technology will be completed. Process modules will be developed, an initial integrated process flow will be demonstrated, a V0.5 PDK will be released, technology capability will be measured through simulation, and yield feasibility will be demonstrated.

TA-2: The first-pass Development Evaluation Circuit (DEC) will be designed, fabricated, and tested.

TA-3: Initial 3DSoC EDA flows will be developed that enable design of combined logic and memory systems that are intimately connected across multiple layers of the 3DSoC. The EDA flow will be released and evaluated by appropriate development teams.

Phase 2: Option 1 (12 months)

TA-1: The process technology integration will be completed and a yield ramp will be started using the DEC designed in Phase 1. A V1.0 PDK will be released to the design community and technology capability will be measured through test chip hardware, including the DEC.

TA-2: The second-pass DEC will be designed, fabricated and tested.

TA-3: 3DSoC EDA flows will be released and supported for a wider, but limited, set of design groups, including groups designing demonstration 3DSoC compute/memory systems.

Phase 3: Option 2 (12 months)

TA-1: Process technology integration will be verified on the DEC/demonstration 3DSoC designs and reasonable yield will be demonstrated. A final PDK will be released to the design community and technology capability will be measured using the DEC and using demonstration

3DSoC designs. 3DSoC MPW runs will be fabricated and made available to the general design community.

TA-3: 3DSoC EDA flows will be released and supported for the general design community, including groups designing 3DSoC compute/memory systems for inclusion on the MPW runs.

C. Foundations Required for Novel Compute (FRANC) Program

1. Background

From the beginning of Moore’s law until its current node, materials and integration technology have been central to its progression. The progression from initial development of high-purity silicon crystal growth and understanding of the utility of its native oxide interface up to today’s engineering of strained silicon and use of high-k dielectrics demonstrates the leading role materials innovation has played in the advance of compute electronics. Similarly, the evolution from Jack Kilby’s recognition of the value of connecting transistors together on chip to the emergence of today’s 2.5D and 3D manufacturing schemes shows the value of integration driven by the ability to realize new architectures. Thus, advanced materials and integration technology is seen as fertile territory in which to explore for new gains in compute capability.

While the potential exists to continue to innovate in this arena within conventional compute paradigms, the key limiting factor is the “memory bottleneck” or the fact that as processing becomes faster, more time and energy is spent transferring data between memory and computing (i.e., data latency) than with the computation itself. This limit is driven by conventional von Neumann architecture, which relies on the separation of memory and logic to facilitate computing through the operation of logic processers on separate memory stores.

The opportunity to overcome this barrier hinges on locating data closer to processors. Here “close” means both temporally close–data can move quickly between processors and memories, and energetically close–only a small amount of power is needed to transmit a bit between processors and memories. For instance, component technologies such as chip scale photonic components offer the opportunity to enable more efficient data movement pathways such that more data can be moved more efficiently, effectively making physically distant memory stores “closer” to processors, though latency can still be an issue. In this direction, industry is pursuing 2.5D integration schemes such as high bandwidth memory (HBM) to address this option.

However, off-chip memory performance gains have lagged processing gains, which still limits performance and adds complexity to design such as hierarchical cache schemes are needed to optimize data availability.

A more aggressive means to reduce latency is to fabricate systems in three dimensions such as anticipated by 3DSoC above. An advantage of this approach is that it preserves the von Neumann compute architecture so it is able to leverage existing approaches to computing and existing investments in software tools. A challenge is the greater complexity needed for 3D processes and risk of the process development required over existing 2D and 2.5D schemes.

While 3DSoC presents a worthy challenge, the fundamental data movement bottleneck of von Neumann compute architectures remain. Hence, the opportunity for a beyond von Neumann compute topology is recognized, especially where circuits can leverage new materials and integration schemes to enable data to be processed in ways that minimize or eliminate data movement.

Compute topologies that minimize data movement would benefit greatly from new material options. For instance, new materials may enable non-volatile memory, which underpins new topologies that do not require the traditional von Neumann separation of memory. However, even conventional logic transistors have resulted in a wide range of elements used in fabrication in a growing number of combinations and larger material sets are likely be needed in future designs. This growing set of options will lengthen an already long process between material discovery and device production. Accelerating the process of material discovery and development will help ensure that new materials and combinations continue to fuel advancements in component technology. Focused exploration of materials that allow for these new compute topologies based on new device physics will help accelerate to specific solutions.

In addition to materials, the advent of 2.5D integration technologies allows for the contemplation of new topologies that shift the location of compute. For instance, processing in memory reduces the need to transfer data from memory by shifting the compute to memory. While most useful algorithms will require the movement of data at some point, there are significant potential for approaches where portions of the computation can be localized. Specialized algorithms and architectures may be able take advantage of processing in memory to improve overall system performance. This approach may be particular advantageous to memory-side learning in support of neuromorphic computing, for example. Other non-von Neumann topologies based on non- Boolean logic have been shown such as, for example, using coupled oscillators.2 Using unconventional components and architectures can target specialized problem classes such as those requiring non-linear optimization.

While the benefits of new materials and circuit topologies drive toward integrated solutions, the advance to denser circuits operating at higher currents places increasing emphasis on the power efficiency needed to mitigate the increased heat load that must be drawn from a chip over each clock cycle. Materials and integration also play key role in power distribution, power management, and heat dissipation. Chips having higher current demands with multiple voltage levels can lead to inefficient approaches to power conversion and distribution. Distribution and optimization of power components on chip, component material optimization, and other such approaches for reducing power overall chip power dissipation have significant opportunity to contribute performance gains. Component technologies such as power management that support this overall topology will be essential to realize this vision for new topologies that leverage new materials and integration technology.

The anticipated increase in computing efficiency is of interest to many entities in addition to the DoD: other defense establishments, internet-of-things developers, and designers of personal electronics including medical electronics. However, the problems have been addressed in the past with conservative approaches such as better memory interfaces including high bandwidth memory (HBM), which has shown near-term benefits, but little chance of achieving great improvements in computing efficiency. The FRANC program aims to realize these revolutionary advancements made possible by new materials, new integration techniques, and new physical architectures.

2 See, for example, A. Parihar, et al., “Vertex coloring of graphs via phase dynamics of coupled oscillatory networks,” Nature Scientific Reports, Article No. 911 (2017); D. Nikonov, et al., “Coupled-Oscillator Associative Memory Array Operation for Pattern Recognition,” IEEE Journal on Exploratory Solid-State Computational Devices and Circuits, Vol. 1 (2015); V. Purahmad, et al., “Nonboolean Pattern Recognition Using Chains of Coupled CMOS Oscillators as Discriminant Circuits,” IEEE Journal on Exploratory Solid-State Computational Devices and Circuits, Vol. 3 (2017).

2. Program Description

The purpose of the FRANC program is to provide the foundation for new materials technology and new integration approaches to be exploited in pursuit of novel compute architectures. The program aims to develop alternate compute topologies that change the compute paradigm from discrete memory and processing to architectures that enable processing to happen where the data is stored with structures that diverge dramatically from conventional digital logic processors, thus allowing for more dramatic gains in compute performance while minimizing the challenges associated with vertical integration.

The FRANC program will seek to demonstrate prototypes that quantify the benefits of this new computing architecture. The benefits of new materials and/or integration will be central to this effort. Monolithic solutions using commercially available processes will be considered non-responsive. Similarly, solutions that encompass design only considerations will be considered non-responsive.

Changes in computing architectures through the FRANC program creates the opportunity for new technology choices across the system hierarchy from new circuit topologies down to the constituent device technologies that drive performance. Examples of technologies of interest include processing chips that incorporate distributed memory, memory chips that incorporate distributed processing, components that exploit device physics to compute approximate solutions to NP-hard problems, innovative power provisioning methods, and wholly new memory designs.

As noted above, only components that provide revolutionary advancements in computing efficiency will be considered. Additionally, a new architecture will have ramifications for ease of programming and broad utility of use.

For that reason, FRANC will focus on two technical areas: TA-1, New Topology circuit prototypes, and TA-2, Building Blocks for new compute.

Proposers to the FRANC program may propose to both TAs if they so choose; however, proposers will be required to submit separate proposals for each individual TA.

3. Technical Areas

As noted above, the FRANC program consists of two technical areas: TA-1: New Topology Circuit Prototypes, and TA-2: Building Blocks. TA-1 contains the sub-thrust TA-1b, Accelerators.

Technical Area 1: New Topology Circuit Prototypes

TA-1 New Topology circuit prototypes will exploit recently emerged materials and/or integration technologies to create an integrated processing and memory systems with revolutionary capabilities. The selection of materials, components, fabrication, and packaging are up to the proposers. However, the proposed circuit prototype must diverge from a conventional Von Neumann compute architecture. Additionally, the completed system must advance over the current state of the art by at least a factor of ten, measured by performer-specified metric such as energy-delay product. Monolithic designs that do not leverage new materials or integration processes will not be considered responsive.

While proposers will provide the benchmarks to evaluate their targeted design, proposers are encouraged to utilize the standard algorithm set associated with a number of relevant missions developed by Pacific Northwest National Lab (PNNL) to facilitate benchmarking; see http://hpc.pnnl.gov/PERFECT/ and http://hpc.pnnl.gov/SEAK/ for more information.

Additionally, the proposers should describe the broad based nature of the performance benefits of any new architecture (i.e. if the topology is particularly suited to a particular workload or broadly applicable).

TA-1 proposers are expected to have a preliminary basis for anticipating a performance benefit to a new topology. TA-1 efforts will then pursue an initial study period to clarify the circuit architecture and provide more robust analysis of expected performance gains. The initial study will also provide estimates of recurring and non-recurring costs to implement components in the proposed new architecture.

The most attractive circuit prototype or prototypes from this initial study period will advance from study phase to a design development phase. The prototype circuits will be matured to final tape-out. Upon successful design review at the end of the second phase, the prototype designs will be fabricated and tested.

TA-1b New Topology: Accelerators is a sub-thrust of this program element. The possibility of non-von Neumann accelerators or processing circuits that participate in an otherwise conventional von Neumann architecture is permissible. The performance benefits of these beyond von Neumann accelerators will need to be quantified while accounting for the overhead of integration into the balance of the compute platform. It is anticipated that accelerator efforts will be significantly more modest that a full New Topology circuit demonstration yet must still abide by the timeline and milestones for the technical area.

Technical Area 2: Building Blocks

TA-2 Building Blocks addresses component and subsystem developments to support future computing systems. This technical area is intended to enable development efforts for component technologies that support novel compute topologies but do not comprise a complete solution.

Because of the potential for a broad interpretation of this technical area, it is essential that proposals address how any proposed component technologies address the key challenges of this call. In particular, successful proposals will address the development of new materials for components or technologies that enable future 2.5D or 3D integrated solutions in the context of beyond von Neumann compute topologies. Additionally, technologies are sought that are close to commercialization so that significant cost share is targeted to make sure there is commitment to develop the technology on the part of the proposer. Proposers are strongly encouraged to participate in cost sharing in this technical area.

http://hpc.pnnl.gov/PERFECT/ http://hpc.pnnl.gov/SEAK/

These efforts may address, but are not limited to the following topics. These listed topics are recommended but not guaranteed to be addressed.

Technologies and approaches to accelerate material discovery, development, and component optimization. The number of material combinations available for component development presents a challenge for rapidly testing new concepts. Thin-film materials need to be deposited and characterized in combination to determine electrical behavior for optimal designs. Thus, a strong need exists to be able to collect a wide range of relevant process, material, electrical, and mechanical data in situ, and to do so both rapidly and inexpensively. These data can then support predictive approaches for optimizing component and systems. Proposals targeting these areas in whole or in part will be considered.

Non-volatile memory (NVM) components. Many approaches to NVM currently exist with no clear leading technology for commercial use. Demonstration promising new NVM technology for specific uses and providing chip-scale demonstrations of their utility will accelerate their acceptance. Technologies of interest go beyond current state of the art, hence, solutions such as ReRAM based on conductive bridge or conventional phase change materials (e.g. GeTe) will not be considered responsive.

Power Management for ICs (PMIC). Solutions to reduce the overall power footprint of ICs is of interest. Heterogeneously integrated power within chips, integrated power components, and islands of power within an IC are potential areas of development.

Photonic components. Innovative approaches to optical transceiver modules for incorporation into state-of-the-art digital SoC solutions are of interest to expand data pathways. Approaches with lower energy consumption and higher aggregate bandwidth beyond that of electrical current interfaces integrated at chip scale are of interest.

Pathways to demonstrate performance in DoD-relevant systems and to realize volume manufacturing, packaging, and system integration should be addressed.

A selectable proposal will fully answer the following: technical merit, potential impact on existing or future computing systems, and feasibility of implementation.

4. Program Structure

Overall, the FRANC program consists of three phases, of 6, 18 and 24 months duration. The initial 6 month phase is intended to allow for a small, focused team to complete the more detailed analysis with a scale consistent with seedling level program.

5. Schedule/Milestones

The FRANC program contains three phases lasting a total of 48 months.

The deliverables/gates of the TA-1, New Topology Circuit Prototypes, are as follows:

Due Milestone End of Phase 1 – 6 months ARO Preliminary design. Performance analysis.

End of Phase 2 – 24 months

ARO

Detailed design.

This is the start of the file's text. The full file is on GovTribe.

File details come from the government source that posted it. Updated .