Appendix B - High Performance Computing Cluster Specifications.pdf

PDF 126 KB Posted

Attached to
Request for Proposal - High Performance Computing Cluster Federal contract opportunity
Solicitation number
A24-1153
Issued by
Department of Energy Ames Laboratory

About this file

This document is Appendix B - High Performance Computing Cluster Specifications for RFP A24-1153 issued by Ames National Laboratory, which is operated by Iowa State University under a prime contract with the U.S. Department of Energy.

The RFP seeks to acquire a High-Performance Computing (HPC) cluster to partially replace an existing cluster at the laboratory. The specifications detail the requirements for a login node, 12 CPU nodes, GPU nodes, and associated network infrastructure. The CPU nodes have options for 64, 96, 128 (preferred), or 192 cores per node. The GPU nodes include an NVIDIA A100 GPU. The HPC cluster requires a non-blocking HDR 200 Infiniband fabric, Ethernet switches for provisioning and management, and other technical requirements. All equipment must be new and sourced to maintain valid manufacturer warranties. The RFP instructs offerors to price each line item individually. The HPC cluster will undergo a 2-week acceptance testing period prior to final acceptance.

View the file

Other files for this federal contract opportunity

Other files attached to Request for Proposal - High Performance Computing Cluster, newest first.
File Type Posted
Addendum 1 - RFP A24-1153 (With Attachments).pdf PDF
A24-1153 HPC Cluster RFP Solicitation.pdf PDF
Appendix A - Term and Conditions for FP Commercial Products and Commercial Services dated 3-12-24.pdf PDF

On GovTribe

Work with this file on GovTribe

  • Download the original file
  • Contacts named in this file
  • Similar government files
  • Ask GovTribe AI about this file

Text version

RFP A24-1153 3/19/2024 Rev 0

Appendix B – High Performance Computing Cluster Specifications

Ames National Laboratory is seeking a High-Performance Computing (HPC) cluster as a partial replacement of an old/preexisting cluster at the laboratory. The following specifications below include a login-node, CPU and GPU nodes, and associated network infrastructure as specified.

A. Head/Login Node, Qty. 1

2U rack-mount server. 32 CPU Cores 3.0GHz, 384GB Memory:

• 2U rack-mount chassis with (24) 2.5" hot-swap drive bays (16 NVMe + 8 NVMe/SATA/SAS) and redundant power supplies.

• Dual socket motherboard for AMD Epyc™ 9004 Processor, 24 DDR5-4800 DIMM slots, 2 M.2 SATA 6Gbps/PCIe Gen 4 x4, 5 PCI-E 5.0 x16.

• 2x AMD Epyc™ 9124 processors (3.0GHz, 16 cores/32 threads, 64MB L3, DDR5-4800, 200W TDP, 200W-240W cTDP).

• 2x High-grade CPU heat sink.

• 24x 16GB DDR5 4800MHz ECC RDIMM (384GB total).

• 2x 960GB Data center U.3 NVMe SSD (2.5", PCIe 4.0 x4, 1DWPD) software RAID1 boot drive.

• 1x 6.4TB Data center U.3 NVMe SSD (2.5", PCIe 4.0 x4, 3DWPD) scratch drive.

• 7x 6.4TB Data center U.3 NVMe SSD (2.5", PCIe 4.0 x4, 3DWPD) software RAID6 local home drive.

• Integrated Gigabit Ethernet RJ45 port for out-of-band remote system management.

• 10G-On Board Integrated dual-port 10GBaseT Ethernet LAN controller (RJ45).

• 10G-Add On Dual-port 10GBaseT network adapter card (RJ45).

• 1G-4 Port Quad-port 1GBaseT network adapter card (RJ45).

• 2x HDR 100 NVIDIA dual-port ConnectX-6 HDR100/200GbE VPI network adapter card with Nvidia registered 3-year warranty.

• Rocky Linux 8 or 9 System OS installed for burn-in and acceptance testing.

Customer will reinstall own OS image.

• 2x 2600W 80+ titanium level power supply, redundant.

• 2x C19 to C14 power cable (2ft).

• Rack-mount rail kit included with chassis.

• 3-Year Warranty with advanced parts replacement.

B. CPU Nodes, Qty. 12

1U rack-mount compute node with 2x AMD Epyc™ (see item C for CPU specs), 768GB memory:

• 1U rack-mount chassis with (12) 2.5" hot-swap NVMe/SATA/SAS drive bays and redundant hot-swap power supplies.

• Dual socket motherboard for AMD Epyc™ 9004 Processor, 24 DDR5-4800 DIMM slots, 2 M.2 SATA 6Gbps/PCIe Gen 4 x4, 3 PCI-E 5.0 x16.

• Processor Please see CPU Options (see item C)

• 2x High-grade CPU heat sink.

• 24x 32GB DDR5 4800MHz ECC RDIMM (768GB total).

• 2x 960GB Data center U.3 NVMe SSD (2.5", PCIe 4.0 x4, 1DWPD) software RAID1 boot drive.

• 1x 6.4TB Data center U.3 NVMe SSD (2.5", PCIe 4.0 x4, 3DWPD) scratch drive.

• Integrated Gigabit Ethernet RJ45 port for out-of-band remote system management including KVM-over-LAN and Serial-over-LAN.

• Integrated quad-port 10Gb Ethernet LAN controller (RJ45).

• 1x HDR 100 NVIDIA single-port ConnectX-6 HDR100/100GbE VPI network adapter card with Nvidia registered 3-year warranty.

• Rocky Linux 8 or 9 System OS installed for burn-in and acceptance testing.

Customer will reinstall own OS image.

• 2x 2600W 80+ titanium level power supply, redundant.

• 2x C19 to C14 power cable (2ft).

C. CPU options (Add to CPU Nodes) please price out individually:

(1) 64 cpu cores per node:

2x AMD Epyc™ 9334 (2.7GHz, 32 cores/64 threads, 128MB L3, DDR5-4800, 210W TDP, 200W-240W cTDP).

(2) 96 cpu cores per node:

2x AMD Epyc™ 9454 (2.75GHz, 48 cores/96 threads, 256MB L3, DDR5-4800, 290W TDP, 240W-300W cTDP).

(3) *128 cpu cores per node:

2x AMD Epyc™ 9534 (2.45GHz, 64 cores/128 threads, 256MB L3, DDR5-4800, 280W TDP, 240W-300W cTDP).

(4) 192 cpu cores per node:

2x AMD Epyc™ 9654 (2.40GHz, 96 cores/192 threads, 384MB L3, DDR5-4800, 360W TDP, 320W-400W cTDP).

*[preferred CPU option #3 (128core per node)]

D. GPU Nodes, Qty. 5

2U rack-mount GPU nodes. 16 CPU Cores 3.0GHz, 1 NVidia A100 (80GB) GPU, 384GB Memory:

• 2U rack-mount chassis with (24) 2.5" hot-swap drive bays (16 NVMe + 8 NVMe/SATA/SAS) and redundant power supplies.

• Dual socket motherboard for AMD Epyc™ 9004 Processor, 24 DDR5-4800 DIMM slots, 2 M.2 SATA 6Gbps/PCIe Gen 4 x4, 5 PCI-E 5.0 x16.

• 1x AMD Epyc™ 9124 processor (3.0GHz, 16 cores/32 threads, 64MB L3, DDR5-4800, 200W TDP, 200W-240W cTDP).

• 1x High-grade CPU heat sink.

• 24x 16GB DDR5 4800MHz ECC RDIMM (384GB total).

• 2x 512GB NVMe M.2 SSD software RAID1 boot drive.

• 1x 3.2TB Data center U.3 NVMe SSD (2.5", PCIe 4.0 x4, 3DWPD) scratch drive.

• Integrated Gigabit Ethernet RJ45 port for out-of-band remote system management.

• Integrated dual-port 10GBaseT Ethernet LAN controller (RJ45).

• 1x HDR 100 NVIDIA single-port ConnectX-6 HDR100/100GbE VPI network adapter card with Nvidia registered 3-year warranty.

• 1x NVIDIA A100 (80GB HBM2e, PCIe 4.0, 300W) with 3-year manufacturer warranty.

• Rocky Linux 8 or 9 System OS installed for burn-in and acceptance testing.

Customer will reinstall own OS image.

• 2x 2600W 80+ titanium level redundant power supply, redundant.

• 2x C19 to C14 power cable (2ft).

E. Non-Blocking HDR 200 Infiniband Fabric

Non-blocking HDR 200 Infiniband fabric:

• NVIDIA 200Gb/s HDR Infiniband switch with redundant 850W power supplies (200Gb/s, 40x QSFP56, unmanaged).

• 2x C13 to C14 power cord.

• 9x NVIDIA 200Gb/s to 2x100Gb/s HDR-2xHDR100 Infiniband passive copper splitter cable (2m).

• 2x NVIDIA 100Gb/s EDR Infiniband active optical cable (30m).

• Fabric configuration and testing.

• NVIDIA HDR Infiniband switch 3-year NVIDIA Silver warranty, firmware support.

F. Cluster Management and IPMI Networks

Ethernet switches for in-band node provisioning and management network (ssh, pdsh, ganglia, …) and out-of-band IPMI network:

• 48-Port 10GbE managed switch (48x RJ45, 4x SFP+).

• 48-Port 1GbE managed switch (48x RJ45, 4x SFP).

• 2x C13 to C14 power cord.

• 20x Cat-6A Ethernet cable (red, 15ft).

• 20x Cat-6A Ethernet cable (white, 15ft).

• 1x Multimode fiber optical cable (50/125 OM3, LC to LC, Aqua, 30m).

• 2x Cisco 10GbE SFP+ transceiver (LC, short range).

• Network configuration and testing.

G. Other

• All memory modules must be from the same batch with the same batch number. Memory brand and part number must be on the approved list from the motherboard manufacturer.

• Supply Chain: All equipment must be new, unused, and sourced such that all manufacturer warranties and support are valid.

• Servers must meet processor manufacturer specifications for supported cTDP. Server internal cooling must conform to process processor manufacturer specifications.

• PXE bios support: BIOS must allow servers to boot without keyboard or monitor; must have temperature & fan monitoring capability; remote power up/down capability and serial over LAN.

• All server nodes must have the uniform version of all embedded firmware (BIOS, BMC, Ethernet, InfiniBand, disk drives, etc.)

• IPMI for integrated cluster management capabilities is required. All IPMI sensor data repositories (SDR’s) must be preloaded so that sensor data (temperatures, fan speeds, voltages) can be read via IPMI from each computer server.

• All servers will go through a 2-week acceptance testing at which time a set of scientific codes per Ames Lab discretion will be run in a loop, independently, across all servers that are a part of the cluster. A server is deemed as accepted if it can record consistent and expected performance on the before mentioned codes.

• MAC addresses for all servers are required to submit to Ames Lab prior shipping.

• Shipping Freight prepaid and insured. Shipping Terms: Destination PPA or DDP AMES PPA.

• Provide applicable Export Classification Control Number (ECCN) designation for all line items as applicable.

• Please price out all line items (A-F) individually.

A. Head/Login Node, Qty. 1
B. CPU Nodes, Qty. 12
C. CPU options (Add to CPU Nodes)
D. GPU Nodes, Qty. 5
E. Non-Blocking HDR 200 Infiniband Fabric
F. Cluster Management and IPMI Networks
G. Other

File details come from the government source that posted it. Updated .