SOW_RFI.docx

DOCX document 208 KB Posted

Attached to
Extreme Scale Spatio-Temporal Learning Federal contract opportunity
Solicitation number
SS_427787
Issued by
Department of Energy Office of Science

View the file

On GovTribe

Work with this file on GovTribe

  • Download the original file
  • Contacts named in this file
  • Similar government files
  • Ask GovTribe AI about this file

Text version

SOW

2/23/2023
Page 1 of 1

Statement of Work Introduction The majority of DOE exascale simulation and experimental applications are spatio-temporal learning challenges, and scaling spatio-temporal learning algorithms on exascale computers is critical to enable scientific breakthroughs and advances in industrial applications.

Spatio-temporal data analysis is prevalent in various exascale prediction and analysis tasks, including those in climate research, DUNE (Deep Underground Neutrino Experiment) simulation, neuroscience analysis including functional magnetic resonance imaging (fMRI) and so on. Such simulations and experiments involve large data volumes, high velocity and multi-modal data streams. When both the data volumes and the model-parameter space are large, it is not practical to train such models in a single node and this is where extreme-scale machine learning can make a significant difference.

Goal: The ultimate goal in extreme scale spatio temporal learning is capturing longer term spatio-temporal dependencies effectively. We have two major directions (neural architecture development and distributed optimization).

In terms of neural architectures, we have finished the proof of principle in climate science (25% better than state of the art forecasting model) and currently using data parallel training at Perlmutter. To scale Transformer with longer-temporal and spatial dependencies, we recently developed a sparse attention kernel approximation method using back-propagation and plan to integrate this feature to our spatio-temporal modeling.

In terms of distributed optimization study, we are interested in distributed data parallel training to incorporate model parallelism either through domain decomposition or model pipelining.

Scope of Work

· Developing hybrid parallelism (data and model) using JAX or alternative platforms for Transformers on DOE supercomputers such as Perlmutter or Polaris at LCF (Leadership Computing Facilities).

· Incorporating scalable algorithms such as sparse attention kernels or approximated attention kernels on Transformer architectures.

Schedule

· We propose the work to be completed by August 31st, 2024.

· Year 1 (FY23): Implement Transformer hybrid parallelism using JAX with the potential support from NVIDIA and NERSC.

· Year 2 (FY24): Scaling study and refining Transformer hybrid parallelism with scalable Transformer algorithms and demonstrate it on at least one science domain.

image1.png

File details come from the government source that posted it. Updated .