Schnase_CMAC2_Proposal.pdf
PDF 3 MB Posted
- Attached to
- Patent Application Federal contract opportunity
- Solicitation number
- 80NSSC18Q0542
About this file
Please see the 19 page Schnase_CMAC2_Proposal
View the file
Other files for this federal contract opportunity
| File | Type | Posted |
|---|---|---|
| RFQ_Open_Market_Template.docx | DOCX document | |
| Project_Slides.pdf | ||
| NASA_Form_1679.pdf | ||
| NASA_Form_1679_(3).pdf | ||
| NASA_Form_1679_(2).pdf | ||
| NASA_Form_1679_(5).pdf | ||
| NASA_Form_1679_(4).pdf | ||
| NASA_Form_1679_(6).pdf | ||
| Questions.docx | DOCX document | |
| RFQ_Open_Market_Template.docx | DOCX document | |
| RFQ_Open_Market_Template.docx | DOCX document | |
| SATPC0009410_Tab_12_RFQ_Open_Market_Template.docx | DOCX document | |
| Questions.docx | DOCX document | |
| RFQ_Open_Market_Template.docx | DOCX document |
Show all 14
On GovTribe
Work with this file on GovTribe
- Download the original file
- Contacts named in this file
- Similar government files
- Ask GovTribe AI about this file
Text version
EXTENDING CLIMATE ANALYTICS-AS-A-SERVICE
TO THE EARTH SYSTEM GRID FEDERATION
John L. Schnase1, Daniel Q. Duffy2, and Glenn S. Tamkin1
1Office of Computational and Information Science and Technology
2NASA Center for Climate Simulation (NCCS) NASA Godard Space Flight Center
Greenbelt, MD 20771
PROPOSAL SUMMARY
We propose to build three extensions to prior-funded work on climate analytics-as-a-service that will benefit the Earth System Grid Federation (ESGF) as it addresses the Big Data challenges of future climate research: (1) We will build a cloud-based, high-performance Virtual Real-Time Analytics Testbed supporting a select set of climate variables from six major reanalysis data sets. This near real-time capability will enable Cloudera Impala-based Structured Query Language (SQL) querying and SciHadoop-based MapReduce analytics over native NetCDF files and will provide a platform for community experimentation with emerging analytic technologies. (2) We will build a full-featured Reanalysis Ensemble Service comprising monthly means data from six reanalysis data sets. The service will provide a basic set of commonly used operations over the reanalysis collections. The operations will be made accessible through NASA's climate data analytics Web service and our client-side Climate Data Services (CDS) API. (3) We will build an Open Geospatial Consortium (OGC) WPS-compliant Web service interface to our climate data analytics service that will enable greater interoperability with next-generation ESGF capabilities. The CDS API will be extended to accommodate the new WPS Web service endpoints as well as ESGF's Web service endpoints. These activities address some of the most important technical challenges for server-side analytics and support the research community's requirements for improved interoperability and improved access to reanalysis data.
ii
TABLE OF CONTENTS
SCIENTIFIC, TECHNICAL, AND MANAGEMENT
1.0 SCIENCE OBJECTIVES
1.1 CMAC Priority Topic Being Addressed
1.2 Background: Big Data Challenges of the Earth System Grid Federation
1.3 Objectives of the Proposed Research
2.0 TECHNICAL APPROACH
2.1 Foundational Technologies from Prior Funded CMAC Work
2.1.1 Climate Analytics-as-a-Service (CAaaS)
2.1.2 MERRA Analytic Services (MERRA/AS)
2.2 Proposed Next-Phase CMAC Advances
2.2.1 Improving Analytic Capabilities
2.2.2 Improving Data Availability
2.2.3 Improving Analytic Interoperability
2.3 Expected Outcomes, Deliverables, and Benefits
3.0 MANAGEMENT PLAN
3.1 Schedule with Milestones
3.2 Investigator Team and Responsibilities
3.2 Licensing Compliance and Software Engineering Practice
REFERENCES
CURRICULUM VITAE
BUDGET JUSTIFICATION
SCIENTIFIC, TECHNICAL, MANAGEMENT
1.0 SCIENCE OBJECTIVES
The Intergovernmental Panel on Climate Change (IPCC) research that provided the basis for 2013's Fifth Assessment Report (AR5) worked with about five petabytes (PB) of data. It is estimated that the IPCC community’s collective work on AR6, which will probably be released around 2020, will generate as much as 100 petabytes of data, a dramatic affirmation that climate science is indeed a “Big Data” domain [1]. The Earth System Grid Federation (ESGF) provides the cyberinfrastructure to support this global scientific collaboration [2]. ESGF also provides the context whereby the science community participates in the design and implementation of this infrastructure. As a result, IPCC’s success depends on ESGF’s ability to scale its capabilities to accommodate the Big Data challenges of AR6.
In prior funded CMAC work, we developed a concept of climate analytics-as-a-service (CAaaS) that advances several ideas and technologies that we believe are of value to the IPCC community [3]. We produced specifications for a climate data analytics system, a Web service interface based on the Open Archival Information System (OAIS) Reference Model, and a Web application programming interface (API) that integrates server-side analytics, digital preservation, and provenance management. We used these specifications to build the Modern Era Retrospective Analysis for Research and Applications Analytic Service (MERRA/AS) and a Persistence Service (PS) that demonstrate an approach to CAaaS, and we implemented the Climate Data Services (CDS) API to make service-oriented client development easier [4]. In the work proposed here, we will build on these results and contribute our technologies and experiences to the extended climate research community through a proactive engagement with the ESGF.
1.1 CMAC Priority Topic Being Addressed
We are responding to the 2.1 Computing Center Based Advanced Data Management and Data Analysis System topic in the ROSES A.40 CMAC solicitation, which focuses on building and expanding data management and data analysis systems that are coupled with the activities of the Earth System Grid Federation and designed to support future climate data-model intercomparison and modeling processes.
1.2 Background: Big Data Challenges of the Earth System Grid Federation
Big Data challenges are often approached from one of two perspectives. They sometimes are viewed as problems of large-scale data management where solutions are offered through an array of somewhat traditional storage and archive theories and technologies. These approaches tend to view Big Data as an issue of storing and managing large amounts of structured data for the purpose of finding particular subsets of interest. Alternatively, Big Data challenges are sometimes viewed as knowledge management problems where solutions are offered through an array of analytic techniques and technologies. These approaches tend to view Big Data as an issue of extracting meaningful patterns from large amounts of unstructured data for the purpose of finding particular insights of interest.
As the ESGF community grapples with its scaling challenges, it seeks to find a balance between these sometimes competing views. This is evident in the charge that the ESGF Compute Working Team (ESGF-CWT) — the group responsible for designing ESGS's "next generation" architecture — has laid out for itself. The Team's overarching goal is to increase the analytical capabilities of the enterprise, primarily by exposing high-performance computing resources and analysis tools to the community through Web services. Ideally, ESGF data from the Federation's distributed collections would be united with the Web-accessible tools and compute resources needed to perform advanced analytics at the scale needed for AR6.
But getting there from here is tough. ESGF's technical heritage is that of a large-scale distributed archive. Its nodes store and distribute data, and they typically support compute resources sufficient only to stream data out of storage onto the network for client consumption, the traditional modus operandi of digital archives. The behaviors implemented and exposed by ESGF's Web service interface are the basic discovery and download operations of an archive. Integrating high-performance computing and high-performance analytics — finding an optimal storage-compute balance in ESGF's ecosystem of distributed resources — is not a trivial exercise.
Our work in this area is driven by the belief that the Big Data challenges of the climate sciences require a fundamental re-working of the architectures for managing petabyte-scale data collections and delivering products to users [4]. Large repositories mean that the data sets themselves cannot be moved.
Instead, following the model of object-oriented design, behavior needs to be co-located with data — analytical operations need to run where the data reside. Complex analyses over large repositories require data to be surrounded by high-performance computing designed for large-scale analytics, not just enough compute power to move files from storage to the network. Large amounts of information coupled with dynamically created analytic products increases the importance of metadata, provenance management, and discovery. Finally, the ability to respond quickly to customer demands for new and often unanticipated uses for climate data requires greater agility in building and deploying applications [4].
In our efforts to address the Big Data challenges of climate science, we are moving toward a notion of Climate Analytics-as-a-Service (CAaaS). We focus on analytics, because we feel that it is the knowledge gained from interactions with Big Data that ultimately produce societal benefits. We focus on CAaaS because we believe it provides a useful way of thinking about the problem: a specialization of the concept of business process-as-a-service, which is an evolving extension of Software-as-a-Service (SaaS) enabled by cloud computing.
CAaaS as we conceive it, and as we have demonstrated in the systems we have built, attempts to harmonize the archive/analytic perspectives on Big Data and mobilizes capabilities in ways we believe can fundamentally change the data practices of this domain. ESGF has been a remarkable resource for the climate research community. But like other science cyberinfrastructures, ESGF's continued success requires continued evolution of its technologies. In the sections below, we identify the major objectives of our proposed work, review accomplishments to date, and then describe how we will build on those accomplishments to benefit ESGF efforts.
1.3 Objectives of the Proposed Research
ESGF development is coordinated by technical working teams that are focusing on the critical elements of a �next generation� ESGF system. Through a series of ongoing team meetings and an annual conference, scores of contributors from institutions throughout the world are identifying key design decisions, conducting proof-of-concept experiments, and, ultimately, implementing capabilities throughout the ESGF network that will yield the scaled capacity needed by IPCC.
The ESGF Compute Working Team (ESGF-CWT) — the group having primary responsibility for architecture and capability development — has identified and are working on many important issues, but there are three areas of technical development where we believe our previous CMAC work can be of value: enhanced server-side analytic capabilities, increased data availability, and improved service interoperability. Our primary objective is to extend the work we have done to address elements in each of these targeted technical areas. In addition, we will work to extend the benefits of these ESGF-driven enhancements to groups beyond the climate science research community that have shown interest in using climate data for research and applications. Our specific objectives include the following:
(1) Improving Analytic Capabilities — Real-time analytics and simplifying the data storage processes associated with high-performance analytics are two of the most important technical challenges for server-side analytics. We will address both of these issues. We will enhance the core MapReduce analytic capabilities developed in our prior CMAC work by integrating a suite of emerging technologies that are moving the Hadoop ecosystem toward real-time analytics. In addition, we will increase the use of cloud computing in our analytic environment and provide methods for direct use of NetCDF files in their native format in analytic processing. Importantly, these capabilities will be made available to the research community in the form of a Virtual Real-Time Analytics Testbed to allow experimentation and to engage the community in further development of server-side analytics.
(2) Improving Data Availability — Using MERRA/AS as a model, we will build a Reanalysis Ensemble Service (RES) that delivers customized NASA/MERRA-2, ECMWF/ERA-Interim, NCEP/CFSR, NCEP/20CR, JMA/JRA-25, and JMA/JRA-55 reanalysis products within an analytic environment. Doing so will enable a fundamentally new capacity to perform reanalysis intercomparison.
In addition, we will develop a set of CDS API utilities to support ensemble analysis, uncertainty quantification, and reanalysis intercomparison. As part of this effort, we will incorporate into our CAaaS framework regridding, georegistration, and formatting services, which are widely used utilities of general value to the community and are specifically needed to implement the RES.
(3) Improving Service Interoperability —The ESGF-CWT is adopting the Open Geospatial Consortium (OGC) Web Processing Service (WPS) interface standard as the communication standard for its next generation architecture. To increase interoperability, we will extend our server-side Web service tier to support the WPS standard. Adding a WPS exposure to our OAIS-based Web service will facilitate integration of NASA's capabilities into the ESGF. To compliment these Web service enhancements, we will extend the client-side CDS API to support the WPS Web service interface standard and ESGF's Web service search interface, and we will incorporate into the CDS API new functions of value to the ESGF community as specified by the community.
2.0 TECHNICAL APPROACH
2.1 Foundational Technologies from Prior Funded CMAC Work
2.1.1 Climate Analytics-as-a-Service (CAaaS)
In our prior funded CMAC work, we developed MERRA Analytic Services. In the course of those activities, we developed a point-of-view on what we believe to be the major factors influencing the general design of climate data analytic services: the need for large-scale data storage to be embedded with high-performance computing and analytic capabilities; the need to support with common server-side operations the natural, hierarchical stratification that exists in most scientific workflows; and the need for client-side Web APIs that support the specialized software development activities of the climate research community. In this section, we describe what we mean by these terms and how we have implemented these concepts in our own work. Additional information can be found in [3].
2.1.1.1 High-performance, server-side computing and analytics
At its core, CAaaS must bring together data storage and high-performance computing in order to perform analyses over data where the data reside. MapReduce has been of particular interest to us, because it provides an approach to high-performance analytics that has been useful to many data intensive problems [5, 6]. MapReduce enables distributed computing on large data sets using high-end computers. It is an analysis paradigm that combines distributed storage and retrieval with distributed, parallel computation, allocating to the data repository analytical operations that yield reduced outputs to applications and interfaces that may reside elsewhere. Since MapReduce implements repositories as storage clusters, data set size and system scalability are limited only by the number of nodes in the clusters. While MapReduce has proven effective for textual data, its use in data intensive science applications has been limited, because scientific data sets are often complex, have high dimensionality, and use binary formats. Much of the work that we have done has focused on bringing scientific data into the MapReduce framework.
MapReduce distributes computations across large data sets using a large number of computers (nodes). In a “map” operation a head node takes the input, partitions it into smaller sub-problems, and distributes them to data nodes. A data node may do this again in turn, leading to a multi-level tree structure. The data node processes the smaller problem, and passes the answer back to a reducer node to perform the reduction operation. In a “reduce” step, the reducer node then collects the answers to all the sub-problems and combines them to form the answer to the problem it was originally trying to solve.
2.1.1.2 Adaptive workflow stratification and canonical operations
In the climate sciences, data intensive analytic workflows generally bridge between a largely unstructured mass of archived scientific data and the highly structured, tailored, reduced, and refined analytic products that are used by individual scientists and form the basis of intellectual work in the domain. The initial steps of an analysis — those operations that first interact with a data repository — tend to be the most general, while data manipulations closer to the client tend to be the most specialized to the individual, to the science question being studied, and to the models being used. The amount of data being operated on also tends to be larger on the repository-side of the workflow, smaller toward the client-side end products.
This stratification can be exploited in order to optimize efficiencies along the workflow chain.
High-performance analytic software, for example, can be used to improve efficiencies of the near-archive operations that initiate workflows. In our work, we have used MapReduce to implement server-side operations that represent a common starting point in many workflows, including average, variance, maximum, minimum, sum, count, and difference operations of the general form:
result <== avg(var, (t0,t1), ((x0,y0,z0),(x1,y1,z1))), that return, in this example, the average value of a variable when given its name, a temporal extent, and a spatial extent. Because of their widespread use, we refer to these simple operations as "canonical ops" — the primitive operations with which more complex expressions can be built. They provide a template for users as they begin their exploration of MapReduce and are useful in their own right as steps in larger analyses. We tend to think of them as a type of assembly language instruction for climate data analytics.
While canonical ops are admittedly low-level capabilities, they are of wide-spread importance. By using high-performance analytic software to implement these canonical ops, the system stands ready to accommodate more sophisticated analytics.
As described in the next section, our goal is to deploy the canonical ops within an adaptive API framework that captures their patterns of use and enables more complex analyses to be assembled and incorporated back into the system. The notion of engaging the broader community to deal with Big Data challenges has been used successfully in other settings, perhaps most notably with GalazyZoo, where a large user community is helping search the Sloan Digital Sky Survey for patterns and observations of potential scientific value [7]. We believe that this type of social networking can play an important role in the future of climate analytics. The approach we are taking sets the stage for the community construction of new capabilities that are adapted to the socially expressed requirements of the system's users.
2.1.1.3 Domain-harmonized, client-side Web APIs
In order to knit these capabilities together and deliver them into practical use, we are developing the Climate Data Services (CDS) application programming interface (API). In building the CDS API, we are trying to provide for climate science a uniform semantic treatment of the combined functionalities of large-scale data management and server-side analytics. We believe that it makes sense to look at an analytics service as being an active archive system where new objects are dynamically created and must be stored, annotated with metadata, and otherwise managed as storage objects � just as the objects in the base collection upon which analytic algorithms operate must be managed. Said another way, the realized objects of an analytics system can be considered elements of a "virtual archive" the existence of which can change the way we think about Big Data storage. And it provides a path whereby existing archive can become more service oriented and offer more sophisticated analytic capabilities to its customers.
In our view, the best way to bring coherence to this archive/analytic Big Data dichotomy is to organize the basic elements of a CAaaS system around a widely accepted archival reference model. For this, we have chosen the International Standards Organization (ISO) Open Archival Information System (OAIS) reference model. OAIS is defined by the Consultative Committee for Space Data Systems (CCSDS). The CCSDS's purview is space agencies, but the OAIS model it developed has proved useful to a wide variety of other organizations and institutions with digital archiving needs. OAIS provides a framework for the understanding and increased awareness of archival concepts needed for long-term digital information preservation and access and provides the concepts needed by non-archival organizations to be effective participants in the preservation process [8].
The OAIS reference model asserts that a scientific archive comprises four data flow interactions:
ingest, query, order, and download. These behaviors act on three types of objects: Submission Information Packages (SIPs), Archive Information Packages (AIPs), and Dissemination Information Packages (DIPs). Each of these types of objects can have associated with it four types of metadata:
Representational Information, Preservation Description Information, Policy Metadata, and Discovered Metadata. With these abstractions, OAIS accommodates the full range of archival preservation functions and defines a minimal set of responsibilities for an archive to be OAIS-compliant [8].
These high-level OAIS abstractions provide a vocabulary that we have adopted for our service functions, Web service protocol, and the CDS API library. Ingest refers to methods that input objects into the system, query methods retrieve metadata relating to data objects in the service, order methods dynamically create data objects, and download methods retrieve objects from the service. We have added execute and status categories to accommodate the dynamic nature of a climate data analytics service archive. Execute methods invoke service-definable extensions, and status methods check on the progress of running operations.
Our outfacing Web service protocol also is based on OAIS's data flow interaction categories.
Within this OAIS-inspired framework, we have created a Python-based client-side CDS API that abstracts low-level inbound and outbound traffic into higher-order functions and methods that are more convenient for software developers to use. The API's basic methods provide a one-to-one mapping of OAIS-classified operations on the client side to corresponding OAIS-classified operations on the server side.
The API's extended methods build on basic methods, placing them under programmatic controls to create more specialized convenience methods and workflows that can be folded back into the API's libraries.
Python scripts and full Python programs can import CDS API library methods to create client software applications that draw on the capabilities of the CAaaS system. The entire CDS API client stack is distributed as a software package that can be used to build applications, a cloud-based service, or distributable virtual cloud images.
By organizing communications and functional capabilities around the OAIS standard, the climate data analytics system behaves like a dynamic archival information system capable of performing full information lifecycle management in an analytics context. Existing archive systems should find it easier to integrate these capabilities, because the interfaces and interactions with the system will be familiar to archive authorities and existing archive systems. Likewise, the behaviors implemented by the climate data analytics system can be organized around traditional archive operational workflows. This approach focuses on the specific analytic requirements of climate science and unites the language and abstractions of collections management with those of high-performance analytics, reflecting at the application level the confluence of storage and computation that is driving Big Data architectures of the future.
2.1.2 MERRA Analytic Services (MERRA/AS)
MERRA Analytic Services (MERRA/AS) pull these elements together in an end-to-end demonstration of CAaaS capabilities. In simple terms, MERRA/AS stores NASA's Modern-Era Retrospective Analysis for Research and Applications (MERRA) data in the Hadoop Distributed Filesystem (HDFS) of a storage cluster and enables MapReduce analytics to be performed over the MERRA data.
2.1.2.1 The MERRA/AS analytics platform
MERRA is produced by NASA's Global Modeling and Assimilation Office (GMAO) using the Goddard Earth Observing System Data Assimilation System Version 5 (GEOS-5). The MERRA reanalysis integrates observational data with numerical models to produce a global temporally and spatially consistent synthesis of 26 key climate variables [9]. Spatial resolution is 1/2 ̊ latitude x 2/3 ̊ longitude x 72 vertical levels extending through the stratosphere. Temporal resolution is 6-hours for three-dimensional, full spatial resolution, extending from 1979-present, nearly the entire satellite era. MERRA data are made available to the general public through the NASA Earth Observing System Distributed Information System (EOS DIS); a subset of the data is made available to the climate research community through ESGF. We focused on MERRA because there is an increasing demand for reanalysis data by an expanding community of consumers, including local governments, federal agencies, and private-sector customers. Reanalysis data are used in models and decision support systems relating to disasters, ecological forecasting, health and air quality, water resources, agriculture, etc.
The Apache Hadoop software library is the classic framework for MapReduce distributed analytics. We are using Cloudera, the 100% open source, enterprise-ready distribution of Apache Hadoop.
Cloudera is integrated with configuration and administration tools and related open source packages. The total size of the MERRA/AS HDFS repository is approximately 480 TB. MERRA/AS is running on a 36-node Dell cluster that has 576 Intel 2.6 GHz SandyBridge cores, 1300 TB of raw storage, 1250 GB of RAM, and a 11.7 TF theoretical peak compute capacity. Nodes communicate through a Fourteen Data Rate (FDR) Infiniband network having peak TCP/IP speeds in excess of 20 Gbps.
The canonical operations that implement MERRA/AS�s average, variance, maximum, minimum, sum, count, and difference calculations are Java MapReduce programs that are exposed as server-side Web service endpoints or as methods in the client-side CDS API Library. There is a substantial code ecosystem behind these simple operations, nearly 6000 lines of Java code being offloaded from the user to the MERRA/AS service in the current implementation. We are using a Representational State Transfer (REST)-style architecture, which is the predominant Web API design model. REST provides scalability of component interactions, accommodates intermediaries like firewalls and proxies without the need to change interfaces, and allows independent deployment of components where implementations can change without the need to change interfaces.
Fig.!1.!Climate!data!analytics!system!architecture.!
2.1.2.2 MERRA/AS in use
Fig. 1 is a diagram showing the overall architecture of our climate data analytics system with key elements identified by number. The system 100 comprises a collection of services that sit atop a high-performance climate data analytics platform (HPDAP) 101. The HDAP compute-storage platform, as described above, currently exists as a storage cluster in our facilities, but it can also be implemented in our Advanced Data Analytics Platform (ADAPT) science cloud or other cloud service.
Our services to date include the MERRA Analytic Service 102 and a Persistence Service 103, which is an Integrated Rule-Oriented Data System (iRODS) storage system that allows MERRA/AS's dynamically created products to be stored and managed as new objects within the system. The system interface 107 comprises an outfacing Representational State Transfer (REST) Web service and an infacing adapter module that maps incoming service requests to specific service methods, thereby exposing capabilities to the outside world.
The Climate Data Services API (CDS API) 108 links to the climate data analytics system through the API's REST interface 109. The API abstracts the system's Web service endpoints into basic 110 and extended 111 utilities as described above. Client software applications 112 have the option of binding directly to Web service endpoints or the CDS API's basic and extended methods.
MERRA/AS is currently in beta testing with about two dozen partners across a wide range of organizations and topic areas [3]. In one application, MERRA/AS's Web service is providing data to NASA's RECOVER wildfire decision support system, which is being used for post-fire rehabilitation planning by Burned Area Emergency Response (BAER) teams within the US Department of Interior and the US Forest Service. This capability has lead to the development of new data products based on climate reanalysis data that until now were not available to the wildfire management community [10].
In our largest deployment exercise to date, the CDS API has been used by the iPlant Collaborative to integrate MERRA data and MERRA/AS functionality into the iPlant Discovery Environment. iPlant is a virtual organization created by a cooperative agreement funded by the US National Science Foundation (NSF) to create cyberinfrastructure for the plant sciences. The project develops computing systems and software that combine computing resources, like those of TeraGrid, and bioinformatics and computational biology software. Its goal is easier collaboration among researchers with improved data access and processing efficiency. Primarily centered in the US, it collaborates internationally and includes a wide range of governmental and private-sector partners [11].
As part of our testing efforts, we have turned to the research literature and interviews with climate scientists for examples of real-world experiments that have MERRA data. By repeating the data-gathering and early-stage data reduction steps of their workflows, we are better understanding the potential benefits of the technology. Initial results have shown that analytic engine optimizations can yield near real-time performance of MERRA/AS's canonical operations and that the total time required to assemble relevant data for many applications can be significantly reduced [3].
2.2 Proposed Next-Phase CMAC Contributions
2.2.1 Improving Analytic Capabilities
The Hadoop ecosystem is notable for the way it embeds computation with data. In the beginning, that ecosystem largely comprised Hadoop's HDFS high-throughput distributed file system and the MapReduce parallel processing system. Today that ecosystem encompasses a burgeoning assemblage of tools and techniques aimed at overcoming earlier limitations and consolidating Hadoop's position in the marketplace. In 2014, for example, Intel invested $740 million in Cloudera, Inc. as part of a $1 billion funding round that brought the company's estimated value to over $4 billion, underscoring investor interest in the Big Data technology movement and the Hadoop software used by Cloudera. For many observers this marked a turning point — there is now a widely held belief that Hadoop and its core con-cepts will be the cornerstone of innovation for the next generation of data management and analytics [12].
A full treatment of this topic is beyond the scope of this proposal. However, what is interesting to note about this transformation is the sheer number of approaches being tried. Borrowing techniques from the DBMS world, innovative indexing, and algorithmic improvements — as seen in HAIL [13], LAIL [14], Hive [15], and SciHadoop [16] — have yielded as much as two orders of magnitude improvement in MapReduce performance [14]. Some approaches — as seen with Impala [17] and Drill [18] — provide near real-time querying services on top of the Hadoop file system that bypass MapReduce altogether. And scalable alternatives to HDFS — as seen with Seagate’s Kinetic Object Store [19], Ceph [20], and MapR [21] — are of increasing interest across the board. In all these activities, what preserves of the original Hadoop philosophy is the notion that storage and compute will be inextricably linked throughout future architectures, even to the level of physical devices.
2.2.1.1 Real-Time Analytics and Native NetCDF Processing
This project and the climate research community are in a position to benefit from these advances. We propose to address two of the most significant impediments to broader use of server-side analytics in our domain: (1) performance — a capability for near real-time analytics to complement the batch processing perspective inherent in traditional, long-running analytic algorithms, and (2) data preparation — the inability of current analytic systems to use raw scientific data without preprocessing, which, among other things, complicates their integration with legacy systems that require POSIX-compliant interfaces to data.
In prior-funded CMAC work, we compared the performance of Cloudera Hadoop, Cloudera Impala, and Apache Hive. Cloudera Hadoop is Cloudera's Apache-licensed open source distribution of Hadoop; Impala and Hive are Structured Query Language (SQL) query engines that operate on the HDFS and offer performance improvements for data warehousing applications. We ran tests on files stored in four formats: comma-separated value (CSV), Hadoop sequence (Seq), and two column-oriented formats, Record Columnar (RC) and Parquet [17]. As shown in Fig. 2, significant reductions in run times for our canonical operations can be realized with Impala/Parquet. In related work, colleagues at George Mason University working on a NASA-funded project have demonstrated improved run times using SciHadoop operating on native NetCDF files stored in a Hadoop file system [22]. These efforts confirm that technologies now exist that could bring us closer to our goals for real-time analytics and native file processing.
2.2.1.2 Implementation Approach
Over the next two years, we intend to tackle the dual challenge of real-time analytics and native file processing through a combination of activities that will advance our basic understanding of the problems and technologies, bring coherence to this chaotic discussion, identify viable solutions for climate data management, and, where solutions are lacking, guide new development.
Virtual Real-Time Analytics Testbed — Using the capabilities of NASA's Advanced Data Analytic Platform (ADAPT) science cloud, we will build Cloudera Impala/Parquet- and SciHadoop/NetCDF-based analytic services to complement MERRA/AS's existing Cloudera Hadoop service. Each service will provide near real-time analytic support for a select set of the most widely used variables (for example, temperature, precipitation, humidity, pressure, winds) from six of the major reanalysis data sets (MERRA-2, ERA-Interim, NCEP-CFSR, ESRL-20CR, JRA-25, and JRA-55). Access to the ADAPT testbed will be provided through a Web service interface and through the Climate Data Services API. We will also instrument the systems to gather performance metrics, and we will provide a social networking interface to capture user feedback about these system.
Climate Data Analytics Working Group and Special Session — In order to more fully engage the community on these issues, we have proposed a special session on this topic to be held as part of the First International Workshop on Spatiotemporal Computing (IWSC '15) to be held at George Mason University July 13-15, 2015 [23]. Invited guests will include technical leads from the SciHadoop and Cloudera communities along with data scientists from NASA JPL and elsewhere who are interested in particular applications of the Hadoop ecosystem to Big Data challenges in the climate sciences. This work will occur during the first year of the project, will engage other CMAC projects as appropriate, and produce a white paper with recommendations to guide further developments in this area.
Fig.!2.!Runtime!performance!for!MapReduce!queries.!!
Multi<year!query!with!various!analytic!file!formats!(top);!!
multi<variable!queries!with!native!NetCDF!files!(bottom).!
Virtual Real-Time Reanalysis Ensemble Service — In the second year of the project, using the insights gained from the first year's Climate Data Analytics Working Group activities, we will place selected technologies into use in the implementation of a Virtual Reanalysis Ensemble Service (vRES).
The vRES will be a high-performance, cloud-based experimental platform supporting an SQL querying capability to complement our standard Web service and CDS API access mechanisms. This environment will allow the climate research community to gain experience with real-time analytics and holistic, long-running batch analytics over data of sufficient complexity to gain useful insights into the technologies and techniques of server-side analytics. Feedback from the community will be used to extend and refine the system. Ensemble reanalysis is also the focus of our second major deliverable as described next.
2.2.2 Improving Data Availability
There is a recognized need within the climate research community for better access to reanalysis data and tools for working with reanalysis data sets [24]. In response, ESGF's Analysis for Model Intercomparison Project (Ana4MIPs) and the Collaborative Reanalysis Technical Environment Intercomparison Project (CREATE-IP) are bringing together various types of commonly-formatted reanalysis data with tools such as the Climate Data Services Visualizer (CDS Visualizer), ESGF's Ultrascale Visualization Climate Data Analysis Tool (UV-CDAT), and ArcGIS to better understand reanalysis differences and uncertainties and increase the usefulness of these data sets.
2.2.2.1 Reanalysis Ensemble Service
We propose to support these efforts and improve the availability of reanalysis data by creating a full-featured Cloudera Hadoop MapReduce-based Reanalysis Ensemble Service (RES). Using MERRA/AS as a model, the RES will deliver customized monthly means products from six of the major reanalysis projects in a uniform analytics environment. In addition, we will develop a set of CDS API utilities to support ensemble analysis, uncertainty quantification, and reanalysis intercomparison.
Fig. 3 provides a demonstration of what we would like to achieve in developing a Reanalysis Ensemble Service. The image on the left shows the departure of 2010 summer temperatures from a 34-year average of summer temperatures as determined by the ECMWF ERA-Interim reanalysis; the image on the right shows the same for the NCEP/CRSR reanalysis. There are obvious difference between the two that could reveal valuable information about the underlying modeling systems, inputs, or atmospheric phenomena. We want to make it easier to do these types of intercomparisons.
A climatological anomaly refers to the positive or negative departure of a climate variable from a long-term average or reference value for the variable. Anomalies can be determined in various ways.
Finding the arithmetic difference between a reference value and an observed or modeled value for the variable in question is perhaps the simplest way to compute an anomaly. Normalized anomalies divide this arithmetic anomaly by the climatological standard deviation to remove dispersive influences and reveal more information about the magnitude of the anomaly. In either case, a set of basic operations are applied to what can be large data sets to make an anomaly calculation: for a given variable, region, and time-span of interest compute sum, count, average, variance, and difference � our canonical ops.
Detecting anomalies is one of climate scientists' most common and useful determinations. The Fig. 3 example was done using traditional manual processes over a period of days. In analogous runs using MERRA/AS and the CDS API, we have written Python workflows that use a single library call to compute similar results in about three minutes [25]. The goal with RES is to do this at an operational scale for our target data sets and to extend the CDS API library to include a collection of anomaly methods and similar calculations of value to the ESGF and extended climate research community.
Fig.!3.!Temperature!anomalies!between!34<year!summer!average!
and!summer!2010!using!the!ERA<Interim!(left)!and!CFSR!(right)!reanalyses."
2.2.2.2. Implementation approach
The Reanalysis Ensemble Service will initially incorporate the following six reanalyses:
(1) NASA Modern Era Reanalysis for Research and Applications Version-2 (MERRA-2): 1980-present � MERRA-2 uses a new version of GEOS-5 that will assimilate observations not available to MERRA. Since there are numerous improvements and updates to the assimilation system, model, and observing systems, the reanalysis will be redone from the beginning.
(2) ECMWF Interim Reanalysis (ERA-Interim): 1979-present � The European Centre for Medium-Range Forecasts (ECMWF) ERA-Interim was originally planned as an 'interim' reanalysis in preparation for the next-generation extended reanalysis to replace ERA-40. It uses a December 2006 version of the ECMWF Integrated Forecast Model (IFS Cy31r2). ERA-Interim is being continued in real time. With some exceptions, ERA-Interim uses input observations prepared for ERA-40 until 2002, and data from ECMWF's operational archive thereafter.
(3) NOAA NCEP Climate Forecast System Reanalysis (CFSR): 1979-present � The National Centers for Environmental Prediction (NCEP) CFSR is a global, high resolution, coupled atmosphere-ocean-land surface-sea ice system that estimates of the state of these coupled domains over this period.
The T382 atmospheric data spans 1979 to 2010. The current T574 analysis is an extension of the CFSR as an operational, real time CFSv2 product from 2011 into the future.
(4) NOAA ESRL 20th Century Reanalysis (20CR): 1980-present � 20CR spans the entire twentieth century. It assimilates surface observations of synoptic pressure, monthly sea surface temperature, and sea ice distribution to produce six-hourly analyses representing the most likely state of the global atmosphere along with uncertainty estimates. This dataset provides the first estimates of global tropospheric and stratospheric variability spanning 1871 to 2012 at six-hourly resolution.
(5) Japanese 25-year Reanalysis (JRA-25): 1979-2004; 2005-Jan.2014 � JRA-25 is the first long-term global atmospheric reanalysis undertaken in Asia. It was completed using the Japan Meteorological Agency (JMA) numerical assimilation and forecast system and assimilates data from many sources, including ECMWF, the National Climatic Data Center (NCDC), and the Meteorological Research Institute (MRI) of JMA. It was continued until 2014, when it was replaced by JRA-55.
(6) Japanese 55-year Reanalysis (JRA-55): 1958-2012, and extended to present � JMA has carried out the Japanese 55-year Reanalysis (JRA-55) using a more sophisticated NWP system, which is based on the operational system as of December 2009, and newly prepared past observations. The analysis period is extended to 55 years starting from 1958, when the regular radiosonde observations became operational on the global basis. JRA-55 has been continued in near real time basis since 2013.
NASA's Climate Model Data Services group has formalized agreements with these reanalysis projects to host full spatial and tempoeral resolution data sets and will, over the course of this project, assemble up to 1 PB of data to support CREATE-IP and Ana4MIPs activities. Climate Model Data Services will keep the data sets current and will add new data sets as they become available.
Most of the work of building the RES involves creating utilities to move data into and out of the Hadoop file system and writing the MapReduce programs that implement the service's canonical ops.
Each file of native NetCDF binary data is converted into separate sequence files — flat files consisting of binary <key, value> pairs that can be operated on by mapper and reducer functions. These sequence files are block-compressed in HDFS and provide direct serialization of binary data types. During sequencing, the data is partitioned by time, so that each record in the sequence file contains the timestamp and name of the parameter (e.g. temperature) as the composite key and the value of the parameter. In operation, mappers filter each sequence file to capture <key, value> pairs that match the variable and time span of interest; reducers perform calculations based on input parameters and create a new subset of sequence files. The resulting sequence files are then transformed back into NetCDF. Removing the costly overhead of sequencing and de-sequencing is a main driver for the use of native NetCDF files in Hadoop.
To integrate the new reanalyses, we will develop the following Java classes for each collection:
(1) a sequencer utility to convert raw data to a MapReduce-consumable input format, (2) a mapper class to filter and combine input sequence records, (3) a reducer class to aggregate and transform filtered input records into sequence file output format, (4) a record reader/writer utility for HDFS input/output used by the mapper/reducer, (5) a driver class to orchestrate the application at runtime, and (6) a de-sequencer utility to convert reduced sequence files to the proper output format (e.g.,NetCDF).
Additional utilities currently used by the Climate Model Data Services group in their work with reanalysis data will also be needed by the Reanalysis Ensemble Service. For convenience, we will integrate these into our CAaaS system. We will provide Web service exposures to the Earth System Modeling Framework (ESMF) ESMPy regridding tool, PCMDI's Climate Model Output Rewriter 2 (CMOR2) reformatting and Climate and Forecasting (CF) compliance tool, and the Geographic Information Systems (GIS) formatting and projection tools available through the Geospatial Data Abstraction Library (GDAL) and ArcGIS. These are widely used utilities: providing them as a general service through NASA's Climate Model Data Services will be a benefit the community.
2.2.3 Improving Analytic Interoperability
The ESGF-CWT is working to increase interoperability within the Earth System Grid Federation in an effort to improve access to distributed resources and to enable scientists to build and share workflows of common interest to the ESGF community. Similar objectives have been achieved, at least in part, by the geospatial community through a series of long-running standards-making activities, one of the most notable being the work of the Open Geospatial Consortium (OGC). OGC is an international industry consortium of over five hundred companies, government agencies, and universities participating in a consensus process to develop publicly available interface standards. OGC�s abstract specifications and implementation standards are designed to support interoperable solutions that "geo-enable" a wide range of hardware platforms and software applications [26].
2.2.3.1 Web services and client-side API enhancements
The ESGF-CWT is adopting OGC’s Web Processing Service (WPS) interface standard for its next generation architecture [27]. WPS is essentially an xml-based remote procedure call (RPC) protocol for invoking processing capabilities as Web services. It has been used in the geospatial community for delivering low-level geospatial processing services. However, WPS can be generalized to other types of applications and data because of its simplicity: WPS uses a single operation (Execute) to invoke remote services; its two other operations (GetCapabilities and DescribeProcess) are used for discovery and to query services for the information necessary to build the signatures needed by Execute operations.
We will extend our Web service and API to support WPS. Adding WPS capabilities to our OAIS-based Web service will facilitate integration of NASA's capabilities into the ESGF by increasing machine-to-machine interoperability; enabling the API to consume WPS Web service endpoints will facilitate client software and workflow development. These enhancements will be of value to the ESGF community because they will improve interoperability at two levels within ESGF�s overall architecture.
Greater system-to-system interoperability improves connectivity and, in the case of WPS, allows the ESGF community to avail itself of WPS-compliant capabilities that exist within the geospatial community; having an API makes it easier to create toolkits, workbenches, workflows, and plug-ins tailored to the ESGF that can improve efficiencies and communication within the community.
2.2.3.2 Implementation approach
In the Web services universe there are two types of interfaces. On the service side, there is a system interface that maps the methods, functions, and programs that implement the service's capabilities to Hypertext Transfer Protocol (HTTP) messages that expose the service's capabilities to the outside world.
Client applications can consume these Web service endpoints to access services. The World Wide Web Consortium (W3C) views Web services as a way to insure machine-to-machine interoperability [28]. The precise messaging format can vary from community to community, often reflecting the specialized functions or audiences they serve. In our case, we have adopted the OAIS standard as described above.
There also are client-side APIs that consume Web service endpoints and abstract them into higher-order functions and methods that software developers find easier to use than the raw endpoints.
These client-side APIs, like the underlying Web services they consume, generally implement abstractions that are specialized to the needs of the audiences they serve. They can be used to build applications, workflows, and domain-specific toolkits, workbenches, and integrated development environments (IDEs).
Because software applications can bind to server-side endpoints as well as client-side libraries, both types of interfaces can be considered…
This is the start of the file's text. The full file is on GovTribe.
File details come from the government source that posted it.