20220518_SOW.pdf

PDF 374 KB Posted

Attached to
NCATS Secure Scientific Platforms Environment Support Federal contract opportunity
Solicitation number
75N95022R00089
Issued by
Department of Health and Human Services National Institutes of Health National Institute on Drug Abuse

About this file

This notice of proposed sole-source acquisition is for professional services and cloud hosting to support the NCATS Secure Scientific Platforms Environment. The incumbent, Palantir Technologies, will provide the commercial software platform, data integration and management services, as well as security, access controls, and governance functionality as defined in the attached Statement of Work. The period of performance is anticipated to be one to three years beginning on September 28, 2022. Interested parties may submit capability statements by the closing date to be considered for a competitive procurement, however the determination to award non-competitively to the incumbent is based on avoiding delays, leveraging their existing insights and FedRAMP certification, and capability to provide the platform as an integrated service.

View the file

On GovTribe

Work with this file on GovTribe

  • Download the original file
  • Contacts named in this file
  • Similar government files
  • Ask GovTribe AI about this file

Text version

Statement of Work

NCATS Secure Scientific Platforms Environment Support

1. Background and Objective

The NCATS Secure Scientific Platform Environment (the “Environment”) is a specialized cloud-based data aggregation and analytics enclave sponsored by the National Center for Advancing Translational Sciences (NCATS) that can integrate, manage, secure, and analyze any kind of scientific data, and provide secure, controlled access to internal and external collaborators.

Within the Environment, multiple institutes and centers (ICs) of the National Institutes of Health (NIH), Federal agencies, and Federal task forces integrate, manage, secure, and analyze all types of scientific data using dedicated platforms, and, equally importantly, make that data available in specific and controlled collaborations with each other and with external collaborators such as research hospitals, academic institutions, and industry stakeholders. The Environment is established through an indefinite delivery, indefinite quantity contract, with task orders issued to support the requirements of individual groups.

The Environment is a data aggregation and analytics asset established and administered by NCATS through the use of contractor support.

2. Requirements

2.1. Independently and not as an agent of the Government, the Contractor shall furnish all the necessary services, qualified personnel, material, equipment, and facilities, not otherwise provided by the Government, as needed to perform the following.

2.2. In accordance with the software requirements described below, the Contract shall provide the following:

2.2.1. The base NCATS Secure Scientific Platforms Environment;

2.2.2. Professional services as needed to support these requirements;

2.2.3. Cloud hosting to support these requirements.

2.3. Software Requirements

2.3.1. A commercial software solution deployable on day one of the project that can be configured within expedited timelines.

2.3.2. An open data architecture, where data always remains under the full control of NCATS and other data owners and can be easily exported in open, non-proprietary data formats via open APIs. The software should be built on an open, distributed microservices architecture with well-documented REST APIs and out-of-the-box connectors that are designed to seamlessly interface with other systems, adapt to meet evolving needs, and avoid system lock-in.

2.3.3. Proven multi-modal data integration capabilities, including the ability to rapidly ingest electronic health/medical record (EHR/EMR) data (including OMOP, TriNetX, ACT, PCORnet, etc.), pathology samples and assay data, unprocessed high-throughput drug screening (HTS) outputs, genomic data (including bulk RNA-seq, scRNA-seq, CITEseq, TCRseq, ChIPseq, Microarray, etc.), imaging data (e.g., MRIs), mass spectrometry, flow cytometry, and other data types used in basic and translational biosciences research and public health, such as administrative, financial, and grants data (e.g., nVision, IMPAC II, I2E, myDCEG, ARS, NIDB, PubMed), and supply chain data. Backed by configurable and interoperable data quality checks and a Git repository for data pipelining.

2.3.4. A multi-tenant secure enclave backed by configurable governance and access policies. Ability to host multiple individual tenants, with subsets of data shareable with different parties in the model as desired and in accordance with access controls. Access to a single view of multi-modal data based on user group and/or role.

2.3.5. Proven granular security controls with the ability for data owners to control all downstream uses of the originating data easily and dynamically, and the ability to conform to NCATS security policies. Ability to request and grant selective access to levels of data sensitivity in-platform based on a user’s intended purpose, implement configurable governance workflows depending on the requirements of data use and data transfer agreements, and audit user behavior after access has been provisioned.

2.3.6. The ability to maintain data and scientific provenance and reproducibility of all integrated data sources. Every resource (dataset, analysis, code, plot, report) contains provenance, metadata, and can be both traced back to the exact version of all upstream dependencies, and where the dependency tree can be easily replayed given new data or updated analysis logic, while still retaining prior versions and branches.

2.3.7. Dynamic data model, object-based search/discoverability, and analysis workflows, allowing easy definition of objects, properties, and links that propagate from a source table. Solution provides natural ways to move between tabular and object-oriented interfaces and data analyses. Proven ability to integrate multi-data model data into a harmonized data model (such as OMOP).

2.3.8. Intuitive, highly configurable user interfaces that have been effectively configured and utilized by technical bioinformaticians, cheminformaticians, data engineers, and data scientists, as well as less technical biologists, chemists, clinicians, analysts, program managers, administrators, and other users.

2.3.9. Ability to perform advanced analytics and informatics (including management of machine learning and other models) in a user's preferred open coding language, as well as in point and click tools, all within the same environment.

Ability to generate no-code analytical templates enabling less technical users to conduct complex analyses and generate visualizations.

2.3.10. A variety of proven, secure, and user-oriented configurable applications and workflows backed by configurable access controls and up-to-date data, including:

2.3.11. Patient digital twin capability for tens of thousands of patients backed by multi-modal data (e.g., clinical, imaging, and tumor sequencing/mutation data).

2.3.12. Laboratory sample and result tracking system.

2.3.13. Streamlined research funding analysis, tracking, and reporting interface for improved funding estimates and budget oversight.

2.3.14. Genomic pipeline code templates for generating analysis and publication-ready visualizations.

2.3.15. Application for sharing and re-use of research outputs. A centralized space where logic, datasets, models, and other research outputs can be securely shared, discovered, and re-used by other researchers. Usage of each artifact should automatically be tracked to ensure attribution for contributing researchers. The application should ensure that use of shared artifacts is compliant with governance rules around data use.

2.3.16. Application for creation and management of code sets. This should allow automatic integration and updates for multiple terminologies and include the ability to version code sets, track their usage, and document them with metadata such as their provenance and intention. Changes to vocabularies should be tracked and users should be alerted when these changes impact existing code sets.

2.3.17. Collaboration capabilities enabling teams comprising of a range of technical and less technical roles to work seamlessly and concurrently on the same data, build on insights, merge similar analytical paths, and track progress in one place.

2.3.18. Ability to scale flexibly with increasing users, data, and pipeline complexity, while providing fine-grained ways to adjust resource consumption—including auto-scalable containerized compute. Proven ability to scale up to thousands of users (including thousands of potential outside collaborators globally), petabytes of raw and processed data, daily updates in the terabytes, and management of thousands of complex bioinformatic pipelines requiring processing components developed in a variety of languages and environments.

2.3.19. Proven interoperability with NCATS’ current IT investment landscape. Includes omnipresent APIs and plugin points that allow the system to keep up with the changing needs of NCATS and partners, and support both third-party software and other analytic applications. NCATS also requires the ability to independently implement new configurations, plugins, integrations, and extensions to meet new and unforeseen needs and interface with external systems. Ability for less technical users to develop workflows and applications in a low-code/no-code suite.

2.3.20. Adherence to Detailed Technical Specifications. The solution shall incorporate the specifications listed in the Detailed Technical Specifications document, attached to this SOW.

2.4. End – nothing follows.

Detailed Technical Specifications Secure Scientific Platform Environment

National Center for Advancing Translational Sciences (NCATS) National Institutes of Health (NIH)

Table of Contents

Technical Capabilities

1. Solution Architecture, Data Integration, and Data Management

1.1 Unified, Commercial Software Data Management Architecture

1.3 Flexible Data Integration Engine & Comprehensive Data Type Experience

1.4 Data Storage, Access, and Catalog

1.5 Full Data Provenance and Schema

1.6 Data Transformation Management

1.7 Data Deposition Back to Established Repositories

1.8 Cohesive User Interfaces for Data Management

2. Data Analysis, Discovery, and Other Workflows

2.1 Analytic Tool Suite

2.2 Search and Exploration

2.3 Web Application and Workflow Builders

3. Security, Administration, and Governance

3.1 Security, Access Controls, and Data Governance

3.2 Platform Management and Administration

3.3 End – nothing follows

Technical Capabilities

1. Solution Architecture, Data Integration, and Data Management

1.1 Unified, Commercial Software Data Management Architecture

1. A core architecture that consists of a suite of distributed microservices that perform fundamental tasks that the rest of the platform, external plugins, and integrations depend on

2. Where a dataset should operate as a high-level conceptual container that manages the full versioned history of a data file (or files) including its history of changes, branches, and access policies

3. Where each dataset maintains a series of transactions, which are the records of the varying logic that has modified the underlying data files over time, with pointers to the specific versions or views of those data files that are physically stored in a distributed file system

4. Where a dataset may contain branches that allow transactions to modify data files in a sandboxed fashion

5. Provides the ability for development and analysis on different versions or subsets of a dataset to occur simultaneously without conflicting

1.2 Scalability and Hosting

1. A FedRAMP Moderate-authorized solution that is offered as software as a service in

AWS US East/West and/or AWS GovCloud

2. The solution’s cloud architecture should provide:

a. Separated cloud, customer, and management networks

b. Isolated virtual private clouds (VPCs) that prevent commingling of data

c. Hardened and filtered entry and egress points between cloud, customer, and management networks

d. Sophisticated provisioning, patching, and monitoring tooling

e. 24/7/365 support and incident response team

f. Automated backups and robust disaster recovery

g. Events logging and log aggregation from services across hosts

h. Interface providing access to consolidated logs and performance metrics allowing standardized monitoring and tuning of services

3. By design, all components of the solution should be horizontally scalable, allowing for high availability and configurable load balancing to optimize performance and facilitate additional user scale, data scale, data complexity, and compute complexity

4. The solution should offer auto-scalable containerized compute, backed by Kubernetes, and distributed parallel-compute engines such as Apache Spark to facilitate dynamic scaling based on compute demand

5. All core services should operate in a high-availability clustered environment using cloud node hosting on Linux platform

6. All data should be stored in one of several highly available and redundant distributed data stores

7. All compute is scheduled and run on open-source cluster-computing frameworks and schedulers, with wrapper libraries and appropriate APIs to allow any plugin or integrated application to run on the same infrastructure

1.3 Flexible Data Integration Engine & Comprehensive Data Type Experience

1. Data ingestion, integration, transformation, and orchestration tools that include:

a. Data Ingestion

i. Support for programmatic (automated and either push/pull) ingestions and ad hoc ingestions (user-initiated, either via user interface, or backend script)

ii. Support for a broad range of source systems, including, but not limited to HDFS, S3, local file systems, RDBMs, streaming systems, JDBC-enabled systems, etc.

iii. Out-of-the-box connectors for accelerated syncs with standard system types (e.g., SAP, Salesforce, etc.) and NIH-specific systems (including Biowulf, HPCDME)

iv. An API for writing new ingestion adaptors from novel or custom data systems

v. Ability to install a daemon in a source system environment entirely within the control of the data owner. This daemon should let the system owner manage credentials, connection details, and files

vi. A central coordinator service that manages the configuration and execution of ingestion jobs through communication with the agent.

The coordinator compiles all job specifications and provides them to the agent, which then performs the system-side aspects of the ingestion task

vii. High resilience to provide robustness in the face of common failure modes during ingestion – poor network connectivity, disk failures, time-outs, and more. Users can choose to configure whether to triage failures in an automated way, or force manual intervention via alerts

viii. A user interface where syncs and jobs can be monitored and configured, and open-ended queries can be performed to test ingestions. Configurations should include the nature of the source system, the relevant destination dataset in the solution, the path that the data should take in transit, and any in-flight modifications to the data

b. Data Transformations

i. Ability to support thousands of data transformations while automatically tracking data lineage

ii. Web-based integrated development environment (IDE) allowing the writing of data transformation code and publishing to GitHub

iii. Code repository capability for data pipelining that is backed by

Git with all features of the protocol expressed (cloning repositories for local development in an IDE of choice, branches, commits, pull requests, etc.)

iv. Continuous integration server to facilitate checks and user-authored unit tests after each code change. Pluggable modules to let users write data transformation code in an open language of their choice (Java, Python, R, SQL, etc.)

v. Templating tools allow users to catalog and parameterize commonly used code

vi. Integrated version control and branching ensure that individuals across an organization can collaborate on logic development without disrupting existing production workflows

vii. Deep branching enables code and underlying data to be experimented with in parallel. Data not requiring any changes while experimenting can be configured to default using a “fallback” branch to avoid unnecessary data duplication

viii. Unique branching of the data itself means that logic branches with disparate transformation approaches can co-exist

ix. Incremental or snapshot transactions available to support regular syncing with a diverse set of data source systems

x. Compilation module that can transform the code into distributed compute operations on the distributed execution environment

xi. Continuous integration backend allowing each change to be committed to a distributed version control system, and versioned artifact to be produced

xii. A distributed environment execution manager that provides simplified access to the distributed cluster, and fully manages necessary dependencies, libraries, and conflicts

xiii. On anticipated roadmap: Configure a user accessible internal container registry where user-authored images can be uploaded, and a build system which can orchestrate jobs in containers instantiated from these images.

xiv. Platform-internal artifact repositories where users can upload code libraries that can be referenced as dependencies in the platform's code authoring environments to perform analysis and develop pipelines.

xv. On anticipated roadmap: Configure a code scanning governance capability for an administrator to review the set of artifacts (container images, user-written code libraries) across the platform that contain vulnerabilities, which artifacts are currently running in production, and who owns them.

c. Data Orchestration

i. Ability to add data health checks at scale (i.e., thousands of checks across thousands of pipelines defined in code or via user interface) including, but not limited to null checks, distribution as expected, no duplicates, reasonably partitioned for optimal performance, foreign-key relationships

ii. Ability to integrate checks with existing (e.g., open source) health check libraries (e.g., OHDSI data quality checks)

iii. Ability to configure data quality metric visualizations and dashboards for individual data sources, groups of data sources, or all data

iv. Ability to add triggers, to run certain data operations upon receiving an alert from a message queue

v. Open APIs that enable outside services to trigger actions and updates within platform

vi. Advanced queuing and resource channels allow for intelligent sharing of resources among simultaneous transformations and prioritization of critical jobs

vii. User interface for context on running jobs and rich metrics on the performance of historical transformation activity to inform scheduling and refactoring

2. Ability to ingest multi-modal data of high-velocity, high-throughput, and any format

a. The solution must be designed to integrate any type of data out-of-the-box including bioscience and non-bioscience data. Additionally, the solution must be preconfigured with specialized tooling to work with the following types of data:

i. Electronic health/medical record data (including OMOP, TriNetX, ACT, PCORnet, etc.)

ii. Pathology samples and assay data

iii. Flow cytometry

iv. Mass spectrometry (MaxQuant, PD, Peaks, etc.)

v. High-throughput drug screening (HTS), including raw plate reader output

vi. Imaging data (e.g., MRI)

vii. Genomic data (bulk RNA-seq, scRNA-seq, CITEseq, TCRseq, ChIPseq, Microarray, etc.)

viii. Financial and budget data (e.g., grant awards)

ix. Supply chain data

3. Ability to model ingested data according to a dynamic ontology that reflects data according to user workflows

a. An ontology management system for the configuration and re-configuration of data modeled as objects and their properties (e.g., “patient” and “COVID+”)

b. The ontology should be configurable to correspond to both existing data models such as OMOP, TriNetX, PCORnet, or ACT/i2b2, and user group-specific workflows—such that data ingested in different data models can be standardized into a single data model (e.g., OMOP)

1.4 Data Storage, Access, and Catalog

1. Distributed and redundant data storage via preferably S3 or Cassandra that provides unique IDs and a REST API, and integrates with the core security and auditing control subsystem

2. Key-value storage that provides:

a. Seamless integration with the security and auditing subsystem so that all resources stored are security-aware throughout their life cycle

b. REST API with bucket/path-based addressing system, allowing applications to easily register, store, and retrieve resources throughout their life cycle

c. Reliable access to billions of records for thousands of users via highly available and scalable architecture

3. Indexing that:

a. Provides a search mechanism as well as a generalized high-scale index layer for the solution and its ecosystem

b. Is security-aware and is closely tied to the core security and auditing subsystem so that search requests – regardless of origin – can be securely managed

c. Makes tasks such as live re-indexing and security accessible

d. By default, will perform search optimizations, such as pre-filtering for paged searches that provide sparse results

e. Can work with the solution's core build service to implement a distributed indexing worker that allows for systematic indexing of jobs as they build and promote new resources

4. A data access API that provides a simplified interface for user ‘PUT’ and ‘GET’ requests for data by managing the underlying File System requests among the solution services in a particular installation

a. The data access API is required to interface with the underlying distributed file store that is being used as part of the solution installation to return the actual resource byte array when requested

5. A data catalog that provides:

a. Highly available metadata store that keeps track of datasets and can handle datasets composed of multiple data files and the associated transactions to maintain ACID properties on datasets

i. Catalog illustrates metadata affiliated with datasets such as size, date of last modification, and branching history

b. Tagging to highlight dataset attributes. Dataset-level tags to delineate raw from derived datasets, with the additional option to mark datasets as “endorsed” by Subject Matter Experts for further downstream analysis

c. Filtering by tags or data usage to rapidly find the most valuable or relevant data

d. Virtual access to a complete dataset that may be composed of many underlying files via many transactions

e. Reconstruction of the full dataset via negotiation of transactions, whether they be a point-in-time exact view of the data, an append or addition to an existing dataset, or a deletion from a dataset

f. A locking mechanism for managing concurrent access to the same dataset, via vector clock and operational transforms (OT)

g. Data and metadata are machine-readable, stored in open formats, and self-describing. All data is be embedded with its own schema for accessibility (e.g., using Apache Parquet) to external systems

h. API access to metadata, complete constructed datasets, or files and transactions therein

1.5 Full Data Provenance and Schema

1. Data in the solution should be inherently schema-less, with additional schema and provenance information provided by dedicated services

2. Each dataset should have a unique and persistent resource identifier—which remains unchanged regardless of dataset version or name

3. Human-readable descriptions of datasets on top of rich metadata—such as schema, author, creation date, and security markings—which the platform should automatically capture alongside data

4. A metadata service that maintains the explicit association between metadata and the data it describes, including the schema definition for each file, so that users can identify data based on its metadata

a. Dataset metadata includes the identifier of the dataset and can be uniquely identified by the primary key of the dataset

b. Automatic recording of intrinsic metadata, such as the complete branched history of each file (e.g., changes made during cleaning and transformation) to enable users to find information about schema changes or dataset statistics

c. Ability for users to assign contextual metadata to resources in an ad hoc manner, such as by adding descriptions to specific columns or tags to datasets

5. A schema management service that:

a. Stores and provides the schema definition for each data file, including variations in the schema that may have occurred across the versioned and branched history of the resource

b. Stores type information that is used in the compute environment. The types can be inferred with an advanced type inference engine; the service should also have a graphical interface for the user to validate, correct, or modify the types manually

c. Tags and “type classes” on tables and columns, allowing semantic meaning to be conveyed and propagated. As an example, a column marked “chemical-compound” will result in all downstream datasets using that column knowing that they also contain a column containing a “chemical-compound,” which can then be associated with specific actions

6. Ability to generate interactive graphical visualizations of the live updating data dependency and provenance tree for all datasets in the solution (while respecting access controls) and provides for any dataset:

a. The ability for users to inspect the transformations that led to it

b. All upstream and downstream datasets to be viewed as well

c. The pipeline graph to be manipulated by users including labeled in various ways

1.6 Data Transformation Management

1. A data transformation management service that provides:

a. A resilient topological graph of dataset dependencies, using metadata captured in the data catalog, provenance, and schema information, consisting of virtual build nodes with job specifications (via a worker process) on how to update a given dataset and dependents

b. A way to dynamically compute staleness of datasets, to determine what needs to be run

c. Advanced topological sorting of datasets to determine the optimal order of job runs, and to allow parallelization of job runs

d. Modular dispatch to dynamically bind and dispatch a worker and an associated executor environment appropriate for the task. This would allow a multi-step build to be run cohesively in a variety of languages in the same run

e. A way to perform efficient selective insert or replacement of data at the row level, pursuant to access controls

1.7 Data Deposition Back to Established Repositories

1. A data transformation management service that provides:

a. A mechanism to intermittently synchronize select data with external relational database management systems (RDMS)

b. A mechanism to intermittently and/or automatically synchronize select data with external services (including public-facing dashboards, analytical solutions, or other software) through APIs

c. Support for the persistence of complex object models through RDMS and API

d. Robust transactional framework to ensure the integrity of the data transfer process

e. Data management process to track staleness and synchronization of datasets, to determine what data needs to be deposited

1.8 Cohesive User Interfaces for Data Management

1. The solution provides intuitive access to all services through the web, and includes:

a. A unified web environment sharing an application frame and common design language

b. Granular permissioning of actions via integrated access controls

c. Interface components that facilitate easy use by both technical and non- technical users alike

2. The solution should focus on the ability for technical users (bioinformaticians, cheminformaticians, data engineers, data scientists, epidemiologists) to be able to collaborate on the same data management and analytic infrastructure as non-technical users (biologists, chemists, clinicians, analysts, program managers, administrators, leadership, etc.), and should provide:

a. The ability to comment on any resource (dataset, analysis, dashboard, compute node, etc.) in a collaborative chat environment, including the ability to tag and notify

b. A fully integrated data issues interface where users can file, discuss, and resolve data quality issues in an interface that can embed flagged datasets, columns, and more

c. Ability to selectively share any analysis, dashboard, or workflow across the platform and choose what permission level to grant (viewer, editor)

d. Shared source of truth with robust change tracking for inventories, protocols, and other data that would otherwise be shared across teams manually by email or spreadsheet

e. Ability for users to discover and reuse analyses and artifacts across applications (e.g., research metrics, conceptual definitions, etc.)

f. Ability in the UI to easily create templates that allow a non- technical user to apply complex analysis workflows and logic paths

2. Data Analysis, Discovery, and Other Workflows

2.1 Analytic Tool Suite

1. A solution that enables individuals across user organizations and authorized partners to work with the same data by providing a suite of analytical tools appropriate for the full range of technical ability. The suite should provide the following functionality:

a. Reporting

i. Ability to unify live views of charts, visualizations, and datasets from various solution services in a single location

ii. Allow users to write collaborative documents, include more context by attaching images or rich text descriptions, and configure long-lived, real-time dashboards of key metrics

iii. Ability for reports to link to live data backing those reports, allowing users to inspect the live provenance of the data in the user interface, as well as all data backing the report

b. Top-Down Analysis at Scale

i. An intuitive and visual path-based, point-and-click environment for tabular manipulation and enrichment

ii. Flexible widgets that allow users to apply a broad range of filters, manipulations, expressions, enrichments, joins, and visualizations (including heatmaps, distributions, histograms, bar charts, scatterplots, regressions, PCA, Venn diagrams, time series, dendrograms, and more)

iii. A path-based model that allows for previous steps to be modified and all subsequent steps to be rebuilt based on those new parameters

iv. The ability for paths to be cloned, branched, and saved as their own objects

v. The ability for iterative steps to be saved automatically while allowing users to undo actions and return to a previous state of the analysis

vi. Have a documented API, allowing developers to add boards as external plugins without requiring downtime

c. Code Workbooks

i. A visual directed acyclic graph, where each node is a raw dataset, a dataset resulting from a code manipulation depending on any of a number of upstream datasets, or a visual plot

ii. The code should be able to be written in Python, R, SQL, or other open common languages, and have direct API bindings to the underlying distributed compute infrastructure (e.g., Spark)

iii. Resulting data from code workbook computation nodes should:

1) Automatically become part of the solution’s data environment

2) Be registered in the data catalog

3) Be immediately available for access in all other applications

d. Modular Code Templates

i. Ability to save common code segments or paths as templates, thereby very quickly enabling any user to express complex analytical ideas with point-and-click tooling

e. Model Management Suite

i. A machine learning framework to enable users to write, test, deploy, manage, and evaluate advanced models on top of live, high-quality data with:

1) A front-end view to expose performance metrics and key information, such as validation statistics, model stages, key parameters, and configurable metadata

2) Core ontology, branching, versioning, and lineage capabilities, including a history browser, a deployment workflow, and a cross-version and branch view for comparing model variants with different parameters

3) Multiple approaches to model training, from open-source architectures (e.g., TensorFlow) to domain-specific solutions via open APIs

4) Models can be trained programmatically, or results can be surfaced with context for human review

5) Models can be constantly validated against and retrained on live data to ensure they remain relevant and useful

2. The analytic tool suite should be easily configurable with a GPU environment as an alternative for standard CPU environments for training of deep learning models (e.g., BERT) or deep learning for feature extraction from medical images

3. The solution should support for multi-tenant usage of code workbooks and top-down analyses, including supporting of full branching log, whereby a user can apply their analytical methods on a branch (with logic and relevant datasets all respecting the branching mechanism), and later propose a merge back to the main branch

4. Resulting data from path-based analyses automatically become part of the solution data environment and immediately available for access in all other applications

5. Ability to selectively share and collaborate throughout the solution’s services, including the ability to add comments, tag and notify users, and report and track issues on identified data columns

2.2 Search and Exploration

1. Beyond simply curating datasets in a central data catalog, the solution must also index analysis and reports such that a user can easily discover data and insights about a topic of interest without knowing beforehand where to navigate

2. Access control settings at the project and individual resource level need to ensure that a user only sees information they are approved to discover

3. Robust data tagging and project hierarchies need to allow users to then narrow global search results down to the information most relevant to their task at hand

4. In addition to global search, the solution should provide a variety of exploratory interfaces to understand the data available, the logic applied to it, and the relationships among different pieces of information:

a. An interface that allows all resources within the solution to be inherently linked to the logic applied to transform the raw information into the subsequent ontology tables

b. An interface that allows search across all raw and transformed datasets so that it is straightforward to search globally by column names to identify where particular information is present

c. An interface (beyond tabular data) to facilitate mapping of information onto conceptual objects with attributes, behavior, and relationships to other objects within the data model

d. An interface that allows users to quickly author a standalone web application which pulls cached data into an interface optimized for carrying out a particular workflow. This interface should thereby streamline workflows for when certain information is accessed repeatedly in the same manner

2.3 Web Application and Workflow Builders

1. Beyond native data exploration, code authoring, and analysis interfaces, the solution should include the following features to enable authorized users to intuitively configure applications, workflows, and dashboards:

a. A development environment which allows any authorized user to create data-driven interactive applications

i. Includes a rich “what you see is what you get” environment, allowing the composing of an interface using extensible “widgets,” including charts, controls, containers, text, images, time series, and others

ii. Each widget can be synchronized to live data with straightforward SQL queries, and complex workflows can be built by defining JavaScript functions that determine interaction between widgets and reaction to user-triggered events

iii. Natively compatible with reading from a SQL connection, via REST API, or anything that provides a JSON response, offering the ability to tie in external systems or APIs

b. A point-and-click application that allows non-technical authorized users to configure step-by-step ontology-backed workflows and dashboards that are integrated with logic-writing capability

i. Enables technical users to define actions that can be taken by users on objects in configured workflows (e.g., automatic assignment of tasks)

2. As all other solution components, applications built in these builders should fully respect the propagated dynamic access controls of the data dependencies

3. Security, Administration, and Governance

3.1 Security, Access Controls, and Data Governance

1. The solution shall have a security and auditing subsystem to enforce security in the solution. The solution must provide authorization of resource access to users who have been authenticated

a. The platform should capture immutable audit logs on all resource access requests throughout the system, including import, read, write, search, export, and deletion attempts along with the user, time, date, and action

2. Security and access controls that include:

a. Multi-source authentication (enterprise LDAP, PKI, Active Directory, etc.) as well as Single Sign-On (SSO) capabilities, including existing integration with NCATS UNA and login.gov

b. Exposed OAuth 2.0 interfaces to facilitate authorization with external applications with either Authorization Code Grant or Client Credentials flows

c. API integrations with:

i. Privacy-preserving record linkage (PPRL) solutions providers to facilitate secure data linkage

ii. External third-party systems for pass-through of credentials to facilitate rapid access to external data and/or the running of jobs on external systems directly from the platform

d. Access controls to secure each piece of information individually in addition to traditional blanket permissions across entire data sources

e. Configurable Projects and Organizations structures backed by access controls

i. The solution should have a collaborative Projects space for organizing users, files, and folders. Each Project should have a default role, which defines the user groups that can access the project and the actions they are allowed to take (e.g., view only permissions). The administrator should be able to grant more granular permissions to specific users and groups in the Project, down to the individual resource level

ii. The solution should have Organization boundaries that supersede Projects. Organizations should provide collaborative spaces that can be limited to a specific set of users and user groups through access controls

iii. Different users should be able to see different lists of Projects and Organizations, as determined by each user’s security permissions

f. Ability to assign individual users granular permissions to data access (including ownership, read, write, access denied, etc.)

g. Role-based access controls (RBAC) which are assignable to individual users or group of users

h. Ability to restrict post-harmonization data access to data owners involved in harmonization for data quality checks

i. Ability to validate harmonized data from across data owners in a controlled environment limited to administrators and release data to broader user base upon validation

j. Ability to assign sensitivity markings to data at ingestion and retain markings through downstream processes (including transformation and analysis) to automatically prevent data from being used by projects or users that do not have the necessary approvals while permitting users to share data and results that are permissible across projects

k. An in-platform data use and data download request and approval process to facilitate purpose-based data access and use

i. The solution should have the ability to enable users to request access to data at a given tier of information and provide justification using an in-platform form

ii. Administrators should be automatically notified of the request as well as be able to grant or decline the request within the solution

iii. The platform should automatically capture any relevant metadata about the request for transparency and auditability

l. Access and role-based access controls that:

i. Can be assigned or revoked by administrators at any time (including data owners) and/or

ii. Can be inherited from source systems and updated at any time

iii. Are based on the principle of least privilege and require privileged access for access administration

m. Propagation and inheritance of access controls to all data and analyses derived directly or indirectly from the originating data, whether those data transformation or analyses are performed:

i. In a standard language (Java, SQL, Scala, Python, etc.)

ii. A technical scripting language (R, SAS, etc.)

iii. A point-and-click analytical environment

n. Data is encrypted at rest and in transit via mechanisms compliant with

Federal Information Processing Standard (FIPS) Publication 140-2

3. FedRAMP-compliant ATO. Comply with FedRAMP Security Assessment and

Authorization (SA&A) requirements and ensure the information system/service under this contract has a valid FedRAMP compliant (approved) authority to operate (ATO) in accordance with FIPS Publication 199 defined security categorization. If a FedRAMP compliant ATO has not been granted, the Contractor shall submit a plan to obtain a FedRAMP compliant ATO as determined by the contract officer

3.2 Platform Management and Administration

1. Continuous integration and monitoring, including:

a. Live downtime-less upgrades enabled by high availability services

b. Centralized monitoring of errors and performance

c. Dynamic scaling of compute and storage to handle changes in usage in a resource-efficient manner

2. Platform administration and a solution configuration manager that allows:

a. Maintenance of a complete and ongoing understanding of the state of the environment:

i. Including the hosts, the services, the nodes for distributed and High Availability services

ii. Where those nodes are deployed

iii. What versions they are on

iv. Version compatibilities, etc.

b. Secure management of cryptographic secrets

c. The concept of ‘roles,’ which declare which services and APIs produce and consume in a declarative manner

d. All services must have a ‘Life Cycle Model’ to indicate the expected or target state at any given time along with the current state (for example, running or upgrading). Service status is a more detailed reporting of the per-host service state that is updated with each agent action

e. Backend command line interfaces (CLIs) that provide full management functionality

f. Front-end graphical interface that provides full management functionalities

3. Standardized microservices:

a. Layout in the filesystem is standardized, allowing a platform administrator to easily find service information and logs

b. Microservices can be deployed via the platform management services directly onto hosts, or, ideally, as lightweight containers using standard container deployment orchestration libraries

c. The orchestration framework should tightly integrate with the security and auditing subsystem

d. These containers allow only a limited number of gateways into the host environment, limiting the potential vectors that could compromise one or more hosts

3.3 End – nothing follows.

20220518_Technical_Specifications.pdf
Technical Capabilities
1. Solution Architecture, Data Integration, and Data Management
1.1 Unified, Commercial Software Data Management Architecture
1.3 Flexible Data Integration Engine & Comprehensive Data Type Experience
1.4 Data Storage, Access, and Catalog
1.5 Full Data Provenance and Schema
1.6 Data Transformation Management
1.7 Data Deposition Back to Established Repositories
1.8 Cohesive User Interfaces for Data Management
2. Data Analysis, Discovery, and Other Workflows
2.1 Analytic Tool Suite
2.2 Search and Exploration
2.3 Web Application and Workflow Builders
3. Security, Administration, and Governance
3.1 Security, Access Controls, and Data Governance
2.3 Web Application and Workflow Builders
3. Security, Administration, and Governance
3.1 Security, Access Controls, and Data Governance
2.3 Web Application and Workflow Builders
3. Security, Administration, and Governance
3.1 Security, Access Controls, and Data Governance
2.3 Web Application and Workflow Builders
3. Security, Administration, and Governance
3.1 Security, Access Controls, and Data Governance

File details come from the government source that posted it. Updated .