Sources Sought Notice_2021-07-21.pdf

PDF 340 KB Posted

Attached to
NCATS Secure Scientific Platforms Environment Support Federal contract opportunity
Solicitation number
75N95021R00053
Issued by
Department of Health and Human Services National Institutes of Health National Institute on Drug Abuse

View the file

On GovTribe

Work with this file on GovTribe

  • Download the original file
  • Contacts named in this file
  • Similar government files
  • Ask GovTribe AI about this file

Text version

SAM.gov Contract Opportunities

SOURCES SOUGHT NOTICE

1. Solicitation Number: 75N95021R00053

2. Title: NCATS Secure Scientific Platforms Environment Support

3. Classification Code: DB10

4. NAICS Code: 511210 Software Publishers, $41.5 million

5. Description:

This is a Sources Sought notice. This is NOT a solicitation for proposals, proposal abstracts, or quotations. The purpose of this notice is to obtain information regarding the availability and capability of all qualified sources to perform a potential requirement.

This notice is issued to help determine the availability of qualified companies technically capable of meeting the Government requirement and to determine the method of acquisition. It is not to be construed as a commitment by the Government to issue a solicitation or ultimately award a contract. Responses will not be considered as proposals or quotes. No award will be made as a result of this notice. The Government will NOT be responsible for any costs incurred by the respondents to this notice. This notice is strictly for research and information purposes only.

The Government is especially interested in determining (1) the availability of qualified small business sources; (2) whether they are small businesses;

HUBZone small businesses; service-disabled, veteran-owned small businesses;

8(a) small businesses; veteran-owned small businesses; woman-owned small businesses; or small disadvantaged businesses; and (3) their size classification relative to the North American Industry Classification System (NAICS) code for the proposed acquisition. Your responses to the information requested will assist the Government in determining the appropriate acquisition method, including whether a set-aside is possible.

Background:

The mission of the National Center for Advancing Translational Sciences (NCATS) is to transform the translational science process in order to get more treatments to more patients more rapidly. Critical to this mission is the ability to manage data, generate insights, and foster collaboration across National Institutes of Health (NIH) Institutes and Centers and with external partners such as research hospitals, academic institutions, other government agencies, and industry stakeholders. NIH, NCATS, and NCATS collaborators are among the world's leading researchers, and they have access to large, cutting-edge data sources. An increasingly complex challenge is to translate these disparate, ever-evolving data sets into actionable insights that accelerate the pace of science and clinical development. Tools are needed to enable teams to ask complex, multi-faceted questions and to collaborate more seamlessly and securely across various teams.

NCATS has established the NCATS Secure Scientific Platform Environment, a specialized cloud-based data aggregation and analytics enclave that can integrate, manage, secure, and analyze any kind of scientific data, and provide secure, controlled access to internal and external collaborators.

Within the Environment, multiple NIH ICs, Federal agencies, and Federal task forces integrate, manage, secure, and analyze all types of scientific data using dedicated platforms, and, equally importantly, make that data available in specific and controlled collaborations with each other and with external collaborators.

The Environment was developed through a pilot effort that validated the use of Palantir Technology' Inc.'s Foundry platform-as-a-service as the basic operating system for the Environment. Foundry addresses the fundamental information security principles of confidentiality, integrity, and availability as described by the National Institute of Standards and Technology (NIST).

The Environment is a mission-critical data management and analysis environment for multiple data management efforts by several NIH ICs and Federal agencies under the leadership of NCATS. The Environment currently uses Palantir Technologies, Inc’s Gotham platform configured with the Palantir Foundry platform.

This platform-as-a-service (PaaS) currently supports NCATS, the National Cancer Institute (NCI), the President's Emergency Plan for AIDS Relief (PEPFAR), and the National COVID Cohort Collaborative (N3C). The Environment has integrated hundreds of live intramural and third-party data sources in support of dozens of ongoing, critical scientific projects that rely on continuous access, data, and analyses within the Platform. The Environment is now the standard means for accessing and collaboratively analyzing NCATS screening data for dozens of investigators at both NCATS and NCI, and for accessing and analyzing RNASeq and several kinds of proteomics data at NCI.

The Platform also supports clinical applications supporting NCATS (the Clinical and Translational Science Awards (CTSA) and Rare Diseases Clinical Research Network (RDCRN)), NCI, and PEPFAR.

In 2020, in response to the COVID-19 emergency response, the Platform was leveraged to support the NCATS National COVID Cohort Collaborative (N3C), https://ncats.nih.gov/n3c. N3C maintains one of the largest collections of clinical data related to COVID-19 symptoms and patient outcomes in the United States.

The N3C is a partnership among more than 70 institutions, including more than 40 NCATS-supported CTSA program hubs, the National Center for Data to Health (Cd2H), and the NIGMS-supported Institutional Development Award Networks for Clinical and Translational Research (IDeA-CTR), with overall stewardship by NCATS. As of July 12, 2021, the N3C Data Enclave maintains over 7.2 billion rows of clinical data covering more than 6.3 million patients, over

2.1 million of them COVID-positive. More than 224 research projects are underway. To protect privacy, this data consists only of limited data sets, de-identified data sets, and synthetic data sets; privacy is ensured through the use of data anonymization through a separate contract with an honest data broker.

The NCATS Secure Scientific Platforms Environment is a managed Platform-as-a-Service environment (PaaS) supported by Palantir Technologies, Inc., under IDIQ Contract number 75N95020D00016, awarded in 2020, with task orders issued to support the requirements of individual groups. The solicitation for that acquisition is archived on SAM.gov Contract Opportunities (inactive opportunities).

Purpose and Objectives: The National Center for Advancing Translational Sciences (NCATS) requires a commercial cloud-based data aggregation and analytics Platform as a Service (PaaS) environment that can integrate, manage, secure, and analyze any kind of scientific data, and provide secure, controlled access to internal and external collaborators.

Project requirements: The vendor must be able to provide its commercial software for enterprise data integration and mangaemtn to meet the following requirements:

1. A commercial software solution deployable on day one of the project and that can be configured within expedited timelines. Respondents must possess FedRAMP authorization so that an agency ATO can be granted upon award of a contract.

2. An open data architecture, where data always remains under the full control of NCATS and can be easily exported in open, non-proprietary data formats via https://ncats.nih.gov/n3c open APIs. The software should be built on an open, distributed microservices architecture with open, well-documented REST APIs that are designed to seamlessly interface with other systems, adapt to meet evolving needs, and avoid system lock-in.

3. Proven data integration capabilities, including the ability to rapidly ingest unprocessed high-throughput drug screening (HTS) outputs, genomic data (including DNA sequencing, RNA-Seq, miRNA, etc.), mass spectrometry, flow cytometry, and other data types used in basic and translational biosciences research.

4. The ability to maintain data and scientific provenance and reproducibility of all integrated data sources where every resource (dataset, analysis, code, plot, report) contains provenance, metadata, and can be both traced back to the exact version of all upstream dependencies, and where the dependency tree can be easily replayed given new data or updated analysis logic, while still retaining prior versions and branches.

5. Dynamic data model, object-based search/discoverability and analysis workflows, allowing easy definition of objects, properties, and links that propagate from a source table, and provide natural ways to move between tabular and object-oriented interfaces and data analyses.

6. Intuitive, highly configurable user interfaces that have been effectively utilized by technical bioinformaticians, cheminformaticians, and data scientists, as well as less technical biologists, chemists, and other scientists.

7. Ability to perform advanced analytics and informatics in a user's preferred coding language, as well as in non-code-based point and click tools, all within the same environment.

8. Collaboration capabilities enabling teams comprising of a range of technical and less technical roles to work seamlessly and concurrently on the same data, build on insights, merge similar analytical paths, and track progress in one place.

9. Ability to scale flexibly with increasing users, data, and pipeline complexity, while providing fine-grained ways to adjust resource consumption. Proven ability to scale up to thousands of users (including thousands of potential outside collaborators globally), petabytes of raw and processed data, daily updates in the terabytes, and complex bioinformatic pipelines requiring processing components developed in a variety of languages and environments.

10. Proven granular, security controls with the ability for data owners to easily control all downstream uses of the originating data, and the ability to conform to NCATS security policies.

11. Proven interoperability with NCATS current IT investment landscape, and includes omnipresent APIs and plugin points that allow the system to keep up with the changing needs of NCATS, and support both third-party software and other analytic applications. NCATS also requires the ability to independently develop new configurations, plugins, integrations, and extensions to meet new and unforeseen needs and interface with external systems.

Anticipated period of performance: The anticipated period of performance is one year.

Other important considerations: In addition to the PaaS environment requirements listed above, detailed technical requirements are provided in the attached Technical Solution Requirements document. Do not address them individually in your capability statement, but be aware that these will be mandatory requirements in any solicitation that is issued.

Capability statement /information sought. Respondents must be registered in the System for Award Management (SAM), and address the following questions:

• Provide your DUNS number, organization name, address, and point of contact. If your business is a small business, include your size and type of business (e.g., 8(a), HubZone, etc., SDVOSB, etc.) pursuant to the applicable NAICS code

• Provide a capability statement that addresses the project requirements listed above, with particular attention to the description of your commercial solution;

information regarding your experience in providing a complex, integrated commercial solution; and general estimates regarding the length of time required to implement the solution. You need not address the attached Technical Solution Requirements in your capability statement; however, as noted, these requirements will be mandatory in any solicitation that is issued.

• Provide any other information that may be helpful in developing or finalizing the acquisition requirements.

The information submitted must be must be in and outline format that addresses each of the elements of the project requirement and in the capability statement /information sought paragraphs stated herein. A cover page and an executive summary may be included but is not required.

The response must include the respondents’ technical and administrative points of contact, including names, titles, addresses, telephone numbers, and e-mail addresses.

All responses to this notice must be submitted by email to Stuart Kern, Contract Specialist, at stuart.kern@nih.gov, and reference sources sought notice number 75N95021R00053.

The response must be received on or before Tuesday, August 3, 2021 at 11:00 a.m.

Eastern time.

Disclaimer and Important Notes: This notice does not obligate the Government to award a contract or otherwise pay for the information provided in response. The Government reserves the right to use information provided by respondents for any purpose deemed necessary and legally appropriate. Any organization responding to this notice should ensure that its response is complete and sufficiently detailed to allow the Government to determine the organization’s qualifications to perform the work.

Respondents are advised that the Government is under no obligation to acknowledge receipt of the information received or provide feedback to respondents with respect to any information submitted. After a review of the responses received, a synopsis and solicitation may be published in SAM.gov Contract Opportunities. However, responses to this notice will not be considered adequate responses to a solicitation.

Confidentiality: No proprietary, classified, confidential, or sensitive information should be included in your response. The Government reserves the right to use any non-proprietary technical information in any resultant solicitation(s).

Detailed Technical Specifications NCATS Secure Scientific Platforms Environment

National Center for Advancing Translational Sciences

(NCATS)

National Institutes of Health

9/2/2020

Table of Contents

Technical Solution Requirements

1. Solution Architecture, Data Integration, and Data Management

1.1 Unified, Commercial Software Data Management Architecture

1.3 Flexible Data Integration Engine & Bioscience Data Experience

1.4 Data Storage, Access, and Catalog

1.5 Full Data Provenance and Schema

1.6 Data Transformation Management

1.7 Data deposition back to established repositories

1.8 Cohesive User Interfaces for Data Management

2. Data Analysis, Discovery, and Bioscience Workflows

2.1 Analytic Tool Suite

2.2 Search and Exploration

2.3 Web Application Builder

3. Security and Administration

3.1 Security and Access Controls

3.2 Platform Management and Administration

3.3 End – nothing follows

Technical Solution Requirements

1. Solution Architecture, Data Integration, and Data Management

1.1 Unified, Commercial Software Data Management Architecture

1. A core architecture that consists of a suite of distributed microservices that perform fundamental tasks that the rest of the platform, external plugins and integrations depend on.

2. Where a dataset should operate as a high level conceptual container that manages the full versioned history of a data file (or files) including its history of changes, branches, and access policies.

3. Where each dataset maintains a series of transactions, which are the records of the varying logic which has modified the underlying data files over time, with pointers to the specific versions or views of those data files that are physically stored in a distributed file system.

4. Where a dataset may contain branches that allow transactions to modify data files in a sandboxed fashion.

5. Provide the ability for development and analysis on different versions or subsets of a dataset to occur simultaneously without conflicting.

1.2 Scalability and Hosting

1. A solution that can be offered in the cloud as software as a service, or as an on-premises installation. The preferred deployment method is in the cloud.

2. The solution's cloud architecture should provide:

a. Separated cloud, customer, and management networks

b. Isolated virtual private clouds (VPCs) that prevent commingling of data

c. Hardened and filtered entry and egress points between cloud, customer, and management networks

d. Sophisticated provisioning, patching, and monitoring tooling

e. 24/7/365 support and incident response team

f. Automated backups and robust disaster recovery

g. Events logging and log aggregation from services across hosts

h. Interface providing access to consolidated logs and performance metrics allowing standardized monitoring and tuning of services

3. By design, all components of the solution should be horizontally scalable, allowing additional user scale, data scale, data complexity, and compute complexity.

4. All core services should operate in a high-availability clustered environment using commodity hardware or cloud node hosting on Linux platform

5. All data should be stored in one of several highly-available and redundant distributed data stores

6. All compute is scheduled and run on open-source cluster-computing frameworks and schedulers, with wrapper libraries and appropriate APIs to allow any plugin or integrated application to run on the same infrastructure

7. The government reserves the right to select/determine the cloud provider of our choice

1.3 Flexible Data Integration Engine & Bioscience Data Experience

1. Data ingestion, integration, transformation, and orchestration tools that include:

a. Data Ingestion

1. Support for programmatic (automated and either push/pull) ingestions and ad-hoc ingestions (user-initiated, either via user-interface, or backend script)

2. Support for a broad range of source systems, including, but not limited to HDFS, S3, local file systems, RDBMs, streaming systems, JDBC-enabled systems, etc.

3. Provide an API for writing new ingestion adaptors from novel or custom data systems

4. Ability to install a daemon in a source system environment entirely within the control of the data owner. This daemon should let the system owner manage credentials, connection details, and files.

5. A central coordinator service that manages the configuration and execution of ingestion jobs through communication with the agent. The coordinator compiles all job specifications and provides them to the agent, which then performs the system-side aspects of the ingestion task.

6. High resilience to provide robustness in the face of common failure modes during ingestion – poor network connectivity, disk failures, time outs, and more. Users can choose to configure whether to triage failures in an automated way, or force manual intervention via alerts.

7. A user interface where syncs and jobs can be monitored and configured, and open-ended queries can be performed to test ingestions.

Configurations should include the nature of the source system, the relevant destination dataset in the solution, the path that the data should take in transit, and any in-flight modifications to the data.

b. Data Transformations

1. Pluggable modules to let users write data transformation code in the language of their choice (Java, Python, R, SQL, etc.).

2. Templating tools allow users to catalog and parameterize commonly used code

3. Integrated version control and branching ensure that individuals across an organization can collaborate on logic development without disrupting existing production workflows

4. Deep branching enables code and underlying data to be experimented with in parallel. Data not requiring any changes while experimenting can be configured to default using a “fall-back” branch to avoid unnecessary data duplication.

5. Unique branching of the data itself means that logic branches with disparate transformation approaches can co-exist

6. Incremental or snapshot transactions available to support regular syncing with a diverse set of source data systems

7. Web-based integrated development environment (IDE) allowing the writing of data transformation code

8. Compilation module that is able to transform the code into distributed compute operations on the distributed execution environment

9. Continuous integration backend allowing each change to be committed to a distributed version control system, and versioned artifact to be produced

10. A distributed environment execution manager that provides simplified access to the distributed cluster, and fully manages necessary dependencies, libraries, and conflicts

c. Data Orchestration

1. Ability to add data health checks (including, but not limited to: null checks, distribution as expected, no duplicates, reasonably partitioned for optimal performance, foreign-key relationships)

2. Ability to add triggers, to run certain data operations upon receiving an alert from a message queue

3. Open APIs that enable outside services to trigger actions and updates within platform

4. Advanced queuing and resource channels allow for intelligent sharing of resources among simultaneous transformations and prioritization of critical jobs

5. User interface for context on running jobs and rich metrics on the performance of historical transformation activity to inform scheduling and refactoring

2. Ability to ingest bioscience data type:

a. The solution must be designed to integrate any type of data out-of-the-box.

Additionally, the solution must be preconfigured with specialized tooling to work with the following types of data:

1. High-throughput drug screening (HTS), including raw plate reader output

2. Mass spectrometry (MaxQuant, PD, Peaks, etc.)

3. Expression data (RNA-Seq, DNAse-Seq, miRNA, etc.)

4. DNA sequencing

5. Flow Cytometry

1.4 Data Storage, Access, and Catalog

1. Distributed data storage via preferably S3 or Cassandra that provides unique IDs and a REST API, and integrates with the core security and auditing control subsystem.

2. Key-value storage that provides:

a. Seamless integration with the security and auditing subsystem so that all resources stored are security-aware throughout their life-cycle.

b. REST API with bucket/path-based addressing system, allowing applications to easily register, store, and retrieve resources throughout their life-cycle.

c. Reliable via highly available and scalable architecture

3. Indexing that:

a. Provides a search mechanism as well as a generalized high-scale index layer for the solution and its ecosystem.

b. Is security-aware and is closely tied to the core security and auditing subsystem so that search requests – regardless of origin – can be securely managed.

c. Makes tasks such as live re-indexing and security accessible.

d. By default, will perform search optimizations, such as pre-filtering for paged searches that provide sparse results.

e. Can work with the solution's core build service to implement a distributed indexing worker that allows for systematic indexing of jobs as they build and promote new resources.

4. A data access API that provides a simplified interface for user ‘PUT’ and ‘GET’ requests for data by managing the underlying File System requests among the solution services in a particular installation.

a. The Data Access API is required to interface with the underlying distributed file store that is being used as part of the solution installation to return the actual resource byte array when requested.

5. A Data Catalog that provides:

a. Highly available metadata store that keeps track of datasets, as well as being able to handle datasets composed of multiple data files, and the associated transactions to maintain ACID properties on datasets.

b. Virtual access to a complete dataset that may be composed of many underlying files via many transactions.

c. Reconstruction of the full dataset via negotiation of transactions, whether they be a point-in-time exact view of the data, an append or to an existing dataset, or a deletion from a dataset.

d. A locking mechanism for managing concurrent access to the same dataset, via vector clock and operational transforms (OT).

e. API access to metadata, complete constructed datasets, or files and transactions therein.

1.5 Full Data Provenance and Schema

1. Data in the solution should be inherently schema-less, with additional schema and provenance information provided by dedicated services.

2. A schema management service that:

a. Stores and provides the schema definition for each data file, including variations in the schema that may have occurred across the versioned and branched history of the resource

b. Stores type information that is used in the compute environment. The types can be inferred with an advanced type inference engine; the service should also have a graphical interface for the user to validate, correct, or modify the types manually.

c. Tags and “typeclasses” on tables and columns, allowing semantic meaning to be conveyed and propagated. As an example, a column marked “chemical-compound” will result in all downstream datasets using that column knowing that they also contain a column containing a “chemical-compound,” which can then be associated with specific actions.

3. Ability to generate interactive graphical visualizations of the live updating data dependency and provenance tree for all datasets in the solution (while respecting access controls) and provides for any dataset:

a. The ability for users to inspect the transformations that led to it.

b. All upstream and downstream datasets to be viewed as well.

c. The pipeline graph to be manipulated by users including labeled in various ways.

1.6 Data Transformation Management

1. A data transformation management service that provides:

a. A resilient topological graph of dataset dependencies, using metadata captured in the Data Catalog, Provenance, and Schema information, consisting of virtual build nodes with job specifications (via a worker process) on how to update a given dataset and dependents.

b. A way to dynamically compute staleness of datasets, in order to determine what needs to be run.

c. Advanced topological sorting of datasets to determine the optimal order of job runs, and to allow parallelization of job runs.

d. Modular dispatch in order to dynamically bind and dispatch a worker and an associated executor environment appropriate for the task. This would allow a multi-step build to be run cohesively in a variety of languages in the same run.

e. A way to perform efficient selective insert or replacement of data at the row level, pursuant to access controls.

1.7 Data deposition back to established repositories

1. A data transformation management service that provides:

a. A mechanism to intermittently synchronize select data with external rational-database management systems (RDMS).

b. A mechanism to intermittently synchronize select data with external services through application programming interfaces (API).

c. Support for the persistence of complex object models through RDMS and API.

d. Robust transactional framework to ensure the integrity of the data transfer process.

e. Data management process to track staleness and synchronization of datasets, in order to determine what data needs to be deposited.

1.8 Cohesive User Interfaces for Data Management

1. The solution provides intuitive access to virtually all services through any modern web browser, and includes:

a. A unified web environment sharing an application frame and common design language

b. Ability to register actions and pluggable capabilities via API plugin points. As an example, a developer can add to the context menu for a dataset in order to analyze it using that custom application.

c. Granular permissioning of actions via integrated access control

d. Interface components that facilitate easy use by both technical and non-technical analysts alike.

2. The solution should focus on the ability for technical users (bioinformaticians, cheminformaticians, data engineers) to be able to collaborate on the same data management and analytic infrastructure as non-technical users (biologists, chemists, analysts), and should provide:

a. The ability to comment on any resource (dataset, analysis, dashboard, compute node, etc.) in a collaborative chat environment, including the ability to tag and notify.

b. A fully integrated data issues interface where users can file, discuss, and resolve data quality issues in an interface that can embed flagged datasets, columns, and more.

c. Ability to selectively share any analysis, dashboard, or workflow across the platform and choose what permission level to grant (viewer, editor).

d. Shared source of truth with robust change tracking for inventories, protocols, and other data that would otherwise be shared across teams manually by email or spreadsheet.

e. Allows analyses and artifacts of analyses to be discoverable and reused across other applications on the platform.

f. Ability in the UI to easily create templates and/or macros that allow a non-technical user to apply complex analysis workflows and code paths.

2. Data Analysis, Discovery, and Bioscience Workflows

2.1 Analytic Tool Suite

1. A solution that enables individuals across NCATS to work with the same data by providing a suite of analytical tools appropriate for the full range of technical ability. The suite should provide the following functionality:

a. Reporting

1. Ability to unify live views of charts, visualizations and datasets from various solution services in a single location.

2. Allow users to write collaborative documents, include more context by attaching images or rich text descriptions, and create long-lived, real-time dashboards of key metrics.

3. Ability for reports to link to live data backing those reports, allowing users to inspect the live provenance of the data in the user interface, as well as all data backing the report.

b. Top Down Analysis at Scale

1. An intuitive and visual path-based, point-and-click environment for tabular manipulation and enrichment.

2. Flexible widgets that allow users to apply a broad range of filters, manipulations, expressions, enrichments, joins, and visualizations (including heatmaps, distributions, histograms, bar charts, scatterplots, regressions, PCA, venn diagrams, time series, dendrograms, and more).

3. A path-based model that allows for previous steps to be modified and all subsequent steps to be rebuilt based on those new parameters.

4. The ability for paths to be cloned, branched, and saved as their own objects.

5. The ability for iterative steps to be saved automatically while allowing users to undo actions and return to a previous state of the analysis.

6. Have a documented API, allowing developers to add boards as external plugins without requiring downtime.

c. Code Workbooks

1. A visual directed acyclic graph, where each node is a raw dataset, a dataset resulting from a code manipulation depending on any of a number of upstream datasets, or a visual plot.

2. The code should be able to be written in Python, R, SQL or other common languages, and have direct API bindings to the underlying distributed compute infrastructure (e.g. Spark).

3. Resulting data from code workbook computation nodes should:

1. Automatically become part of the solution's data environment,

2. Be registered in the data catalog, and

3. Immediately available for access in all other applications.

d. Modular Code Templates

1. Ability to save common code segments or paths as templates or macros, thereby very quickly enabling any user to express complex analytical ideas with point-and-click tooling.

2. The solution should support for multi-tenant usage of Code Workbooks and Top Down Analysis, including supporting of full branching log, whereby a user can apply their analytical methods on a branch (with code and relevant datasets all respecting the branching mechanism), and later propose a merge back to the master branch.

3. Resulting data from path-based analyses automatically become part of the solution data environment and immediately available for access in all other applications.

4. Ability to selectively share and collaborate throughout the solution's services, including the ability to add comments, tag and notify users and report and track issues on identified data columns.

2.2 Search and Exploration

1. Beyond simply curating datasets in a central data catalog, the solution must also index analysis and reports such that a user can easily discover data and insights about a topic of interest without knowing beforehand where to navigate.

2. Access control settings at the project and individual resource level need to ensure that a user only sees information they are approved to discover.

3. Robust data tagging and project hierarchies need to allow users to then narrow global search results down to the information most relevant to their task at hand.

4. In addition to global search, the solution should provide a variety of exploratory interfaces to understand the data available, the logic applied to it, and the relationships among different pieces of information:

a. An interface that allows all resources within the solution to be inherently linked to the logic applied to transform the raw information into the subsequent ontology tables.

b. An interface that allows search across all raw and transformed datasets so that it is straightforward to search globally by column names to identify where particular information is present

c. An interface (beyond tabular data) to facilitate mapping of information onto conceptual objects with attributes, behavior, and relationships to other objects within the data model.

d. An interface that allows users to quickly author a standalone web application which pulls cached data into an interface optimized for carrying out a particular workflow. This interface should thereby streamline workflows for when certain information is accessed repeatedly in the same manner.

2.3 Web Application Builder

1. Beyond native data exploration, code authoring, and analysis interfaces, the solution should include the following features:

a. A development environment which allows any user to create data-driven interactive applications.

b. Includes a rich “what you see is what you get” environment, allowing the composing of an interface using extensible “widgets,” including charts, controls, containers, text, images, time series, and others.

c. Each widget can be synchronized to live data with straightforward SQL queries, and complex workflows can be built by defining JavaScript functions that determine interaction between widgets and reaction to user triggered events.

d. Is natively compatible with reading from a SQL connection, via REST API, or anything that provides a JSON response, offering the ability to tie in external systems or APIs.

2. As all other solution components, applications built in this web interface should fully respect the propagated dynamic access controls of the data dependencies.

3. Security and Administration

3.1 Security and Access Controls

1. The solution shall have a security and auditing subsystem to enforce security in the solution. The solution must provide authorization of resource access to users who have been authenticated.

2. Security and access controls that include:

a. Multi-source authentication (enterprise LDAP, PKI, Active Directory, etc.) as well as Single Sign-On (SSO) capabilities.

b. Access controls to secure each piece of information individually in addition to traditional blanket permissions across entire data sources.

c. Ability to assign individual users granular permissions to data access (including ownership, read, write, access denied, etc.).

d. Role-based access controls (RBAC) are assignable to individual users or group of users.

e. Allow for conditional policies based upon defined criteria. For example, applications and individual actions (termed “resources”) can each be given a policy, taking the form of statements that express a (Condition, Verb, Operation) tuple, such as “if User X is a member of Group Y (condition) they are Allowed (verb) to Edit (operation) resource Z.”

f. Access and role-based access controls that can be:

1. assigned by administrators and/or

2. inherited from source systems and updated at any time.

3. Role-based access control based on least privilege and privileged access for access administration.

g. Propagation and inheritance of access controls to all data and analyses derived directly or indirectly from the originating data, whether those data transformation or analyses are performed:

1. In a standard coding language (Java, SQL, Scala, Python, etc.)

2. As technical scripting language (R, SAS, etc.), or

3. A point-and-click analytical environment

h. Data is encrypted at rest via compliant mechanisms (e.g., full disk encryption, LUKS, TrueCrypt).

i. Data in transit from client-to-server or server-to-server is end-to-end encrypted (via SSL/TLS).

3. FedRAMP Compliant ATO. Comply with FedRAMP Security Assessment and Authorization (SA&A) requirements and ensure the information system/service under this contract has a valid FedRAMP compliant (approved) authority to operate (ATO) in accordance with Federal Information Processing Standard (FIPS) Publication 199 defined security categorization. If a FedRAMP compliant ATO has not been granted, the Contractor shall submit a plan to obtain a FedRAMP compliant ATO as determined by the contract officer.

3.2 Platform Management and Administration

1. Continuous Integration and Monitoring, including:

a. Live downtime-less upgrades enabled by high availability services.

b. Centralized monitoring of errors and performance.

c. Dynamic scaling of compute and storage to handle changes in usage in a resource efficient manner.

2. Platform Administration and a solution Configuration Manager that allows:

a. Maintaining a complete understanding of the state of the environment always:

1. Including the hosts, the services, the nodes for distributed and High

Availability services

2. Where those nodes are deployed

3. What versions they are on

4. Version compatibilities, etc.

b. Secure management of cryptographic secrets.

c. The concept of ‘roles’, which declare which services and APIs produce and consume in a declarative manner.

d. All services must have a ‘Life Cycle Model’ to indicate the expected or target state at any given time along with the current state (for example, running or upgrading). Service status is a more detailed reporting of the per-host service state that is updated with each agent action.

e. Backend command line interfaces (CLIs) that provide full management functionality.

f. Front-end graphical interface that provide full management functionalities

3. Standardized microservices:

a. Layout in the filesystem is standardized, allowing a platform administrator to easily find service information and logs.

b. Microservices can be deployed via the platform management services directly onto hosts, or, ideally, as lightweight containers using standard container deployment orchestration libraries.

c. The orchestration framework should tightly integrate with the security and auditing subsystem.

d. These containers allow only a limited number of gateways into the host environment, greatly limiting the potential vectors that could compromise one or more hosts.

3.3 End – nothing follows.

AP.03.B.SOW_Detailed Technical Specifications.pdf
Technical Solution Requirements
1. Solution Architecture, Data Integration, and Data Management
1.1 Unified, Commercial Software Data Management Architecture
1.3 Flexible Data Integration Engine & Bioscience Data Experience
1.4 Data Storage, Access, and Catalog
1.5 Full Data Provenance and Schema
1.6 Data Transformation Management
1.7 Data deposition back to established repositories
1.8 Cohesive User Interfaces for Data Management
2. Data Analysis, Discovery, and Bioscience Workflows
2.1 Analytic Tool Suite
2.2 Search and Exploration
2.3 Web Application Builder
3. Security and Administration
3.1 Security and Access Controls
3.2 Platform Management and Administration
3.3 End – nothing follows.

File details come from the government source that posted it. Updated .