FSA Data Strategy v1.0 - September 25 2020 (Final).pdf

PDF 1 MB Posted

Attached to
Data Modernization Federal contract opportunity
Solicitation number
DATAMODERN2024
Issued by
Department of Education Contracts and Acquisition Management

About this file

This document is a Request for Information (RFI) from the Department of Education Contracts and Acquisition Management division related to the Federal Student Aid (FSA) Data Modernization effort. FSA is evaluating the need for future solicitations to modernize its Title IV Financial Aid Origination and Disbursement (TIVOD) technology ecosystem. The focus areas include improving data management and governance, modernizing the data architecture and engineering, accelerating data integration for insights and oversight, and decreasing operational costs while increasing speed to market. FSA is interested in understanding available technologies, sources, and provider capabilities to deliver these capabilities. This RFI is for planning purposes only and does not constitute a solicitation or obligate FSA to contract for any items or services discussed. Interested parties are invited to provide voluntary submissions, which should be marked for any proprietary or competition-sensitive information.

View the file

Other files for this federal contract opportunity

Other files attached to Data Modernization, newest first.
File Type Posted
FSA Canonical Model 2024_1016.pdf PDF
Amendment_0003_Revised_RFI_Data_Modernization 101024.pdf PDF
Question and Answers 101024.xlsx XLSX spreadsheet
Amendment_0003_Revised_RFI_Data_Modernization 101024.pdf PDF
Data Modernization_RFI.pdf PDF

On GovTribe

Work with this file on GovTribe

  • Download the original file
  • Contacts named in this file
  • Similar government files
  • Ask GovTribe AI about this file

Text version

FSA Data Strategy Version 1.0 ● 9/25/2020

Final Version

FSA Data Strategy Revision History

Version: 1.0 2 9/25/2020

Revision History

VERSION DATE AUTHOR DESCRIPT ION

d1.0 07/23/2020 Shyam Pai James Puggi Homayoon Khalili

Internal draft document offered to FSA’s stakeholders from the Enterprise Data Directorate for review and comments

1.0 09/25/2020 Shyam Pai James Puggi Homayoon Khalili

Initial baseline which includes updates from stakeholders’ feedback

FSA Data Strategy Contents

Version: 1.0 3 9/25/2020

Contents

SECTION 1. DOCUMENT INFORMATION

1.1. Purpose and Scope

1.2. Audience

1.3. Reference Documents

SECTION 2. EXECUTIVE SUMMARY

2.1. Data Strategy Framework

2.2. Data Strategy Guiding Principles

2.3. State of Implementation

SECTION 3. DATA QUALITY

3.1. Establish a Data Quality Team

3.2. Identify Key Data Fields

3.3. Establish Data Quality Standards

3.4. Set Metadata and Business Glossary Baseline

3.5. Implement Reference and Master Data

SECTION 4. DATA SECURITY

4.1. Examples of Data Breach

4.2. Data in the Cloud

4.3. Roles at FSA (for Cloud deployments)

4.4. Shortcomings and Our Strategy To Close The Gap

4.5. Data at Rest

4.6. Data in Motion

4.7. Access

4.7.1. Privileged Access Management (PAM)

4.8. Privacy

4.8.1. Roles and Responsibilities

SECTION 5. DATA REPORTING

5.1. Operational Reporting

5.2. Analytics Reporting

5.3. Data Mining

5.4. Machine Learning

5.5. Reporting Software

SECTION 6. COMPLIANCE

6.1. Department of Education Data Maturity Assessment

6.1.1. Data Management Strategy and Oversight

6.1.2. Data Governance

6.1.3. Data Quality

6.1.4. Data Operations

6.1.5. Data Management

6.1.6. Platform and Architecture

6.1.7. Knowledge and Skills

6.1.8. Customer Support and Engagement

6.1.9. Supporting Processes

6.2. Federal Data Strategy

6.2.1. Federal Principles

6.2.2. Federal Practices

6.2.3. Federal Action Steps: Year 1 Actions by Practice

FSA Data Strategy Contents

Version: 1.0 4 9/25/2020

6.2.4. Cross-Agency Priorities (CAP) Goals

6.3. Evidence Act

6.3.1. Enhance the data infrastructure

6.3.2. Improve access to government data

6.3.3. Develop common understanding

6.4. FUTURE Act

SECTION 7. DATA ARCHITECTURE

7.1. Data Management Framework

7.1.1. Business Systems (OLTP)

7.1.2. Operational Data Store (OLTP)

7.1.3. GSS 1075 (OLTP)

7.1.4. Data Lake (OLAP)

7.1.5. Data Warehouse (OLAP)

7.1.6. Enterprise Data Mart (OLAP)

7.2. Data Reporting Framework

7.2.1. Data Lineage

7.2.2. Data Protection

7.2.3. Data Flows

7.3. Sandbox for Training and Tool Evaluation

7.4. Data Storage

7.4.1. AWS Data Lake Storage

7.4.2. AWS Database Storage

7.5. Data Retention, Archiving, and Backup

7.6. Master, Reference, and Meta- Data Management

7.7. Service-Based Architecture

7.7.1. Service Registry

7.8. Data Modeling

7.8.1. Enterprise Data Model

7.8.2. Data Model Management

7.8.3. Data Model Registry

7.8.4. Data Stewardship

SECTION 8. DATA GOVERNANCE

8.1. Establish Data Governance Committee

8.2. Manage Baseline and Changes

8.3. Publish Standards, Policies, and Processes

8.4. Data Governance Guiding Principles

8.5. Data Governance Charter

8.5.1. Role of the Enterprise Data Directorate

SECTION 9. DATA USE CASES

9.1. Use Case Template

9.2. Bulk Data Ingest Use Case

9.3. Data Request Use Case

9.4. Analytics Use Case

9.5. Research and Modeling

9.6. 360-Degree Views and Reporting Use Case

9.7. Transactional Data, Processing, and Reporting Use Case

9.8. Master and Reference Data Use Case

9.9. External Data Delivery Use Case

SECTION 10. DATA MANAGEMENT OPERATIONS AND MAINTENANCE

OVERSIGHT

10.1. Data Requirements Definition

FSA Data Strategy List of Tables

Version: 1.0 5 9/25/2020

10.2. Data Lifecycle Management

10.3. Guiding principles for data operations are:

10.4. Change Management

10.5. Maintenance

10.6. Change Request Management

10.7. Configuration Management

10.8. Data Storage and Operations

10.9. Technology Stack and Tools

10.10. Capacity Management

SECTION 11. DATA MANAGEMENT INVESTMENT OVERSIGHT

11.1. Objectives

11.2. Business Case

11.3. Program Funding

11.4. Financial Oversight

11.5. Data Management Function

APPENDIX A: ACRONYMS AND ABBREVIATIONS

APPENDIX B: GLOSSARY

APPENDIX C: TOPICS FOR THE DATA GOVERNANCE COMMITTEE

List of Tables Table 1. Reference Documents Table 2. Data Quality Measures Table 3. Model Characteristics Table 4. Key Data Strategy Roles Table 5. Acronyms and Abbreviations Table 6. Glossary

List of Figures Figure 1: Data Strategy Framework Figure 2: FSA Data Architecture Figure 3: FSA Reporting Architecture Figure 4: FSA Data Governance Committee Structure

FSA Data Strategy Document Information

Version: 1.0 6 9/25/2020

Section 1. Document Information

1.1. Purpose and Scope

This Federal Student Aid (FSA) Data Strategy focuses on actionable approaches and guidelines to achieve the objectives identified in the Data Strategy Framework, described in the Executive Summary section below. The sections are structured with a brief description of the component being discussed, followed by how the component aligns to one of the four major objectives, and description of an approach, guidelines, and metrics (if applicable) that can be used by project teams and the Data Governance Committee for managing that component.

There is detail here about key entities and fields for mastering data, mapping out reference data, data architecture, governance, and data management processes.

Upon publishing this Strategy, the Enterprise Data Directorate (EDD) intends that project teams and their contractors will use it for data quality, security, and architecture guidance, and internal teams to understand it as they develop contract requirements and participate in the governance process. EDD intends to update this Strategy annually as additional discoveries are made, or new requirements develop.

1.2. Audience

The FSA Data Strategy is intended for use by FSA’s stakeholders and their contractors.

1.3. Reference Documents

The following documents provide either governance or guidance for this document.

Table 1. Reference Documents

DO C U ME N T T IT LE DO C U ME N T LO C A T IO N

Architecture Documents:

Enterprise Data Architecture https://fsa.share.ed.gov/to/architecture/ea/SitePages/Data%20Architecture.aspx

Strategic Documents:

Federal Data Strategy FSA Strategic Plan https://strategy.data.gov/assets/docs/2020-federal-data-strategy-action-plan.pdf https://studentaid.gov/strategicplan

Compliance Documents:

DATA Act Evidence Act FUTURE Act https://www.congress.gov/113/plaws/publ101/PLAW-113publ101.pdf https://www.congress.gov/115/plaws/publ435/PLAW-115publ435.pdf https://www.congress.gov/116/plaws/publ91/PLAW-116publ91.pdf https://fsa.share.ed.gov/to/architecture/ea/SitePages/Data%20Architecture.aspx https://strategy.data.gov/assets/docs/2020-federal-data-strategy-action-plan.pdf https://strategy.data.gov/assets/docs/2020-federal-data-strategy-action-plan.pdf https://studentaid.gov/strategicplan https://www.congress.gov/113/plaws/publ101/PLAW-113publ101.pdf https://www.congress.gov/115/plaws/publ435/PLAW-115publ435.pdf https://www.congress.gov/116/plaws/publ91/PLAW-116publ91.pdf

FSA Data Strategy Executive Summary

Version: 1.0 7 9/25/2020

Section 2. Executive Summary FSA’s business systems are a rich source of customer, school, and financial partner data that can be harnessed for enabling data-driven decisions and policies, managing fraud and default risks, and improving customer service and partner engagement.

Both the Next Gen FSA program and the EDD share these broad objectives. To accomplish these, however, FSA must establish controls for data quality, data security, and privacy, and develop reporting architectures to make data accessible. FSA must also implement these control processes through a comprehensive strategy and approach that complies with Federal and Department of Education guidelines.

2.1. Data Strategy Framework

This Data Strategy (Strategy) is built around four major objectives (in blue below): Data Quality, Data Security, Data Reporting, and Compliance. These are based on FSA’s Strategic Plan and the Department’s Data Maturity Assessment requirement. The objectives are enabled by a strong Data Architecture and managed by Data Governance.

Shown below is the framework for the Strategy.

Figure 1: Data Strategy Framework

These objectives can only be accomplished with strong executive support and sponsorship, funding for initiatives that might be proposed by the EDD, and active participation by data owners and stewards in the management of this Strategy through the proposed Data Governance Committee.

Below is a brief description of the six components of the Framework:

1. Data Quality: This objective will be accomplished by identifying and focusing on key data entities and fields from across FSA’s business systems, but this does not mean all available fields. Key fields will be identified by the Data Quality Team, which will be part of the Data Governance Committee.

The team will then define and establish data quality standards and measurements (Section 3: Data Quality provides more detail on potential metrics), assess quality, and assist with establishing a Metadata and business glossary baseline. That baseline will then be used to implement reference and master data that will be made available to FSA business systems for use.

2. Data Security: This objective will be accomplished by a proposed Data Security team under the Data Governance Committee and will assess the current state of Data Security across FSA’s

Version: 1.0 8 9/25/2020 business systems for fields (such as Personally Identifiable Information (PII) and Sensitive Personally Identifiable Information (SPII)), databases and other storage (data encryption), and data in transit (flows). The team will develop a dashboard with metrics and measurements and propose initiatives to tighten data security where needed. The assessment will also be captured in the Metadata tool for use as a baseline by the Data Governance Committee.

3. Data Reporting: This objective will be accomplished by a separate team under the Data Governance Committee, with a focus on data organization — Operational, Analytic, Mining, and Machine Learning

— for delivery and online access. This team will work with various internal user groups to determine the type of reporting needs to organize the data around and will then work with the Data Architecture team to develop data models.

4. Compliance: This objective will be met by ensuring that the other components in the Framework adhere to the Federal Data Strategy and the Department of Education’s Data Maturity Model. The Compliance team under the Data Governance Committee will focus on assessing compliance, developing a dashboard, and proposing next steps or initiatives as needed.

5. Data Architecture: This objective will be developed by a team that will address batch and real-time data flows, data storage platforms and their purpose, archiving platform and requirements, data models to organize data for various needs, and the technology stack that will standardize and streamline software tools.

6. Data Governance: This objective will manage the baseline set by each of the teams above, as well as control and manage change to the baseline. The Data Governance Committee will be comprised of a Data Steering Committee, which will approve the baseline, major changes to the baseline, course-correct as necessary, and approve data initiatives, and the individual focus area teams described above. These focus area teams will meet as often as needed to develop assessments, statuses, changes, and new initiatives for the Data Steering Committee. The Data Governance Committee will queue up issues for the FSA Board and FSA Council.

2.2. Data Strategy Guiding Principles

The FSA Data Strategy is aligned to the following guiding principles:

• Maintain control of key data entities and fields throughout their lifecycles, including creation, flow, transformation, storage, user and system access, publishing, and long-term storage (archive)

• Maintain data quality using uniformity, accuracy, completeness, consistency, uniqueness, timeliness, and auditability as measures

• Maintain data security and privacy beyond cybersecurity requirements

• Reduce the number of copies of similar data for key data entities and fields, so that data is more easily controlled and secured, more consistent, and easier to access and use in reporting

• Engage a cross-section of data owners, stewards, and security and technical architects to help set the baseline for data quality, security, reporting, compliance, and standards, and then manage changes to the baseline

• Changes to data must be auditable, and must adhere to audit and compliance requirements, based on FSA’s business policies and Federal laws and regulations

• Define who is responsible, accountable, consulted, and informed (RACI) for data managed through the governance process

• Move toward real-time data exchanges between business systems

• Move toward independence from external servicer systems

Version: 1.0 9 9/25/2020

2.3. State of Implementation

FSA is in the process of implementing elements of the Data Strategy objectives within Enterprise Data Management & Analytics Platform Services (EDMAPS). Some objectives have been achieved fully, some partially, and some have not yet been addressed.

• Completed: Enterprise Data Warehouse & Analytics (EDWA) is a fully functioning Analytics Reporting capability and is actively used by internal teams for a wide range of analytics uses, from fulfilling over 600 Data Requests from outside groups annually, when combined with data from business systems, to Program Compliance. EDD is recommending enhancements to EDWA in the Data Reporting section below.

• Partially Completed: The concept behind Master Data Management (MDM) is to make available mastered / reference data to business systems (e.g., Digital and Customer Care (DCC) and Partner Participation and Oversight (PPO)), so that they can reference high-quality mastered data from a single source. This reduces the number of duplicate copies, and moves us toward reliable, complete, consistent, and current data. Person MDM is complete; Loan, Grant, School, and Financial Partner MDMs are yet to be designed and implemented. The Data Strategy calls for co-development with PPO and Interim Servicing Solution (ISS) systems because they have the need and use cases and are involved in re-engineering their systems. MDM development and use will fall into place naturally as part of these projects.

• Partially Completed: Metadata and Business Glossary Baseline has been completed for EDWA, Person MDM, and portions of the Data Lake that are in Production. This is a baseline of what has been implemented within EDMAPS. In addition, this Data Strategy calls for capturing metadata and business glossary information from business systems, so that FSA can have a complete baseline for key data entities and fields for the enterprise. The idea is to set a comprehensive baseline for use with governance — for establishing and managing to quality controls, security controls, reporting access controls, and compliance.

• Not Yet Addressed: Creating a comprehensive data quality assessment across the FSA business systems for critically important business data (e.g., a person’s basic data, loan balances, due dates, etc.), and subsequent proposals and activities to cleanse, master, and set quality standards.

• Not Yet Addressed: Re-starting the Data Governance Committee and establishing a Data Architecture Board to drive Strategy components; establishing standards, policies, and controls;

managing changes through change control processes; and setting and managing to data metrics.

• Not Yet Addressed: Establishing an enterprise Operational Data Store (ODS) to centralize key operational metrics within a database system to enable 360-degree views of customers, loans, grants, schools, and financial partners. The ODS can be a powerful and comprehensive database but requires extraordinary care and attention to ensure the data is consistent and reliable. This Data Strategy proposes to regularly receive data from business systems, synthesize and standardize the data, store it, and make it easily available for reporting. The ODS, in turn, can be used as a key feeder system to EDWA.

• Not Yet Addressed: Creating functional-area data marts that are smaller versions of EDWA, but expressly focused on delivering function-specific analytics reporting capability to internal teams, such as Program Compliance or Customer Experience.

FSA Data Strategy Data Quality

Version: 1.0 10 9/25/2020

Section 3. Data Quality Achieving control over data quality is the cornerstone to becoming a data-driven organization. Without consistently reliable data, internal teams will continue to spend time cleansing, synthesizing, and verifying data. Additionally, business systems will continue to adjust incoming and outgoing data to accommodate for data format and quality problems. Bad data should never be used for decision or policymaking; it is bad for customer service and bad for partner engagement.

Data quality is defined by the following specific factors:

Table 2. Data Quality Measures

QU A LIT Y ME A S U R E DE S C R I P T I ON

1. Uniformity Data field formats must be uniform across business and reporting systems

2. Accuracy Data should match reality

3. Completeness Data fields should contain complete information, with no missing elements

4. Consistency De-conflicted data within fields and databases across systems

5. Uniqueness Duplicative data should be reduced wherever possible

6. Timeliness Data must be as fresh and current as possible where it is needed

7. Auditability Changes to data, fields, or database tables must be auditable

Repeatability is another quality measure that means producing the same result every time a report is run on the same data set. This measure is not directly related to the data quality measures above but can be achieved by setting date and time stamps on fields on the database side and reporting queries to fetch the same data set.

It is important to note that only key data entities and fields will be subject to these quality controls, not all available fields in every database of every system. These key entities and fields will be identified by the Data Quality Team (referred to below).

3.1. Establish a Data Quality Team

The Data Quality Team is a cross-organizational team of data owners, stewards, and architects that will define specifics of which data entities and fields are important to put through the quality control process, and subsequent Metadata, Business Glossary, and Master Data Management (MDM) implementations.

The data owners and stewards have a functional understanding, while the architects have an understanding of the best practices for implementation. As such, this team will also participate in the development of the Metadata and Business Glossary, creating reference data as part of Master Data Management, and with the overall governance process through the Data Governance Committee.

3.2. Identify Key Data Fields

Among the thousands of fields that are available across business systems, most of them are only needed for the system’s own internal processing and are not useful or important to the enterprise. The Data Quality Team will identify key, enterprise-level data entities and fields. These fields are typically shared between systems and used in data analysis and reporting. The Quality Team will establish a Metadata and Business Glossary baseline to track these fields through all of FSA’s business systems, based on a data lineage

Version: 1.0 11 9/25/2020 feature of the governance tool. The initial list of key entities are Person, Loan, Grant, School, and Financial Partner. Specific fields within these entities will be identified and fleshed out by the Data Quality Team.

The following are a proposed set of initial key entities for MDM. These will need to be finalized and approved by the Data Governance Committee, based on FSA’s enterprise requirements and direction. MDM entities consist of relatively static, non-transactional data that can be mastered and provided to business systems as reference data. MDM is the single, centralized source of basic reference data for each entity.

o Person Entity: Categories of data fields within this entity include a person’s identifiers, including SPII, addresses, phone numbers, emails, and employers. This entity may also include income, tax information, Expected Family Contribution (EFC), and loan or grant calculations.

o Loan Entity: Categories of data fields include loan type, identifiers, origination and disbursement information, owed balances, due dates, and statuses.

o Grant Entity: Categories of data fields include grant types, identifiers, and disbursement information.

o School Entity: Categories of data fields include school identifiers, including PII, addresses, phone numbers, emails, contact person’s phone numbers and emails.

o Financial Partner Entity: Categories of data fields might be partner identifiers, including PII, addresses, phone numbers, emails, contact person’s phone numbers and emails. (Financial Partners include banks, guarantee agencies, and private collection agencies.)

These entities will ultimately be “related” in database terms so that data users can know which person received what loans or grants, which college they attended and when, how much of the loan or grant they used, and the status and balance of their repayment. In addition, data users can also know which guarantee agency or private collection agency is handling delinquencies and defaults.

3.3. Establish Data Quality Standards

The Data Quality Team will identify or confirm the data entities as described above, and drill down to specific fields that should be controlled for quality. The team will also use the quality areas identified above and define specific measurable metrics for each data field. Lastly, the team will develop a quality status dashboard that shows quality percentages met or improved, that can be used by the team and, more broadly, the Data Governance Committee to monitor progress.

3.4. Set Metadata and Business Glossary Baseline

The next step is to capture and set a baseline for metadata and the business glossary. Metadata captures database level details for fields, such as formats, data types, its data source, and validation and quality rules. Business Glossary captures the business names of fields and their functional meaning, usage, calculations, rules, and policy. Each field will be marked as SPII, PII, or derivative that, when combined with other non-PII fields, can become PII. The Data Quality Team will drive the population of the Metadata dictionary and Business Glossary, and related standards and policies.

FSA currently has IBM’s Information Governance Catalog (IGC) tool that contains metadata and glossary information for EDWA, Person Master Data Management (PMDM), and the Data Lake that are in production. The Data Quality team will also set the baseline for data lineage, which tracks fields from their source all the way to their final system. This will give the Data Governance Committee a full view of how the data moves and where it may get changed and distributed. The team will review and derive any revisions based on this Strategy document and subsequent quality efforts.

Version: 1.0 12 9/25/2020

3.5. Implement Reference and Master Data

Following the effort of identifying key data fields, establishing quality standards, and capturing metadata and business glossary, Reference and Master data systems must be implemented. Reference data maps shared fields (e.g., names, formats, etc.). Master data uses hierarchy and business rules to determine, store, and present mastered data of the key entities such as Person. Both reference and master data are for use by business systems to support their transactions, and for reporting systems. When business and reporting systems use the same reference and master data, it will improve data quality by reducing the number of copies of the same data in multiple systems, by always using the same standard data set, and by reducing data exchanges for these data entities between systems. Reference and master data are part of this Data Strategy’s Master Data Management concept and should be implemented within EDMAPS.

PMDM is already deployed and supports DCC. Other MDM’s, including Organization MDM (OMDM) will begin development in the near future.

FSA Data Strategy Data Security

Version: 1.0 13 9/25/2020

Section 4. Data Security Data security refers to digital privacy measures to maintain confidentiality in which data is accessible by authorized users while still ensuring that data is accessible to fulfill business needs. Data security also includes auditing of data to align with organizational security standards and comply with regulations, contractual agreements, and business requirements. Continuous monitoring of data ensures identification of vulnerabilities, lapses (misconfigurations), or threats. Data can be protected in several ways such as data encryption, data back up, and data masking.

A framework must be created that will secure data across multiple environments and architecture tiers, and regulate the usage of the shared services. This framework provides FSA with the confidence that their data is being stored in a safe manner and the data is only accessed through secure channels.

Data security principles are:

1. Data shall be kept secure.

2. Data security is to be universally applied throughout FSA.

3. The Privacy Advocate shall be consulted on the principles for data architecture and privacy protection of data in operational systems and data warehouses.

4. Data security will be enforced at both the application level and at the data level.

5. Security at the data level will adequately provide open access to information within a secure environment.

6. All access to authoritative data is achieved via common data access services hosted by an Enterprise tenant.

7. Protecting the privacy of Customer and Partner data is the most important criteria of data security.

8. All Authoritative FSA data will be accessed through a data security layer that will be configurable at the user, location, workstation, and time level.

9. Security mechanisms must not impede the day-to-day operations of FSA and Department of Education to an unacceptable degree.

The three critical components of the data security strategy are to secure data in the following states:

1. Data at Rest: Data at rest in information technology means inactive data that is stored physically in any digital form (e.g., databases, data warehouses, spreadsheets, archives, tapes, off-site backups, mobile devices, etc.).

2. Data in Motion: Data is most vulnerable when it is in motion. The risk to data is often underestimated because organizations believe all the sensitive data is contained within a few secure systems.

3. Data in Use: Data in use is more vulnerable than data at rest because, by definition, it must be accessible to those who need it. The key here is to secure data in use by controlling access and incorporating authentication/authorization mechanisms.

Organizations also need to be able to track and report suspicious activity, diagnose potential threats, and improve security by addressing the following questions:

• What type of sensitive data does the organization store, use, and transmit?

• Who has access to this data?

• Where, when, and why are they using it?

• How is the data stored when it is not in use?

Version: 1.0 14 9/25/2020

• How is access to the database controlled?

• What mechanisms are used to transport the data?

4.1. Examples of Data Breach

A data breach is an incident where information is stolen or taken from a system without the knowledge or authorization of the data owner. Stolen data may involve sensitive, proprietary, or confidential information such as credit card numbers, customer data, trade secrets, or matters of national security.

Among the organizations that have had harmful data breach and loss events are Target, Home Depot, Anthem, the Federal Office of Personnel Management, and the National Security Agency.

• Guidance from Government entities: National Institute of Standards and Technology (NIST), Office of Management and Budget (OMB), Department of Education, and FSA.

• Data Security addressed in NIST 800-53 to protect the confidentiality, integrity, and availability of information.

• Secure Data (storage and transit):

• Protection of Data at Rest, 800-53 section SC-28 (includes guidance for off-line storage as well).

Information that requires protection include configurations for devices and systems (firewalls, routers, ITS/IPS, etc.).

• Protect Data in Transit (communications outside the boundary of the system is exposed to the potential of interception and modification).

• Guidance provided in 800-53, SC-8 and other referenced sections therein, provide guidance on protecting the confidentiality and integrity of the transmitted information.

• Cryptographic Key establishment and Management, 800-53 section SC-12 (specific guidance in 800-57 series).

Other controls in this area address things such as: separation of production and non-production environments, removal and destruction of media, capacity to ensure availability, protection against data leaks, etc.

FSA needs a comprehensive approach to Data Security that extends from the infrastructure layer to the compute, network, and storage layers. It is not EDD’s intent to cover all those areas in this Data Strategy, but it’s worth mentioning that the silo structure of our organization with different groups responsible for discrete systems does not lend itself to a comprehensive strategy across the enterprise. Many of the standards and guidelines have been developed and employed by Technology Office teams with sole emphasis on the systems they control or the NGDC (Next Generation Data Center). There is a lack of documentation for COD (Common Origination and Disbursement) and EDWA (Enterprise Data Warehouse and Analytics), and other Next Gen systems and applications, especially given their deployment into the Amazon Cloud.

4.2. Data in the Cloud

There has been a recent growth in Cloud computing deployments (both in private and public sectors), mostly because Cloud offers greater flexibility and availability in obtaining computing resources at lower cost. Security and privacy concerns for organizations considering new or legacy transitions to the Cloud, should be at the forefront of these deployments.

Version: 1.0 15 9/25/2020

NIST publication SP 800-145 defines Cloud computing as a model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services) that can be rapidly provisioned and released with minimal management effort or Cloud provider interaction.

The Security and Privacy Downside Public Cloud computing also brings with it potential areas of concern, when compared with computing environments found in traditional data centers. Some of the more fundamental concerns include the following:

• System Complexity: A public Cloud computing environment is extremely complex compared with that of a traditional data center. Many components make up a public Cloud, resulting in a large attack surface which translates to the need for better oversight and more careful consideration.

Challenges also exist in understanding and securing application programming interfaces (APIs) that are often proprietary to a Cloud provider.

• Shared Multi-tenant Environment: Client organizations typically share components and resources with other consumers that are unknown to them. Rather than using physical separation of resources as a control, Cloud computing places greater dependence on logical separation at multiple layers of the application stack. While not unique to Cloud computing, logical separation is a non-trivial problem that is exacerbated by the scale of Cloud computing.

• Network Threats: Threats to network and computing infrastructures continue to increase each year and become more sophisticated. Having to share an infrastructure with unknown outside parties can be a major drawback for some applications and require a high level of assurance pertaining to the strength of the security mechanisms used for logical separation.

• Internet-facing Services: Public Cloud services are delivered over the Internet, exposing the administrative interfaces used to self-service and manage an account, as well as non-administrative interfaces used to access deployed services. Applications and data that were previously accessed from the confines of an organization’s intranet, but moved to a public Cloud, must now face increased risk from network threats that were previously defended against at the perimeter of the organization’s intranet and from new threats that target the exposed interfaces.

• Remote Administrative Access: Relying on remote administrative access as the means for the organization to manage assets that are held within the Cloud also increases risk, compared with a traditional data center, where administrative access to platforms can be restricted to direct or internal connections.

• Loss of Control: While security and privacy concerns in Cloud computing services are like those of traditional non-Cloud services, they are amplified by external control over organizational assets and the potential for mismanagement of those assets. Transitioning to a public Cloud requires a transfer of responsibility and control to the Cloud provider over information as well as system components that were previously under the organization’s direct control. Loss of control over both the physical and logical aspects of the system and data diminishes the organization’s ability to maintain situational awareness, weigh alternatives, set priorities, and effect changes in security and privacy that are in the best interest of the organization.

4.3. Roles at FSA (for Cloud deployments)

At FSA, Cloud deployments have been the focus of most new developments. There are currently several major systems and applications operating in the Amazon Web Services (AWS) Cloud and others are being considered for migration to AWS. Most of FSA’s data, which includes a substantial amount of PII, is part of this shift in system deployment strategy from on-prem data center systems to the Amazon Cloud.

The roles and responsibilities listed here follow the traditional data center deployment models and are integral for the migration and continual operation of these systems and applications:

Version: 1.0 16 9/25/2020

• Department of Education Chief Information Officer (CIO) and Chief Information Security Officer (CISO), FSA CISO, Contracts Team, System Owner, Information System Security Officer, Security Control Assessor, Information Technology Risk Management Group, and Vendor.

The specific Responsibilities for these groups can be found in the Cloud Security Guide Document Information (Version: 1.8 6 4/26/2019).

It is crucial for the security of FSA’s data that all these separate individuals and groups carry out their responsibilities with the right level of expertise. Given the continuous new technologies and changes to the Cloud infrastructure and services, an overall strategy to keep the teams current and working together is central to the data security strategy.

Some of the general risk areas for FSA include the following:

• Lack of expertise at critical oversight positions within FSA

• Overwhelming number of systems both in operation and development reporting to the same individual

• Ever expanding scope of new systems migrating to the Cloud

• Leveraging the same Vendor throughout the systems under development

• Lack of separation between Cloud operator Vendor and System Developer/Integrator Vendor

• Lack of an API tier to uniformly address the data tier. That would go a long way towards ensuring a holistic security approach to the data tier.

4.4. Shortcomings and Our Strategy To Close The Gap

1. Subject Matter Expert (SME): Review of the above roles and responsibilities for analysis as part of securing FSA’s data and managing the risk. The roles have to be lined up with the current Cloud capabilities and Risks, with emphasize on the right level of subject matter expertise and adherence to documentation review by Federal government employees responsible for the critical positions such as Information System Owner (ISO), Information System Security Officer (ISSO), and Information Technology Risk Management (ITRM) team. The roles identified at FSA do not engage in a comprehensive view of the security because everyone is focused in their own area and having multiple layers of management does not allow for a critical look at where and how something can fall through the gap(s). For example, how systems are accessed by consumers of FSA data should correlate with what kind of data (PII vs non-PII) and the amounts of data they are accessing. A single solution can’t address all the possible combinations of data use (for example, analytics and reporting vs publicly available data).

2. Governance: Developing the control and oversight by the organization over policies, procedures, and standards for application development and information technology service acquisition, as well as the design, implementation, testing, use, and monitoring of deployed or engaged services. With the wide availability of Cloud computing services, lack of organizational controls over employees or the vendors engaging such services can be a source of problems. The lack of visibility into how the Cloud services and capabilities are employed by the vendor is of great concern. This is further exasperated by the lack of expert level knowledge at those critical government positions which approve and “go along” with the vendor recommendations to meet deadlines. While Cloud computing simplifies platform acquisition, it does not alleviate the need for governance; instead, it has the opposite effect, amplifying that need.

3. Data Ownership: The organization’s ownership rights over the data must be firmly established in the service contract to enable a basis for trust and privacy of data. If everything, including data ownership, is left to the ISO then FSA is vulnerable to the possibility of the ISO being more focused on the technical controls and not necessarily familiar with how the data is used.

Version: 1.0 17 9/25/2020

4. Visibility: Continuous monitoring of information security requires maintaining ongoing awareness of security controls, vulnerabilities, and threats to support risk management decisions. Collecting and analyzing available data about the state of the system should be done regularly and as often as needed by the organization to manage security and privacy risks, as appropriate for each level of the organization involved in decision making. The ongoing security findings at FSA, related to the systems deployed in the Cloud, raises the risk of a possible breach or data loss due to lack of forward thinking on the part of the vendor and/or the critical roles associated with the systems(s).

5. Architecture: Applications are built on the programming interfaces of Internet-accessible services, which typically involve multiple Cloud components communicating with each other over APIs. It is important to understand the technologies the Cloud provider uses to provision services and the implications the technical controls involved have on security and privacy of the system throughout its lifecycle. With such information, the underlying system architecture of Cloud can be decomposed and mapped to a framework of security and privacy controls that can be used to assess and manage risk.

6. Identity and Access Management: The organizational identification and authentication framework may not naturally extend into a public Cloud and extending or changing the existing framework to support Cloud services may prove difficult. The alternative of employing two different authentication systems, one for the internal organizational systems and another for external Cloud-based systems, is a complication that can become unworkable over time. Identity federation, popularized with the introduction of service-oriented architectures, is one solution. FSA’s Identity and Access Management systems (AIMS and PAS) are highly customized and were implemented long before Cloud deployments became the norm. Some of the additional requirements due to Cloud deployments are:

• Clear separation of the managed identities of the Cloud consumer from those of the Cloud provider must also be ensured to protect the consumer’s resources from provider-authenticated entities and vice versa.

• Easy and cost-effective integration with other Cloud Services (and service providers) with minimal customization.

• Automated provisioning of Cloud Identities without manual intervention (to minimize human errors).

7. Authentication and Access Control: Authentication is the process of establishing confidence in user identities. Authentication assurance levels should be appropriate for the sensitivity of the application and information assets accessed and the risk involved. Security Assertion Markup Language (SAML) alone is not sufficient to provide Cloud-based identity and access management services. The capability to adapt Cloud consumer privileges and maintain control over access to resources is also needed.

8. Security configuration management baseline (hardening standards):

• FSA provides guidance for hardening all information systems in accordance with Federal Information Security Modernization Act (FISMA) regulations, Department of Education requirements, and Defense Information Systems Agency’s Security Technical Implementation Guides (DISA STIGs).

• The NGDC has started work on complying with the above STIG guidelines, and for Next Gen systems under development (or in production) the compliance and assessment is scheduled for June 2021.

• What are the Next Gen Program Office efforts to ensure that FSA can meet those guidelines?

• Who are the responsible personnel in charge of these systems, and do they have the right level of expertise?

• Is the acquisition for these systems including the right language to ensure compliance?

• Given the current scope of work for the leads on these systems, is there enough bandwidth to ensure a timely compliance with these requirements?

Establishing a data security team may be the right answer to solving this and similar challenges.

Version: 1.0 18 9/25/2020

4.5. Data at Rest

Encrypting hard drives is one of the best ways to ensure the security of data at rest. Sensitive data stored either on GFE or non-GFE (contractor-owned) equipment must be safeguarded in accordance with NIST guidance and OCIO Policy, including but not limited to the following standards:

• Folders/files containing sensitive PII or other sensitive data stored in a shared drive must be encrypted and the folders configured to restrict access on a need-to-know basis.

• Data backups must be encrypted and securely transported/filed/archived.

• Cryptographic mechanisms must be employed to protect the integrity of audit information related to high-value assets (e.g., log and audit tools).

• OMB M-07-16 specifies that Federal agencies must encrypt, using only NIST certified cryptographic modules, all data on mobile computers/devices carrying agency data unless the data is determined not to be sensitive, in writing, by the agency’s Deputy Secretary or their designee.

The primary security controls for restricting access to sensitive information stored on end user devices are encryption and authentication. Encryption can be applied granularly, such as to an individual file containing sensitive information, or broadly, such as encrypting all stored data. The appropriate encryption solution for a particular situation depends primarily upon the type of storage, the amount of information that needs to be protected, the environments where the storage will be located, and the threats that need to be mitigated.

There is a need for a detailed analysis of the storage encryption (across databases) and key storage and management (across the enterprise), with clear attention to segregation of duties and considering data availability and continuity of business activities in case of a disaster. FSA needs an analysis of its Encryption Management and must consider uniformly applying a policy across all its systems and applications, with considerations around availability and security of the data.

4.6. Data in Motion

Certificate, Public Key Infrastructure (PKI), and Transport Layer Security (TLS)

FSA Certificate Guidance defines the minimum acceptable standards for implementing secure communications across networks and describes the current FSA procedures regarding secure certificates.

In accordance with the Federal Information Processing Standards (FIPS) guidance, NIST guidelines, and FISMA compliance, FSA system development focuses on the use of the network and back-end computer processing systems to:

• Improve security for citizens’ access to government services and information.

• Facilitate the flow of government information between branches and agencies.

• Reduce operating costs using electronic business processes.

FSA Scanning and Analysis Teams (well established processes at FSA, applicable to both on-premises and Cloud-based systems):

• Determines correct levels of certificates and communication standards during routine and specific audits/scans.

• Executes vulnerability and configuration compliance scans that FSA requests as part of ongoing risk management activities.

• Analyzes raw scan results to determine the accuracy of identified deficiencies.

The following applies to FSA systems, applications, appliances, or websites:

Version: 1.0 19 9/25/2020

1. Employ the current U.S. Department of Education, BOD 18-01, Transport Layer Security (TLS) Memorandum, December 20, 2019 (Appendix D).

2. Employ the U.S. Department of Homeland Security (DHS), Binding Operational Directive (BOD) 18- 01, Enhance Email and Web Security, October 2017 (Appendix C).

3. Ensure every certificate includes a fully qualified domain name.

4. Remove expired certificates.

5. All FSA systems must comply with the requirements established in the OMB Memorandum M-15-13, Policy to Require Secure Connections across Federal Websites and Web Services, June 8, 2015 (Appendix B), which provides guidance to Federal government agencies making the transition to

HTTPS.

6. Deploy HTTPS upon launch on new websites and services on FSA domains or subdomains.

7. Ensure existing websites and services are accessible through a secure connection (HTTPS-only, with HTTP Strict Transport Security (HSTS)).

8. Use HTTPS on FSA intranets.

FSA’s strategy around the above processes, even though well established, need to involve better planning on the part of the vendors to minimize the risk on the data. The continued scan findings on AWS hosted systems currently managed by Accenture, indicates a shortcoming which needs to be addressed. FSA needs to develop a plan on how to better address these findings going forward and therefore minimize the risk.

4.7. Access

Enterprises rely upon strong access control mechanisms to ensure that corporate resources (e.g., applications, networks, systems, and data) are not exposed to anyone other than authorized users.

Data sensitivity and privacy of information have become increasingly an area of concern for organizations.

The identity proofing and authentication aspects of identity management entail the use, maintenance, and protection of PII collected from users. Preventing unauthorized access to information resources in the Cloud is also a major consideration. Some of the main principles of Access Control are shown below (with the NIST 800-53 area indicated):

• Access Enforcement (AC-3): Organizations can control access to PII through access control policies and access enforcement mechanisms (e.g., access control lists). This can be done in many ways. One example is implementing role-based access control (RBAC) and configuring it so that each user can access only the pieces of data necessary for that user‘s role. Another example is Attribute Based Access Control (ABAC). ABAC enables the appropriate permissions and limitations for each user’s access request based on individual attributes and allows for the management of those permissions by multiple systems from a single platform, reducing administrative burden. FSA needs to perform an analysis of the possibility of leveraging any ABAC to further enhance its security posture.

• Separation of Duties (AC-5): Organizations can enforce separation of duties for duties involving access to PII. For example, the users of de-identified PII data would not also be in roles that permit them to access the information needed to re-identify the records.

• Least Privilege (AC-6): Organizations can enforce the most restrictive set of rights/privileges or accesses needed by users (or processes acting on behalf of users) for the performance of specified tasks. Concerning PII, the organization can ensure that users who must access records containing PII only have access to the minimum amount of PII, along with only those privileges (e.g., read, write, execute) that are necessary to perform their job duties.

Version: 1.0 20 9/25/2020

FSA needs to focus on the following areas for Access Control to its data:

• FSA does not have the visibility into the enforcement of the above principles (Access Enforcement, Separation of Duties, and Least Privileges), especially for the systems operated by Accenture in the AWS Cloud. Given the role that Accenture plays in both operating the Cloud and developing the systems, the need for a clear understanding of these principles is urgent.

• There is a definite need for an analysis on the access control principles across all FSA systems, starting with those systems under development, to ensure these concepts are integrated before going to production.

• Additionally, given the number of systems being deployed and developed in the AWS Cloud by Accenture as separate…

This is the start of the file's text. The full file is on GovTribe.

File details come from the government source that posted it. Updated .