ATR Data Analytics Solution RFI QA.docx

DOCX document 60 KB Posted

Attached to
ATR IT MODERNIZATION: DATA ANALYTICS SOLUTION Federal contract opportunity
Solicitation number
TMF5-1bRFI
Issued by
Department of Justice Offices Boards and Divisions

About this file

This is a Questions and Answers (Q&A) document for a Request for Information (RFI) issued by the Department of Justice Antitrust Division (ATR) regarding an IT modernization initiative for a Data Analytics Solution (DAS). The RFI seeks vendor capabilities for implementing a scalable, AI-enabled data analytics platform to support advanced data processing, analytics, and mission-driven decision making for ATR's litigation and investigative operations.

Technical and Operational Requirements: ATR requires a cloud-based analytics solution capable of processing 5 KB to 400 TB of data volumes, with batch ingestion of at least 500 GB per hour and streaming throughput of 50,000 records per second. The platform must support structured, unstructured, and semi-structured data ingestion from multiple sources including case management systems, litigation databases, and court records. Peak concurrent users are estimated at 10, though the user base may expand to over 30 users including system administrators, data scientists, economists, attorneys, paralegals, and data processing specialists. The solution must achieve 99.95% availability with a Recovery Time Objective (RTO) of 1 hour or less and Recovery Point Objective (RPO) of 5 minutes, with multi-availability zone deployment. Critical requirements include support for Python, R, and Shell scripting; natural language querying capabilities for non-technical users; Git-based source control integration; and advanced features such as automated data classification, hallucination detection, PII redaction, and toxicity filtering for AI-generated outputs.

Security, Compliance, and Deployment: The solution must be deployed in Azure Government Cloud with FedRAMP High Authorization and operate exclusively within government cloud boundaries with no data leaving the environment. All AI models must originate from US-based companies, with data-in-use compliance required during inference operations. The platform must support CUI (Controlled Unclassified Information) classification and comply with NIST SP 800-53 Rev 5, NIST AI RMF, executive orders on AI (EO 14147, 14179, 14283, 14275, 14319), and OMB directives (M-25-21, M-25-22, M-26-04). ATR is open to multi-product solutions or single-vendor suites provided all Appendix A requirements are met; teaming agreements complying with FAR 52.219-14 are acceptable but will not be treated as pass-through contracts. The anticipated procurement vehicle is GSA Schedule (Multiple Award Schedule) or NASA SEWP, with an expected contract structure of a base year plus four option years. The RFI response deadline was May 1, 2026, with the RFP anticipated for release within fiscal year 2026. ATR has not determined contract type (Firm Fixed Price, Time and Materials, or Cost Plus Fixed Fee) and requests vendor proposals include both platform licensing and professional services pricing, with ROMs provided for both pilot deployment and full enterprise rollout based on vendor industry experience.

View the file

Other files for this federal contract opportunity

Other files attached to ATR IT MODERNIZATION: DATA ANALYTICS SOLUTION, newest first.
File Type Posted
DAS Evaluation Criteria Appendix A.xlsx XLSX spreadsheet
Data Analytics Solution (DAS) RFI.docx DOCX document

On GovTribe

Work with this file on GovTribe

  • Download the original file
  • Contacts named in this file
  • Similar government files
  • Ask GovTribe AI about this file

Text version

4/24/2026 Subject: ATR Data Analytics Solution RFI QA document

· …had one question regarding this statement: Assess the ability to integrate with the Agency’s existing cloud infrastructure (Azure-based environment).Would the government also entertain vendors who are deployed in AWS Government hosting platforms? Yes

· Is ATR open to a multi-product platform response, or are they seeking a single-vendor analytics suite (e.g., pure Databricks/Snowflake)? Yes Open (they must be in Gov. Cloud with FedRAMP High Authorization)

· What existing data sources does ATR need to ingest from — case management systems, litigation databases, court records, financial filings, third-party data providers? Any additional detail is appreciated. Initial data ingestion comes from many sources, including data produced by parties during investigations and second requests. This will include structured, unstructured and semi-structured data.

· What is the current Azure environment (Azure Government or Azure Commercial)? Will the solution require FedRAMP High or Moderate authorization? Currently Government Cloud requiring FedRAMP High Authorization

· Is there an existing ATO (Authority to Operate) process in place at ATR, and what is the expected timeline? This process is initiated as part of an Authority to Test and will be required to meet all Authority to Operate controls prior to being production ready

· Will structured and unstructured data be stored in the government environment, or can it data be processed in a FedRAMP-authorized cloud? The data must be processed within government cloud boundaries. No data can leave government environment.

· What is the expected data volume range (TBs of data, number of concurrent users)? 5kbs to as much as 400TBs can be produced at any one time. Volume depends on data use type. Economic data comes in much larger sizes.

· How many users will be using the tool? Will the users primarily be Data Scientists, Economists and/or Lawyers? Any breakdown related to user demographics would be appreciated. Please provide peak concurrent users expected to access the Data Analytics Solution as well. Current needs: >30, System Admins, Data Scientists and Economists mainly, 10 users peak. It is possible that other types of users could be involved, but that is unknown at this time. They are likely to be Attorneys, Paralegals or Data Processing specialists that help with the loading and validation of production data.

· Is the Government open to layered or component solutions in which a specialized vendor addresses a defined capability bundle (for example, investigative analytics over structured and unstructured litigation data, with native entity extraction, relationship graphing, and source-traceable AI), as a complement to a separate enterprise platform addressing data warehousing, MLOps, and broad analytics infrastructure? Or does ATR intend a single-prime platform award covering the full Appendix A scope? (sidenote from vendor: We ask because our platform delivers a set of capabilities the named reference vendors do not — specifically around investigative entity analysis, relationship discovery across disparate sources, and AI-generated answers with structural source citations preserved by the data architecture rather than reconstructed by the model. We are not, however, an end-to-end data warehousing or model lifecycle management platform, and we would not want to consume ATR’s evaluation time submitting against requirements that lie outside our intended scope.) Open to discussion, but all requirements for Appendix A must be met.

· Can you confirm if this is related to a recompete or if it's new work? New

· NFR 1.18 requires FedRAMP compliance. Does the Agency require FedRAMP Moderate or FedRAMP High authorization? Please confirm the minimum required impact level. Yes FedRAMP High Authorization

· What is the highest data classification level the DAS must handle? (e.g., Unclassified, Controlled Unclassified Information (CUI), For Official Use Only (FOUO), Personally Identifiable Information (PII)) CUI

· BR 1.02 states that the solution must be flexible to work with other cloud platforms. Does the Agency anticipate the DAS operating solely within its existing Azure infrastructure, or is a multi-cloud or hybrid-cloud deployment (where some components operate in a separate FedRAMP-authorized cloud environment) within scope for the resulting procurement? We are willing to explore other possible architecture options, including multi-cloud or hybrid-cloud environments because our data currently exists within multiple cloud environments, and there may be a need to connect to those data sources. All environments must support FedRAMP High Authorization

· NFR 1.17 specifies RTO <= 1 hour, RPO <= 5 minutes, and 99.95% uptime. Is the disaster recovery site expected to be within the same cloud provider and region, or must it be geographically separated (e.g., East/West US)? DR should be in same cloud provider, but we will accept a recovery site within the same region, but different availability zones.

· NFR 1.17 specifies RTO <= 1 hour, RPO <= 5 minutes, and 99.95% uptime. Is an active-active or active-passive failover architecture acceptable to satisfy this requirement? Both are acceptable, but are application dependent

· Approximately how many total named users will require access to the DAS platform? >30 initially. We do not have an estimate for future needs.

· Can the Agency provide an approximate breakdown of DAS users by role type (e.g., system administrators, data scientists/engineers, economists/analysts, read-only/business users)? Yes. Please provide tiered licensing structure if that is the intent here. Users may include: >30, System Admins, Data Scientists, Economists, Attorneys, Paralegals and Data Processing specialists.

· What is the approximate total data volume currently under management at ATR that would be ingested or migrated into the DAS? > 450 TB already in Azure Cloud. This information does not include the size of any production data which can range from 500TB to 1PB.

· What is ATR's estimated annual data growth rate? At least 10% based off the last 3 years and the current rate of ingestion for this year of economic data. However, when we include the data produced during investigations and litigation, that amount grows at approximately 20-30% per year.

· FR 1.01 requires support for unstructured data ingestion. Approximately what percentage of ATR's total data assets are unstructured (e.g., legal documents, deposition transcripts, contracts, correspondence, multimedia files)? > 90%

· What is the approximate total volume of unstructured data currently managed by ATR?

Approximately 450 TB of economic data and between 500TB to 1PB of production data.

· FR 1.02 requires support for MPEG and MP3 formats, and FR 1.15 requires automated ingestion and classification. Approximately what volume of audio or video files (e.g., recorded depositions, wiretaps, hearing recordings) does ATR currently manage? The volume of video and audio files varies significantly by case and is not fixed. If your solution has any limitations processing or storing this type of data, please include that in the comments section of the evaluation criteria.

· What are the typical file sizes and formats of ATR's audio and video evidence files? Varies extensively, can be in any audio or video format type, but most common is MP4.

· BR 1.18 specifies a streaming throughput requirement of at least 50,000 records per second. What are the anticipated streaming data sources that drive this requirement (e.g., financial data feeds, regulatory filing APIs, internal system event logs, real-time monitoring data)? Typically, data is ingested via physical media or from a staged NAS, developing an automated ingestion pipeline is desired as part of another project related to this request

· Approximately how many automated data pipelines, ETL workflows, or scheduled jobs currently exist in ATR's environment that would need to be migrated, replicated, or rebuilt in the new DAS? Currently we have about 5 pipelines and about 200 jobs per week, but they’re not automated and will have to be created as part of the new solution. No need to migrate anything

· Will ETL pipeline development, configuration, and ongoing maintenance be performed by contractor staff, government staff, or a combination of both? Combination

· Will AI/ML model training, development, and deployment be performed by contractor data scientists or government economists/analysts? Possibly all of the above

· Is there an existing data analytics platform, data warehouse, or incumbent contractor currently supporting ATR's data analytics needs? Yes

· What data currently resides in the existing system, and will the incumbent contractor be expected to cooperate with transition activities? Case data, there is no incumbent contract currently supporting this work as part of the product implementation. All resources are both federal employees and onsite contract support from other Division contracts. We do not expect support beyond tier 2 or 3 maintenance support to be required as part of any RFP.

· Does the Agency currently hold active enterprise licenses for PowerBI, Tableau, and/or IBM Cognos? Yes, we own PowerBI licenses. We do not own Tableau or IBM Cognos.

· Will the existing BI visualization tools (PowerBI, Tableau, IBM Cognos) continue to operate alongside the new DAS, or is the DAS expected to replace any of these tools? Continue to use existing visualization tools

· FR 1.10 requires Git-based source control integration. What specific Git platform does ATR currently use (e.g., GitHub Enterprise, Azure DevOps, GitLab)? We do not have one standard.

· Does ATR have a preferred or existing CI/CD automation platform? No preferred /Yes we have existing CI/CD platforms (BitBucket, Nexus, Ansible Tower, Jenkins, Terraform)

· Does ATR currently have AI or ML models in production that must be migrated to or supported by the new DAS? We are exploring OpenAI RAG models but have not created any AI models in production.

· What frameworks were used to build ATR's existing AI/ML models (e.g., scikit-learn, TensorFlow, PyTorch, R statistical models)? We do not have any active models in production.

· Is a parallel operation period required during transition from the existing environment to the new DAS? Yes If a parallel operation period is required, what is the minimum acceptable duration before the legacy system can be decommissioned? Approximately 90 days

· What contract type is anticipated for the DAS acquisition (e.g., Firm Fixed Price, Time and Materials, Cost Plus Fixed Fee)? ATR has not made of decision on contract type.

· What period of performance is anticipated for the DAS acquisition (e.g., base year plus option years)? Base, plus 4 options

· What is the anticipated timeline from RFI response (May 1, 2026) to RFP release? Does ATR anticipate releasing the RFP within FY2026? Yes, we expect to release the RFP within this FY.

· What enterprise productivity, cost management and analytics tools are currently being utilized by ATR? Cloud based Dashboard and Tools

· Will it be necessary to integrate current solutions into the new solution, or will the existing data need to be migrated to the new solution? TBD, depending on solution selection

· Can the government elaborate on the currently deployed visualization tools (such as PowerBI, Tableau, or others), any preferred or mandated tools, the specific Azure services in use (e.g., Azure Data Lake Gen2, Azure SQL, Azure Synapse, Azure Databricks), and the immediate versus future support requirements for other cloud platforms (like AWS or GCP) to understand integration touchpoints? We have data visualization through PowerBI. As for tools and Cloud platforms the purpose of this RFI is to find out the industry tools available. Propose the tool you think is best based on the RFI information and your industry experience.

· Does ATR require existing ETL workflows to be migrated, or will this be a completely new ETL implementation? TBD, depending on solution. We have tested some solutions.

· Is the government's objective to provide analytic services or to establish an analytic support system for ATR users? Both

· Can the government provide details about ATR's current training infrastructure? Specifically, do you have a Learning Management System (LMS) in place? Yes

· What kind of access and functionalities are available to non-technical users? Does their access include things like self-service query builders, pre-built dashboards, natural language interfaces, and other similar features? Pre-built Dashboards

· Does ATR require support for specific data classification levels, such as Controlled Unclassified Information (CUI) or For Official Use Only? Yes, CUI

· Does ATR have specific data retention policies for different data types? Additionally, what is the established process for data disposal? Yes. Data disposal is dependent upon record schedules, litigation holds, and case disposition.

· What are ATR's specific requirements for cross-matter information barriers in litigation contexts? Criminal data is subject to Grand Jury secrecy requirements. Conflicts of interest may also be relevant, so RBAC roles are important. Data security should allow for flexibility based on roles and groups.

· What specific ROI metrics does ATR leadership currently track? Are there established baselines for time-to-decision or TCO? We currently have no tracking mechanisms but would be interested in metrics that would provide valuable insights to senior leadership, including CIO, about the speed to review and system responsiveness.

· What are the current wait times for compute resource availability that need to be improved? Ability to analyze >1TB of data per day

· Are there specific DOJ, ATR or federal AI transparency requirements we should be aware of beyond general explainability? Systems must meet all requirements as listed in EOs and OMB Directives on AI such as EO 14147, EO 14179, EO 14283, EO 14275, EO 14319, M-25-21, M-25,22, and M-26-04 as well as any new EO and OMB Directives on the subject.

· How many active litigation matters or investigations does ATR typically manage concurrently? The RFI is intended for high-level capability assessment, we are unable to provide additional details beyond what is included in the RFI materials. If your solution has any limitations, please include it into the comments section of the evaluation criteria.

· Does the "US-based companies" requirement for AI models apply to open-source models, or only proprietary/commercial models? Yes to all. It must not have any foreign source within its supply chain pursuant to EO 13873.

· Please specify the main external data sources that ATR should connect to? The RFI is intended for high-level capability assessment, we are unable to provide additional details beyond what is included in the RFI materials.

· Are there particular third-party data providers or commonly utilized public datasets that are regularly integrated for antitrust analysis? The RFI is intended for high-level capability assessment, we are unable to provide additional details beyond what is included in the RFI materials.

· What source control and CI/CD tools are ATR currently using or prefer (e.g., GitHub, GitLab, Azure DevOps)? The RFI is intended for high-level capability assessment, we are unable to provide additional details beyond what is included in the RFI materials.

· Are there any existing infrastructure as code (IaaC) implementations within ATR that need to be integrated or migrated? Yes

· What are the typical file sizes and formats that are most frequently encountered by ATR? Everything, 5 KBs up to 400TBs

· What kinds of data classification or categorization schemes are required? Are there particular legal or regulatory categories that specifically apply to antitrust data? Various classifications and categories are required, including CUI. The RFI is intended for high-level capability assessment, we are unable to provide additional details beyond what is included in the RFI materials.

· What degree of data lineage tracking is necessary? Is it sufficient to track at the field level, or is more detailed transformation tracking required? More detailed

· What kinds of natural language queries do different user types need to execute? NLQ

· Are there predefined Service Level Agreements (SLAs) in place? If yes, what do they entail? SLAs agreements will be spelled out in the RFP. Please provide the standards that your solution is able to meet.

· NFR 1.18 requires the solution to comply with FedRAMP security standards. For vendors who deploy their solution on a FedRAMP-authorized cloud platform (such as Microsoft Azure Government) but do not themselves hold a FedRAMP Authority to Operate (ATO), will the Government accept compliance demonstrated through operation within the authorized cloud provider's FedRAMP boundary — provided the vendor can supply documented boundary scope, system security plan (SSP) references, and a written declaration from the cloud provider confirming the vendor's services operate within the authorized environment? Or does the Government require the vendor itself to hold an independent FedRAMP ATO? They have to have their own ATO or be on an active path to having FEDRAMP authorization. We are unable to sponsor a product. For those operating within a FEDRAMP cloud boundary they have to satisfy inherited controls at the infrastructure level, provide SSP appendices from the CSP, and lastly produce a written attestation stating they operate within our guidelines

· The RFI requests that vendors indicate eligibility under GSA Schedule or SEWP contract vehicles. Will the Government accept responses from vendors who participate as a named subcontractor or teaming partner under a Prime contractor that holds the qualifying GSA Multiple Award Schedule (MAS) / IT Category (formerly IT Schedule 70) or SEWP V vehicle — where the subcontractor performs a defined, substantial portion of the technical scope? If so, is a formal executed Teaming Agreement sufficient documentation, or does the Government require the Prime to be the responding entity with the subcontractor named in the response? The service should be compliant with standards such as FAR 52.219-14. This should not be treated as a pass-through contract.

· The RFI requires at least three case studies demonstrating 'direct experience supporting government agencies in implementing data analytics solutions.' Will the Government consider case studies from regulated private sector engagements — such as financial fraud analytics, legal discovery platforms, or healthcare data governance — where the vendor can demonstrate a direct and explicit analogy to ATR's stated requirements? If private sector case studies are considered, what level of mission-mapping documentation is expected to establish relevance to ATR's investigative and litigation analytics needs? We will consider any relevant and recent experience, but strong government experience is highly valued.

· NFR 1.17 specifies an RPO of ≤ 5 minutes. Achieving a 5-minute RPO for analytical workloads — particularly those involving large Delta Lake or data lakehouse architectures — requires synchronous or near-synchronous geo-replication, not standard backup-based recovery. Does ATR's 5-minute RPO requirement apply to: (a) all data within the DAS platform, including raw ingested datasets and processed analytical outputs; All data (b) active compute state and in-progress pipeline jobs; Both or (c) only committed, query able data assets in the primary data store? Yes Additionally, will the Government accept an RPO demonstrated through synchronous zone-redundant storage combined with asynchronous geo-replication to a paired region as compliant with NFR 1.17? Yes

· FR 1.15 requires automated classification, sorting, and labelling of ingested data, and FR 1.16 requires throughput of ≥1TB in under 90 minutes. For the purposes of evaluating FR 1.15 and FR 1.16: (a) Does the 1TB/90-minute throughput requirement apply to all file types including binary media files, or to structured/semi-structured data specifically? Yes (b) Does automated classification include content-based classification (e.g., AI-driven document categorization) or schema/metadata-based classification only? Yes (c) Are there specific file format categories — such as audio (MP3) or video (MPEG) as referenced in FR 1.02 and FR 1.03 — that must be ingested and classified within this throughput threshold? Yes

· BR 1.28 requires that the solution 'only use AI models from US-based companies.' Does this requirement extend to: (a) data residency during inference — i.e., does government data processed by the AI model need to remain within US-based infrastructure during the inference call, not merely at rest? Yes, government data processed during inference must remain within US-based infrastructure throughout the inference call. data-at-rest compliance alone is insufficient. The requirement extends to data-in-transit and data-in-use during model inference operations (b) model fine-tuning or retraining on ATR data — must any training or fine-tuning activity involving government data occur exclusively within FedRAMP-authorized US-based infrastructure? Yes, any training, fine tuning or retraining activity involving ATR government data must occur exclusively within FEDRAMP-authorized, US-based infrastructure. Vendors must be able to demonstrate and document that no government data transits outside this boundary during such activities (c) third-party AI components embedded within primary US-based platforms — for example, if a US-based platform incorporates a non-US open-source model, does that constitute non-compliance? The primary platform must be US-based and FEDRAMP authorized. Embedded third-party or open-source model components of non-US origin would require additional scrutiny and vendor disclosure. DOJ reserves the right to evaluate such components on a case-by-case basis and may determine non US embedded models as non-compliant due to the security and supply chain risk.

· Appendix A contains 114 requirements across nine categories spanning Business, Functional, and Non-Functional types. Will the Government apply equal weighting to all requirements and categories in scoring vendor responses, or will certain categories — such as Security & Governance (Category IV), AI Capabilities (Category IX), or Performance & Scalability (Category VI) — carry greater weighting? Equal weighting to all categories Additionally, will 'Fully Supported / Out-of-the-Box' responses be scored materially higher than 'Fully Supported / Configurable' responses, or are both treated as equivalent demonstrations of capability? Equivalent demonstrations of capability. This information is being gathered for informational purposes as part of market exploration.

· BR 1.04 requires native support for Python, R, and Shell scripting. (Tools are provided as an example) ATR's economists perform advanced quantitative analysis including regression modelling, market simulation, and econometric analysis — workflows commonly implemented in R (using packages such as fixest, lfe, and sandwich) and in Python statistical libraries. Will the Government prioritize or score more favorably solutions that provide native, managed execution environments for these specialized economic modelling workflows — including RStudio Server, Jupyter-based R kernels, and statistical library management — over solutions that support R and Python only through generic scripting interfaces? Possibly Additionally, does ATR anticipate workflows requiring integration with Stata or SAS datasets? Yes

· BR 1.10 requires the solution to 'improve accessibility of data and analytics tools for non-technical users,' and FR 1.46 references 'natural language queries.' ATR's attorneys represent the primary non-technical user constituency. Does the Government expect the solution to include a Natural Language Querying (NLQ) capability that allows non-technical users — such as attorneys — to ask questions of data in plain English without writing SQL or code? Yes If so, what is the expected scope: (a) structured data querying only, (b) unstructured document search and retrieval, or (c) both? Both Will demonstrations of attorney-facing NLQ interfaces carry meaningful weight in the evaluation of BR 1.10 and FR 1.46? Yes

· The RFI requests a pricing model overview including pricing approach, key drivers, licensing structure, and a Rough Order of Magnitude (ROM) for professional services. For government budgeting and procurement planning purposes, does ATR have a preference between: (a) subscription-based pricing with fixed annual fees per user tier, (b) consumption-based pricing tied to compute usage and data volume, or (c) a hybrid model with a fixed platform fee and variable consumption charges? Additionally, for the professional services ROM, should vendors scope implementation as a single deployment engagement, or should the ROM reflect a phased multi-year implementation aligned with ATR's modernization roadmap? The government is looking for a solution that offers the best value to the government.

· Please clarify the anticipated acquisition vehicle for this requirement. Specifically, will ATR procure through the General Services Administration (GSA) Multiple Award Schedule (MAS), the NASA Solutions for Enterprise-Wide Procurement (SEWP), another existing Indefinite Delivery, Indefinite Quantity (IDIQ) vehicle, or open market competition under Federal Acquisition Regulation (FAR) Part 15? GSA is the preferred source, but NASA SEWP is an alternative vehicle available. No decision has been made yet as to which vehicle will be used for final purchase.

· Will ATR pursue a software-centric acquisition, an implementation and integration services acquisition, or a combined platform-and-services procurement? This information is being gathered for informational purposes as part of market exploration. We are asking the vendors to provide the best solution based on their industry experience.

· Does ATR anticipate a single-award contract, a multiple-award arrangement, or a task order against an existing IDIQ vehicle? Understanding the anticipated award structure will help respondents frame their teaming posture and pricing narrative appropriately. We will be using either GSA MAS or NASA SEWP for the purchase order. It is likely that will be awarded to a single vendor solution, but teaming agreements that meet standards such as FAR 52.219-14 can be considered. This should not be treated as a pass-through contract.

· Will the eventual solicitation require respondents to hold a specific small business designation, such as SBA 8(a) or SBA Small Business set-aside, or will the competition be unrestricted? This has not been decided. This will determine whether small business prime or subcontracting posture is strategically relevant.

· Please confirm whether ATR requires deployment exclusively within Azure Government (GovCloud) regions or whether a commercial Azure environment with appropriate FISMA and FedRAMP controls is also acceptable. This distinction directly affects platform feature availability and compliance architecture. No

· Does ATR have an existing Azure tenant, an Azure Data Lake Storage account, or other provisioned Azure infrastructure that the solution must integrate with, or will the DAS be deployed into a net-new environment? Understanding this will shape integration complexity estimates and pricing.

· Please identify which enterprise identity provider is currently in use at ATR. Specifically, whether Entra ID (formerly Azure Active Directory), on-premises Active Directory, or a third-party identity provider such as Okta is the authoritative identity source for SSO and MFA integration. We have multiple authentication mechanisms that can be used for SSO and MFA, including Okta.

· Are there existing data pipeline orchestration tools, ETL platforms, or data catalog services currently in use at ATR that the DAS must interoperate with or replace? Respondents need this information to accurately assess integration scope and avoid duplicative pricing. Yes

· Does ATR currently operate an enterprise Git-based source control platform such as Azure DevOps or GitHub Enterprise, and if so, is the DAS expected to integrate with that existing system or provide its own version control and CI/CD capabilities? Yes, we expect Git-based systems to integrate with existing systems.

· Please provide the approximate user population expected to access the DAS, broken down by persona where possible, for example: attorneys, economists, data scientists, analysts, and leadership or executive users. User count and role distribution directly affect licensing, workspace design, and training scope. >30 Users may include: System Admins, Data Scientists, Economists, Attorneys, Paralegals and Data Processing specialists.

· Please clarify which data types and sources represent the highest-priority ingestion candidates for the initial deployment. For example, structured litigation databases, unstructured document repositories, streaming market data feeds, or multimedia evidence files. Prioritization will allow respondents to construct more accurate sizing estimates. Prioritization varies depending on litigation schedule

· Does ATR anticipate generative AI or large language model use cases that involve agency-owned data, publicly available data, or both? Both Please also confirm whether geographic restrictions on model hosting or model provenance apply beyond the requirement in Appendix A BR 1.28 specifying US-based AI model companies. Yes, all components used for any part of the system must be US based.

· Are there specific litigation or investigation workflow tools currently in use, such as Relativity, Nuix, or DOJ case management systems, that the DAS must integrate with or that respondents should assume as adjacent systems in their architecture? Yes. RelativityOne and others may apply.

· Appendix A BR 1.18 specifies batch ingestion of at least 500 GB per hour and streaming throughput of at least 50,000 records per second. Please confirm whether these are current sustained operational requirements or projected future-state targets, as this distinction affects cluster sizing, cost modeling, and infrastructure commitment. All are Current capability

· FR 1.16 requires data ingestion throughput of 1 TB or more in under 90 minutes. Please confirm whether this represents a peak sustained ingestion requirement, an average daily load, or an episodic burst scenario, such as large evidence file ingestion at the start of a major investigation. Peak or episodic scenario dependent on data volumes and other factors.

· NFR 1.17 specifies a Recovery Time Objective (RTO) of 1 hour or less and a Recovery Point Objective (RPO) of 5 minutes or less, with 99.95% availability. Please confirm whether these thresholds apply to the full DAS platform or specifically to defined mission-critical workloads, and whether ATR requires multi-region or cross-zone disaster recovery architecture. Full DAS, with Multi-AZ deployment

· Does ATR have an anticipated total data storage footprint for the DAS at initial deployment and at projected full operating capability? Understanding baseline and growth expectations will allow respondents to provide more accurate storage cost estimates and avoid presenting a ROM that is materially under or over the government's planning range. >450 TB in current economist environment with a predicted growth rate of 10% annually. For production annual data volume is between 200-300TB of new data with a growth rate of approximately 20-30% per year.

· NFR 1.18 requires FedRAMP compliance. Please confirm whether ATR requires a FedRAMP High authorization, a FedRAMP Moderate authorization, or whether a FISMA High system operating under an existing Agency Authorization to Operate (ATO) framework is also acceptable. This distinction eliminates a significant category of platforms that hold Moderate but not High authorizations. Yes, Requires FedRamp High Compliance

· BR 1.22 requires that human review of Government Data by the solution provider be restricted, logged, justified, and visible to the Government. Please clarify whether this requirement applies to all vendor personnel including support staff and site reliability engineers, or specifically to AI system operators and data engineers with direct access to litigation data. Vendors would be required to get cleared to access the government data, facilities and systems

· BR 1.28 states the solution shall only use AI models from US-based companies. Please confirm whether this restriction applies to foundation models accessed through API, fine-tuned or custom models deployed within the ATR environment, or both. Additionally, please clarify whether open-source models hosted on US-based infrastructure qualifies under this requirement. The requirement is not limited to one deployment model, any AI model used as part of the solution, regardless of how it is accessed or hosted, must originate from a US-based company. Regarding open source models hosted on US-based infrastructure: hosting location alone does not satisfy this requirement. The model must be developed and maintained by a US-based company or organization. An open source model of foreign origin hosted on US infrastructure would not qualify under BR.1.28

· Please identify any agency-standard visualization, reporting, records retention, legal hold, or Section 508 accessibility tools that respondents must assume in their solution architecture. Several Appendix A requirements reference visualization tools such as Power BI, Tableau, and Cognos. Does ATR have a mandated enterprise standard, or are respondents free to propose the most appropriate tooling? ATR currently owns PowerBI.

· Should vendors provide pricing ROM for pilot deployment only, full enterprise rollout, or both? Providing this context will allow respondents to present a more useful cost comparison rather than pricing an unanchored scope. Both

· Please confirm whether ATR intends to procure platform licenses, professional services, or both as part of this acquisition. Understanding whether professional services will be scoped separately from the platform subscription will allow respondents to present cleaner pricing structures. You should be prepared to answer for both.

· Appendix A includes FR 1.02 and FR 1.03 with identical requirement text. Please confirm whether this is a duplication error or whether two distinct requirements were intended. Respondents cannot provide differentiated responses to requirements that are textually identical. Duplication mistake

· Appendix A contains three requirements numbered BR 1.22 with distinct requirement text. Please confirm the correct numbering for these three requirements so respondents can address each individually and without ambiguity in the response matrix. These are three distinct requirements

· Two requirements are numbered BR 1.28 with distinct text: one addressing operational AI use for real-time and batch decision-making, and one requiring US-based AI model sourcing. Please confirm the correct requirement number for each so respondents can address them distinctly in the response matrix. These are two distinct requirements

· For requirements that respondents assess as Partially Supported, does ATR have a preference for how respondents describe the gap, specifically whether the gap closure should be described as a configuration activity, a professional services engagement, or a roadmap commitment with a projected availability date? Provide gap support evidence in the comments section of the evaluation criteria

· FR 1.43 references hallucination detection and PII and toxicity checks as AI output validation controls. Please confirm whether these requirements apply to all AI-generated outputs within the DAS or specifically to outputs surfaced to end users such as attorneys and analysts. This affects the scope of AI governance tooling respondents must include in their solution. AI-generated outputs

· What is the estimated budget range or Rough Order of Magnitude (ROM) for this initiative? ATR will not provide a response. This proposal should be based on your industry experience and present the best value to the government.

· What is the anticipated Period of Performance (base and option years) for this effort? Base of 12 months from date of award, plus 4 options.

· Is there an existing system or contractor currently supporting data analytics within ATR? If yes, please provide details. The RFI is intended for high-level capability assessment, we are unable to provide additional details beyond what is included in the RFI materials.

· Does the Government have an estimate of required resources, labor categories, or level of effort (LOE)? We are relying on your expertise as industry professionals to provide your LOE and labor categories that will provide the best value to the government for the solution requirements.

· Can the Government provide more details on the existing Azure-based infrastructure and current data architecture? The RFI is intended for high-level capability assessment, we are unable to provide additional details beyond what is included in the RFI materials.

· Are there any preferred or mandated platforms/tools (e.g., Databricks, Snowflake, Synapse Analytics)? No

· Can the Government provide representative use cases and expected data types/volumes (structured/unstructured)? The RFI is intended for high-level capability assessment, we are unable to provide additional details beyond what is included in the RFI materials.

· What are the expectations for post-implementation support, including ownership (Government vs Contractor) and operational responsibilities? It is expected that the professional services team would work with existing federal and contract resources during the engagement to allow them to work alongside the team to understand the support and maintenance required.

· Can ATR provide representative examples of litigation datasets (e.g., emails, chat logs, transaction data, PDFs, economic models) and typical volumes encountered during major merger or criminal investigations? The RFI is intended for high-level capability assessment, we are unable to provide additional details beyond what is included in the RFI materials.

· Are there DOJ specific evidentiary requirements for data lineage, immutability, or audit trails that must be preserved for litigation for Chain‑ of‑ Custody‑ Requirements ? Yes

· What econometric tools, libraries, or modeling frameworks (e.g., Stata, R packages, SAS) must be supported for antitrust economic analysis? All of the above

· Are there defined turnaround expectations for data processing during fastmoving merger reviews or criminal investigations‑? Not specifically, but speed is of high importance.

· What percentage of ATR’s data is structured vs. unstructured, and does the ingestion pipeline need to prioritize one over the other? ~90% Unstructured, No prioritization

· Does ATR anticipate real-time‑ ingestion (e.g., logs, API feeds), or is streaming primarily for internal telemetry? Real-time ingestion is a desired functionality

· What is the expected maximum number of concurrent ingestion jobs during peak litigation periods? 5-15

· What level of explainability is required for AI-generated ‑insights? We do not have this information.

· Could ATR please provide a list of approved or prohibited AI model providers? AI models must have all components and references in the United States and should be operated in Gov-cloud environments.

· Which validation controls are mandatory among Hallucination detection, PII redaction, Toxicity filtering, Policy compliance checks, Source grounding ‑validation? We are looking for vendors to provide their recommendations based on industry experience.

· What is the required format, retention period, and review process for logs documenting human access to Government Data? These requirements should meet all FISMA standards

· Does ATR require ZeroTrust‑ enforcement at the: Workspace level, Dataset level, Notebook level, Model level, API level? Yes, at all levels

· What are the performance, scalability requirement? See previous answers. Are certain workloads (e.g., ingestion, AI inference) required to exceed 99.95% uptime? Yes

· Could ATR please provide expected user counts by personnels such as economists, attorneys, data scientists, paralegals? >30 Users may include:System Admins, Data Scientists, Economists, Attorneys, Paralegals and Data Processing specialists.

· How many environments such as dev,test,prod do we need access to setup? Minimum of 2

· What platform/tools are the DOJ Antitrust Division currently using for data analytics? ATR is currently using various tools for data analytics.

· Is Azure predominantly used as the sole cloud platform for DOJ's other mission support requirements? ATR retains data in on prem and SaaS environments across multiple clouds.

· Would the government provide additional information on their current data analytics infrastructure? The RFI is intended for high-level capability assessment, we are unable to provide additional details beyond what is included in the RFI materials.

· Would the government provide additional information on their current data analytics challenges? The RFI is intended for high-level capability assessment, we are unable to provide additional details beyond what is included in the RFI materials.

· Will there be a requirement for other domain classifications or cross domain support? Possibly

· Would the government provide additional information on their current challenges ingesting and parsing documents? The RFI is intended for high-level capability assessment, we are unable to provide additional details beyond what is included in the RFI materials.

· Would the government provide more details on the current SLAs in regards to real time inference latency? The RFI is intended for high-level capability assessment, we are unable to provide additional details beyond what is included in the RFI materials.

· Would the government provide a range for acceptable response time limits for AI executed validation processes? The RFI is intended for high-level capability assessment, we are unable to provide additional details beyond what is included in the RFI materials.

· Would the government provide the security policy references for AI enforced compliance? NIST SP 800-53 Rev 5, NIST AI RMF, and FEDRAMP auth requirements where applicable. Systems must meet all requirements as listed in EOs and OMB Directives on AI such as EO 14147, EO 14179, EO 14283, EO 14275, EO 14319, M-25-21, M-25,22, and M-26-04 as well as any new EO and OMB Directives on the subject.

· Would the government provide the security levels required (protecting sensitive data)? Data processed within this environment is classified at the moderate impact level in accordance with FIPS 199 and NIST SP 800-60. All solution components must meet FEDRAMP moderate baseline control requirements at a minimum. Data containing CUI must be handled in accordance with NIST SP 800-171 and applicable DOJ CUI policy. In addition, Grand Jury secrecy data must be handled in according to DOJ Justice Manual 9-11.000. Encryption in-transit and at rest should fall within FIPS 140-2/140-3 modules.

· Is there a high limit for scalability the government would like to attain? Based on data volumes and growth. Current volumes are >450 TB in current economist environment with a predicted growth rate of 10% annually. For production annual data volume is between 200-300TB of new data with a growth rate of approximately 20-30% per year.

· Would the government provide a high-end limit for larger datasets? 100TB

· The term "efficient" is used in several requirement areas. Could the government provide a realistic high range for data sets and a time range defining "efficient"? FR 1.16 requires throughput of ≥1TB in under 90 minutes.

· To achieve the Division's goal of accelerating time-to-insight during fast-paced merger reviews and criminal prosecutions, how will the evaluation weigh a platform's native ability to instantly search and semantically cluster unstructured evidentiary data compared to its handling of traditional structured metrics? This will be carefully considered, but evaluation factors will be defined in the RFP.

· To ensure the highest level of prosecutorial integrity and evidentiary defensibility, how will the Division evaluate a solution's native capability to strictly ground generative AI outputs in case-specific evidence and mitigate hallucinations, while also providing visibility into context provided and reasoning for LLM auditability and defensibility? Evaluation by SMEs through testing. AI tools will go through a rigorous evaluation before they will be approved for use.

· To protect highly sensitive economic data and prevent spillage between restricted investigations, does the Division require the solution to enforce granular, document-level and field-level access controls natively within the data store index, rather than relying solely on application-level security? Yes

· To maximize platform adoption among non-technical investigative staff, does the Division prioritize solutions that provide intuitive, out-of-the-box natural language exploration of complex litigation data over platforms that require extensive coding or advanced querying expertise? No

· Can ATR clarify the expected scope vendors should assume for the requested professional services ROM, including whether it should cover discovery, pilot, production implementation, migration, training, and/or ongoing support? Additionally, does ATR anticipate a follow-on RFP/RFQ, demonstrations, oral presentations, or proof-of-concept activities after review of RFI responses? ATR will not provide a response for this RFI. This proposal should be based on your industry experience and present the best value to the government.

· Beyond Power BI and Tableau, which specific enterprise productivity, collaboration, or analysis tools are prioritized for integration or interoperability (e.g., Microsoft 365 applications, document management systems, case management, or litigation support tools)? No other tools have been prioritized

· What kind of databases or systems are you using currently for structured and unstructured data? The RFI is intended for high-level capability assessment, we are unable to provide additional details beyond what is included in the RFI materials.

File details come from the government source that posted it. Updated .