Amendment 0003_Attachment 1_Tool Threshold Assessment.docx

DOCX document 27 KB Posted

Attached to
U. S. Department of Treasury_ Artificial Intelligence (AI) Tools Federal contract opportunity
Solicitation number
2032H325Q00063
Issued by
Department of the Treasury Departmental Offices

About this file

Amendment 0003_Attachment 1_Tool Threshold Assessment is a detailed technical evaluation document for the U.S. Department of Treasury's AI Tools solicitation (2032H325Q00063). The document outlines rigorous technical and security criteria for two AI tool categories: a Coding Assistant Tool and an AI Chat Tool, with 22 and 14 specific threshold requirements respectively. Key evaluation criteria include standalone operation, context understanding, large context windows, SOC2 compliance, network efficiency, role-based access controls, security features, and comprehensive administrative controls.

The assessment mandates strict security measures such as data isolation, authentication compatibility, audit log compliance, and robust session management protections. Vendors must demonstrate capabilities like supporting multiple language models, providing detailed documentation, enabling admin consoles, implementing granular access controls, and addressing complex security scenarios such as preventing unauthorized session continuation and data exfiltration. The document requires vendors to prove their tools can meet sophisticated technical and security standards for use across Treasury and its bureaus, with specific thresholds for performance, integration, transparency, and compliance.

View the file

Other files for this federal contract opportunity

Other files attached to U. S. Department of Treasury_ Artificial Intelligence (AI) Tools, newest first.
File Type Posted
II_01_2032H325Q00063_Attachment 2_Price Matrix_Amendment 0004.xlsx XLSX spreadsheet
2032H325Q00063_Amendment 0004_.pdf PDF
Amendment 0004_ 2032H325Q00063_Revised Terms and Conditions.pdf PDF
Amendment 0003_2032H325Q00063.pdf PDF
2032H325Q00063_Amendment 0002_Final.pdf PDF
II_01_2032H325Q00063__Amendment 1.pdf PDF
II_01_2032H325Q00063_Attachment 2_Price Matrix_Amend1_20250709.xlsx XLSX spreadsheet
II_01_2032H325Q00063_Attachment 1_SOW_Amend 1.pdf PDF
II_QA_2032H325Q00063.xlsx XLSX spreadsheet
DRAFT RFQ_Questions and Answers_AI Tools.xlsx XLSX spreadsheet
II_01_2032H325Q00063_Attachment 1_SOW.pdf PDF
II_01_2032H325Q00063_Attachment 2_Price Matrix.xlsx XLSX spreadsheet
II_01_2032H325Q00063_TCs.pdf PDF
Show all 13

On GovTribe

Work with this file on GovTribe

  • Download the original file
  • Contacts named in this file
  • Similar government files
  • Ask GovTribe AI about this file

Text version

Solicitation Number: 2032H325Q00063 Vendor Name:

Coding Assistant Tool Criteria

#
Criteria
Threshold
Tool Threshold Assessment

(Meets / Does Not Meet) Additional Information / Supporting Information

1
Standalone Operation
Operates locally; no external server or cloud dependency beyond IDE.
2
Context Understanding (Large Code Bases)
100 files or 1M lines indexed
3
Agentic Coding and Multi-File Editing
Autonomous tasks across 5 files
4
Large Context Window
200k tokens
5
SOC2 Compliance
SOC2 Type II or equivalent
6
Vector Database Embedding Option
Indexing 500k files
7
Speed of Inferencing
Latency 1s on standard hardware
8
Autocomplete Relevance (95%)
95% relevance, 10 languages
9
Network Traffic Efficiency
10 MB/min for typical session
10
CPU Processing Efficiency
20% CPU on standard hardware
11
Access to Provisioned Models
Custom/fine-tuned models via APIs
12
SSO SAML 2.0 Support
Explicit SAML 2.0 support (ADFS-compatible)
13
Support for Top-Tier Models
The agentic IDE must support multiple LLMs, with at least on having >= 10 billion parameters or a >=200k token context window, verified by official documentation, to provide functionality for code generation, debugging, multi-file navigation, and compatibility with a range of models.
14
Documentation Creation
Accurate docs for 5 files or modules. Documentation must be generated in Markdown format to support easy version control and conversion to other formats. Generated documentation should include function/method descriptions, parameter details, return values, usage examples, inline code comments, and API references.
15
Support for Prompting
Natural language prompting for tasks
16
Admin Console Command Enable/Disable
Admin console for command control
17
Whitelist/Blacklist Terminal Commands
Whitelist/blacklist terminal commands. Terminal command controls must be configurable at both global and role-based levels, with project-level overrides where appropriate. Vendors should provide default safe/unsafe command lists based on industry best practices, while allowing Treasury administrators to customize these lists for specific security requirements.
18
Usage Analytics
Analytics for 10 users. Key metrics must include active usage time, feature utilization, completion acceptance rates, error frequencies, and performance metrics. Both interactive dashboards and raw data exports (CSV/JSON) are required to support internal reporting and integration with Treasury's existing analytics systems.
19
Role-Based Access Controls (RBAC)
3 distinct roles with permissions. The solution must provide both predefined roles (e.g., Administrator, Developer, Reviewer) and support for fully customizable roles. RBAC must integrate with Treasury's identity management systems through standard protocols like SAML 2.0, with the ability to map external groups to internal roles.
20
Custom API Integration
APIs/SDKs for ≥5 integration types. Integration types must include CI/CD pipelines (e.g., GitHub Actions, Azure DevOps), issue tracking systems (e.g., Jira, Azure DevOps), code repositories, knowledge management systems, and security scanning tools. RESTful APIs are required.
21
Testing and Validation Framework
≥2 industry standard languages tested with automated methods (e.g., unit or integration tests). The framework must support common testing patterns such as mocking, assertions, and test coverage reporting.
22
Model Transparency and Reasoning Disclosure
1 documented example showing how the model selects or prioritizes code suggestions, with explanation of underlying logic

AI Chat Tool Criteria

#
Criteria
Threshold
Tool Threshold Assessment

(Meets / Does Not Meet) Additional Information / Supporting Information

1
Data isolation
Treasury data not used for model training.
2
Data residency
Data retention within U.S.
3
Data Control
Admin can delete/purge data. The system must support both immediate deletion and configurable soft delete with recovery windows (up to 30 days). Comprehensive audit trails of all deletion actions are required, including what was deleted, by whom, when, and under what authority. These logs must be tamper-proof and accessible for compliance verification.
4
Access isolation
Data segregation and RBAC.
5
Explainability
Transparency and user-level control over AI outputs.
6
Security hardening
Guardrails for adversarial input protection.
7
Content filtering
Guardrails against prohibited terms/content.
8
Authentication compatibility
Mandatory integration with authentication (CAIA/ESIM).
9
Audit/log compliance
Log retention must comply with M-21-31 requirements, with a minimum retention period of 12 months for high-value audit log data. Access controls must limit log access to authorized security personnel and support the principle of least privilege. Logs are readily available for regular internal audits and security investigation purposes.
10
Testing and validation
3 adversarial or prompt injection test cases with expected vs. actual outputs
11
Model Transparency and Reasoning Disclosure
1 documented example of chat output with explanation of response generation and filtering logic. Transparency documentation must include confidence scores for suggestions, information about input factors that influenced the response, and explanation of any rule-based overrides or guardrails that modified the output.
12
Prompt Validation
Filters user prompts for toxicity, PII, malicious intent, and off-topic content.
13
Response Validation
Validates model responses for bias, hallucination, sensitive data, and includes transparency or explanation features.
14
User Onboarding/Offboarding Support
Supports automated user provisioning and

de-provisioning per government approval workflow.

Security Scenario Narrative Vendors must also address the following real-world scenario to demonstrate control against unauthorized session continuation and exfiltration for proposed tools: A user logs into the AI tool using their official Treasury credentials and performs work tasks. They then copy a session cookie or equivalent authentication artifact. Later, from a personal (non-Treasury) device, they use that session to access the workspace and download restricted materials—bypassing network controls normally enforced within Treasury’s secure environment. Vendors must describe how their tool detect, prevent, or mitigate this session hijack/exfiltration. Responses should detail: • Session authentication lifecycle and token handling • IP/device-based session restrictions • Token expiration and revalidation • Admin override or alert mechanisms Security Controls Evaluation

#
Criteria
Threshold
Tool Threshold Assessment

(Meets / Does Not Meet) Additional Information / Supporting Information

1
Session authentication lifecycle and token handling
Tool must demonstrate control over session lifecycle and secure token management. Tool must detect, prevent, or mitigate session hijack using secure session lifecycle and token management.

Session tokens within the system must have a maximum lifetime of 12 hours, with sliding expiration supported for active sessions. Refresh tokens should be limited to 24 hours maximum. All tokens must be securely stored, transmitted, and revocable through administrative action. Token validation must occur with each significant action, not just at session initiation.

2
IP/device-based session restrictions
Tool must restrict session use to known IPs/devices. Tool must prevent unauthorized reuse of sessions by enforcing IP/device-based access restrictions.
3
Token expiration and revalidation
Tool must enforce short-lived tokens with revalidation. Token revalidation within the system shall occur at minimum every 30 minutes of active use. Idle timeout must be configurable, with a default maximum of 15 minutes of inactivity before requiring reauthentication.
4
Admin override or alert mechanisms
The tool must provide administrative override capabilities, including forced logout and session termination for individual users or groups. Real-time alerts must be generated for suspicious activities such as multiple failed authentication attempts, access from unusual locations, or attempted session hijacking.

File details come from the government source that posted it. Updated .