Toth Technologies on VA SOW REVISED4-22-2026.docx
DOCX document 35 KB Posted
- Attached to
- DA10--Intent to sole source Data Aggregator and Viewing Tool Federal contract opportunity
- Solicitation number
- 36C24E26Q0042
About this file
This is a Statement of Work (SOW) for a Data Aggregator and Viewing Tool to be developed for the Department of Veterans Affairs Research team. The vendor will build a web-based interface and queryable database to aggregate and index VA-housed healthcare data, including radiologic imaging (21,201 CT studies comprising 5.3 terabytes, 64,068 prostate MRI studies totaling 14.5 terabytes, and 2,488 head and neck imaging studies representing 0.6 terabytes), 41,895 digitized pathology slides (42 terabytes), and 3,703 additional pathology slides (3.0 terabytes). The system will allow VA researchers and data scientists to easily access, filter, visualize, and export anonymized data subsets for research purposes.
The SOW specifies 13 primary tasks organized into four categories: ingestion and preprocessing (multi-modality ingest engine, data harmonization), database layer (schema definition, direct-to-database loader), web-based exploration UI (browseable patient index, data visualizations, patient record viewer), and data export capabilities (export interface, network drive access, downloader tool). Additional requirements include integration testing, deployment on a dedicated Research Scientific Computing Device outside the VA base network, and comprehensive user documentation with video instructions. The contract is firm fixed price with a 120-day period of performance from contract start, with weekly progress meetings required. The vendor will not access actual VA data, using only publicly available data for development and testing. Deliverables include functional software passing three evaluation steps, a deployed web interface, and successfully demonstrated data download functionality. The backend will utilize Python, SQLAlchemy, and SQLite/PostgreSQL; the frontend will use Flask, Bootstrap, and Plotly; and deployment will use Docker containerization or Podman. Place of performance is the contractor's site.
View the file
Other files for this federal contract opportunity
| File | Type | Posted |
|---|---|---|
| 36C24E26Q0042.docx | DOCX document |
On GovTribe
Work with this file on GovTribe
- Download the original file
- Contacts named in this file
- Similar government files
- Ask GovTribe AI about this file
Text version
STATEMENT OF WORK (SOW)
1. Contracting Officer’s Representative (COR).
2. Contract Title.
Data Aggregator and Viewing Tool
3. Background.
The goal for the Research team procuring this service is to aggregate and index VA-housed healthcare data available for Research, and allow in-house researchers and data scientists to easily utilize this data for their work. All data specified below are part of IRB/R&D approved VA studies and are also VA Office of Research and Development (ORD) funded. Radiologic data for our lung cancer research currently includes 21,201 full-resolution Digital Imaging and Communications in Medicine (DICOM) imaging studies, corresponding to approximately 5.3 terabytes of CT data curated from multiple VA medical centers. Prostate cancer imaging aggregation currently includes 64,068 prostate magnetic resonance imaging (MRI) studies, totaling approximately 14.5 terabytes of radiologic data. In addition, 41,895 digitized pathology slides have been generated from biopsy and surgical specimens, comprising approximately 42 terabytes of whole slide imaging data. The head and neck cancer dataset includes 2,488 radiologic imaging studies, totaling approximately 0.6 terabytes, and approximately 3,703 digitized pathology slides, comprising 3.0 terabytes of whole slide imaging data. This data has also been anonymized using a custom tool already delivered to the VA. This goal will be achieved by building an easy to use web interface backed by a queryable database that contains a manifest of all the data available (identifiable and anonymized/coded). This will be built on a Research Scientific Computing Device approved through enterprise risk assessment (ERA) and located outside of the base VA network. This system, once complete, will allow VA scientists to see what data is available, and provide an easy ability to get that data into their workspace, whether through downloading or pointing to a network drive. The ultimate goal is to reduce the friction between scattered datasets and novel healthcare breakthroughs within the VA. The datasets are too scattered and there’s too much inefficiency in being able to utilize the data for medical research. This tool will accelerate the VA’s efforts to develop new precision medicine tools for Veterans.
4. Scope.
The procurement for computing services to enable the automated building of a database containing a manifest of available data with a deployed web interface for selecting datasets or subsets, and the ability to download selected data for research within the VA. The software is custom and only useful our research operations and will only be used on our dedicated RSCD. The vendor will not be given access to VA data but use publicly data available for creation. All testing and evaluation of the database will be conducted by VA personnel in consultation with the vendor.
5. Specific Tasks.
Specification Overview:
Ingestion & Preprocessing
1. Multi-Modality Ingest Engine
· Accept manifests (CSV) of pre-anonymized data.
· Accept manifests of data with PHI, extract common patient identifiers (SSN+DOB), and generate internal randomized patient keys.
· Normalize date formats, field names, and data types across modalities.
2. Data Harmonization Layer
· Merge patient data from all modalities into unified data structure.
· Resolve discrepancies across modalities (e.g., demographic mismatches).
· Flag or log ambiguous merges for review.
· Link associated medical imagery to the patient/study.
· Add fields like PatientAge, ModalityFlags, LatestLabDate, etc.
· Condense one-to-many lab/test records into structured arrays or summaries.
Database Layer
3. Schema Definition and Migrations
· Define SQL schema for each of: Patient, CT, MRI, Lab_Test, etc.
· Define foreign key constraints.
· Include indices on PatientID, StudyDate, etc.
4. Direct-to-DB Loader
· Load processed manifest records into SQL database.
· Provide hooks for batch ingestion, validation, and rollback on error.
· Web-Based Exploration UI
5. Browseable Patient Index
· Build a Flask UI to browse by patient fields (e.g., modality, indication, body part, age, lab results).
· Support filters, pagination, and search history.
6. Data Visualizations
· For both filtered and cumulative data, display charts:
· Modality counts (bar/pie chart)
· Age distributions (histogram, violin)
· Number of studies per year (line/bar)
· Longitudinal time length (years of data available per patient, histogram)
7. Patient Record Viewer & Selector
· Table or card-based interface for viewing patient summaries.
· Include checkbox-style selection with "Select All" and "Deselect All" functions.
· Persist selections via session or browser storage.
· Data Export
8. Export Interface
· Allow the user on the web app to select either all or a subset of filtered data.
· Show a “checkout summary” screen, that shows statistics of the selected dataset, such as total size (e.g. in gigabytes) and estimated download time.
· Show options to access data via network drive (preferred) or download the data.
· Export a manifest file in either JSON or CSV format of the selected dataset showing metadata about the data.
9. Network Drive Access
· If the user has access to a network drive with the source data, allow them to download a location manifest to access the source data.
· Have the docker container running the backend to create a new directory for the user, with a UUID subfolder name, that contains symlinks to selected data.
10. Downloader Tool
· Create a tool that inputs a manifest file and downloads the data locally.
· Containerize it.
· Secure it with best practices so only authorized users can download the data.
· Show progress bar, ETA, graphical elements as needed.
Deployment & Testing
11. Integrations and Unit Tests
· Validate merge logic, anonymization, and derived fields.
· Simulate user actions across UI: ingestion → filtering → selection → export.
· Confirm that selected records are correctly exported in requested format and anonymized.
12. Deployment
· Let researchers stand up the web interface within RSCD.
· System should regularly check it has access to the network drive of source data.
· Have single source of truth (on network drive) of database of available data.
· Easy link to downloader tool.
13. User Documentation & Usage Examples
· Write clean, public-facing README with:
· Overview of the tool
· How to ingest and explore data
· How to select, anonymize, and export for modeling
· Screenshots of the web interface
· Create video instructions and examples
System Requirements:
Modalities:
CT, MRI, Digital Pathology, Bloodwork, Test Results Data Pipeline:
Anonymize → Ingest → Preprocess → Store → Explore → Select → Export Backend:
Python + SQLAlchemy + SQLite/PostgreSQL https://www.oit.va.gov/Services/TRM/ToolPage.aspx?tid=5702 https://www.oit.va.gov/Services/TRM/ToolPage.aspx?tid=6406
Frontend:
Flask + Bootstrap + Plotly https://www.oit.va.gov/Services/TRM/ToolPage.aspx?tid=10316 https://www.oit.va.gov/Services/TRM/ToolPage.aspx?tid=7795
Deployment:
Docker containerization/ Podman - https://www.oit.va.gov/Services/TRM/ToolPage.aspx?tid=14897&tab=2
Evaluation by Research Team:
Step 1 will be to take an existing public data study, ingest it, and then query known data (as a sanity check) on RSCD (non-GFE device).
Step 2 will be to run this on RSCD for several VA data sets, then, someone from the Research team will manually check for accuracy of aggregation and visualization of data.
Step 3 will be to work with the vendor to present aggregate results and set up punch list to fix any errors.
Deliverables:
· Delivery of functional software that passed the evaluation steps shown above.
· Database containing manifest of available data.
· Deployed web interface for selecting datasets or subsets.
· Successfully downloading selected data for research.
6. Performance Monitoring The software is expected to be completed within 120 days and the requestor will have weekly meetings to discuss progress on the software. Once the evaluation steps are completed, the code will be run on a VA device and ensured to operate as described.
7. Contract Type: Firm fixed price
8. Government-Furnished Equipment (GFE)/Government-Furnished Information (GFI).
None
11. Place of Performance.
Contractor’s site
12. Period of Performance.
Until delivery, estimated to be done within 120 days of contract start.
13. Delivery Schedule.
Within 120 days of contract start.
14. Security Requirements Security Requirements apply.
File details come from the government source that posted it. Updated .