IARPA-BAA-24-04_VideoLINCS_Amendment 004_01 AUG_ 24_C.pdf
PDF 931 KB Posted
- Attached to
- Video Linking and Intelligence From Non-Collaborative Sensors (Video LINCS) Program Federal contract opportunity
- Solicitation number
- IARPA-BAA-24-04
About this file
This document is a Broad Agency Announcement (BAA) from the Intelligence Advanced Research Projects Activity (IARPA) for the Video Linking and Intelligence from Non-Collaborative Sensors (Video LINCS) program. The Video LINCS program aims to develop algorithms to autonomously re-identify objects across diverse video sensor data and map re-identified objects to a common reference frame. IARPA is seeking innovative solutions for this 48-month effort, which is divided into three phases. The program will pursue rigorous independent testing and evaluation, with performers required to provide containerized software and share datasets across the program. Key objectives include person, vehicle, and generic object re-identification and geo-localization, with performance metrics defined for each phase. Proposals are due August 9, 2024, with multiple awards anticipated. The Government will evaluate proposals based on technical merit, proposed work plan, program relevance, experience, and resource realism.
View the file
Other files for this federal contract opportunity
| File | Type | Posted |
|---|---|---|
| Video LINCS BAA 24 Round 1 and 2 updated responses_ 7-29-24_C .pdf | ||
| IARPA-BAA-24-04_VideoLINCS_Amendment 003_26 JUL 24_C.pdf | ||
| VideoLINCS BAA questions Round 2_26 JUL 24_C.pdf | ||
| IARPA-BAA-24-04_VideoLINCS_Amendment 002_11 JUL 24_C.pdf | ||
| VideoLINCS BAA questions Round 1_11 JUL 24_C (1).pdf | ||
| IARPA-BAA-24-04_VideoLINCS_Amendmement 001_C.pdf | ||
| IARPA-BAA-24-04_VideoLINCS_20240521_C.pdf |
On GovTribe
Work with this file on GovTribe
- Download the original file
- Contacts named in this file
- Similar government files
- Ask GovTribe AI about this file
Text version
UNCLASSIFIED
IARPA
Broad Agency Announcement
IARPA-BAA-24-04
Video Linking and Intelligence From Non-Collaborative Sensors
(Video LINCS) Program
BAA Release Date:
May 15, 2024
Amendment 001: June 25, 2024
ALL CHANGES FOR AMENDMENT 001 ARE HIGHLIGHTED IN YELLOW
Amendment 002: July 9, 2024
ALL CHANGES FOR AMENDMENT 001 ARE HIGHLIGHTED IN GREEN
Amendment 003: July 26, 2024
ALL CHANGES FOR AMENDMENT 001 ARE HIGHLIGHTED IN RED
Amendment 004: August 1, 2024
ALL CHANGES FOR AMENDMENT 001 ARE HIGHLIGHTED IN BLUE
Table of Contents
1 Table of Contents
SECTION 1: FUNDING OPPORTUNITY DESCRIPTION
1.A. Program Overview
1.A.1.a TA-1 – ReID Challenges
1.A.1.b TA-2 – Object Geo-localization Challenges
1.A.2. Program Phases
1.A.2.a Phase 1
1.A.2.b Phase 2
1.A.2.c Phase 3
1.B Team Expertise
1.C Program Scope and Limitations
1.D Program Data
1.D.1 Development Data
1.D.2 Evaluation Data
1.D.3 External Data Sources
1.E Testing & Evaluation (T&E)
1.F Program Metrics
1.G Program Waypoints, Milestones and Deliverables
1.G.1 Waypoints
1.G.2. Program Waypoints, Milestones and Deliverables Timeline
1.G.3. Software Deliverable Formatting
1.G.3.1 Program API and Framework
1.H Meeting and Travel Requirements
1.H.1. Technical Exchange Meetings and Workshops
1.H.2. Site Visits
1.I Period of Performance
1.J Place of Performance
1.K References
SECTION 2: AWARD INFORMATION
SECTION 3: Eligibility Information
UNCLASSIFIED
3.A Eligible Applicants
3.A.1 Organizational Conflicts of Interest (OCI)
3.A.2 Multiple Submissions to the BAA
3.B U.S. Academic Institutions
3.C Other Eligibility Criteria
3.C.1 Collaboration Efforts
SECTION 4: PROPOSAL AND SUBMISSION INFORMATION
4.A Proposal Information
4.B. Proposal Format and Content
4. B.1 Volume 1: Technical and Management Proposal
4. B.1.a Section 1: Cover Sheet & Transmittal Letter
4. B.1.b Section 2: Summary of Proposal (see below for page limit)
4. B.1.c. Section 3: Detailed Proposal Information
4.B.1.d. Section 4: Attachments
4.B.2. Volume 2: Cost Proposal (No Page Limit)
4. B.2.a. Section 1: Cover Sheet
4. B.2.b. Section 2: Estimated Cost Breakdown
4.B.2.c. Section 3: Supporting Information
4.C Submission Details
4.C.1. Due Dates
4.C.2. Proposal Delivery
4.D Funding Restrictions
SECTION 5: PROPOSAL REVIEW INFORMATION
5.A Technical and Funding Availability Evaluation Factors
5.A.1. Technical Evaluation Factor (technical criteria listed below)
5.A.1.a. Overall Scientific and Technical Merit
5.A.1.b. Effectiveness of Proposed Work Plan
5.A.1.c. Contribution and Relevance to the IARPA Mission and Program Goal
5.A.1.d. Relevant Experience and Expertise
5.A.1.e Resource Realism
5.A.2. Funding Availability Factor
5.A.2.a. Budget Constraints
5.A.2.b. Program Balance
UNCLASSIFIED
5.B Method of Evaluation and Selection Process
5.C Negotiation and Contract Award
5.D Proposal Retention
SECTION 6: AWARD ADMINISTRATION INFORMATION
6.A Award Notices
6.B Administrative and National Policy Requirements
6.B.1. Proprietary Data
6.B.2. Intellectual Property
6.B.3. Human Use
6.B.4. Animal Use
6.B.5. Publication Approval
6.B.6. Export Control
6.B.7. Subcontracting
6.B.8. Reporting
6.B.9. System for Award Management (SAM)
6.B.10. Representations and Certifications
6.B.11. Lawful Use and Privacy Protection Measures
6.B.12. Public Access to Results
6.B.13. Other Contract Requirements
6.B.13.a. Provisions
6.B.13.b. Clauses
SECTION 7: APPENDIX
APPENDIX A - Templates for Volume 1: Technical Proposal
Appendix A.1: Cover Sheet for Volume 1: Technical and Management Proposal
Appendix A.2: Academic Institution Acknowledgment Letter
Appendix A.3: Intellectual Property and Data Rights
Appendix A.4: Organizational Conflicts of Interest Certification Letter
Appendix A.5: Three Chart Summary of the Proposal
Appendix A.6: Research Data Management Plan (RDMP) BAA 24-04
Appendix A.7 Statement of Work (SOW) Template: (RESERVED)
APPENDIX B - Templates for Volume 2: Cost Proposal
Appendix B.1: Cover Sheet for Volume 2: Cost Proposal
Appendix B.2: Contractor/Subcontractor Cost Element Sheet for Volume 2 Cost Proposal
UNCLASSIFIED
Appendix B.3: Software and IP Costs
Appendix B.4: Travel Costs Trip breakdown
Appendix B.5 Contract Deliverables Table
GENERAL INFORMATION
This notice constitutes a Broad Agency Announcement (BAA) and sets forth research of interest in the area of video-based object re-identification and geo-localization The solicitation process will follow Federal Acquisition Regulation (FAR) Part 35, Research and Development Contracting, as supplemented with additional information included in this notice. Awards based on responses to this BAA will be considered the result of full and open competition.
1. Federal Agency Name – Office of the Director of National Intelligence (ODNI) / Intelligence
Advanced Research Projects Activity (IARPA)
2. Funding Opportunity Title – Video Linking and Intelligence from Non-Collaborative
Sensors (Video LINCS) Program
3. Announcement Type – Initial
4. Funding Opportunity Number – IARPA-BAA-24-04
5. Catalog of Federal Domestic Assistance Numbers (CFDA) – Not applicable
6. Questions
Submit questions on administrative, technical, or contractual issues by email to DNI-IARPA- BAA-24-04@iarpa.gov. All requests must include the full name and affiliation of a point of contact. Do not send questions with proprietary content. A consolidated Question and Answer response will be posted on SAM.gov for Contract Opportunities website (https://SAM.gov/) and linked from the IARPA website (https://www.iarpa.gov/research-programs/video-lincs).
No answer will go directly to the submitter. IARPA will accept questions until July 17, 2024
@ 4:00 PM ET.
7. Dates
7.1 Posting Date: May 21, 2024
7.2 Questions: July 17, 2024 @ 4:00PM EST
7.3 Proposal Due Date for Initial Round of Selections: August 9, 2024 @ 4:00 PM ET
7.4 BAA Closing Date: August 13, 2024, 4:00 PM ET
8. Anticipated individual awards – Multiple awards anticipated mailto:DNI-IARPA-BAA-24-04@iarpa.gov mailto:DNI-IARPA-BAA-24-04@iarpa.gov https://sam.gov/ https://www.iarpa.gov/research-programs/video-lincs
UNCLASSIFIED
9. Types of instruments that may be awarded – Procurement Contracts and Other Transactions
10. Agency Points of Contact
ATTN: IARPA-BAA-24-04
Office of the Director of National Intelligence Intelligence Advanced Research Projects Activity Washington, DC 20511 Electronic mail: DNI-IARPA-BAA-24-04@iarpa.gov
11. Program Manager (PM) – Dr. Reuven Meth
12. Program Website – https://www.iarpa.gov/research-programs/video-lincs
13. BAA Summary – This BAA (IARPA-BAA-24-04) is for the Video Linking and Intelligence from Non-Collaborative Sensors (Video LINCS) Program. IARPA is seeking innovative solutions for the Video LINCS Program in this BAA, which is envisioned to be a 48-month effort across all technical areas. The Video LINCS program aims to develop novel capabilities to autonomously re-identify objects across diverse video sensor collections and map all objects to a common reference frame.
1 Procurement Contract: This is a standard government contract that follows the processes, format and terms and conditions as outlined in the Federal Acquisition Regulations (FAR) and supplementing Agency specific regulations.
Other Transaction Authority: Agreements generally are not subject to the federal laws and regulations governing procurement contracts and thus are not required to comply with the Federal Acquisition Regulation (FAR), its supplements, or laws that are limited in applicability to procurement contracts. They may be used with non-traditional contractors under certain circumstances.
https://www.iarpa.gov/research-programs/video-lincs
UNCLASSIFIED
SECTION 1: FUNDING OPPORTUNITY DESCRIPTION
The Intelligence Advanced Research Projects Activity (IARPA) often selects its research efforts through the Broad Agency Announcement (BAA) process. The use of a BAA solicitation allows a wide range of innovative ideas and concepts. The BAA shall appear first under Contract Opportunities on https://sam.gov, then the IARPA website at http://www.iarpa.gov/. The following information is for those wishing to respond to this Program BAA.
This BAA (IARPA-BAA-24-04) is for the Video Linking and Intelligence from Non-Collaborative Sensors (Video LINCS) Program. IARPA is seeking innovative solutions for the Video LINCS Program in this BAA, which is envisioned to be a 48-month effort across all focus areas, beginning approximately May 2025.
1.A. Program Overview
In today’s world, video is a ubiquitous sensor. Voluminous video data is collected (e.g., security, closed circuit television (CCTV), aerial, etc.), but the manpower and available attention to provide desired security is limited. Operators are overloaded with many video feeds with limited attention for the available data volumes or ability to discern threats. Tragic incidents trigger forensic analyses that require significant manpower to pore through troves of imagery.
Diverse video sensors, even when installed at a single facility, often do not collaborate with each other. Some automated analytics have been developed, ranging from motion detection to object classification to summarization, providing valuable tools for filtering video content. These capabilities assist operators in focusing their attention. Ultimately, the burden on operators is high, requiring significant expertise, recall and resources to associate content across footage and discern threats.
The goal of the Video LINCS program, it to develop re-identification (reID) algorithms to autonomously associate objects across diverse, non-collaborative, video sensor footage and map re-identified objects to a unified coordinate system (geo-localization). The reID and geo-localization algorithms will distill raw pixel data into spatio-temporal motion vectors, providing the ability to analyze these patterns for anomalies and threats. While the ultimate goal will be to re-identify general objects, the program will start with person reID, progress to vehicle reID, and culminate with reID of generic objects across a video collection.
The overall program design is depicted in Figure 1, where a diverse collection of video imagery from different sensors is input into the Video LINCS system that reidentifies objects across the collection and geo-localizes those objects, providing tracks for each object in both the initial camera coordinate and common geo-coordinate reference frames. Offerors to the Video LINCS program shall propose solutions for both the re-identification and geo-localization objectives.
https://sam.gov/ http://www.iarpa.gov/
Display all object locations on a single map
Figure 1: Video LINCS Program Design
It is anticipated that annotated videos of people, vehicles and generic objects over a range of sensor types and collection conditions will be central to both Research and Development (R&D) and Test and Evaluation (T&E) objectives. Offerors who propose use of such R&D data will be required to collect (and/or simulate) and annotate R&D datasets throughout the program lifecycle and share them across the program with other performers. The Government will collect and annotate data primarily for T&E objectives and may potentially provide some data for R&D purposes. All data collections with potential Human Subject Research (HSR) will require Institutional Review Board (IRB) approval, and all performers will be required to deliver a privacy plan that comprehensively describes how any potential personally identifiable information (PII) will be safeguarded.
The Video LINCS program will pursue rigorous and comprehensive independent T&E to ensure that research outcomes are well characterized, deliverables are aligned with program objectives, and performance is measured across a wide range of conditions. T&E activities will not only inform Government stakeholders on Video LINCS research progress but will also serve as valuable feedback to the performers to improve their research approaches, algorithm training practices, and system development. The Video LINCS program will continually refine and improve T&E methodologies throughout the lifecycle of the program.
Technical Challenges and Objectives The Video LINCS program consists of two Technical Areas (TAs):
• TA-1 Re-identification (ReID): Autonomously and automatically associating the same object (person, vehicle, or generic object) across a video corpus.
• TA-2 Object Geo-localization: Geo-locating objects to provide positions for all objects in a common world reference frame.
The objective is for Video LINCS to be effective in the open-world setting where there is no advance knowledge of the sensor types, sensor collection geometries, scenes, objects to be re-identified, etc. There will be no a priori set of people, vehicles or objects to be queried or gallery/library to serve as a reference for matching. The goal will be to re-identify arbitrary content within the video set, to determine the presence and location of each object throughout the video collection, without any auxiliary or prior knowledge of that object. This presents multiple challenges for both reID and object geo-localization.
UNCLASSIFIED
1.A.1.a TA-1 – ReID Challenges
The Video LINCS reID objective is to reidentify objects across an arbitrary video collection without any prior access to the representative video or the objects to be reidentified. This presents a number of challenges.
Video in a collection may come from many different sources and collection geometries, containing significantly different characteristics that are difficult to associate. The variations in appearance within and across video sensor footage, which may range across many collection types such as stationary and mobile; ground-based, building/roof/pole mounted and aerial; outdoor and indoor;
day and night; varying weather conditions; temporally proximate and distant;
noise/artifacts/corruption, etc., present significant challenges to reidentifying objects across such collections.
The reID objective varies significantly from identification (ID). For ID, specific identities are known a priori and queries are matched to a gallery/library of candidates. For reID on an arbitrary video collection, there are no a priori galleries. The lack of an a priori gallery of objects to be re-identified will require systems to autonomously determine when to expand system generated galleries to include additional objects vs. expanding matches to existing objects. For people, recognizable faces will predominantly be unavailable throughout the video collection. For vehicles, license plates will not be available for association.
While general categories of video will be known (e.g. aerial drone, building mounted security and CCTV, etc.) no representative data will be provided for system development and training.
Furthermore, Video LINCS will focus on correct association/reID across all appearances of the same object in collection – not merely retrieving a single match such as that measured by the Cumulative Match Characteristic (CMC).
Video LINCS is seeking an end-to-end reID solution, where video is input into the system and reidentified object locations are output. The system needs to automatically locate the objects and associate them - across scale, aspect, density, crowding, obscuration, etc. - without introducing false detections and false matches.
Video LINCS will take a staged approach, beginning with cases where the class of the object to be re-identified is known, such as people and vehicles, and ultimately extend to all objects in a video collection when there is no a priori knowledge of the object class.
1.A.1.b TA-2 – Object Geo-localization Challenges
The Video LINCS geo-localization objective is to provide spatio-temporal tracks of re-identified objects, which requires mapping from camera coordinates to a common reference frame – nominally, a geo reference frame.
Object geo-localization presents a number of challenges. Homographies are often used to transform from image-plane to ground-plane coordinates, but the homographies are often computed manually, and are based on the assumption/approximation of a planar scene. Geometric inverse mapping using camera pose and auxiliary information may be attempted, but camera pose
UNCLASSIFIED
is not always available, and scale must be resolved. Even when camera pose is available, it may not be sufficiently accurate to facilitate inverse mapping.
Video LINCS TA-2 objectives will include autonomous coordinate remapping for a wide variety of video collections and conditions, to facilitate autonomous operation under broad conditions where different levels of metadata and reference imagery are available.
1.A.2. Program Phases
The Video LINCS program is planned as a 48-month Research and Development (R&D) effort and is divided into three phases. In an 18-month Phase 1, Video LINCS teams will demonstrate the feasibility of reID for people in a video corpus and object geo-localization to provide all subject motion in a common reference frame. During this phase, each person’s clothing will remain the same (short-term / temporally proximate reID) and provided metadata (e.g. time-stamps and camera pose, to the extent available) will be noise free. During an 18-month Phase 2, reID will expand to include people with clothing changes (long-term / temporally distant reID), include vehicles, require functionality on generic objects, additional sensor types and collection geometries will be included in the evaluations, and noise will be introduced into provided metadata. During a 12-month Phase 3, evaluation will have a stronger focus on reID of generic objects, temporally distant reID of vehicles will be performed, additional sensor types and collection geometries will be included, and there will be greater uncertainty in camera pose. The progression of objectives for both TAs across the program phases is shown in Figure 2.
18 Months 18 Months 12 Months
Phase 1 Feasible
18 Months
TA1 - ReID
• People
• Temporally proximate
(people - same clothing)
• Within collection type
• Accurate time-stamps
TA2 - Geo-localization
• Camera pose provided
(location & orientation)
Phase 2 Robust
18 Months
TA1 - ReID
• People, vehicles, generic objects
• Temporally distanced (people - attribute changes)
• Across collection types
• Time-stamp errors
TA2 - Geo-localization
• Camera location provided
(no orientation)
Phase 3 Generalize
12 Months
TA1 - ReID
• Generic objects
• Temporally distanced
(vehicles)
• Across broader collection types
TA2 - Geo-localization
• Camera general vicinity provided
Figure 2: Phase Objectives
Offerors are required to propose to all phases (Phases 1, 2, & 3) and all Technical Areas (TA-1 & TA-2) of the Program under this BAA. Proposals that omit any of phases or Technical Areas will be considered non-compliant.
Test set conditions will expand throughout the program phases to increase the number of
UNCLASSIFIED
people/vehicles to be re-identified, increase the number of distinct identities within the datasets, increase the number of sensor collection classes/types, and lower the level of auxiliary metadata available. A listing of nominal test-set conditions throughout the phases is provided in Table 1.
The exact conditions are subject to potential update during program execution.
Table 1: Nominal test set conditions throughout program phases.
Phase 1 Phase 2 Phase 3 People # Distinct truthed people 750 1,500 2,500 % Small (< 150 pixels high) ≥ 50% all phases % Medium (150 < pixels high <
250) < 30% all phases
% Large (250 < pixels high) < 20% all phases # Additional confusers 1,000 5,000 10,000 # Attribute changes per person N/A ≥ 5 ≥ 5 # Sensor collection types 3 5 6 Vehicles # Distinct truthed vehicles N/A 1,000 2,000 Video sensor pose
Collection sensor location and orientation when available
Full pose provided
Location, no orientation
Approximate location within 100 m
Leaderboards are planned to provide ongoing measurement of performance with rankings for all performers. IARPA will provide access to self-test datasets for performers to run algorithms on their own hardware with results submitted to a T&E hosted server for scoring and leaderboard posting. A separate sequestered data leaderboard will be used for independent evaluation where performer algorithms will be run by the T&E team on a sequestered dataset hosted on T&E hardware. Monthly leaderboard submissions (initially to the self-test leaderboard, and subsequently to the sequestered data leaderboard) will be required to provide ongoing insight into performer progress. Semi-annual waypoints will be incorporated where performance will be reported to IARPA management and software will be delivered to T&E to facilitate independent builds and runs from source code. Performer data collections and annotations must be shared with other performers after each semi-annual performance reporting waypoint. Key program activities and milestones are shown in Figure 3.
Figure 3: Video LINCS performer key activities and milestones throughout the program
Performer teams will be responsible for obtaining (collecting, curating, simulating) and annotating data in support of proposed R&D objectives. The Government will collect, curate and annotate data for evaluation. Performers will submit containerized (nominally Docker) software using/conforming to the framework and APIs specified by the T&E team for independent runs on
UNCLASSIFIED
sequestered evaluation datasets.
Varying levels of reference data may be made available to systems at test time, especially for later phases when less camera information is provided. For outdoor scenes, such reference data may include georeferenced orthographic satellite imagery such as that available from Google Earth or Maxar Vivid Basemaps. Such imagery and maps may be dated and not accurately reflect the scene layout as imaged by the video sensors. For indoor scenes, such reference data may include floorplans that have accurate proportions but not necessarily absolute scale, and potentially snapshots or short video clips of some areas included in the floorplan. Offerors’s may propose solutions that rely on other reference data, provided such data is readily available and/or may be easily obtained for an arbitrary location. In the event that such reference data is made available at test time, it will be provided to all performer teams.
1.A.2.a Phase 1
During an 18-month Phase 1, Video LINCS performers will research and develop the capability and demonstrate the feasibility for reID of people in a video corpus and geo-localization of those people to provide all subject motion in a common reference frame. Monthly submissions will be required to the program leaderboard throughout the program (initially to the self-test leaderboard and subsequently to the sequestered data leaderboard, as shown in Figure 3), with runtimes that are no longer than five times the test video length (and quicker in subsequent phases as shown in Table 2) when run on the T&E hardware for sequestered dataset runs.
TA-1 – ReID
During this phase, each person’s clothing will remain the same (short-term / temporally proximate reID) and provided metadata (e.g., time-stamps and camera pose, to the extent available) will be noise free. ReID will be performed across videos within the same collection type (e.g. drone, CCTV, etc.). Multiple collection types will be included, but reID will only be performed within each collection type and not across the different types during this phase. Accurate time stamps will be provided for all video data. Performer systems will be required to re-identify people across all videos in the test set from the same collection type.
The Phase 1 test set will have approximately 750 distinct people which will be measured for reID performance (see Table 2 for metrics), and there will be an additional 1000 people present in the imagery to serve as potential confusers. People to be reidentified will largely have common external attributes (e.g. clothing) throughout the collection, but some attributes that could change over short periods (e.g. donning or removal of glasses or a cap, lifting or setting down a purse or bag, etc.) may be present.
TA-2 – Object Geo-localization
All people in the test set will need to be mapped to a common geo-coordinate frame and accuracy of geo-localizations will be measured. Full sensor pose will often be available to performer systems during test runs but pose measurements might not be synchronized with the video frames for cameras with variable pose (e.g. panning and moving cameras) and will contain uncertainty.
Geo-localization will be tested both as an independent algorithm as well as for the end-to-end reID system. For independent algorithm testing, camera coordinates of subjects in the test video will be
UNCLASSIFIED
provided at test time and performer systems will remap to geo-coordinates which will be measured for accuracy. For end-to-end reID, geolocation accuracy of subjects identified by the system will be measured. Videos that do not contain full sensor pose will not require person geo-localization in this phase.
1.A.2.b Phase 2
TA-1 – ReID During an 18-month Phase 2, Video LINCS teams will demonstrate the feasibility of reID for vehicles, extend and develop new reID capabilities for people, and initiate reID for generic objects.
For people, reID will extend to temporally distant intervals where temporary attributes (e.g.
clothing) may vary and metadata (e.g. time-stamps and camera pose) may have errors or might not be available. ReID will be performed across videos from different collection types (e.g. drone, CCTV, etc.), and the number of collection types will expand beyond those used in Phase 1. The number of distinct people to be re-identified in the test collection will increase to 1500, and there will be 5000 additional people present in the imagery to serve as potential confusers.
For vehicles, the reID conditions will be similar to those addressed for people in phase 1. During this phase, vehicle reID will be performed over short-term / temporally proximate intervals and provided metadata (e.g. time-stamps and camera pose, to the extent available) will be noise free.
ReID will be performed across videos within the same collection type (e.g. drone, CCTV, etc.).
Multiple collection types will be included, but reID will only be performed within each collection type and not across the different types during this phase. Accurate time stamps will be provided for all video. Performer systems will be required to re-identify vehicles across all videos in the test set from the same collection type. The Phase 2 vehicle test set will have approximately 1000 distinct vehicles which will be measured for reID performance, and there will be additional vehicles present in the imagery to serve as potential confusers.
For generic objects, systems will need to autonomously detect and reID foreground objects throughout the collection, without any prior information regarding the types or characteristics of objects to be re-identified. To demonstrate the concept, consider video at a port of entry that contains ships. The system should autonomously reID ships in such a collection, without any a priori knowledge that ships are present. Similarly, if the system is provided videos of marbles, the system should autonomously reID the marbles throughout the collection. During Phase 2, performers will need to develop and implement a generic object reID capability which will be applied to test data and performance will be measured on a self-test dataset. For Phase 2, generic objects will include all foreground objects that move within at least one video in the collection, in which case they must be re-identified throughout the collection for all appearances, including instances where that object is stationary.
TA-2 – Object Geo-localization
For vehicle reID, all vehicles in the test set will need to be mapped to a common geo-coordinate frame and accuracy of geo-localizations will be measured. Sensor location often will be available to performer systems during test runs, but orientation may not be present, and location measurements might not be synchronized with the video frames for moving cameras and will
UNCLASSIFIED
contain uncertainty. Geo-localization will be tested both as an independent algorithm as well as for the end-to-end reID system. For independent algorithm testing, camera coordinates of vehicles in the test video will be provided at test time and performer systems will remap to geo-coordinates which will be measured for accuracy. For end-to-end reID, geolocation accuracy of vehicles identified by the system will be measured. Videos that do not contain sensor location will not require vehicle geo-localization in this phase.
For person reID, all people in the test set will need to be mapped to a common geo-coordinate frame and accuracy of geo-localizations will be measured. Camera location will often be available to performer systems during test runs but orientation may not be available, and location measurements might not be synchronized with the video frames for moving cameras and will contain uncertainty. Geo-localization will be tested both as an independent algorithm as well as for the end-to-end reID system. For independent algorithm testing, camera coordinates of subjects in the test video will be provided at test time and performer systems will remap to geo-coordinates which will be measured for accuracy. For end-to-end reID, geolocation accuracy of people reidentified by the system will be measured. Videos that do not contain sensor location will not require person geo-localization in this phase.
1.A.2.c Phase 3
TA-1 – ReID During a 12-month Phase 3, Video LINCS teams will demonstrate the feasibility of reID for generic objects and extend and develop new reID capabilities for people and vehicles.
For people and vehicles, reID will be applied to temporally distant intervals where temporary attributes may vary and metadata may have errors and/or might not be available. ReID will be performed across videos from different collection types, and the number of collection types will expand beyond those used in Phase 2. The number of distinct people to be re-identified in the test collection will increase to 2,500, and there will be 10,000 additional people present in the imagery to serve as potential confusers. The Phase 3 vehicle test set will have approximately 2000 distinct vehicles which will be measured for reID performance, and there will be additional vehicles present in the imagery to serve as potential confusers.
For generic objects, systems will need to autonomously detect and reID foreground objects throughout the collection, without any prior information regarding the types or characteristics of objects to be re-identified. Objectives will extend beyond Phase 2 to all foreground objects, even if they are stationary throughout the collection and are never observed to move. Evaluation will progress during this phase to sequestered datasets with independent T&E system runs (the sequestered data leaderboard).
TA-2 – Object Geo-localization
All people and vehicles in the test set will need to be mapped to a common geo-coordinate frame and accuracy of geo-localizations will be measured. Approximate camera location often will be available to performer systems during test runs providing the general vicinity and no orientation information will be provided. Geo-localization will be tested both as an independent algorithm as well as for the end-to-end reID system. For independent algorithm testing, camera coordinates of
UNCLASSIFIED
people, vehicles and objects in the test video will be provided at test time and performer systems will remap to geo-coordinates which will be measured for accuracy. For end-to-end reID, geolocation accuracy of people, vehicles and objects identified by the system will be measured.
Videos that do not contain any location information will not require object geo-localization in this phase.
1.B Team Expertise
Collaborative efforts and teaming across candidate Offerors are highly encouraged. It is anticipated that teams will be multidisciplinary and may include expertise in one or more of the disciplines listed below. This list is included only to provide notional potential expertise areas for the Offerors;
satisfying all the areas of technical expertise below is not a requirement for selection and unconventional or innovative team expertise may be needed based on the proposed research.
Proposals shall include a description and the mix of skills and staffing that the Offeror determines will be necessary to carry out the proposed research and achieve program metrics.
• Artificial Intelligence
• Computer vision, including object detection, tracking, person/vehicle/object modeling, generic vision learning
• Deep learning
• Geometric camera projections and inverse projections
• Image and video geo-localization
• Machine learning
• Modeling and simulation
• Open set classification
• Re-identification
• Soft biometrics
• Software engineering
• Software integration
• Systems integration
• Truthed video data collection and annotation (truthing in both anonymized identities and geo-location) including potential HSR
• Vehicle fingerprinting
• Video data generation (including simulation, generative modeling)
1.C Program Scope and Limitations
Proposals shall explicitly address:
• Underlying theory: Proposed strategies to meet program-specified metrics must have firm theoretical bases that are described with enough detail that reviewers will be able to assess the viability of the approaches. Proposals shall properly describe and reference previous work upon which their approach is founded.
• Research & Development approach: Proposals shall describe the technical approach to meeting program objectives and metrics.
• Technical risks: Proposals shall identify technical risks and proposed mitigation strategies for each identified risk.
UNCLASSIFIED
• Software development: Proposals shall describe the approach to software architecture, modularization, and integration.
The following areas of research are out of scope for the Video LINCS program:
• Research that does not have strong theoretical and experimental foundations
• Approaches that rely solely on integration of existing technologies without experimental or scientific improvements
• Development of hardware or sensors
• Approaches that consist merely of integrating currently existing software systems
• Research of non-visible image/video modalities (e.g. audio, emissive infrared)
• Research on face recognition. Existing face recognition technologies may be employed, but there should be no expectation that recognizable faces will be present in the video.
• Research on license plate recognition. Existing license plate recognition technologies may be employed, but there should be no expectation that recognizable license plates will be present in the video.
• Video from space-based platforms
• Methods that require a human in the loop
Operation on footage from event-based cameras is not an objective of the program, and event-based camera footage is not planned for the test set.
Wide angle video sensor footage is within scope. Footage may contain noise and background variations as is typical for real-world collections such as compression artifacts, dirt or moisture on the camera lens, daytime, twilight, night, rain, snow conditions, etc.
Expansion of video collections, where new videos are incrementally added to a collection after reID has already been performed, must be supported by Video LINCS systems, without requiring re-ingestion of previously processed video.
Performers are permitted to use data from external sources provided that the data is consistent with an approved privacy plan, may be shared across all performer and T&E teams, and is not proprietary. Please include a listing of any datasets that you wish to share within your proposal to the effort. Any data used for training shall be made available to all selected performer and T&E teams on Video LINCS and be provided with Government Purpose Rights. Datasets need not be made public beyond the program teams.
Delivered solutions must include all resources and execute autonomously, not relying on internet access during execution. There may also be the ability to accumulate external data for system reference. If a solution would require external data sources, proposals should describe the necessary data and what would be required to obtain such data for offline use.
1.D Program Data
Video data of a wide range of objects (people, vehicles and generic objects) from many video sensor types and collection conditions with accurate ground truth will be necessary for characterizing performance of Video LINCS developments (T&E). Similar data will likely be used
UNCLASSIFIED
by performers for their R&D as they develop Video LINCS solutions.
The Government will provide a list of nominal video types to be included in test video collections at program kickoff. Nominal video types to be included in evaluations include stationary and mobile sensors that are air, building, and ground based (e.g. drone/UAV, security/CCTV, ground mobile).
Performer teams will be responsible for obtaining (collecting, curating, simulating) and annotating data in support of R&D objectives. All data collected and/or annotated by performer teams under Video LINCS funding, as well as any data used by performer teams for Video LINCS R&D must be shared program wide with other performer and T&E teams. Such data may be used exclusively by the team who collected and curated the data for the initial evaluation reporting period in which it was used, and must be shared program wide after performance results for that period are reported to IARPA management.
All performers will be required to develop and submit a Privacy Protection Plan (PPP) that comprehensively describes the efforts the teams will take to protect personally identifiable information and safeguard the security of any personal data collected or services involved in collection, transmission, processing, and storage of this data. The PPP should cover data that will be collected by the performer as well as datasets obtained from other sources and used on the program (such as self-test datasets provided by T&E), and clearly describe procedures that will be used to safeguard any data resources that require protection throughout the program and the data life cycle. Any claims that data is anonymous must be based on evidence and supported with sufficient information regarding how the data has been anonymized.
Version 1.0 of the Video LINCS PPP shall be included in the Offeror’s proposal as an appendix that covers all datasets to be leveraged as part of the proposed research approaches. This appendix will not count against page count limits. The Video LINCS PPP shall be updated at the beginning of each Phase and when new sources of data or datasets are proposed for use within a Performer’s Video LINCS research activities.
1.D.1 Development Data
Offerors have the freedom to propose any solution to the Video LINCS challenges that will meet program objectives. There is no requirement that such solutions must require training or validation data, but truthed data collection/curation and annotation is highly recommended. For solutions that propose the use of training or validation data, the following is required:
1. A detailed plan for collection/curation/annotation to demonstrate scope and anticipated contribution to the proposed research.
2. All data collections that may potentially contain personally identifiable information must receive Institutional Review Board (IRB) approval.
3. All data collected, curated, annotated and/or simulated by performer teams under Video LINCS funding, as well as any data used by performer teams for Video LINCS R&D must be shared program wide with other performer and T&E teams. Such data may be used exclusively by the team who collected, curated, annotated and/or simulated the data for the initial evaluation period in which it was used, and must be shared program wide after each
UNCLASSIFIED
evaluation waypoint (nominally 6 month waypoints throughout the program - see Figure 3 and Table 5).
4. If a proposed solutions includes data collection, curation, annotation and/or simulation, a minimum of two collection, curation, annotation and/or simulation events must be performed within each phase (aside from Phase 3 which requires at least one such data generation event), with program-wide sharing of at least one of those datasets before the phase is 2/3 complete.
5. Data that is proposed/used must have data use rights that facilitate sharing across the Video LINCS program with the Government and performer teams. A data sharing plan that includes details of data usage rights and any potential data usage restrictions for data that will be used and shared across the program shall be included in the Offeror’s proposal as an appendix.
1.D.2 Evaluation Data
The T&E team will collect, curate and annotate data for evaluation. The T&E team will provide access to self-test datasets for performers to run algorithms on their own hardware throughout the program lifecycle, and performers will submit containerized (nominally Docker) software using/conforming to the framework and APIs (Section 1.G.3.1) specified by the T&E team for independent runs on sequestered evaluation datasets. It is possible that T&E provided self-test data may not be authorized for cloud processing and may require a local processing infrastructure.
1.D.3 External Data Sources
Performers may utilize external datasets for development. External data is data obtained by performers that is available from third parties or that has been collected by the performer outside of the Video LINCS program. An example of a third-party dataset would be a reID dataset collected by a university under an IRB-approved HSR protocol, that has been approved by the dataset owner for release to the research community. Data collected by a performer under a different program is considered external data, even if the other program’s data collection was Government-sponsored.
All external datasets must be approved for use in the Video LINCS program by IARPA, in accordance with HSR and applicable privacy policies, statutes, and regulations. IRB approval or a ruling of IRB HSR exemption may be required before IARPA approval of the use of data in Video
LINCS.
Performers may not use proprietary datasets unless these datasets are made available to all R&D performers and Government T&E in the Video LINCS program without restriction. Public release of proprietary datasets is not a requirement; however, release for use within the Video LINCS program is required. Moreover, for any dataset not collected under the scope of the Video LINCS program, performers must provide the Government with an accounting of all resources used and sources from which data is drawn and describe how the data will be used for development, testing, and training of algorithms.
All external datasets that are part of an offeror’s proposed research approach must be summarized in the proposal with the following minimum information:
• Dataset name
• Short description
• Data owner
• License or use rights
• Method or URL link to obtain data
1.E Testing & Evaluation (T&E)
T&E will be conducted by an independent team of Government and contractor staff carrying out evaluation and analyses of Performer research Deliverables using program test datasets and protocols. The Video LINCS Program will pursue rigorous and comprehensive T&E to ensure that research outcomes are well characterized, deliverables are aligned with program objectives, and that algorithm performance is measured across diverse conditions. T&E activities will inform IARPA and Government stakeholders on Video LINCS research progress and serve as invaluable feedback to the Performers to improve their research approaches, algorithm training practices, and system development. Figure 4 shows the nominal Video LINCS T&E workflow.
T&E Data Collection
Performer Self-Test
T&E Infrastructure
Instrumented Data Collection, Curation
& Annotation
Common Framework
& APIs Sequestered Data Runs
Manual Annotation
Baseline
Visualization Framework
Performer TA-1 Performer TA-2Self-Test
Dataset
Sequestered Dataset
Leaderboards
TA1 Test Harness
TA2 Test Harness
Baseline TA-1 Baseline TA-2
Performer TA-1 Performer TA-2
Performance Analysis
T&EPerformer
Figure 4: Test and Evaluation. The T&E team will collect datasets for performer self-test and T&E sequestered evaluation. Performers will develop systems conforming to the program common framework and APIs and submit containerized software for independent T&E runs on sequestered data. Scores will be posted to program-wide leaderboards.
Leaderboards are planned to facilitate ongoing evaluation and ranking of performer systems against T&E datasets. Both self-test and sequestered data leaderboards will be instantiated by T&E. For the self-test leaderboard, performers will be provided test video datasets to facilitate system runs on performer hardware. System output from performer runs on the self-test dataset will be submitted to a server hosted by the T&E team which will score the submission and performance results will be posted on a system wide leaderboard. Performers will not be permitted to view or probe the self-test data. The self-test leaderboard will facilitate rapid evaluation cycles to assist with development and provide an unofficial ranking across performers. Monthly submissions to the self-test leaderboard will be required in the period before the sequestered data
UNCLASSIFIED
leaderboard is operational (see Figure 3).
For official program evaluations, performer systems will be submitted to run independently on T&E hardware which will contain sequestered video test data. Monthly submissions will be required, with semi-annual evaluation period waypoints where the systems will be officially ranked and reported to IARPA management. T&E results from all Performers will be shared with all teams to establish an understanding of the current state and progress of Video LINCS research.
T&E results will also be shared with USG external stakeholders, including their contractors, for Government purposes. IARPA may conduct other supplemental evaluations or measurements at its sole discretion to evaluate the Performers’ research and Deliverables.
For both leaderboards, the T&E team will post finer grained performance measures beyond the official program metrics to provide insights into where systems require improvement. Such feedback from T&E will assist teams with determining conditions under which the systems require improvement without exposing the test set.
1.F Program Metrics
Achievement of metrics is a performance indicator under IARPA research programs. IARPA has defined Video LINCS program metrics to evaluate effectiveness of the proposed solutions in achieving the stated program goals and objectives, and to determine whether satisfactory progress is being made. The metrics described in this BAA are shared with the intent to scope the effort, while affording maximum flexibility, creativity, and innovation to Offerors proposing solutions to the stated problem.
The final Video LINCS T&E protocols and evaluation methodology will be provided at program kickoff. Program metrics may be refined during the various phases of the Video LINCS program;
if metrics change, revised metrics will be communicated in a timely manner to Performers. The evaluation methodology may be revised by the Government at any time during the program lifecycle to better meet program needs.
For TA-1 - ReID, the goal is to re-identify the same object throughout a video collection. Ideally, all appearances of the same object should be declared to be the same object by Video LINCS systems, without any misses and without matches to different objects (false positives). This is well encapsulated by the F1 measure as described by Ristani et al [2]
F1 = TPSys
(TPGT + (TP + FP)Sys) 2⁄
2 TPSys 2 TPSys + FPSys + FNSys where
TP = True Positive, FP = False Positive, FN = False Negative, GT = Ground Truth, Sys = System which is the number of correct system detections scaled by the average of the number of ground truth and system computed detections. Thus, this metric measures how often the correct object is re-identified, with even penalties for false positives and false negatives.
UNCLASSIFIED
The processing time speedup for reID will also be measured, where the processing time for automated system reID will be computed relative that required for manual reID.
For TA-2 – Geo-localization, the ground-plane accuracy will be measured by computing the position error in meters
||(long, lat, alt)Sys - (long, lat, alt)GT||2 averaged over all object appearances.
Geo-localization will only be evaluated over object appearances in videos which include “video sensor pose” information as described in Table 1. When no sensor location is available, objects in the video will only be evaluated for reID.
ReID and geo-localization metrics will be computed for objects that are visible in a given video frame, potentially including instances with partial occlusion, but will not include instances where an object is not visible in the frame (e.g., if it is fully occluded), even if presence of the object in the scene might be inferred using other evidence.
Runtime for both TA-1 and TA-2 will be measured relative to test-set video length/runtime. The primary limiting factor for runtime requirements will be to ensure sufficient capacity for all performer submissions on the T&E server infrastructure2.
Table 2 provides anticipated primary program performance milestones throughout the program phases.
Table 2: Primary program metrics and performance milestones by Phase.
Objective Conditions Phase 1 Phase 2 Phase 3 TA1: ReID accuracy:
(F1)
Person Temporally proximate
0.6 0.75 0.9
Temporally distanced N/A 0.6 0.75 Vehicle Temporally proximate N/A 0.6 0.75
Temporally distanced N/A N/A 0.5 Generic Object N/A N/A 0.5
TA1: ReID speedup vs. manual annotation TA2: Geo-localization (m) (accuracy in meters)
Full…
This is the start of the file's text. The full file is on GovTribe.
File details come from the government source that posted it. Updated .