HRT-23-015_MIMIC.pdf

PDF 2 MB Posted

Attached to
Developing Pedestrian Realistic Artificial Datasets (RADs) Federal contract opportunity
Solicitation number
693JJ325R000014
Issued by
Department of Transportation Federal Highway Administration

About this file

This document is a Federal Highway Administration (FHWA) research report titled "MIMIC—Multidisciplinary Initiative on Methods to Integrate and Create Artificial Realistic Data" published in January 2023. The research project focused on developing synthetic datasets for highway safety research, specifically targeting two types of crash scenarios at diamond interchanges: left-turn crashes at ramp terminals and speed change lane (SCL) crashes. The project developed a three-step framework to generate realistic artificial data (RAD) by identifying contributing factors, establishing cause-effect relationships, and quantifying crash frequencies using data from Washington and Missouri.

The researchers created two primary outputs: tabular datasets for statistical and machine learning modeling, and virtual reality (VR) simulation testbeds using Strategic Highway Research Program Naturalistic Driving Study data. A web-based software was developed to provide access to 196 pregenerated datasets and enable custom data requests. The VR testbeds serve dual purposes of driver education and human factors research, offering interactive crash scenario visualizations that can be used across different visualization platforms. The project demonstrated a generic framework for generating synthetic data that could potentially be applied to other roadway facilities like work zones, alternative intersections, and bicycle facilities.

View the file

Other files for this federal contract opportunity

Other files attached to Developing Pedestrian Realistic Artificial Datasets (RADs), newest first.
File Type Posted
693JJ325R000014-Amendment 0003.pdf PDF
693JJ325R000014-Amendment 0002.pdf PDF
RFP 693JJ325R000014 Questions and Governments Response_Additional Questions.pdf PDF
FHWA-HRT-23-122.pdf PDF
Task 1212_Report_FHWA-HRT-23-121_CLEAN_I.pdf PDF
RFP 693JJ325R000014 Questions and Governments Response.pdf PDF
RADFactSheet.pdf PDF
693JJ325R000014-Amendment 0001.pdf PDF
693JJ325R000014.pdf PDF

On GovTribe

Work with this file on GovTribe

  • Download the original file
  • Contacts named in this file
  • Similar government files
  • Ask GovTribe AI about this file

Text version

Research, Development, and Technology Turner-Fairbank Highway Research Center 6300 Georgetown Pike McLean, VA 22101-2296

PUBLICATION NO. FHWA-HRT-23-015 JANUARY 2023

MIMIC—Multidisciplinary Initiative on Methods to Integrate and Create Artificial Realistic Data

FOREWORD

Data-driven safety analysis models help State and local agencies quantify safety data, identify high-risk roadway features, and predict the effects of proposed safety measures. However, even when a model performs well overall, it may not accurately represent the interactions between variables for a specific location or crash because the underlying relationships in the real world are unknown. One proposed solution is to generate realistic artificial datasets (RADs) with predetermined safety relationships built into them. Because these relationships are known, the RAD can serve as a testbed, revealing how well a model reflects those underlying cause-and-effect relationships.

This study describes the development of RAD for ramp terminals and speed change lanes at diamond interchanges. A web-based software was developed under the Federal Highway Administration’s Exploratory Advanced Research Program. The software provides the ability to generate RAD for multiple years and locations as well as access to pregenerated datasets. This report will be of interest to academics and researchers developing crash modification functions and statistical models to determine how the models best represent real-world relationships.

Brian P. Cronin, P.E.

Director, Office of Safety and Operations.

Research and Development

Notice This document is disseminated under the sponsorship of the U.S. Department of Transportation (USDOT) in the interest of information exchange. The U.S. Government assumes no liability for the use of the information contained in this document.

The U.S. Government does not endorse products or manufacturers. Trademarks or manufacturers’ names appear in this report only because they are considered essential to the objective of the document.

Quality Assurance Statement The Federal Highway Administration (FHWA) provides high-quality information to serve Government, industry, and the public in a manner that promotes public understanding. Standards and policies are used to ensure and maximize the quality, objectivity, utility, and integrity of its information. FHWA periodically reviews quality issues and adjusts its programs and processes to ensure continuous quality improvement.

TECHNICAL REPORT DOCUMENTATION PAGE

1. Report No. 2. Government Accession No. 3. Recipient’s Catalog No.

FHWA-HRT-23-015

4. Title and Subtitle 5. Report Date MIMIC—Multidisciplinary Initiative on Methods to Integrate and Create Artificial Realistic Data.

January 2023

6. Performing Organization Code

7. Author(s) 8. Performing Organization Report No.

Edara, P. (0000-0003-2707-642X), Sun, C. (0000-0002-8857-9648), Brown, H. (0000-0003-1473-901X), Savolainen, P. (0000-0001-5767-9104), Shankar, V. (0000-0002-6671-2268), Balakrishnan, B. (0000-0002-0994-0213), Shang, Y. (0000-0001-7771-4034), Chakraborty, S. (0000-0003-2022-1735), Adu-Gyamfi, Y. (0000-0002-1924-9792), Li, C. (0000-0002-3237-1477), Aati, K. (0000-0001-8834-7735), Lima, S., Huang, Y. (0000-0002-7346-5293), Mussah, A. (0000-0002-1084-5598), Hopfenblatt, J.

9. Performing Organization Name and Address 10. Work Unit No. (TRAIS) University of Missouri-Columbia E2509 Lafferre Hall Columbia, MO 65211

11. Contract or Grant No.

693JJ31950023

12. Sponsoring Organization Name and Address 13. Type of Report and Period Covered United States Department of Transportation Federal Highway Administration HSA Room #E71-324 1200 New Jersey Avenue SE Washington, DC 20590

Final Report; September 2019–October 2022

14. Sponsoring Agency Code

HRSO-2

15. Supplementary Notes This project was supported by the Exploratory Advanced Research (EAR) Program. EAR Program oversight was provided by David Kuehn and Jim Shurbutt. Yusuf Mohamedshah served as Federal Highway Administration (FHWA) Contracting Officer’s Technical Manager. Additional project guidance was provided by FHWA staff members Carol Tan and Ana Maria Eigen. The project’s technical advisory committee consisted of Dean Kanitz (Michigan Department of Transportation (DOT)), Ida van Schalkwyk (Washington State DOT), Ray Shank (Missouri DOT), and John Miller (FHWA Missouri Division).

16. Abstract Traditional safety modeling efforts primarily focus on accurately estimating crash frequencies or rates. The true relationships between crashes and potential causal factors are not always easily discernible from safety models. While a model consisting of multiple causal factors may produce accurate estimates of crash measures, it may not accurately explain all causal relationships.

Knowing the true cause-and-effect relationships is important while choosing countermeasures to address safety problems. This Exploratory Advanced Research Program project developed a framework to generate realistic artificial data (RAD) datasets that mimic the known causal relationships between contributing factors and crashes. The proposed framework is generic and can be used to generate RAD for other facilities, such as work zones, bicycle/pedestrian facilities, innovative geometric designs, etc. The framework was applied to generate RAD for ramp terminals and speed change lane facilities at diamond interchanges. A web-based software was developed to provide easy access to the RAD dataset. The software provides 196 pregenerated datasets and the option to request custom datasets. Sample RAD datasets were used to test negative binomial and a suite of machine learning models. A model evaluation rubric was developed to evaluate and compare the performance of different models.

Additionally, this project developed a second type of RAD dataset—the virtual reality (VR) simulation testbeds for crashes and near-crashes occurring at interchanges. Driving simulator studies offer another source of RAD for evaluating new behavioral and roadway countermeasures. The testbeds were developed using safety critical events recorded in the Strategic Highway Research Program 2 Naturalistic Driving Study data. VR offers an engaging visualization platform to educate the public about interchange crashes and to evaluate different countermeasures. These interventions are well aligned with the USDOT’s National Roadway Safety Strategy’s Safe System Approach of considering an overlapping set of safety measures—roadway countermeasures, behavioral interventions, enforcement, vehicle safety features, and emergency medical care—to achieve zero roadway fatalities.

17. Key Words 18. Distribution Statement Crashes, realistic artificial data, safety, synthetic data, visualization, simulator testbeds

No restrictions. This document is available to the public through the National Technical Information Service, Springfield, VA 22161.

http://www.ntis.gov

19. Security Classification (of this report)

20. Security Classification (of this page)

21. No. of Pages

22. Price

NA

Unclassified. Unclassified.

Form DOT F 1700.7 (8-72) Reproduction of completed page authorized http://www.ntis.gov/ ii

SI* (MODERN METRIC) CONVERSION FACTORS

APPROXIMATE CONVERSIONS TO SI UNITS

Symbol When You Know Multiply By To Find Symbol

LENGTH

in inches 25.4 millimeters mm ft feet 0.305 meters m yd yards 0.914 meters m mi miles 1.61 kilometers km

AREA

in2 square inches 645.2 square millimeters mm2 ft2 square feet 0.093 square meters m2 yd2 square yard 0.836 square meters m2 ac acres 0.405 hectares ha mi2 square miles 2.59 square kilometers km2

VOLUME

fl oz fluid ounces 29.57 milliliters mL gal gallons 3.785 liters L ft3 cubic feet 0.028 cubic meters m3 yd3 cubic yards 0.765 cubic meters m3

NOTE: volumes greater than 1,000 L shall be shown in m3

MASS

oz ounces 28.35 grams g lb pounds 0.454 kilograms kg T short tons (2,000 lb) 0.907 megagrams (or “metric ton”) Mg (or “t”)

TEMPERATURE (exact degrees) °F Fahrenheit 5 (F-32)/9 Celsius °C or (F-32)/1.8

ILLUMINATION

fc foot-candles 10.76 lux lx fl foot-Lamberts 3.426 candela/m2 cd/m2

FORCE and PRESSURE or STRESS lbf poundforce 4.45 newtons N lbf/in2 poundforce per square inch 6.89 kilopascals kPa

APPROXIMATE CONVERSIONS FROM SI UNITS

Symbol When You Know Multiply By To Find Symbol

LENGTH

mm millimeters 0.039 inches in m meters 3.28 feet ft m meters 1.09 yards yd km kilometers 0.621 miles mi

AREA

mm2 square millimeters 0.0016 square inches in2 m2 square meters 10.764 square feet ft2 m2 square meters 1.195 square yards yd2 ha hectares 2.47 acres ac km2 square kilometers 0.386 square miles mi2

VOLUME

mL milliliters 0.034 fluid ounces fl oz L liters 0.264 gallons gal m3 cubic meters 35.314 cubic feet ft3 m3 cubic meters 1.307 cubic yards yd3

MASS

g grams 0.035 ounces oz kg kilograms 2.202 pounds lb Mg (or “t”) megagrams (or “metric ton”) 1.103 short tons (2,000 lb) T

TEMPERATURE (exact degrees) °C Celsius 1.8C+32 Fahrenheit °F

ILLUMINATION

lx lux 0.0929 foot-candles fc cd/m2 candela/m2 0.2919 foot-Lamberts fl

FORCE and PRESSURE or STRESS N newtons 2.225 poundforce lbf kPa kilopascals 0.145 poundforce per square inch lbf/in2 *SI is the symbol for International System of Units. Appropriate rounding should be made to comply with Section 4 of ASTM E380.

(Revised March 2003) iii

TABLE OF CONTENTS

EXECUTIVE SUMMARY

CHAPTER 1. BACKGROUND

Project Motivation Project Overview Study Scope

CHAPTER 2. LITERATURE REVIEW

Synthetic Data

Generation of Synthetic Data Evaluation of Synthetic Data Applications of Synthetic Data

Interchange Safety Modeling State of the Practice Review Crash Modeling Methods and CMFs

CHAPTER 3. METHODOLOGY

Selection of Independent Variables Establishment of Cause–Effect Relationships Generation of Crash Data

CHAPTER 4. LT CRASHES AT INTERCHANGE RAMP TERMINALS

CHAPTER 5. SCL CRASHES AT INTERCHANGES

CHAPTER 6. MODEL EVALUATION RUBRIC

Descriptive Analysis Model Selection Training and Testing Overall Model Performance Model Inference

CHAPTER 7. MODEL TESTING USING RAD DATASETS

Team 1 Models Team 2 Models Team 3 Models Evaluation of Models

CHAPTER 8. SOFTWARE DEVELOPMENT

RAD Software Development Web Development

User Authentication Custom RAD Request Handler Running Time Estimator Multijob Scheduler Email Notification Download Pregenerated RAD VR Animation and Simulator Testbeds iv

CHAPTER 9. SIMULATOR TESTBED DEVELOPMENT

CHAPTER 10. CONCLUSIONS AND CONSIDERATIONS FOR FUTURE RESEARCH

APPENDIX. DESCRIPTIVE STATISTICS OF SAMPLE RAD DATA

ACKNOWLEDGMENTS

REFERENCES

v

LIST OF FIGURES

Figure 1. Screenshot. Landing page with main menu options in the RAD software Figure 2. Graphic. Software development workflow Figure 3. Screenshot. Main landing page of the simulator testbed user interface Figure 4. Screenshot. User menu showing three visualization options Figure 5. Map. Example diamond interchange with two ramp terminals Figure 6. Graphic. Entrance SCLs.(5) Figure 7. Map. Real-world example of an entrance SCL Figure 8. Graphic. Components of SCLs at an interchange.(7) Figure 9. Screenshot. Example of process for spatially locating crashes Figure 10. Graphs. Generated crash data distribution and observed distribution Figure 11. Screenshot. Files in RAD folder Figure 12. Graphic. Software development workflow Figure 13. Screenshot. User authentication page Figure 14. Screenshot. Landing page with main menu options in the RAD software Figure 15. Screenshot. Custom RAD request handler Figure 16. Screenshot. Running time estimator for custom query Figure 17. Screenshot. Window to download pregenerated RAD Figure 18. Screenshot. Web page for VR animation and simulator testbeds Figure 19. Graphic. Four-step process to develop testbeds Figure 20. Graphic. LT crash event Figure 21. Graphic. LT near-crash event Figure 22. Graphic. Near-crash event within entrance SCL.(5) Figure 23. Graphic. Crash event within exit SCL on freeway lane.(5) Figure 24. Graphic. Crash event within exit SCL on deceleration lane.(5) Figure 25. Graphic. Drawing vehicle trajectories in a computer-aided design program Figure 26. Graphic and Photographs. Extracting roadway signs from NDS videos Figure 27. Graphic. Example of a highway and an overpass structure Figure 28. Screenshot. Simulating vehicles in a simulation-optimized runtime build Figure 29. Screenshot. Main landing page of the simulator testbed user interface Figure 30. Screenshot. User menu showing three visualization options Figure 31. Screenshot. Aerial view of an SCL crash Figure 32. Screenshot. SCL crash shown in 360-degree view vi

LIST OF TABLES

Table 1. Criteria and scores for model evaluation Table 2. Variables used in the safety and operational performance analysis of interchange facilities Table 3. Interchange-related studies listed in CMF Clearinghouse Table 4. Contributing factors for LT crashes at entrance ramp Table 5. Contributing factors for SCL crashes Table 6. Variables in the RAD dataset for LT crashes at ramp terminals Table 7. Additional variables specific to individual crashes for LT crashes at ramp terminals Table 8. Variables in the RAD dataset for SCL facilities Table 9. Additional variables specific to individual crashes Table 10. Criteria and scores for model evaluation Table 11. Evaluation of the statistical and machine learning models developed using RAD datasets Table 12. Descriptive statistics of crash frequency (RAD with 400 LT crash sites) Table 13. Descriptive statistics of crossroad AADT (RAD with 400 LT crash sites) Table 14. Descriptive statistics of left-turn AADT (RAD with 400 LT crash sites) Table 15. Descriptive statistics of independent variables (RAD with 400 LT crash sites) Table 16. Descriptive statistics of crash frequency (RAD with 400 SCL sites) Table 17. Descriptive statistics of freeway AADT (RAD with 400 SCL sites) Table 18. Descriptive statistics of ramp AADT (RAD with 400 SCL sites) Table 19. Descriptive statistics of ramp truck AADT (RAD with 400 SCL sites) Table 20. Descriptive statistics of roadway geometric variables (RAD with 400 SCL sites) Table 21. Descriptive statistics of other independent variables (RAD with 400 SCL sites) vii

LIST OF ABBREVIATIONS

3D three dimensional AADT annual average daily traffic AASHTO American Association of State Highway and Transportation Officials ADT average daily traffic AIC Akaike information criterion ARD artificial realistic data BIC Bayesian information criterion CMF crash modification factor CNN convolutional neural network DDSA data-driven safety analysis DOT department of transportation EDC Every Day Counts FHWA Federal Highway Administration EAR Exploratory Advanced Research GAN generative adversarial networks HSIS Highway Safety Information System HSM Highway Safety Manual KNN k-nearest neighbor LT left turn MSE mean squared error NCHRP National Cooperative Highway Research Program NDS Naturalistic Driving Study PDO property damage only RAD realistic artificial data RSE relative squared error SCL speed change lane SHRP Strategic Highway Research Program SPF safety performance function SVM support vector machine USDOT U.S. Department of Transportation VR virtual reality

EXECUTIVE SUMMARY

Data-driven methods are an important component of transportation safety decisionmaking. The U.S. Department of Transportation’s (USDOT) Strategic Plan (2022–2026) stresses the importance of using data-driven methods as part of the overall Safe System Approach toward achieving zero roadway fatalities.(1,2) These methods typically require analytical evaluation of predicted and expected crashes based on geometric and traffic characteristics and other contributing factors. One tool that could facilitate this evaluation is realistic artificial data (RAD).(3,1) RAD can be beneficial to highway safety research in several ways: assessing a new crash estimation method, comparing methods to analyze alternatives, and conducting human factors evaluation of behavioral and roadway countermeasures. Advancing RAD will also enhance the Federal Highway Administration’s (FHWA) efforts to encourage practitioners to apply data-driven methods to safety decisionmaking through programs such as Every Day Counts by expanding the number of tools available for safety analysis.(4)

Although artificial data have been used in many diverse applications, such as security, image processing, surveys, cancer genomics, infrared spectroscopy, and geography, their use in transportation has been limited. The main goal of this Exploratory Advanced Research (EAR) Program project was to develop a framework to generate RAD for interchange facilities and to generate datasets using that framework. Even though interchanges are ubiquitous in our transportation network and carry significant traffic volumes, accurate crash data for such facilities are lacking throughout the United States. The scope of this project involves generating RAD for two types of crashes occurring at diamond interchanges—ramp terminal left-turn (LT) crashes and speed change lane (SCL) crashes.

The data generation framework consists of three main steps. The first step identifies a set of contributing factors at the selected interchange facility (e.g., SCL, ramp terminal). Roadway, traffic, and driver contributing factors were synthesized from the literature from each selected facility. Sampling distributions were generated for each of the factors using observed data from Washington and Missouri. The RAD for these factors were then generated by repeatedly sampling the distributions for a given sample size (e.g., 500 sites). Data from other States were also considered. Highway Safety Information System data for interchange crashes were obtained for a 5-yr period. Data from Washington were the most complete and recent for the purposes of developing RAD, although data from California, Illinois, Maine, and Minnesota were also reviewed. In addition to Washington, interchange safety data from Missouri were also used. The Missouri data were acquired from Missouri DOT’s Transportation Management System as part of a recently completed Highway Safety Manual (HSM) calibration project.(5,6,7)

The second step of the data generation process establishes the effect of each contributing factor on crash frequency. This information was also synthesized from published literature, HSM, and the Crash Modification Factors Clearinghouse.(8) When no reliable information was available for a particular variable, assumptions were made based on analyzing observed data from Washington and Missouri.

The third step of the data generation process quantifies the combined effect of all contributing factors on crash frequency. This quantification was done in two stages. First, the research team estimated the composite crash measure for a given site based on its roadway and traffic characteristics. They considered both individual effects of each factor and interaction effects between two or more factors in generating the composite measure. A site with a higher composite crash measure was likely to experience a higher crash frequency. In the second stage, the researchers converted the composite crash measure to realistic crash frequency (i.e., counts) using observed crash data. This conversion was done using a hierarchical Poisson approach, with parameters optimized for each level of the hierarchy using observed crash data. The research team adjusted the generated crash data distribution parameters to match the overall distributional shape and crash counts at individual sites. Once the crash counts were generated, they used the crash severity distributions to subdivide the overall crash counts into fatal, injury, and property damage only crashes. In addition to crash severity, crash-specific factors pertaining to the driver (e.g., distraction, age, gender), vehicle type, and roadway (e.g., road condition at the time of crash) were also generated for each crash.

The researchers developed a model evaluation rubric to evaluate the performance of models developed using RAD with a scoring system of 0–100. Table 1 shows the six criteria and the maximum points assigned to each model. Because the primary goal of RAD is to evaluate the ability of models to accurately estimate the cause–effect relationships, model inference is weighed more than other criteria.

Table 1. Criteria and scores for model evaluation.

Criteria Points Descriptive analysis of data 10 Model selection 10 Training and testing data 10 Overall prediction accuracy 20 Model inference 50 Total score 100

The RAD datasets generated for LT and SCL facilities were used to test crash prediction models.

Two teams estimated statistical models, while one team developed a series of machine learning models. Statistical models include various forms of negative binomial regression, whereas the machine learning models ranged from a simple ridge regression model to a complex deep learning model TabNet.(9) The model evaluation rubric was applied to the models developed by the three teams. All teams provided basic descriptive statistics of the RAD datasets. Overall scores (out of 100) ranged between 72 and 91, with the main difference in performance appearing in the model inference criteria.

To facilitate the use of RAD, a web-based software was developed to provide access to RAD datasets. Figure 1 is a screenshot of the RAD website’s homepage showing the three types of data that are available for each of the two facility types. Figure 2 shows the workflow of the software. The users submit a RAD data request to the web server. The web server will call the RAD generator to produce a set of RAD datasets. Depending on the type of request, either a pregenerated dataset or a custom dataset will be produced. For custom datasets, an email notification with the download URL for the generated data will be sent to the user-provided email.

Source: FHWA.

SHRP = Strategic Highway Research Program; NDS = Naturalistic Driving Study.

Figure 1. Screenshot. Landing page with main menu options in the RAD software.

Figure 2. Graphic. Software development workflow.

The second type of RAD datasets developed in this project are the virtual reality (VR) simulation testbeds for crashes and near-crashes occurring at interchanges. The testbeds were developed using a four-step process:

1. Analyze Strategic Highway Research Program (SHRP)2 Naturalistic Driving Study (NDS) videos.

2. Develop crash diagrams.

3. Create three-dimensional (3D) modeling of roadway and environment.

4. Reconstruct crash in VR.

Step 1 involves obtaining and analyzing videos of safety-critical events occurring at interchanges. This step was accomplished using SHRP2 NDS data. A total of 114 crash and near-crash events involving left-turning vehicles and 310 events occurring on SCLs were evaluated to develop the testbeds. Both video and kinematic data for safety-critical events were used for reconstruction.

The second step of the crash reconstruction process involves crash diagramming. This task entails drawing detailed trajectories of vehicles involved in the crash event. After drawing the trajectories, road signs are generated. Signs similar to those observed in the crash videos were generated because the actual crash locations were withheld due to privacy concerns.

After extracting the vehicle trajectories and basic signage from the NDS videos, the third step involves creating the roadway and the environment using 3D modeling tools. Coded roadway elements include travel lanes, shoulders, medians, barriers, terrain, overpasses, pavement markings, etc. Environment elements include signage, overall lighting, and foliage next to the highway.The fourth and final step in the testbed development process involves creating a crash simulation. Vehicle information and the trajectories extracted in the first step were overlaid on top of the roadway and environment elements created in the second and third steps to reconstruct the crash. A commonly used simulation engine was used to create the testbeds.

The researchers created a graphical user interface to facilitate the use of simulator testbeds created for LT and SCL crashes. Figure 3 shows a screenshot of the homepage of the user interface. A user has three visualization options (as shown in figure 4):

• An aerial view is a recreated animation of a crash.

• The 360-degree view places the user in the driver’s seat of the subject vehicle and provides the driver’s perspective of the crash.

• The test-drive view is similar to the 360-degree mode, with the exception that the user actively controls the vehicle.

Although aerial and 360-degree views are not interactive, the test drive mode gives control to the user to drive through the scenario and react to the conditions that led to a crash.

Figure 3. Screenshot. Main landing page of the simulator testbed user interface.

Figure 4. Screenshot. User menu showing three visualization options.

In summary, this EAR Program project developed synthetic datasets for interchange facilities for the first time. A three-step framework was developed to generate RAD. The developed datasets were then used by state-of-the-art statistical and machine learning approaches for modeling crash frequency and to ascertain the cause-and-effect relationships. A web-based software was developed to provide easy access to the RAD datasets. The software provides 196 pregenerated datasets and the option to submit custom data requests. The RAD dataset is provided in a spreadsheet format similar to the safety datasets obtained from State DOTs. The proposed framework is generic and can be used to generate RAD for other facilities, such as work zones, bicyclist/pedestrian facilities, innovative geometric designs, etc.

This project also extended the idea of RAD by generating realistic simulation testbeds using NDS data. The VR RAD testbeds were developed with two intended purposes. First, the VR animations of crashes and near-crashes can be used for driver education. Because the testbeds were developed using actual crashes documented in the NDS, they provide a realistic experience that is more engaging. For example, outreach activities targeted at teen drivers can use the animations to help provide a realistic, immersive experience of a crash and to encourage safe driving practices in such circumstances. Lower hardware costs, a younger workforce, and investments from technology companies in improving VR experience are all significant reasons to believe that the transportation industry will increasingly embrace VR-enabled training and education.

A second purpose served by the VR testbeds is to assist with human factors research. For example, a driving simulator platform can readily use the testbeds to test the performance of safety countermeasures, such as in-vehicle driver information systems, roadside dynamic message signs, collision avoidance systems, etc. This purpose is well aligned with the USDOT’s National Roadway Safety Strategy’s Safe System Approach, which considers an overlapping set of safety measures: roadway countermeasures, behavioral interventions, enforcement, vehicle safety features, and emergency medical care.

CHAPTER 1. BACKGROUND

Data are critical to understanding crash causation, which leads to the optimization of safety countermeasures. The Federal Highway Administration (FHWA) has been an active proponent of data-driven methods for safety decisionmaking. Through the Every Day Counts (EDC) program (EDC-3 and EDC-4), FHWA has been encouraging practitioners to apply a data-driven safety analysis (DDSA) approach to safety decisionmaking.(4) The DDSA advocates for a new line of thinking that relies on predicted and expected safety values using statistical methods.

Many States now use DDSA approaches (75 percent per FHWA EDC website) to strategically invest in systemic treatments that target specific crash types rather than chasing high-crash locations and addressing them piecemeal. The U.S. Department of Transportation’s (USDOT) Research, Development, and Technology Strategic Plan (2018–2022) further reinforced the importance of reliable data and effective analytical tools in achieving the strategic safety goal of zero fatalities.(1) Specifically, the safety data initiative of the systemic safety approach “seeks to develop new and integrated data sources, analysis, and visualization techniques to enhance our understanding of crash risk and our ability to mitigate it.”(1)

PROJECT MOTIVATION

As described by Hauer, artificial realistic data (ARD) dataset can be a useful tool for research on highway safety.(3) Specifically, three areas where the use of ARD could be beneficial are:

determining a sample size or assessing a new estimation method, comparing methods to analyze alternatives, and evaluating ways to generate multivariate models. Some of the common challenges in developing safety models include difficulties in identifying causal relationships in the data; the use of average values for variables, variable errors, missing variables, and complex variable dependencies, and the use of simple mathematical functions. ARD offers an innovative way to address these challenges by developing models that not only accurately predict crash frequency but also accurately explain the cause-and-effect relationships between crashes and the independent variables.

Council et al. led the first ARD effort sponsored by FHWA to examine the performance of different modeling methods for cross-sectional studies.1 Using Highway Safety Information System (HSIS) data from Washington, a dataset was created consisting of 2,400 mi of homogenous segments of 0.02 mi each. Crashes were assigned to each of the segments based on certain known causal relationships. The case study examined single-vehicle-lane-departure crashes occurring on rural two-lane roadways. The crash and roadway data were then provided to a modeler not privy to the assumed causal relationships. The modeler was tasked with estimating regression models and deriving the causal relationships. The model results were then checked against the assumed relationships. This effort is one of the few endeavors in transportation to create synthetic data for safety modeling.

1Council, F., E. Hauer, B. Lan, D. Harwood, and R. Srinivasan. 2017. Use of “Artificial Realistic Data” (ARD) to Assess the Performance of Cross-Sectional Analysis Methods in Capturing Causal Relationships Between Individual Roadway Attributes and Safety. Unpublished Report. Washington, DC: Federal Highway Administration.

PROJECT OVERVIEW

In this Exploratory Advanced Research (EAR) Program project, this initial effort by Council et

al. was extended by developing synthetic datasets for interchange facilities.2 Specifically, datasets were generated for crashes occurring at ramp terminals and speed change lanes (SCLs) of diamond interchanges. Diamond interchanges are one of the highly prevalent designs in the United States. These realistic artificial data (RAD) datasets were then used by state-of-the-art statistical and machine learning approaches for modeling crash frequency and to ascertain the cause-and-effect relationships. A web-based software was developed to easily access the RAD datasets. The software provides 196 pregenerated datasets and the option to submit custom data requests. The RAD dataset is provided in a spreadsheet format similar to the safety datasets obtained from State DOTs.

RAD datasets can be used to test the performance of different safety modeling approaches. For example, if a modeler estimated different forms of statistical models using a crash dataset from a particular State DOT, the different models can only be compared using overall goodness-of-fit measures (e.g., prediction accuracy, likelihood value). Since the ground truth cause–effect relationships between independent and dependent variables are seldom known, the models cannot be compared by their ability to extract the true cause–effect relationships. RAD datasets, on the other hand, are created using cause–effect relationships that were established using literature reviews, subject matter expert interviews, and observed safety data from Washington and Missouri. Thus, different models can be compared based on overall goodness of fit as well as on model inference, that is, the model estimated cause–effect relationships versus the ground truth (assumed) relationships. If a model can satisfactorily extract these relationships from RAD data, the user can confidently apply that model to real data (e.g., from a State DOT) and generate reliable crash modification factors (CMFs).

This project also expands the idea of RAD by generating realistic simulation testbeds using Strategic Highway Research Program (SHRP)2 Naturalistic Driving Study (NDS) data. The virtual reality (VR) RAD testbeds were developed with two intended purposes. First, the VR animations of crashes and near-crashes can be used for driver education. Since the testbeds were developed using actual crashes documented in the NDS, they provide an immersive, realistic experience that is engaging. For example, outreach activities targeted at teen drivers can use the animations to help provide a realistic, immersive experience of a crash and to encourage safe driving practices in such circumstances. Lower hardware costs, a younger workforce, and investments from technology companies in improving the VR experience are all significant reasons to believe that the transportation industry will increasingly embrace VR-enabled training and education.

The second purpose of developing the VR testbeds of safety-critical events is to assist with human factors research to improve interchange safety. For example, a driving simulator platform can readily use the testbeds to test the performance of safety countermeasures, such as in-vehicle driver information systems, roadside dynamic message signs, collision avoidance systems, etc.

2Council, F., E. Hauer, B. Lan, D. Harwood, and R. Srinivasan. 2017. Use of “Artificial Realistic Data” (ARD) to Assess the Performance of Cross-Sectional Analysis Methods in Capturing Causal Relationships Between Individual Roadway Attributes and Safety. Unpublished Report. Washington, DC: Federal Highway Administration.

This purpose is well aligned with the USDOT’s National Roadway Safety Strategy’s Safe System Approach, which considers an overlapping set of safety measures: roadway countermeasures, behavioral interventions, enforcement, vehicle safety features, and emergency medical care.

STUDY SCOPE

This project focused on two facilities of a diamond interchange. The first facility is the ramp terminal. Figure 5 shows a diamond interchange with two ramp terminals (one on the south side and one on the north side). Each site refers to one ramp terminal. The crash type of interest is the multivehicle crash occurring between vehicles turning left onto the entrance ramp of the freeway (shown as a left-pointing, curved arrow in the diagram) and the oncoming through vehicle on the crossroad (shown as a downward-pointing, straight arrow). In the RAD dataset, there is no spatial correlation between consecutively numbered sites.

Original photo: Imagery © 2020 Maxar Technologies, map data © 2020 Google®. Modifications by FHWA (see acknowledgments section).

Figure 5. Map. Example diamond interchange with two ramp terminals.

The second interchange facility is the freeway SCL. An SCL facility is an uncontrolled terminal between a ramp and a freeway. The schematic in figure 6 shows an entrance SCL measured from the gore point to the end of the taper.(5) Figure 7 shows a real-world example of an entrance SCL segment.

© 2014 American Association of State Highway and Transportation Officials.

Figure 6. Graphic. Entrance SCLs.(5)

Original photo: Imagery © 2022 Maxar Technologies, Map data © 2022 Google®. Modifications by FHWA (see acknowledgments section).

Figure 7. Map. Real-world example of an entrance SCL.

CHAPTER 2. LITERATURE REVIEW

The research team reviewed literature from two different domains. First, studies documenting the development and use of synthetic data were reviewed. Second, due to the focus of this project on interchange safety, literature pertaining to the understanding of crash causation at interchanges was examined. While the synthetic data review provided information on available data generation methods, the interchange safety review helped to obtain information on the key independent variables, their impact on crash frequency, and the state-of-the-practice crash prediction models.

SYNTHETIC DATA

Although the concept of artificial or synthetic data is new in transportation, its use has been demonstrated in other disciplines. A literature review revealed studies have demonstrated the successful development, evaluation, and application of artificial datasets.

Generation of Synthetic Data

Probabilistic models and deep generative models are two main methods for synthetic data generation. Probabilistic models focus on mimicking the structure and distribution of real data, whereas deep generative models focus on replicating the structure of real data and the information it contains. The Bayesian network and the Markov model are two popular probabilistic models, and deep generative models include variational autoencoder and generative adversarial networks (GAN). In Ping et al., a synthetic data generation tool called DataSynthesizer was proposed using the Bayesian network.(10) Zhang et al. proposed a Bayesian network to generate synthetic high-dimensional data.(11) The Markov chain approach can be applied to generate temporal synthetic data, such as solar states for a smart grid.(12) Islam et al.

presented a data augmentation technique to reproduce crash data.(13) GAN have been used to generate synthetic health data and sensor data.(14,15) A method called TGAN was proposed by Xu and Veeramachaneni to synthesize tabular data using GAN.(16)

Ichim provided an overview of several methods of generating synthetic data, such as probability distribution, Latin hypercube sampling, information preserving statistical obfuscation, data shuffling, and multiple imputations.(17) The author proposed the use of a quantile-based bootstrap strategy and tested it using survey data. Bootstrapping has also been used in other studies by Barth et al., Jia and Culver, and Thanathamathee and Lursinsap (2013).(18–20) Other methods that have been used include Hadoop, convolutional neural network (CNN), genetic algorithm, Bayesian Hierarchical model, and n-spheres.(21–25)

Evaluation of Synthetic Data

The quality of synthetic data is critical to its widespread adoption. One straightforward method to evaluate the quality of synthetic data is to compare the distribution of each variable with the original dataset, but this method does not consider joint distributions of variables. To overcome the limitations, the synthetic and original datasets can be compared by visualizing the joint distribution of high-dimension data with dimension-reduction techniques. The relative performance of two machine learning algorithms on the synthetic dataset and the original dataset can be used to measure the quality of synthetic data. A good synthetic dataset should preserve the same relative performance as the original dataset.(26)

In this project, a rubric was developed to evaluate the synthetic datasets that were developed. A rubric contains three main components: evaluative criteria; quality definition for those criteria at different levels, and a scoring strategy.(27) The goal is to create a rubric grading system to rank different models based on their performance, which will be helpful for modelers to revise and improve their models. Standard rubrics are created based on expert review and general rules of thumb.(28)

Applications of Synthetic Data

Patki et al. developed and used a synthetic data vault to create synthetic data for five datasets.(29) A crowdsourced experiment was then performed in which data scientists were asked to create predictive models with both the synthetic data and real data. The results showed that there was no significant difference in the models developed from the synthetic data and real data. Carlucci et al. created a synthetic database of depth images, and experiments to test the database on two publicly available object datasets showed that the features obtained from processing the synthetic data were stronger.(22) In a research study, Soltana et al. developed and tested an approach to generate synthetic test data using synthetic data for citizens’ records in a public administration system.(30) The case study demonstrated that the results met the criteria for both logical validity and statistical representation.

von Neumann-Cosel et al. demonstrated the successful application of synthetic datasets in a simulated environment through research in which a lane-tracking algorithm was tested.(31) The use of synthetic images was innovative as lane-tracking algorithms are usually evaluated with real camera data and then verified using ground truth data. The study found that specific output parameters of image processing algorithms could be tested using synthetic images. The implementation of the simulation environment to test the lane-tracking algorithm allowed the process for investigating various scenarios to be automated, thus reducing the required extent of testing on actual roads.

Synthetic datasets have also been implemented successfully in various applications in civil engineering. Jia and Culver investigated several synthetic flow generation methods to develop flow records for hydrological calibration to address the challenges created by the limited availability of historical flow data.(19) The methods were tested using a case study at Buck Mountain Run in Albemarle, VA. The results showed that the best combination of methods was the bootstrapped artificial neural network for low- and medium-flow predictions and a modified drainage area ratio for the highest 10 percent of synthetic flows. In another study, Naess and Claussen used synthetic data to assess the performance of different estimators for the prediction of values for long return periods such as wind loads.(32) Sakshaug and Raghunathan used a Bayesian Hierarchical method to generate synthetic datasets for estimating small areas.(24)

INTERCHANGE SAFETY MODELING

State of the Practice Review

Freeway interchanges consist of freeway segments, SCLs, entrance and exit ramp segments, and ramp terminals (intersections). Given the important role freeways play in carrying high traffic volumes, the safety of interchanges is an important concern. American Association of State Highway and Transportation Officials’ (AASHTO) Highway Safety Manual (HSM) was updated in 2014 to provide safety performance functions (SPFs) and CMFs for different freeway facilities, including interchanges.(5) However, there are several important limitations related to the use of SPFs and CMFs.(33) The first practical limitation is the amount of data required for calibration. For example, the predictive method for basic freeway segments in the HSM requires a total of 14 data elements, whereas the method for ramps requires 10 elements. Unfortunately, collecting and analyzing data for several of these elements require a high level of effort, as noted in National Cooperative Highway Research Program (NCHRP) Project 17-45 (e.g., length and radii of horizontal curves, length of/offset to median barriers, clear zone width), which can inhibit the effective utilization of these tools.(34) The authors noted challenges in transferability of models to other contexts, which may be reflective of differences in geometric characteristics in the areas where these studies have been conducted. In addition, a recent meta-analysis of SPFs for freeway merge and diverge areas found that existing research in this area has been largely inconsistent.(35) For example, the effect of deceleration length on safety was reported to be significant in some studies and insignificant in others. Collectively, these results reinforce another limitation noted in the HSM in that the SPFs must be calibrated to reflect local driver populations, conditions, and environments.

Another limitation to the use of SPFs and CMFs is specific to the ramp terminal facility type. A review of the literature for on-ramp terminal crashes shows that most studies are constrained by the lack of crash data specific to this type of freeway geometric element. As such, several studies rely on simulations to address the situations of analyzing scenarios that help improve safety and traffic flow through these sections.(36,37)

Determining whether a crash should be located on a ramp terminal is not as straightforward as it seems. The NCHRP 17-45 project, which influenced the production of the HSM chapter on freeway interchanges, includes extensive discussions of the process and criteria for identifying interchange ramp-terminal crashes.(34) In the analysis of this study to generate RAD, the crashes utilized were collected from Missouri and Washington, which have both taken extensive measures to locate and identify such crashes. Washington, in their reporting of freeway crashes for the HSIS database, included an intersection-related variable that makes the process more manageable. Missouri is also able to provide on-ramp terminal crash data at specifically selected locations due to a previously completed research study.(6,7) This project overcame a tremendous data challenge by using Washington and Missouri ramp terminal crash data.

A few studies have attempted to create localized SPFs for interchange ramp terminals by locating and analyzing crashes within specific study areas for specific States.(38–40) An SPF is a calibrated relationship between collision frequency, traffic volume, and other characteristics of a site.(40) Typically, there are several variables employed in these analyses, which are justified by extensive literature in the transportation safety field. These variables include those components influenced by factors such as traffic volume (exposure), roadway geometry, and traffic signal timing, as described in the literature.

Some of the most common variables utilized in the existing literature (table 2) include annual average daily traffic (AADT) and average daily traffic (ADT) values, clearance intervals as a factor for traffic signal timing, and roadway geometric elements such as roadway surface type, number of lanes, median type and width, and shoulder widths.

Table 2. Variables used in the safety and operational performance analysis of interchange facilities.

Independent variable Studies that feature the selected variable Segment length Parajuli et al.;* Le and Porter; Park, Fitzpatrick, and Lord; Claros, Edara, and Sun (2017).(40–43) Speed limit Wang, Qin, and Noyce;* Chen et al. (2011a); Bonneson and

Zimmerman; Fang, Elefteriadou, and Elias; Claros, Edara, and Sun (2016).(38,44–47)

Number of lanes Elefteriadou et al.; Parajuli et al.;* Le and Porter; Claros, Edara, and Sun (2017); Chen at al. (2011a); Claros, Edara, and Sun (2016); Wang et al.; Chen et al. (2011b)(37,40,41,43,44,47–49)

ADT Elefteriadou et al.; Torbic et al.; Le and Porter; Park, Fitzpatrick, and Lord; Chen et al. (2011b), Liu et al.(37,39,41,42,49,50)

AADT Wang, Qin, and Noyce;* Parajuli et al.;* Claros, Edara, and Sun (2017); Chen et al. (2011a); Claros, Edara, and Sun (2016); Wang et al.(38,40,43,44,47,48)

Surface type Chen et al. (2011b); Liu et al.(49,50)

Median width Park, Fitzpatrick, and Lord; Claros, Edara, and Sun (2017); Wang et al.(42,43,48)

Lane width Fang, Elefteriadou, and Elias(46)

Shoulder width Park, Fitzpatrick, and Lord(42)

Terminal spacing Wang, Qin, and Noyce;* Claros, Edara, and Sun (2016)(38,47)

Signal timing Elefteriadou et al.; Wang, Qin, and Noyce;* Bonneson and Zimmerman; Fang, Elefteriadou, and Elias(37,38,45,46)

Traffic control type Torbic et al.;* Claros, Edara, and Sun (2016 and 2017)(39,43,47)

Interchange configuration

Torbic et al.;* Parajuli et al.;* Fang , Elefteriadou, and Elias;

Claros, Edara, and Sun (2016).(39,40,46,47)

*Studies used in the generation of SPFs for interchange ramp terminals.

In an extensive study on the safety performance of ramp terminals, Parajuli et al. collected data from 380 ramp terminals in Ontario, QC, Canada.(40) In that study, six different ramp-terminal-specific SPFs were developed, accounting for the factors of geometry type, traffic control type, and crash severity level.

Wang, Qin, and Noyce presented interesting insights into the effects of yellow or all-red interval timing on crash safety at interchange ramp terminals, as well as terminal spacing and exclusive right turn phases.(38) Their study suggested that the crash frequency will increase with the deficient yellow or all-red intervals and with the increase in terminal spacing.

Crash Modeling Methods and CMFs

Many statistical, machine learning, and deep learning methods have been applied in the study of crash modeling. For statistical methods, generalized linear model, Poisson, and zero inflated negative binomial are often employed to model crash frequency.(51) The Poisson distribution has the advantage of simulating the unobserved heterogeneity for a smaller dataset, whereas the negative binomial distribution can simulate a dataset with a lot of zeros and a long tail.(52,53) The Poisson-Gamma model can reflect the skewness of the data and be more tunable with a gamma prior.(54) In this EARP project, a modified version of the Poisson-Gamma model is applied to generate the synthetic data.

Iranitalab and Khattak applied various machine learning methods for predicting crash severity, including k-nearest neighbor (KNN), support vector machine (SVM), decision tree, and random forest.(55) For deep learning methods, Huang et al. used CNN for highway crash detection and risk estimation.(56) A long short-term memory–CNN based model was proposed to predict real-time crash risk on arterials.(57)

The CMF Clearinghouse is an online repository of CMFs for various types of facilities. The research team queried CMFs for interchange-related facilities.

This is the start of the file's text. The full file is on GovTribe.

File details come from the government source that posted it. Updated .