<?xml version="1.0" encoding="utf-8"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>TRB Publications Index</title><link>http://pubsindex.trb.org/</link><atom:link href="http://pubsindex.trb.org/common/TRIS Suite/feeds/rss.aspx?s=PHNlYXJjaD48cGFyYW1zPjxwYXJhbSBuYW1lPSJzcGVjaWZpY3Rlcm1zIiB2YWx1ZT0iOTI2wqE1MTA1NcKhNTEwODTCoTI5NznCoTExMDUwN8KhOTQ0NzXCoTUwNzDCoTU0MjQiIC8%2BPHBhcmFtIG5hbWU9ImxvY2F0aW9uIiB2YWx1ZT0iMiIgLz48cGFyYW0gbmFtZT0ic3ViamVjdGxvZ2ljIiB2YWx1ZT0ib3IiIC8%2BPHBhcmFtIG5hbWU9InRlcm1zbG9naWMiIHZhbHVlPSJvciIgLz48L3BhcmFtcz48ZmlsdGVycyAvPjxyYW5nZXMgLz48c29ydHM%2BPHNvcnQgZmllbGQ9InB1Ymxpc2hlZCIgb3JkZXI9ImRlc2MiIC8%2BPC9zb3J0cz48cGVyc2lzdHM%2BPHBlcnNpc3QgbmFtZT0icmFuZ2V0eXBlIiB2YWx1ZT0icHVibGlzaGVkZGF0ZSIgLz48L3BlcnNpc3RzPjwvc2VhcmNoPg%3D%3D" rel="self" type="application/rss+xml" /><description></description><language>en-us</language><copyright>Copyright © 2015. National Academy of Sciences. All rights reserved.</copyright><docs>http://blogs.law.harvard.edu/tech/rss</docs><managingEditor>tris-trb@nas.edu (Bill McLeod)</managingEditor><webMaster>tris-trb@nas.edu (Bill McLeod)</webMaster><image><title>TRB Publications Index</title><url>http://pubsindex.trb.org/Images/PageHeader-wTitle.png</url><link>http://pubsindex.trb.org/</link></image><item><title>Exploring the Potential of Remote Sensing and Machine Learning for Scalable Sidewalk Condition Assessment</title><link>http://pubsindex.trb.org/view/2726634</link><description><![CDATA[Sidewalk condition plays a critical role in ensuring pedestrian safety, accessibility, and compliance with regulatory standards. Conventional assessment methods typically involve manual inspections using categorical ratings, which are labor-intensive, subjective, and limited in spatial coverage. This study evaluates the use of satellite imagery and machine learning to support sidewalk condition assessments. A classification model was developed using synthetic aperture radar (SAR) imagery combined with sidewalk physical attributes, including width, slope, and material type. A random forest classifier was trained to predict four condition categories: good, fair, poor, and severe. To address the substantial imbalance in the distribution of classes, a binary formulation was also tested by grouping segments into defective and nondefective classes. Data resampling techniques combining under- and oversampling were applied to improve model performance. The results indicated that the binary model with combined sampling achieved the best performance, with a recall of 0.85 and G-mean of 0.81. Models trained on the original four classes showed lower performance owing to underrepresentation of the poor and severe categories. Feature-importance analysis highlighted SAR amplitude as the most influential predictor across all scenarios. The findings demonstrated the potential of SAR imagery to support scalable and data-driven evaluation of sidewalk conditions. This approach offers a viable complement to traditional inspection methods by enabling targeted resource allocation and broader spatial coverage in pedestrian infrastructure management.]]></description><pubDate>Mon, 13 Jul 2026 17:05:41 GMT</pubDate><guid>http://pubsindex.trb.org/view/2726634</guid></item><item><title>Machine-Learning-Based Traffic State Prediction in Car–Bicycle Mixed Traffic Using Synthetic Data</title><link>http://pubsindex.trb.org/view/2724773</link><description><![CDATA[This study explores the use of machine learning models to predict traffic conditions in mixed car–bicycle traffic environments. A synthetic dataset was developed from numerical evaluations of traffic flow theory, capturing a wide range of multimodal traffic scenarios. Random forest (RF), multi-layer perceptron (MLP), and linear regression models were trained to estimate key traffic metrics, including output flow, delay, and density. The analysis focuses on model performance under different data splits, especially when sorting by variables such as initial car flow and bicycle flow. Results show that, while RF performs well for previously observed traffic conditions, MLP offers stronger generalization to unseen traffic conditions, particularly in high-flow and high-density regimes. However, prediction performance varies depending on the input variable used for sorting and the distribution of training data. These findings underscore the importance of balanced, diverse datasets and support the use of data-driven models for traffic state estimation in multimodal urban networks.]]></description><pubDate>Fri, 10 Jul 2026 12:12:17 GMT</pubDate><guid>http://pubsindex.trb.org/view/2724773</guid></item><item><title>Residential Price Modeling Using Spatially Validated Machine Learning Methods: A Comparison Across Geographical Contexts</title><link>http://pubsindex.trb.org/view/2724620</link><description><![CDATA[Land use (i.e., buildings) and transportation infrastructure are tightly coupled systems, such that residential property prices play a critical role in transportation planning. In a similar manner, transportation infrastructure influences property prices, such that accurate forecasts of both systems are related research problems. The objective of this study is to examine these interactions and evaluate the effectiveness of machine learning methods in modeling residential real estate prices across different urban contexts. Specifically, we first examine the impact of land and transportation infrastructure on residential real estate prices using Extreme Gradient Boosting (XGBoost) and Random Forest (RF) machine learning methods, and compare results between two cities representing diverse geographic and socioeconomic contexts. Second, we investigate the application of machine learning methods on spatial data and provide a comparison of non-spatial and spatial cross-validation on the performance of machine learning methods. We use SHapley Additive exPlanations (SHAP) values to study the impact of land use and transportation infrastructure on real estate prices. The models are applied, and results are compared between the Rawalpindi and Islamabad Metropolitan Area in Pakistan and the City of Toronto in Canada. We find that, despite differences in demographics and economic development, the two cities exhibit similarities in the effect of transportation infrastructure and local amenities on dwelling prices. Proximity to the major central city cores (i.e., downtowns) increases sale price. The effect of transportation infrastructure is differentiated, with high quality transit (e.g., subway and BRT) increasing and conventional bus stop proximity decreasing sale price, respectively. We confirm the previous finding in other fields that non-spatial cross-validation over-estimates the prediction accuracy of machine learning algorithms on spatially referenced datasets. We find that the XGBoost model has slightly higher performance than the RF model. We recommend careful use of machine learning methods in the case of spatial data specifically in modeling of land prices.]]></description><pubDate>Thu, 09 Jul 2026 14:05:01 GMT</pubDate><guid>http://pubsindex.trb.org/view/2724620</guid></item><item><title>Enhancing Travel Mode-Choice Modeling with Route-Based Attributes and Explainable Machine Learning</title><link>http://pubsindex.trb.org/view/2724619</link><description><![CDATA[Modeling mode choice is essential for designing efficient and sustainable mobility systems. Revealed-preference surveys provide valuable information, but they rely on self-reported data, which can be biased and are typically unavailable for unchosen alternatives. This study proposes an integrated analytical process that combines revealed-preference survey data, route-level attributes derived from a digital trip planner, machine-learning classifiers, and explainable artificial intelligence (XAI) methods to evaluate predictive performance and behavioral interpretation jointly. Using a dataset of 1,372 trips collected in a university commuting context as an illustrative application, survey responses were enriched with mode-specific travel times and geometric characteristics of planner-recommended routes obtained from Google Maps’ application programming interface. Four tree-based classifiers were evaluated in a leak-free validation framework, and model behavior was interpreted using SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-agnostic Explanations) at both the global and local levels. The results indicate that when route-level geometric attributes were combined with revealed-preference survey data, predictive accuracy remained comparable, whereas interpretability improved substantially. XAI analyses revealed that route characteristics such as straightness, sinuosity, and angular deviation emerged as significant predictors that modulated perceived travel effort, particularly for walking and public transport, despite their limited impact on aggregate performance.]]></description><pubDate>Thu, 09 Jul 2026 14:05:01 GMT</pubDate><guid>http://pubsindex.trb.org/view/2724619</guid></item><item><title>Using Machine Learning to Understand Electric and Hybrid Vehicles Ownership in Burdened and Nonburdened Communities</title><link>http://pubsindex.trb.org/view/2724616</link><description><![CDATA[Transitioning to electric and hybrid vehicles (EHVs) for all communities is a pivotal step toward sustainable transportation and environmental conservation. This paper aims to understand the adoption of EHVs, focusing on burdened communities (BCs) in the United States. The EHV ownership-based analysis combines two datasets—behavioral data from the Puget Sound Regional Travel Survey integrated with BCs (Justice40) data covering transportation insecurity, environmental burden, social vulnerability, health vulnerability, and climate and disaster risk burden. After creating this unique database, descriptive analysis and modeling are used to analyze the data and predict EHV ownership in the future. Specifically, we use a new method that combines particle swarm optimization (PSO) with a stacking model named PSO-Stacking, which incorporates heterogeneous base learners of machine learning and deep learning. PSO applies a customized objective function to select the optimal hyperparameters for heterogeneous learners within the stacking model, effectively addressing challenges such as multicollinearity, data imbalance, nonlinearity, and overfitting. The proposed solution covers more accurate results than standard benchmark models for EHV ownership in BCs and non-BCs. In addition, the results of the PSO-Stacking method are explained using the local interpretable model-agnostic explanations technique. Results show a negative correlation between the BCs indicators, that is, higher transportation insecurity associated with lower EHV ownership. Furthermore, BCs have higher future climate risk scores, diesel particulate matter levels, and PM₂.₅ in the air than non-BCs because of higher conventional vehicle ownership. These communities are at higher risk and can benefit from electrification, EV infrastructure, and EV policies to address environmental challenges.]]></description><pubDate>Thu, 09 Jul 2026 14:05:01 GMT</pubDate><guid>http://pubsindex.trb.org/view/2724616</guid></item><item><title>Large Language Model-Enhanced General Transportation Agent Framework for Human Mobility Forecasting and Synthetic Travel Survey Data Generation</title><link>http://pubsindex.trb.org/view/2721769</link><description><![CDATA[This work presents the LLM-enhanced general transportation agent, a novel framework that leverages large language models (LLMs) to simulate individual-level human mobility behavior. Unlike prior approaches that treat LLMs as generic predictors or planners, this framework employs role-play prompting and synthetic sociodemographic profiles to position LLMs as simulated individuals responding to household travel surveys. The system integrates population synthesis, persona-rich prompting, structured response tools, a dynamic survey engine, and an LLM-as-a-judge evaluation method to generate and vet context-aware, realistic behavioral data. Case studies in Chicago, USA and Lyon, France demonstrate cross-linguistic adaptability and behavioral fidelity, while highlighting challenges in non-English contexts. Quantitative evaluation shows that midsized and smaller models most accurately reproduce empirical travel survey distributions, whereas larger instruction-tuned models generate more coherent and naturalistic responses. Smaller models exhibit greater susceptibility to logical inconsistencies and stereotyped language, emphasizing potential bias propagation in synthetic datasets. The framework enables scalable, model-agnostic synthetic mobility data generation and supports applications such as scenario prototyping, discretionary activity simulation, travel diary construction, and integration into agent-based and microsimulation transportation models.]]></description><pubDate>Thu, 09 Jul 2026 14:05:01 GMT</pubDate><guid>http://pubsindex.trb.org/view/2721769</guid></item><item><title>KYA: Vision–Language Assistant for Emotional Reactions to Risky Driving</title><link>http://pubsindex.trb.org/view/2720345</link><description><![CDATA[This study introduces a vision–language pipeline that detects risky driving behaviors and generates emotionally expressive responses to support driver awareness and comfort. Although vision–language models have advanced perception and reasoning in autonomous driving, existing systems rarely consider the emotional dimension or real-world user experience. Keep Yelling Assistant (KYA) detects high-risk driving maneuvers in real time, such as sudden cut-ins. It then produces emotional responses through a large language model tailored to driver preferences. The framework comprises two core modules. The vision module uses You Only Look Once (YOLO) v8 variants to detect nearby vehicles and identify risky behaviors such as sudden cut-ins. Key driving metrics, including relative distance, speed, and projected reach time, are extracted and normalized to produce a structured behavior log. The language module processes this log with user-defined emotional tone settings (e.g., neutral, humorous, analytical) and generates verbal reactions using state-of-the-art large language models (LLMs) (ChatGPT-4o, Claude 3, Gemini 2.5, and Copilot). We evaluated the proposed system using dashcam videos containing risky driving behaviors and a user study involving 108 participants. Participants selected preferred response styles, and LLMs were evaluated based on emotional alignment. All models received favorable ratings, though preferences varied across personas. Notably, the combination of YOLOv8s and ChatGPT-4o achieved the highest score, 4.29 out of 5.00. By integrating real-world perception with emotionally adaptive dialogue, KYA advances emotionally intelligent in-vehicle artificial intelligence. It highlights new opportunities to improve safety, trust, and driver comfort in conventional and autonomous vehicles.]]></description><pubDate>Thu, 02 Jul 2026 16:00:31 GMT</pubDate><guid>http://pubsindex.trb.org/view/2720345</guid></item><item><title>Integrating Microscopic Traffic Simulation and Machine Learning for Evaluating Road Diet Impacts on Traffic Performance and Intersection Safety</title><link>http://pubsindex.trb.org/view/2720341</link><description><![CDATA[Road diets represent a cost-effective strategy for enhancing roadway safety; however, comprehensive frameworks integrating operational assessment with surrogate safety evaluation remain limited. This study develops an end-to-end microscopic simulation and machine learning framework to quantify operational and safety impacts of an urban road diet. A 0.22 mi segment of Frazier Avenue in Chattanooga, Tennessee, U.S.—recently converted from a four-lane undivided cross-section to a two-lane facility with a center left-turn lane—is modeled in SUMO using multi-month GridSmart trajectory counts, speed measurements, annual average daily traffic, and OpenStreetMap-derived geometry. A one-way analysis-of-variance-based sensitivity analysis identifies nine influential driver and vehicle parameters, which are calibrated via a genetic algorithm that minimizes speed root mean square error; validation against turning-movement counts using the Geoffrey E. Havers statistic satisfies FHWA acceptance criteria. Three 24 h scenarios are simulated (pre-diet, post-diet, and post-diet geometry with pre-diet demand), complemented by a demand sensitivity analysis up to +30% average daily trips. Corridor-wide, the road diet yields modest mobility penalties—average travel time increases by 4.3%, average speed decreases by 5.2%, and average waiting time increases by 40.7%—with more pronounced degradation westbound and intersection-specific queue growth. Surrogate safety measures at the major signalized intersection indicate substantial safety gains: total vehicle conflicts with time-to-collision (TTC) &lt;3 s decrease by 45.7%, and critical conflicts with TTC &lt;1 s decrease by 62.3%, with conflict hotspots markedly attenuated. Finally, an XGBoost regressor using pre-diet queue length, density, travel time, and speed predicts post-diet waiting times with R2 = 0.905, demonstrating the potential of simulation-driven predictive models to support proactive corridor management and design of future road diet projects.]]></description><pubDate>Thu, 02 Jul 2026 16:00:30 GMT</pubDate><guid>http://pubsindex.trb.org/view/2720341</guid></item><item><title>Enhancing Wind Field Prediction and Reconstruction around Windbreak Walls along High-Speed Railway by Advanced Neural Network Architectures: Accuracy and Stability Assessment</title><link>http://pubsindex.trb.org/view/2717202</link><description><![CDATA[Accurate prediction of wind fields around high-speed railway (HSR) infrastructure is critical for operational safety and energy efficiency. This study evaluates neural network approaches for predicting wind fields around HSR windbreak walls, focusing on transformer models. Field measurements were conducted using 15 anemometer masts arranged inside and outside windbreak walls on the Lanzhou–Xinjiang railway. We compared multiple deep learning architectures (multilayer perceptron, long-short-term memory, temporal convolutional network and transformer) for predicting interior wind conditions based on exterior measurements. The key findings are summarized as follows: (1) prediction accuracy improved substantially with longer historical contexts (10–60 timesteps); (2) significant spatial variability exists in wind predictability across measurement locations; (3) feature importance analysis identified critical measurement points, enabling cost-effective maintenance strategies and optimized sensor deployment; and (4) sequence mean filling performed best among the three tested strategies for handling missing data, maintaining positive predictive power even with substantial sensor loss. Among these models, the transformer model achieved the best overall performance (𝘙² = 0.9665 at 𝘛 = 60), with its advantage becoming most pronounced at longer historical contexts. These findings have important implications for railway safety and wind energy applications, enabling more efficient monitoring networks and robust forecasting systems. The demonstrated effectiveness of transformer models represent a significant advancement in applying attention-based architectures to infrastructure monitoring challenges.]]></description><pubDate>Wed, 24 Jun 2026 10:29:07 GMT</pubDate><guid>http://pubsindex.trb.org/view/2717202</guid></item><item><title>Large Language Model-Based Realistic Safety-Critical Driving Video Generation</title><link>http://pubsindex.trb.org/view/2712066</link><description><![CDATA[Designing diverse and safety-critical driving scenarios is essential for evaluating autonomous driving systems. In this paper, we propose a novel framework that leverages large language models (LLMs) for few-shot code generation to automatically synthesize driving scenarios within the CARLA simulator, which has flexibility in scenario scripting, efficient code-based control of traffic participants, and enforcement of realistic physical dynamics. Given a few example prompts and code samples, LLM generates safety-critical scenario scripts that specify the behavior and placement of traffic participants, with a particular focus on collision events. To bridge the gap between simulation and real-world appearance, we integrate a video generation pipeline using Cosmos-Transfer1, which converts rendered scenes into realistic driving videos. Our approach enables controllable scenario generation and facilitates the creation of rare but critical edge cases, such as pedestrian crossings under occlusion or sudden vehicle cut-ins. Comprehensive quantitative evaluations across multiple environments demonstrate a favorable balance between visual realism, perceptual quality, and structural consistency in the generated videos. Experiment results demonstrate the effectiveness of our method in generating a wide range of realistic, diverse, and safety-critical scenarios, offering a promising tool for simulation-based testing of autonomous vehicles.]]></description><pubDate>Wed, 10 Jun 2026 09:06:00 GMT</pubDate><guid>http://pubsindex.trb.org/view/2712066</guid></item><item><title>A Multimode Cooperative Control Architecture for Connected and Automated Vehicle Platoon Splitting and Merging in Mixed Traffic</title><link>http://pubsindex.trb.org/view/2712065</link><description><![CDATA[Connected and automated vehicle (CAV) platoons provide significant advantages in enhancing traffic efficiency and safety through vehicle-to-vehicle cooperative driving. However, owing to the uncertainty of human-driven vehicles in mixed traffic environments, platoons must frequently split to avoid potential collisions and merging is required to maintain platoon following. To address this challenge, this paper proposes a cooperative control architecture for CAV platoons that includes a single-vehicle cruising control mode and a platoon-following control mode, enabling independent operation of each mode and discrete event transitions around split and merge maneuvers. In single-vehicle mode, a driving safety potential field model is proposed for collision-avoidance trajectory planning, and a distributed model predictive control algorithm is designed to achieve the distinct control objectives of the two modes. Then, a long short-term memory (LSTM) neural network and fuzzy logic are combined to predict collision risk and determine platoon split events. A cooperative control system is implemented to ensure continuous control and flexible switching between the two modes. Finally, joint simulations in PreScan, CarSim, and MATLAB/Simulink were conducted to evaluate the performance of the system across various obstacle scenarios. The results demonstrate that the proposed control architecture effectively coordinates vehicle maneuvers and adapts platoon formation to changes in traffic conditions.]]></description><pubDate>Tue, 09 Jun 2026 14:35:55 GMT</pubDate><guid>http://pubsindex.trb.org/view/2712065</guid></item><item><title>Enhancing Rail Obstacle Detection Systems: Optimizing Accuracy in Adverse Weather Conditions</title><link>http://pubsindex.trb.org/view/2711639</link><description><![CDATA[Railway safety is paramount, especially with the increasing reliance on rail transport and the potential for catastrophic consequences from train colliding with obstacles. This paper introduces a novel obstacle detection methodology using Convolutional Neural Networks (CNNs) to enhance detection accuracy, particularly for diverse and unforeseen obstacles, including wildlife intrusion, under challenging environmental conditions. We employ the state-of-the-art (You Only Look Once) YOLOv11-Seg algorithm for simultaneous rail segmentation and obstacle detection, defining a critical safety margin around the tracks. A key contribution of this work is a novel synthetic image generation algorithm designed to address the critical scarcity of real-world obstacle data, particularly for rare and unpredictable hazards such as animals and uncharacterized debris. This algorithm strategically places various obstacles, extracted from diverse sources, at random locations on the rail or within the safety margin. Crucially, it incorporates diverse and realistic environmental conditions, such as train vibrations, rain, snow, dust, fog, and varying light intensities to augment the training data and improve the model’s robustness against these highly transient events. Experimental results demonstrate the effectiveness of the YOLOv11-Seg network, trained on our synthetically augmented data set, in accurately performing both segmentation and obstacle detection in a single step, paving the way for improved railway safety systems.]]></description><pubDate>Fri, 05 Jun 2026 11:27:53 GMT</pubDate><guid>http://pubsindex.trb.org/view/2711639</guid></item><item><title>Traffic Breakdown Prediction Beyond Stochastic Capacity Models: Machine Learning Approach</title><link>http://pubsindex.trb.org/view/2709302</link><description><![CDATA[This study focuses on the application of machine learning models for a short-term prediction of traffic breakdowns on freeways. Traffic breakdowns, which occur when demand exceeds the momentary capacity, are typically predicted using probabilistic methods, but these approaches do not fully capture the short-term variability inherent in traffic flow. In this work, the methodology is advanced by employing machine learning techniques to predict traffic flow conditions, relying exclusively on lane-by-lane analysis of current detector data without utilizing any upstream or downstream information. Traffic conditions are classified into distinct categories, including breakdowns, and a neural network is employed to predict them, providing a robust method for identifying intervals in which the momentary capacity of a freeway is reached. Capacity estimates from the neural network are then compared with those from widely accepted statistical methods, revealing minimal differences, and thereby validating the effectiveness of the neural network approach in capacity analysis. Moreover, comparing the short-term flow conditions predicted based on the two approaches revealed superiority of neural network in providing significantly more accurate classifications. These findings highlight the significant potential of machine learning methods as powerful tools for momentary capacity estimation, with applications across various transportation systems management and operations strategies.]]></description><pubDate>Wed, 03 Jun 2026 09:07:22 GMT</pubDate><guid>http://pubsindex.trb.org/view/2709302</guid></item><item><title>A Data-Driven Simulation and Machine Learning Framework for Shopping Trip Forecasting with Spatial Clustering</title><link>http://pubsindex.trb.org/view/2709131</link><description><![CDATA[Retailing plays a pivotal role in the functioning of urban systems. While upstream supply chain activities such as manufacturing and distribution primarily affect freight movement, the retail interface translates consumer demand into individual travel behavior, shaping local traffic conditions and feeding back into upstream logistics. Despite its importance, shopping-related travel remains under-modeled in urban mobility research. To address this gap, this study develops a purpose-specific travel forecasting and simulation framework for predicting shopping trip demand in urban areas. The forecasting model integrates commercial-environment attributes, trip characteristics, and sociodemographic factors. A suite of machine learning (ML) models is evaluated, and the best-performing model is selected for the proposed simulation. Microlevel predictions are then scaled to the full urban region, followed by zonal aggregation and k-means spatial clustering to identify distinct retail-demand patterns and support scenario testing. Numerical results show that the random forest model outperforms alternative ML classifiers and, when implemented in the simulation, generates a citywide estimate indicating that shopping trips represent 14.3% of all weekday travel, in line with external regional benchmarks. The combined ML–simulation framework demonstrates strong predictive performance and reveals meaningful spatial and behavioral insights relevant to policymaking and planning applications. Although applied to Halifax, the modular structure of the framework makes it transferable to other urban regions and adaptable to additional trip purposes, supporting future extensions involving multiactivity modeling, causal impact analysis, and integration with passive mobility datasets.]]></description><pubDate>Tue, 02 Jun 2026 11:01:49 GMT</pubDate><guid>http://pubsindex.trb.org/view/2709131</guid></item><item><title>Toward Asphalt Pavement Construction Safety Improvement with Generative Artificial Intelligence</title><link>http://pubsindex.trb.org/view/2706339</link><description><![CDATA[Recent advancements in generative artificial intelligence (GenAI), particularly large language models (LLMs), have shown promise in enhancing safety analysis within the construction industry. This study explores the integration of structured information from the U.S. Pennsylvania Department of Transportation’s job safety analysis (JSA) documents with unstructured accident narratives from the U.S. Occupational Safety and Health Administration’s Integrated Management Information System (IMIS), focusing on asphalt pavement construction—a sector marked by complex operations and hazardous equipment. Multiple LLMs were employed to classify accident narratives across four dimensions: construction type, relevant JSA job and step, environmental or operational influence, and hazard type. This approach aims to identify frequently cited work activities, assess gaps in safety documentation, and improve future hazard recognition. While general job classification achieved moderate success, performance declined for step-level and contextual classifications, largely because of ambiguous language and overlapping job responsibilities in the narratives. Despite these limitations, LLMs uncovered critical patterns. Paver and roller operations emerged as high-risk activities, often influenced by environmental factors such as traffic or weather. Furthermore, exploratory hazard analysis revealed that model-suggested hazard labels were sometimes more contextually appropriate than those in the original IMIS database, indicating opportunities for data refinement. By aligning structured safety documentation with real-world incidents through GenAI, this study highlights a novel pathway for data-driven safety planning in highway construction. While expert oversight remains essential, the results demonstrate the potential of LLMs to support more adaptive, targeted, and proactive approaches to risk assessment and mitigation.]]></description><pubDate>Thu, 28 May 2026 10:47:37 GMT</pubDate><guid>http://pubsindex.trb.org/view/2706339</guid></item></channel></rss>