1
Return

Does road diversity really matter in testing automated driving systems?

delete2026-07-25
delete0
delete
OA
AI
S
Stefan Klikovits *
V
Vincenzo Riccio
E
Ezequiel Castellano
A
Ahmet Cetinkaya
A
Alessio Gambi
P
Paolo Arcaini
DOI:10.1007/s10664-026-10920-5delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The use of automated driving systems (ADSs) in the real world requires rigorous testing to ensure safety. To increase trust, ADSs should be tested on a large set of diverse road scenarios. Literature suggests that if a vehicle is driven along a set of geometrically diverse roads—measured using various diversity measures (DMs)—it will react in a wide range of behaviours, thereby increasing the chances of observing failures, or strengthening the confidence in its safety, if no failures are observed. However, this assumption has never been tested before, nor have road DMs been assessed for their properties. Our goal was to perform an exploratory study on 53 currently used and new, potentially promising road DMs. Specifically, our research questions looked into the road DMs themselves, to analyse their properties (e.g. monotonicity, computation efficiency), and to test correlation between DMs. Furthermore, we investigated the use of road DMs to determine whether the assumption that diverse test suites of roads expose diverse driving behaviour holds. Our empirical analysis relies on a state-of-the-art, open-source ADS testing infrastructure and uses a data set containing over 97,000 individual road geometries and matching simulation data that were collected using two driving agents. By considering test suites of various sizes and measuring their roads’ geometric diversity, we studied road DM properties, the correlation between road DMs, and the correlation between road DMs and the observed behaviour. Our findings reveal a strong correlation between road diversity and behavioural diversity, confirming that geometrically diverse test suites systematically exercise diverse driving behaviours. We identified Dist. Entropy and Summing aggregations as most effective, with Feature Map achieving the strongest correlation of 0.95 while requiring minimal computation time. The analysed measures maintain robust correlation with behavioural diversity across test suites containing roads of varying lengths, eliminating the need for length normalisation. These results empirically validate the fundamental assumption underlying diversity-driven ADS testing: road geometry diversity serves as a reliable proxy for behavioural diversity. For practitioners, we recommend Feature Map or Dist. Entropy as optimal choices, whilst Averaging-based measures should be avoided entirely. The near-identical correlation patterns observed across architecturally different driving agents indicate that our findings generalise beyond specific ADS implementations, providing a solid foundation for diversity-driven test generation and selection.
Keywords:
Road diversity measure
Autonomous driving systems
Behaviour diversity
Scenario-based testing

Journal

Empirical Software Engineering cover
Empirical Software Engineering
IF:
3.6
Papers:
1.9K
Citations:
5.3K

Organization

J
johannes kepler university linz
Scholars:
701
Papers: 299
Citations: 0
A
Austrian Institute of Technology
Scholars:
36
Papers: 20
Citations: 2.2K
N
national institute of informatics
Scholars:
34
Papers: 33
Citations: 0
S
shibaura institute of technology
Scholars:
234
Papers: 138
Citations: 1
U
university of udine
Scholars:
1.0K
Papers: 454
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers