Measuring semantic similarity between words or phrases is central to natural language processing, information retrieval and computational linguistics. Despite their importance, similarity and distance measures are typically chosen by default (e.g., cosine similarity) or in an ad hoc fashion, with little empirical justification. This lack of systematic evaluation creates two gaps: first, the absence of a comprehensive taxonomy of measures that spans set-based, vector-based and information-theoretic approaches; second, the lack of task-aware benchmarking that quantifies how these measures perform across different models and applications. In this paper, we address these gaps by comparing 15 similarity and distance measures on four NLP tasks (sentence similarity, kNN classification, correlation analysis and visualisation) using multiple benchmark datasets and embedding models. Our results reveal that the effectiveness of similarity measures varies substantially depending on the task and model, challenging the assumption that cosine similarity is universally optimal. These findings highlight the practical risk of relying on default measures and provide a principled basis for selecting similarity functions in NLP, information retrieval and related fields.
Cambria, E., Nobani, N., Pallucchini, F., Mercorio, F. (2026). Benchmarking Distributional Vector Similarity Measures: A Survey. EXPERT SYSTEMS, 43(8) [10.1111/exsy.70354].
Benchmarking Distributional Vector Similarity Measures: A Survey
Nobani N.;Pallucchini F.;Mercorio F.
2026
Abstract
Measuring semantic similarity between words or phrases is central to natural language processing, information retrieval and computational linguistics. Despite their importance, similarity and distance measures are typically chosen by default (e.g., cosine similarity) or in an ad hoc fashion, with little empirical justification. This lack of systematic evaluation creates two gaps: first, the absence of a comprehensive taxonomy of measures that spans set-based, vector-based and information-theoretic approaches; second, the lack of task-aware benchmarking that quantifies how these measures perform across different models and applications. In this paper, we address these gaps by comparing 15 similarity and distance measures on four NLP tasks (sentence similarity, kNN classification, correlation analysis and visualisation) using multiple benchmark datasets and embedding models. Our results reveal that the effectiveness of similarity measures varies substantially depending on the task and model, challenging the assumption that cosine similarity is universally optimal. These findings highlight the practical risk of relying on default measures and provide a principled basis for selecting similarity functions in NLP, information retrieval and related fields.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


