Automatic systems with word sense induction for lexical semantic change detection
Authors
-
Denis V. Kokosinskii
-
Nikolay V. Arefyev
Keywords:
lexical semantic change detection
mathematical modeling in linguistics
natural language processing
clustering
Abstract
Lexical Semantic Change Detection (LSCD) is the task of identifying words that have changed their meaning over time. The semantic change can be formalized using different mathematical models. The Jensen–Shannon Distance model quantifies semantic change at the level of word senses, while the COMPARE model works directly at the level of individual word usages. Automatic LSCD systems reflect this split: some of them explicitly induce word senses by means of clustering, while others operate at the level of individual word usages without resorting to clustering. We argue that the automatic LSCD systems with word sense induction are preferable due to their interpretability, since they answer not only the question of the degree to which a word’s meaning has changed, but also the question of which word senses caused the lexical semantic change. The main counter-argument against such systems was their inferior quality. Recent studies found that the usage-level systems outperform the systems with word sense induction, even on benchmarks that themselves rely on induced senses. In this paper, we propose an automatic LSCD system with word sense induction that surpasses the existing usage-level systems in quality. We systematically evaluate each component of the systems on benchmarks for eight languages. It is shown that specialized compact models (∼350 million parameters) achieve the quality on par with human annotators in the semantic similarity estimation problem (the Word-in-Context subtask). A new metric for evaluating clustering quality is proposed, which makes it possible to find a more optimal system configuration and thereby eliminate the quality gap. As a result, we build a sense-level LSCD system that combines interpretability, quality, and computational efficiency. In this paper, we perform a systematic evaluation of each component of automatic LSCD systems on benchmarks across eight languages. In our evaluation, we find that specialized compact models (neural networks with 350M parameters) achieve human-level quality in pairwise similarity estimation (the task known as Word-in-Context, or WiC). We also separately evaluate the clustering step in WSI-based systems across multiple clustering algorithms and configurations using a custom metric tailored to the specifics of the existing datasets. By carefully tuning the individual components and analyzing computational efficiency trade-offs, we build a WSI-based LSCD system that combines interpretability, quality, and computational efficiency.
Section
Methods and algorithms of computational mathematics and their applications
References
- A. Blank, Prinzipien des Lexikalischen Bedeutungswandels am Beispiel der Romanischen Sprachen(Niemeyer, Tübingen, 1997).
doi 10.1515/9783110931600
- F. Periti and N. Tahmasebi, “A Systematic Comparison of Contextualized Word Embeddings for Lexical Semantic Change,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Mexico City, Mexico, June 16–21, 2024, Vol. 1: Long Papers(Association for Computational Linguistics, 2024), pp. 4262–4282.
doi 10.18653/v1/2024.naacl-long.240
- F. Periti and S. Montanelli, “Lexical Semantic Change through Large Language Models: a Survey,” ACM Computing Surveys 56 (11), 1–38 (2024).
doi 10.1145/3672393
- A. Kutuzov, L. Øvrelid, T. Szymanski, and E. Velldal, “Diachronic word embeddings and semantic shifts: a survey,” in Proceedings of the 27th International Conference on Computational Linguistics, Santa Fe, New Mexico, USA, August 20–26, 2018(Association for Computational Linguistics, 2018), pp. 1384–1397.
https://aclanthology.org/C18-1117/ Cited July 17, 2026.
- D. Schlechtweg, B. McGillivray, S. Hengchen, et al., “SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection,” in Proceedings of the Fourteenth Workshop on Semantic Evaluation, Barcelona, Online, December 12–13, 2020(International Committee for Computational Linguistics, 2020), pp. 1–23.
doi 10.18653/v1/2020.semeval-1.1
- D. Schlechtweg, S. Schulte im Walde, and S. Eckmann, “Diachronic Usage Relatedness (DURel): A Framework for the Annotation of Lexical Semantic Change,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, New Orleans, Louisiana, June 1–6, 2018, Vol. 2: Short Papers(Association for Computational Linguistics, 2018), pp. 169–174.
doi 10.18653/v1/N18-2027
- M. T. Pilehvar and J. Camacho-Collados, “WiC: the Word-in-Context Dataset for Evaluating Context-Sensitive Meaning Representations,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, Minnesota, June 2–7, Vol. 1: Long and Short Papers(Association for Computational Linguistics, 2019), pp. 1267–1273.
doi 10.18653/v1/N19-1128
- J. H. Lau, P. Cook, D. McCarthy, et al., “Word Sense Induction for Novel Sense Detection,” in Proceedings of the 13th Conference of the European Chapter of the Association for Computational Linguistics, Avignon, France, April 23–27, 2012(Association for Computational Linguistics, 2012), pp. 591–601.
https://aclanthology.org/E12-1060/ Cited July 17, 2026.
- M. Martinc, S. Montariol, E. Zosa, and L. Pivovarova, “Capturing Evolution in Word Usage: Just Add More Clusters?’’ in WWW’20: Companion Proceedings of the Web Conference, Taipei, Taiwan, April 20–24, 2020(ACM, New York, 2020), pp. 343–349.
doi 10.1145/3366424.3382186
- F. Periti, A. Ferrara, S. Montanelli, and M. Ruskov, “What is Done is Done: an Incremental Approach to Semantic Shift Detection,” in Proceedings of the 3rd Workshop on Computational Approaches to Historical Language Change, Dublin, Ireland, May 26–27, 2022(Association for Computational Linguistics, 2022), pp. 33–43.
doi 10.18653/v1/2022.lchange-1.4
- A. Kutuzov and M. Giulianelli, “UiO-UvA at SemEval-2020 Task 1: Contextualised Embeddings for Lexical Semantic Change Detection,” in Proceedings of the Fourteenth Workshop on Semantic Evaluation, Barcelona, Online, December 12–13, 2020(International Committee for Computational Linguistics, 2020), pp. 126–134.
doi 10.18653/v1/2020.semeval-1.14
- M. Giulianelli, M. Del Tredici, and R. Fernández, “Analysing Lexical Semantic Change with Contextualised Word Representations,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online, July 5–10, 2020(Association for Computational Linguistics, 2020), pp. 3960–3973.
doi 10.18653/v1/2020.acl-main.365
- F. D. Zamora-Reina, F. Bravo-Marquez, and D. Schlechtweg, “LSCDiscovery: A shared task on semantic change discovery and detection in Spanish,” in Proceedings of the 3rd Workshop on Computational Approaches to Historical Language Change, Dublin, Ireland, May 26–27, 2022(Association for Computational Linguistics, 2022), pp. 149–164.
doi 10.18653/v1/2022.lchange-1.16
- A. Kutuzov, S. Touileb, P. Mæhlum, et al., “NorDiaChange: Diachronic Semantic Change Dataset for Norwegian,” in Proceedings of the Thirteenth Language Resources and Evaluation Conference, Marseille, France, June 20–25, 2022(European Language Resources Association, 2022), pp. 2563–2572.
https://aclanthology.org/2022.lrec-1.274/ Cited July 18, 2026.
- J. Chen, E. Chersoni, D. Schlechtweg, et al., “ChiWUG: A Graph-based Evaluation Dataset for Chinese Lexical Semantic Change Detection,” in Proceedings of the 4th Workshop on Computational Approaches to Historical Language Change, Singapore, December 6, 2023(Association for Computational Linguistics, 2023), pp. 93–99.
doi 10.18653/v1/2023.lchange-1.10
- A. Kutuzov and L. Pivovarova, “Three-part diachronic semantic change dataset for Russian,” in Proceedings of the 2nd International Workshop on Computational Approaches to Historical Language Change, Online, August 6, 2021(Association for Computational Linguistics, 2021), pp. 7–13.
doi 10.18653/v1/2021.lchange-1.2
- D. Schlechtweg, N. Tahmasebi, S. Hengchen, et al., “DWUG: A large Resource of Diachronic Word Usage Graphs in Four Languages,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Punta Cana, Dominican Republic, Online, November 7–11, 2021(Association for Computational Linguistics, 2021), pp. 7079–7091.
doi 10.18653/v1/2021.emnlp-main.567
- N. Bansal, A. Blum, and S. Chawla, “Correlation Clustering,” Machine Learning 56 (1), 89–113 (2004).
doi 10.1023/B: MACH.0000033116.57574.95.
- A. Hätty, D. Schlechtweg, and S. Schulte im Walde, “SURel: A Gold Standard for Incorporating Meaning Shifts into Term Extraction,” in Proceedings of the 8th Joint Conference on Lexical and Computational Semantics, June 2019, Minneapolis, Minnesota, USA(Association for Computational Linguistics, 2019), pp. 1–8.
doi 10.18653/v1/S19-1001
- S. Kurtyigit, M. Park, D. Schlechtweg, J. Kuhn, and S. Schulte im Walde, “Lexical Semantic Change Discovery,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)(Association for Computational Linguistics, Online, 2021), pp. 6985–6998.
doi 10.18653/v1/2021.acl-long.543
- A. Aksenova, E. Gavrishina, E. Rykov, and A. Kutuzov, “RuDSI: Graph-based Word Sense Induction Dataset for Russian,” in Proceedings of TextGraphs-16: Graph-based Methods for Natural Language Processing, October 2022, Gyeongju, Republic of Korea(Association for Computational Linguistics, Gyeongju, Republic of Korea, 2022), pp. 77–88.
https://aclanthology.org/2022.textgraphs-1.9 Cited July 29, 2026.
- M. Martinc, P. K. Novak, and S. Pollak, “Leveraging Contextual Embeddings for Detecting Diachronic Semantic Shift,” in Proceedings of the Twelfth Language Resources and Evaluation Conference, Marseille, France, May 11–16, 2020(European Language Resources Association, 2020), pp. 4811–4819.
https://aclanthology.org/2020.lrec-1.592/ Cited July 17, 2026.
- C. Beck, “DiaSense at SemEval-2020 Task 1: Modeling Sense Change via Pre-trained BERT embeddings,” in Proceedings of the Fourteenth International Workshop on Semantic Evaluation, Barcelona, Online, December 12–13, 2020(Association for Computational Linguistics, 2020), pp. 50–58.
doi 10.18653/v1/2020.semeval-1.4
- J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, Minnesota, June 2–7, 2019, Vol. 1: Long and Short Papers(Association for Computational Linguistics, 2019), pp. 4171–4186.
doi 10.18653/v1/N19-1423
- S. Laicher, S. Kurtyigit, D. Schlechtweg, et al., “Explaining and Improving BERT Performance on Lexical Semantic Change Detection,” in Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop, Online, April 19–23, 2021(Association for Computational Linguistics, 2021), pp. 192–202.
doi 10.18653/v1/2021.eacl-srw.25
- A. Conneau, K. Khandelwal, N. Goyal, et al., “Unsupervised Cross-lingual Representation Learning at Scale,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online, July 5–10, 2020(Association for Computational Linguistics, 2020), pp. 8440–8451.
doi 10.18653/v1/2020.acl-main.747
- M. Giulianelli, A. Kutuzov, and L. Pivovarova, “Do Not Fire the Linguist: Grammatical Profiles Help Language Models Detect Semantic Change,” in Proceedings of the 3rd Workshop on Computational Approaches to Historical Language Change, Dublin, Ireland, May 26–27, 2022(Association for Computational Linguistics, 2022), pp. 54–67.
doi 10.18653/v1/2022.lchange-1.6
- M. Rachinskiy and N. Arefyev, “GlossReader at SemEval-2021 Task 2: Reading Definitions Improves Contextualized Word Embeddings,” in Proceedings of the 15th International Workshop on Semantic Evaluation (SemEval-2021), Online, August 5–6, 2021(Association for Computational Linguistics, 2021), pp. 756–762.
doi 10.18653/v1/2021.semeval-1.100
- N. Arefyev, M. Fedoseev, V. Protasov, et al., “DeepMistake: Which Senses are Hard to Distinguish for a Word-in-Context Model,” in Komp’juternaja Lingvistika i Intellektual’nye Tehnologii(RGGU, Moscow, 2021), pp. 16–30.
https://dialogue-conf.org/media/5491/arefyevnplusetal133.pdf Cited July 29, 2026.
- P. Cassotti, L. Siciliani, M. DeGemmis, et al., “XL-LEXEME: WiC Pretrained Model for Cross-Lingual LEXical sEMantic changE,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, Toronto, Canada, July 9–14, 2023, Vol. 2: Short Papers(Association for Computational Linguistics, 2023), pp. 1577–1585.
doi 10.18653/v1/2023.acl-short.135
- M. Fedorova, T. Mickus, N. Partanen, et al., “AXOLOTL’24 Shared Task on Multilingual Explainable Semantic Change Modeling,” in Proceedings of the 5th Workshop on Computational Approaches to Historical Language Change, Bangkok, Thailand, August 15, 2024(Association for Computational Linguistics, 2024), pp. 72–91.
doi 10.18653/v1/2024.lchange-1.8
- D. Kokosinskii, M. Kuklin, and N. Arefyev, “Deep-change at AXOLOTL-24: Orchestrating WSD and WSI Models for Semantic Change Modeling,” in Proceedings of the 5th Workshop on Computational Approaches to Historical Language Change, Bangkok, Thailand, August 15, 2024(Association for Computational Linguistics, 2024), pp. 168–179.
doi 10.18653/v1/2024.lchange-1.16
- D. Schlechtweg, T. Choppa, W. Zhao, and M. Roth, “CoMeDi Shared Task: Median Judgment Classification & Mean Disagreement Ranking with Ordinal Word-in-Context Judgments,” in Proceedings of Context and Meaning: Navigating Disagreements in NLP Annotation, Abu Dhabi, UAE, January 19, 2025(International Committee on Computational Linguistics, 2025), pp. 33–47.
https://aclanthology.org/2025.comedi-1.4/ Cited July 19, 2026.
- A. Raganato, T. Pasini, J. Camacho-Collados, and M. T. Pilehvar, “XL-WiC: A Multilingual Benchmark for Evaluating Semantic Contextualization,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online, November 16–20, 2020(Association for Computational Linguistics, 2020), pp. 7193–7206.
doi 10.18653/v1/2020.emnlp-main.584
- F. Martelli, N. Kalach, G. Tola, and R. Navigli, “SemEval-2021 Task 2: Multilingual and Cross-lingual Word-in-Context Disambiguation (MCL-WiC),” in Proceedings of the 15th International Workshop on Semantic Evaluation (SemEval-2021), Online, August 5–6, 2021(Association for Computational Linguistics, 2021), pp. 24–36.
doi 10.18653/v1/2021.semeval-1.3
- Q. Liu, E. M. Ponti, D. McCarthy, et al., “AM2iCo: Evaluating Word Meaning in Context across Low-Resource Languages with Adversarial Examples,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Punta Cana, Dominican Republic, Online, November 7–11, 2021(Association for Computational Linguistics, 2021), pp. 7151–7162.
doi 10.18653/v1/2021.emnlp-main.571
- G. A. Miller, C. Leacock, R. Tengi, and R. T. Bunker, “A Semantic Concordance,” in Human Language Technology: Proceedings of a Workshop Held at Plainsboro, New Jersey, March 21–24, 1993
https://aclanthology.org/H93-1061/ Cited July 19, 2026.
- D. Homskiy and N. Arefyev, “DeepMistake at LSCDiscovery: Can a Multilingual Word-in-Context Model Replace Human Annotators?’’ in Proceedings of the 3rd Workshop on Computational Approaches to Historical Language Change, Dublin, Ireland, May 26–27, 2022(Association for Computational Linguistics, 2022), pp. 173–179.
doi 10.18653/v1/2022.lchange-1.18
- W. M. Rand, “Objective Criteria for the Evaluation of Clustering Methods,” Journal of the American Statistical Association 66 (336), 846–850 (1971).
doi 10.1080/01621459.1971.10482356
- K. Erk, D. McCarthy, and N. Gaylord, “Measuring Word Meaning in Context,” Computational Linguistics 39 (3), 511–554 (2013).
doi 10.1162/coli_a_00142
- A. Kilgarriff, “How Dominant Is the Commonest Sense of a Word?’’ in International Conference on Text, Speech and Dialogue Lecture Notes in Computer Science. Vol. 3206.(Springer, Berlin Heidelberg, 2004), pp. 103–111.
doi 10.1007/978-3-540-30120-2_14
- D. Kokosinskii and N. Arefyev, “Multilingual Substitution-based Word Sense Induction,” in Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Torino, Italia, May 20–25, 2024(ELRA and ICCL, Torino, 2024), pp. 11859–11872.
https://aclanthology.org/2024.lrec-main.1035/ Cited July 19, 2026.
- Oxford English Dictionary(Oxford University Press, Oxford, online edition, 2026).
https://www.oed.com Cited July 17, 2026.
- T. Blevins and L. Zettlemoyer, “Moving Down the Long Tail of Word Sense Disambiguation with Gloss Informed Bi-encoders,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, July, 2020(Association for Computational Linguistics, Online, 2020), pp. 1006–1017.
doi 10.18653/v1/2020.acl-main.95
- T. Pasini, A. Raganato, and R. Navigli, “XL-WSD: An Extra-Large and Cross-Lingual Evaluation Framework for Word Sense Disambiguation,” in Proceedings of the AAAI Conference on Artificial Intelligence, February 2-9, 2021, Online35 (15), 13648–13656 (2021).
doi 10.1609/aaai.v35i15.17609
- A. Grattafiori, A. Dubey, A. Jauhri, et al., “The Llama 3 Herd of Models,” arXiv: 2407.21783 (2024).
doi 10.48550/arXiv.2407.21783