Modélisation lexicale et intelligence artificielle : une ressource bilingue pour l’architecture contemporaine
Okładka czasopisma Studia Romanica Posnaniensia, tom 53, nr 3, rok 2026, tytuł Recensement et description des données en vue des applications lexicographiques
PDF (Français (France))

Słowa kluczowe

specialized lexicography
bilingual resource
contemporary architecture
TEI-LMF modeling
artificial intelligence

Jak cytować

Bartolomé-Díaz, Z. (2026). Modélisation lexicale et intelligence artificielle : une ressource bilingue pour l’architecture contemporaine. Studia Romanica Posnaniensia, 53(3), 7–28. https://doi.org/10.14746/strop.2026.53.3.1

Abstrakt

This article presents a methodology for the design of a bilingual French-Spanish lexical resource dedicated to contemporary architecture, a field characterized by a constantly evolving terminology and a lack of specialized dictionaries. Our proposal is based on a model that integrates the entire lexicographic workflow – from corpus building to final encoding, including extraction, alignment, definition, and enrichment – while leveraging the contributions of artificial intelligence. Statistical and neural models are used to automate steps that are traditionally lengthy and demanding – such as term identification, semantic clustering, draft definition generation, or the detection of variants. However, this automation remains framed by systematic human validation, which is indispensable to ensure conceptual accuracy and terminological consistency. More broadly, our methodology provides a roadmap that can be applied to other specialized domains and to under-represented languages.

https://doi.org/10.14746/strop.2026.53.3.1
PDF (Français (France))

Bibliografia

Artetxe, M., Labaka, G. & Agirre, E. (2018). A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings. In I. Gurevych & Y. Miyao (éds.), Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (vol. 1 : Long Papers), 789-798. Melbourne : Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/P18-1073

Bartolomé-Díaz, Z. & Trujillo-González, V.C. (2023). Les termes architecturaux. Recommandations officielles et réalités des usages. Çédille, revista de estudios franceses, 24.

Becchi, A. et al. (2008). Les livres d’architecture : Leurs éditions de la Renaissance à nos jours. Perspective. Actualité en histoire de l’art, 2. DOI: https://doi.org/10.4000/perspective.3396

Bergenholtz, H. & Tarp, S. (1995). Manual of Specialised Lexicography. Amsterdam (Pays-Bas) : John Benjamins Publishing Company. https://benjamins.com/catalog/btl.12 DOI: https://doi.org/10.1075/btl.12

Bojanowski, P. et al. (2017). Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5, 135-146. DOI: https://doi.org/10.1162/tacl_a_00051

Camacho-Collados, J. & Pilehvar, M.T. (2018). From word to sense embeddings: A survey on vector representations of meaning. Journal of Artificial Intelligence Research, 63 (1), 743-788. DOI: https://doi.org/10.1613/jair.1.11259

Cañete, J. et al. (2020). Spanish Pre-trained BERT Model and Evaluation Data. Practical ML for Developing Countries Workshop : learning under limited/low resource scenarios.

De Schryver, G.-M. (2023). Generative AI and lexicography: The current state of the art using ChatGPT. International Journal of Lexicography, 36. DOI: https://doi.org/10.1093/ijl/ecad021

Devlin, J. et al. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, (vol. 1 : Long and Short Papers) pp. 4171-4186). Minneapolis, Minnesota (États-Unis) : Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/N19-1423

Durand, B. (2023). La fabrication d’une « architecture durable » en France (2000-2010) [thèse de doctorat]. Université Paris-Est. https://theses.hal.science/tel-04353097

Fung, J. et al. (2024). Human-in-the-loop technical document annotation: Developing and validating a system to provide machine-assistance for domain-specific text analysis. National Institute of Standards and Technology. https://www.nist.gov/publications/human-loop-technical-document-annotation-developing-and-validating-system-provide DOI: https://doi.org/10.6028/NIST.TN.2287

Grabowski, Ł. (2023). Statistician, programmer, data scientist? Who is, or should be, a corpus linguist in the 2020s? Journal of Linguistics/Jazykovedný Casopis, 74 (1), 52-59. DOI: https://doi.org/10.2478/jazcas-2023-0023

Kilgarriff, A. et al. (2004). The Sketch Engine, 17.11.2016. 105-115. https://euralex.org/publications/the-sketch-engine/

Lew, R. (2023). ChatGPT as a COBUILD lexicographer. Humanities and Social Sciences Communications, 10 (1). DOI: https://doi.org/10.1057/s41599-023-02119-6

Lew, R. (2024). Dictionaries and lexicography in the AI era. Humanities and Social Sciences Communications, 11 (1), 1-8. DOI: https://doi.org/10.1057/s41599-024-02889-7

L’Homme, M.-C. (2019). Lexical semantics for terminology. Amsterdam : John Benjamins Publishing Company.

Li, Q. & Tarp, S. (2024). Using generative AI to provide high-quality lexicographic assistance to Chinese learners of English. Lexikos, 34, 397-418. DOI: https://doi.org/10.5788/34-1-1944

Martin, L. et al. (2020). CamemBERT: A tasty French language model. In D. Jurafsky et al. (éds.), En Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 7203-7219). Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/2020.acl-main.645

Mikolov, T. et al. (2013). Efficient Estimation of Word Representations in Vector Space. In 1st International Conference on Learning Representations. Scottsdale (AZ, États-Unis). https://arxiv.org/abs/1301.3781.

Miller, G.A. (1995). WordNet : A lexical database for English. In Communications of the ACM, vol. 38, (pp. 39-41). San Francisco (CA, États-Unis) : Morgan Kaufmann. DOI: https://doi.org/10.1145/219717.219748

Navigli, R. & Ponzetto, S. (2012). BabelNet: The automatic construction, evaluation and application of a wide-coverage multilingual semantic network. Artificial Intelligence, 193, 217-250. DOI: https://doi.org/10.1016/j.artint.2012.07.001

OpenAI et al. (2023). GPT-4 Technical Report. https://cdn.openai.com/papers/gpt-4.pdf

Ortega-Martín, M. et al. (2023). Spanish built factual freectianary (Spanish-BFF): The first AI-generated free dictionary.

Ortega-Martín, M. et al. (2024). Building another Spanish dictionary, this time with GPT-4.

Periti, F., Alfter, D., & Tahmasebi, N. (2024). Automatically generated definitions and their utility for Modeling Word Meaning. In Y. Al-Onaizan, M. Bansal & Y.-N. Chen (éds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 14008-14026). Miami, FL : Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/2024.emnlp-main.776

Romary, L. (2010). Using the TEI framework as a possible serialization for LMF. Rendering endangered languages lexicons interoperable through standards harmonization. RELISH Workshop Nijmegen (Pays-Bas). https://hal.inria.fr/inria-00511769

Romary, L. (2013). TEI and LMF crosswalks. Journal for Language Technology and Computational Linguistics 30 (1), 47-70. DOI: https://doi.org/10.21248/jlcl.30.2015.195

Romary, L. et al. (2019). LMF reloaded. In Proceedings of the 13th International Conference of the Asian Association for Lexicography: Past, Present and Future (pp. 533-539). Istanbul : Asos Publisher. https://inria.hal.science/hal-02118319v1

Scarlini, B., Pasini, T. & Navigli, R. (2020). With more contexts comes better performance: Contextualized sense embeddings for all-round word sense disambiguation. In B. Webber et al. (éds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (pp. 3528-3539). Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/2020.emnlp-main.285

Toprak, A. & Turan, M. (2025). Automated thematic dictionary creation using the web based on WordNet, Spacy, and Simhash. Data and Information Management, 9 (3), 100088. DOI: https://doi.org/10.1016/j.dim.2024.100088

Yu, W. et al. (2022). Dict-BERT : Enhancing language model pre-training with dictionary. Findings of the Association for Computational Linguistics: ACL 2022, 1907-1918. DOI: https://doi.org/10.18653/v1/2022.findings-acl.150