Abstract
This article presents a methodology for the design of a bilingual French-Spanish lexical resource dedicated to contemporary architecture, a field characterized by a constantly evolving terminology and a lack of specialized dictionaries. Our proposal is based on a model that integrates the entire lexicographic workflow – from corpus building to final encoding, including extraction, alignment, definition, and enrichment – while leveraging the contributions of artificial intelligence. Statistical and neural models are used to automate steps that are traditionally lengthy and demanding – such as term identification, semantic clustering, draft definition generation, or the detection of variants. However, this automation remains framed by systematic human validation, which is indispensable to ensure conceptual accuracy and terminological consistency. More broadly, our methodology provides a roadmap that can be applied to other specialized domains and to under-represented languages.
References
Artetxe, M., Labaka, G. & Agirre, E. (2018). A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings. In I. Gurevych & Y. Miyao (éds.), Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (vol. 1 : Long Papers), 789-798. Melbourne : Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/P18-1073
Bartolomé-Díaz, Z. & Trujillo-González, V.C. (2023). Les termes architecturaux. Recommandations officielles et réalités des usages. Çédille, revista de estudios franceses, 24.
Becchi, A. et al. (2008). Les livres d’architecture : Leurs éditions de la Renaissance à nos jours. Perspective. Actualité en histoire de l’art, 2. DOI: https://doi.org/10.4000/perspective.3396
Bergenholtz, H. & Tarp, S. (1995). Manual of Specialised Lexicography. Amsterdam (Pays-Bas) : John Benjamins Publishing Company. https://benjamins.com/catalog/btl.12 DOI: https://doi.org/10.1075/btl.12
Bojanowski, P. et al. (2017). Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5, 135-146. DOI: https://doi.org/10.1162/tacl_a_00051
Camacho-Collados, J. & Pilehvar, M.T. (2018). From word to sense embeddings: A survey on vector representations of meaning. Journal of Artificial Intelligence Research, 63 (1), 743-788. DOI: https://doi.org/10.1613/jair.1.11259
Cañete, J. et al. (2020). Spanish Pre-trained BERT Model and Evaluation Data. Practical ML for Developing Countries Workshop : learning under limited/low resource scenarios.
De Schryver, G.-M. (2023). Generative AI and lexicography: The current state of the art using ChatGPT. International Journal of Lexicography, 36. DOI: https://doi.org/10.1093/ijl/ecad021
Devlin, J. et al. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, (vol. 1 : Long and Short Papers) pp. 4171-4186). Minneapolis, Minnesota (États-Unis) : Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/N19-1423
Durand, B. (2023). La fabrication d’une « architecture durable » en France (2000-2010) [thèse de doctorat]. Université Paris-Est. https://theses.hal.science/tel-04353097
Fung, J. et al. (2024). Human-in-the-loop technical document annotation: Developing and validating a system to provide machine-assistance for domain-specific text analysis. National Institute of Standards and Technology. https://www.nist.gov/publications/human-loop-technical-document-annotation-developing-and-validating-system-provide DOI: https://doi.org/10.6028/NIST.TN.2287
Grabowski, Ł. (2023). Statistician, programmer, data scientist? Who is, or should be, a corpus linguist in the 2020s? Journal of Linguistics/Jazykovedný Casopis, 74 (1), 52-59. DOI: https://doi.org/10.2478/jazcas-2023-0023
Kilgarriff, A. et al. (2004). The Sketch Engine, 17.11.2016. 105-115. https://euralex.org/publications/the-sketch-engine/
Lew, R. (2023). ChatGPT as a COBUILD lexicographer. Humanities and Social Sciences Communications, 10 (1). DOI: https://doi.org/10.1057/s41599-023-02119-6
Lew, R. (2024). Dictionaries and lexicography in the AI era. Humanities and Social Sciences Communications, 11 (1), 1-8. DOI: https://doi.org/10.1057/s41599-024-02889-7
L’Homme, M.-C. (2019). Lexical semantics for terminology. Amsterdam : John Benjamins Publishing Company.
Li, Q. & Tarp, S. (2024). Using generative AI to provide high-quality lexicographic assistance to Chinese learners of English. Lexikos, 34, 397-418. DOI: https://doi.org/10.5788/34-1-1944
Martin, L. et al. (2020). CamemBERT: A tasty French language model. In D. Jurafsky et al. (éds.), En Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 7203-7219). Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/2020.acl-main.645
Mikolov, T. et al. (2013). Efficient Estimation of Word Representations in Vector Space. In 1st International Conference on Learning Representations. Scottsdale (AZ, États-Unis). https://arxiv.org/abs/1301.3781.
Miller, G.A. (1995). WordNet : A lexical database for English. In Communications of the ACM, vol. 38, (pp. 39-41). San Francisco (CA, États-Unis) : Morgan Kaufmann. DOI: https://doi.org/10.1145/219717.219748
Navigli, R. & Ponzetto, S. (2012). BabelNet: The automatic construction, evaluation and application of a wide-coverage multilingual semantic network. Artificial Intelligence, 193, 217-250. DOI: https://doi.org/10.1016/j.artint.2012.07.001
OpenAI et al. (2023). GPT-4 Technical Report. https://cdn.openai.com/papers/gpt-4.pdf
Ortega-Martín, M. et al. (2023). Spanish built factual freectianary (Spanish-BFF): The first AI-generated free dictionary.
Ortega-Martín, M. et al. (2024). Building another Spanish dictionary, this time with GPT-4.
Periti, F., Alfter, D., & Tahmasebi, N. (2024). Automatically generated definitions and their utility for Modeling Word Meaning. In Y. Al-Onaizan, M. Bansal & Y.-N. Chen (éds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 14008-14026). Miami, FL : Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/2024.emnlp-main.776
Romary, L. (2010). Using the TEI framework as a possible serialization for LMF. Rendering endangered languages lexicons interoperable through standards harmonization. RELISH Workshop Nijmegen (Pays-Bas). https://hal.inria.fr/inria-00511769
Romary, L. (2013). TEI and LMF crosswalks. Journal for Language Technology and Computational Linguistics 30 (1), 47-70. DOI: https://doi.org/10.21248/jlcl.30.2015.195
Romary, L. et al. (2019). LMF reloaded. In Proceedings of the 13th International Conference of the Asian Association for Lexicography: Past, Present and Future (pp. 533-539). Istanbul : Asos Publisher. https://inria.hal.science/hal-02118319v1
Scarlini, B., Pasini, T. & Navigli, R. (2020). With more contexts comes better performance: Contextualized sense embeddings for all-round word sense disambiguation. In B. Webber et al. (éds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (pp. 3528-3539). Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/2020.emnlp-main.285
Toprak, A. & Turan, M. (2025). Automated thematic dictionary creation using the web based on WordNet, Spacy, and Simhash. Data and Information Management, 9 (3), 100088. DOI: https://doi.org/10.1016/j.dim.2024.100088
Yu, W. et al. (2022). Dict-BERT : Enhancing language model pre-training with dictionary. Findings of the Association for Computational Linguistics: ACL 2022, 1907-1918. DOI: https://doi.org/10.18653/v1/2022.findings-acl.150
License
Copyright (c) 2026 Zaida Bartolomé-Díaz

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
- The Author hereby warrants that he/she is the owner of all the copyright and other intellectual property rights in the Work and that, within the scope of the present Agreement, the paper does not infringe the legal rights of another person.The owner of the copyright work also warrants that he/she is the sole and original creator thereof and that is not bound by any legal constraints in regard to the use or sale of the work.
- The Author grants the Purchaser non-exclusive and free of charge license to unlimited use worldwide over an unspecified period of time in the following areas of exploitation:
2.1. production of multiple copies of the Work produced according to the specific application of a given technology, including printing, reproduction of graphics through mechanical or electrical means (reprography) and digital technology;
2.2. marketing authorisation, loan or lease of the original or copies thereof;
2.3. public performance, public performance in the broadcast, video screening, media enhancements as well as broadcasting and rebroadcasting, made available to the public in such a way that members of the public may access the Work from a place and at a time individually chosen by them;
2.4. inclusion of the Work into a collective work (i.e. with a number of contributions);
2.5. inclusion of the Work in the electronic version to be offered on an electronic platform, or any other conceivable introduction of the Work in its electronic version to the Internet;
2.6. dissemination of electronic versions of the Work in its electronic version online, in a collective work or independently;
2.7. making the Work in the electronic version available to the public in such a way that members of the public may access the Work from a place and at a time individually chosen by them, in particular by making it accessible via the Internet, Intranet, Extranet;
2.8. making the Work available according to appropriate license pattern Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) as well as another language version of this license or any later version published by Creative Commons. - The Author grants the Adam Mickiewicz University in Poznan permission to:
3.1. reproduce a single copy (print or download) and royalty-free use and disposal of rights to compilations of the Work and these compilations.
3.2. send metadata files related to the Work, including to commercial and non-commercial journal-indexing databases. - The Author declares that, on the basis of the license granted in the present Agreement, the Purchaser is entitled and obliged to allow third parties to obtain further licenses (sublicenses) to the Work and to other materials, including derivatives thereof or compilations made based on or including the Work, whereas the provisions of such sub-licenses will be the same as with the attributed license pattern Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0)or another language version of this license, or any later version of this license published by Creative Commons. Thereby, the Author entitles all interested parties to use the work, for non-commercial purposes only, under the following conditions:
4.1. acknowledgment of authorship, i.e. observing the obligation to provide, along with the distributed work, information about the author, title, source (link to the original Work, DOI) and of the license itself.
4.2. the derivatives of the Work are subject to the same conditions, i.e. they may be published only based on a licence which would be identical to the one granting access to the original Work. - The University of Adam Mickiewicz in Poznań is obliged to
5.1. make the Work available to the public in such a way that members of the public may access the Work from a place and at a time individually chosen by them, without any technological constraints;
5.2. appropriately inform members of the public to whom the Work is to be made available about sublicenses in such a way as to ensure that all parties are properly informed (appropriate informing messages).
Other provisions
- The University of Adam Mickiewicz in Poznań hereby preserves the copyright to the journal as a whole (layout/stylesheet, graphics, cover design, logo etc.)
