Exportar Publicação
A publicação pode ser exportada nos seguintes formatos: referência da APA (American Psychological Association), referência do IEEE (Institute of Electrical and Electronics Engineers), BibTeX e RIS.
Moura, P., Gama, I., Batista, F. & Lopes, A. L. (2026). Optimising Retrieval for Linguistic Question-Answering in European Portuguese: A Benchmark on Ciberdúvidas Da Língua Portuguesa. 15th Symposium on Languages, Applications and Technologies (SLATE 2026).
P. A. Moura et al., "Optimising Retrieval for Linguistic Question-Answering in European Portuguese: A Benchmark on Ciberdúvidas Da Língua Portuguesa", in 15th Symp. on Languages, Applications and Technologies (SLATE 2026), Lisboa, 2026
@misc{moura2026_1788026670968,
author = "Moura, P. and Gama, I. and Batista, F. and Lopes, A. L.",
title = "Optimising Retrieval for Linguistic Question-Answering in European Portuguese: A Benchmark on Ciberdúvidas Da Língua Portuguesa",
year = "2026",
howpublished = "Digital",
url = "https://slate-conf.org/2026/home"
}
TY - CPAPER TI - Optimising Retrieval for Linguistic Question-Answering in European Portuguese: A Benchmark on Ciberdúvidas Da Língua Portuguesa T2 - 15th Symposium on Languages, Applications and Technologies (SLATE 2026) AU - Moura, P. AU - Gama, I. AU - Batista, F. AU - Lopes, A. L. PY - 2026 CY - Lisboa UR - https://slate-conf.org/2026/home AB - Information retrieval for question-answering remains underexplored in specialised domains and under-resourced language variants such as European Portuguese. Existing benchmarks largely target general-domain English data and document-centric retrieval, failing to capture the semantic alignment required for linguistic consultation tasks over curated question–answer (QA) pairs. We address this gap by introducing a controlled evaluation framework for retrieval over the "Ciberdúvidas da Língua Portuguesa" corpus, comprising 29,145 expert-validated QA entries. Our approach systematically analyses the interaction between indexing strategies, encoder models, and retrieval paradigms, while modelling real-world query variability through a paraphrase-based benchmark of 600 queries across five user profiles, manually validated by a professional linguist to ensure semantic fidelity. Experiments show that dense retrieval with an IR-optimised monolingual encoder significantly outperforms both sparse (BM25) and hybrid methods, achieving a Mean Reciprocal Rank (MRR) of 0.93. Notably, hybrid retrieval underperforms due to lexical mismatch interference, challenging prevailing assumptions in the literature. Our contributions include a novel benchmark framework for linguistic QA retrieval, empirical evidence supporting monolingual IR-specialised models, and insights into retrieval robustness under paraphrastic variation, enabling improved QA systems for specialised and low-resource environments. ER -
English