Registo de Dados de Investigação
This dataset contains a large scale paraphrase-based query benchmark built from 1000 original questions randomly sampled from Ciberdúvidas da Língua Portuguesa, an expert-curated Portuguese-language consultation service (https://ciberduvidas.iscte-iul.pt).
For each of the 1000 original questions, five paraphrased reformulations were generated using the mistral-large-2512 model via the Mistral API, according to five distinct query profiles designed to simulate the range of phrasings a real user might submit:
- Synthetic: short, direct, search-engine style queries
- Formal: grammatically rigorous and highly detailed queries
- Informal: casual, conversational tone queries
- Professor: queries using technical pedagogical/linguistic terminology
- Student: direct, comprehension-focused queries
This scaled benchmark served as a primary evaluation suite for the retrieval and reranking components of a Conversational Agent system. It was used to: compare embedding models, retrieval strategies, evaluate a domain-adapted cross-encoder reranker and support an automatic Ragas-based comparison of candidate generation models.
Each entry includes:
- id: unique identifier of the original Ciberdúvidas question
- url: direct link to the original question on the Ciberdúvidas website
- paraphrases: object with the five paraphrased query variants
(synthetic, formal, informal, teacher, student)
| Nome do Ficheiro | Tamanho Ficheiro |
|---|---|
| ciberduvidas_5000_paraphrases.json | 1605 KB |
| Nome | Afiliação | ORCID |
|---|---|---|
| Moura, Pedro | Iscte – Instituto Universitário de Lisboa | 0009-0006-7389-5759 |
| Batista, Fernando | Iscte – Instituto Universitário de Lisboa | 0000-0002-1075-0177 |
| Lopes, António | Iscte – Instituto Universitário de Lisboa | 0000-0003-3045-0304 |
Ainda não há publicações associadas a este registo
Ainda não há projetos associados a este registo
Ainda não há ODS associados a este registo
Ainda não há Áreas FoS associadas a este registo
English