Exportar Publicação
A publicação pode ser exportada nos seguintes formatos: referência da APA (American Psychological Association), referência do IEEE (Institute of Electrical and Electronics Engineers), BibTeX e RIS.
Rebecca Cunha, Nunes, Luis & de Almeida, A. (2026). Synthetic data for machine learning: a study on quality and evaluation. INFORUM 2026,.
R. D. Cunha et al., "Synthetic data for machine learning: a study on quality and evaluation", in INFORUM 2026, 2026
@misc{cunha2026_1790668522497,
author = "Rebecca Cunha and Nunes, Luis and de Almeida, A.",
title = "Synthetic data for machine learning: a study on quality and evaluation",
year = "2026",
url = "https://inforum.pt/index"
}
TY - CPAPER TI - Synthetic data for machine learning: a study on quality and evaluation T2 - INFORUM 2026 AU - Rebecca Cunha AU - Nunes, Luis AU - de Almeida, A. PY - 2026 UR - https://inforum.pt/index AB - Generative Adversarial Networks (GANs) have emerged as a promising solution for synthetic tabular data generation in class–imbalanced financial datasets. However, evaluation in the literature has remained predominantly focused on downstream machine learning utility, leaving statistical fidelity, domain feasibility, and privacy dimensions largely unquantified. This paper presents a multidimensional empirical evaluation of CTGAN and CTAB-GAN+, two state-of-the-art tabular GAN architectures, applied to credit card default prediction under class imbalance. Using the UCI Default of Credit Card Clients dataset, we compare synthetic minority oversampling against two reference conditions, an unaugmented imbalanced baseline and a SMOTE baseline across four evaluation pillars: ML utility, per-variable statistical fidelity, a seven-property semantic imperceptibility audit, and distance-based pri- vacy risk. Our results reveal a consistent pattern: the architecture that appears superior under standard utility benchmarks simultaneously exhibits severe distributional distortions, widespread domain violations, and substantial loss of inter-variable correlation structure, none of which are captured by the Train on Synthetic Test on Real paradigm, yet each representing a critical obstacle to adoption in regulated environments. The architecturally more sophisticated model, despite lower predictive scores, achieves substantially superior fidelity and zero domain violations. Both architectures pass all primary privacy safety criteria. These findings demonstrate that multidimensional evaluation is not optional but essential for the responsible deployment of synthetic data. ER -
English