Comunicação em evento científico
Synthetic data for machine learning: a study on quality and evaluation
Rebecca de Oliveira Cunha (Rebecca Cunha); Luís Nunes (Nunes, Luis); Ana de Almeida (de Almeida, A.);
Título Evento
INFORUM 2026
Ano (publicação definitiva)
2026
Língua
Inglês
País
--
Mais Informação
Web of Science®

Esta publicação não está indexada na Web of Science®

Scopus

Esta publicação não está indexada na Scopus

Google Scholar

Esta publicação não está indexada no Google Scholar

Esta publicação não está indexada no Overton

Abstract/Resumo
Generative Adversarial Networks (GANs) have emerged as a promising solution for synthetic tabular data generation in class–imbalanced financial datasets. However, evaluation in the literature has remained predominantly focused on downstream machine learning utility, leaving statistical fidelity, domain feasibility, and privacy dimensions largely unquantified. This paper presents a multidimensional empirical evaluation of CTGAN and CTAB-GAN+, two state-of-the-art tabular GAN architectures, applied to credit card default prediction under class imbalance. Using the UCI Default of Credit Card Clients dataset, we compare synthetic minority oversampling against two reference conditions, an unaugmented imbalanced baseline and a SMOTE baseline across four evaluation pillars: ML utility, per-variable statistical fidelity, a seven-property semantic imperceptibility audit, and distance-based pri- vacy risk. Our results reveal a consistent pattern: the architecture that appears superior under standard utility benchmarks simultaneously exhibits severe distributional distortions, widespread domain violations, and substantial loss of inter-variable correlation structure, none of which are captured by the Train on Synthetic Test on Real paradigm, yet each representing a critical obstacle to adoption in regulated environments. The architecturally more sophisticated model, despite lower predictive scores, achieves substantially superior fidelity and zero domain violations. Both architectures pass all primary privacy safety criteria. These findings demonstrate that multidimensional evaluation is not optional but essential for the responsible deployment of synthetic data.
Agradecimentos/Acknowledgements
This work was partially supported by national funds through Fundação para a Ciência e a Tecnologia, I.P. (FCT), ISTAR-Iscte Pluriannual Projects UID/04466/2025 (https://doi.org/10.54499/UID/04466/2025).
Palavras-chave
Class imbalance,GANs,Multidimensional evaluation,Privacy,Statistical fidelity,Synthetic tabular data