Talk
Synthetic data for machine learning: a study on quality and evaluation
Rebecca de Oliveira Cunha (Rebecca Cunha); Luís Nunes (Nunes, Luis); Ana de Almeida (de Almeida, A.);
Event Title
INFORUM 2026
Year (definitive publication)
2026
Language
English
Country
--
More Information
Web of Science®

This publication is not indexed in Web of Science®

Scopus

This publication is not indexed in Scopus

Google Scholar

This publication is not indexed in Google Scholar

This publication is not indexed in Overton

Abstract
Generative Adversarial Networks (GANs) have emerged as a promising solution for synthetic tabular data generation in class–imbalanced financial datasets. However, evaluation in the literature has remained predominantly focused on downstream machine learning utility, leaving statistical fidelity, domain feasibility, and privacy dimensions largely unquantified. This paper presents a multidimensional empirical evaluation of CTGAN and CTAB-GAN+, two state-of-the-art tabular GAN architectures, applied to credit card default prediction under class imbalance. Using the UCI Default of Credit Card Clients dataset, we compare synthetic minority oversampling against two reference conditions, an unaugmented imbalanced baseline and a SMOTE baseline across four evaluation pillars: ML utility, per-variable statistical fidelity, a seven-property semantic imperceptibility audit, and distance-based pri- vacy risk. Our results reveal a consistent pattern: the architecture that appears superior under standard utility benchmarks simultaneously exhibits severe distributional distortions, widespread domain violations, and substantial loss of inter-variable correlation structure, none of which are captured by the Train on Synthetic Test on Real paradigm, yet each representing a critical obstacle to adoption in regulated environments. The architecturally more sophisticated model, despite lower predictive scores, achieves substantially superior fidelity and zero domain violations. Both architectures pass all primary privacy safety criteria. These findings demonstrate that multidimensional evaluation is not optional but essential for the responsible deployment of synthetic data.
Acknowledgements
This work was partially supported by national funds through Fundação para a Ciência e a Tecnologia, I.P. (FCT), ISTAR-Iscte Pluriannual Projects UID/04466/2025 (https://doi.org/10.54499/UID/04466/2025).
Keywords
Class imbalance,GANs,Multidimensional evaluation,Privacy,Statistical fidelity,Synthetic tabular data