Exportar Publicação

A publicação pode ser exportada nos seguintes formatos: referência da APA (American Psychological Association), referência do IEEE (Institute of Electrical and Electronics Engineers), BibTeX e RIS.

Exportar Referência (APA)
Rebecca Cunha, Nunes, Luis & de Almeida, A. (2026). Synthetic data for machine learning: a study on quality and evaluation. INFORUM 2026,.
Exportar Referência (IEEE)
R. D. Cunha et al.,  "Synthetic data for machine learning: a study on quality and evaluation", in INFORUM 2026, 2026
Exportar BibTeX
@misc{cunha2026_1790668522497,
	author = "Rebecca Cunha and Nunes, Luis and de Almeida, A.",
	title = "Synthetic data for machine learning: a study on quality and evaluation",
	year = "2026",
	url = "https://inforum.pt/index"
}
Exportar RIS
TY  - CPAPER
TI  - Synthetic data for machine learning: a study on quality and evaluation
T2  - INFORUM 2026
AU  - Rebecca Cunha
AU  - Nunes, Luis
AU  - de Almeida, A.
PY  - 2026
UR  - https://inforum.pt/index
AB  - Generative Adversarial Networks (GANs) have emerged as a promising solution for synthetic tabular data generation in class–imbalanced financial datasets. However, evaluation in the literature has remained predominantly focused on downstream machine learning utility, leaving statistical fidelity, domain feasibility, and privacy dimensions largely unquantified. This paper presents a multidimensional empirical evaluation of CTGAN and CTAB-GAN+, two state-of-the-art tabular GAN architectures, applied to credit card default prediction under class imbalance. Using the UCI Default of Credit Card Clients dataset, we compare synthetic minority oversampling against two reference conditions, an unaugmented imbalanced baseline and a SMOTE baseline across four evaluation pillars: ML utility, per-variable statistical fidelity, a seven-property semantic imperceptibility audit, and distance-based pri-
vacy risk. Our results reveal a consistent pattern: the architecture that appears superior under standard utility benchmarks simultaneously exhibits severe distributional distortions, widespread domain violations, and substantial loss of inter-variable correlation structure, none of which are captured by the Train on Synthetic Test on Real paradigm, yet each representing a critical obstacle to adoption in regulated environments. The architecturally more sophisticated model, despite lower predictive scores, achieves substantially superior fidelity and zero domain violations. Both architectures pass all primary privacy safety criteria. These findings demonstrate that multidimensional evaluation is not optional but essential for the responsible deployment of synthetic data.
ER  -