Exportar Publicação
A publicação pode ser exportada nos seguintes formatos: referência da APA (American Psychological Association), referência do IEEE (Institute of Electrical and Electronics Engineers), BibTeX e RIS.
Vanda Barata, de Almeida, A. & Nunes, Luis (2026). Beyond Downstream Accuracy: A Multi-Axis Framework for Evaluating Synthetic EEG in Seizure Detection. EPIA Conference on Artificial Intelligence (EPIA 2026),.
V. Barata et al., "Beyond Downstream Accuracy: A Multi-Axis Framework for Evaluating Synthetic EEG in Seizure Detection", in EPIA Conf. on Artificial Intelligence (EPIA 2026), 2026
@misc{barata2026_1790107204845,
author = "Vanda Barata and de Almeida, A. and Nunes, Luis",
title = "Beyond Downstream Accuracy: A Multi-Axis Framework for Evaluating Synthetic EEG in Seizure Detection",
year = "2026"
}
TY - CPAPER TI - Beyond Downstream Accuracy: A Multi-Axis Framework for Evaluating Synthetic EEG in Seizure Detection T2 - EPIA Conference on Artificial Intelligence (EPIA 2026) AU - Vanda Barata AU - de Almeida, A. AU - Nunes, Luis PY - 2026 AB - Seizures are rare events in electroencephalogram (EEG) recordings, and clinical data are scarce and hard to access, making it difficult to train reliable classifiers. Synthetic data augmentation is increasingly used to address this, yet most studies evaluate generators using only utility metrics like accuracy or F1 score, without validating the generated data itself. Such metrics may be misleading under extreme class imbalance and do not assess whether synthetic signals are physiologically plausible or clinically relevant. This paper proposes a multi-axis evaluation framework that evaluates synthetic EEG beyond utility alone, across four complementary dimensions: (1)fidelity, the degree to which generated signals preserve spectral, temporal, and spatial structure; (2)discriminability, how easily a classifier can distinguish real from synthetic data; (3)utility, measured by per-patient Area Under the Precision-Recall Curve (AUPRC) and sensitivity at 95\% specificity, ideally under leave-one-patient-out cross-validation; and (4) privacy, the extent to which generators memorise subject-specific patterns and whether augmentation reduces or amplifies the detector's reliance on patient identity. We apply the framework to three generators (TimeGAN, Conditional Variational Autoencoder, and Latent Diffusion Model) on the CHB-MIT dataset, evaluating only three of the four axes due to computational constraints. We show that even three axes reveal failure modes invisible to utility metrics alone - the generator with the most uniform spectral fidelity does not improve detection, while one with better fidelity in seizure-relevant bands does. Only the joint assessment across axes explains why and correctly identifies which generator yields a meaningful improvement. ER -
English