Ciência_Iscte
Comunicações
Descrição Detalhada da Comunicação
Efficient Frequency-Aware Multiscale Vision Transformer for Event-to-Video Reconstruction
Título Evento
2025 33rd European Signal Processing Conference (EUSIPCO)
Ano (publicação definitiva)
2025
Língua
Inglês
País
Itália
Mais Informação
Web of Science®
Esta publicação não está indexada na Web of Science®
Scopus
Esta publicação não está indexada na Scopus
Google Scholar
Esta publicação não está indexada no Google Scholar
Esta publicação não está indexada no Overton
Abstract/Resumo
Event-to-video (E2V) reconstruction is a critical task in event-based vision, benefiting from the advantages of event cameras, such as high dynamic range and low latency. However, existing deep learning reconstruction methods often prioritize temporal consistency and over-emphasize low-frequency features, leading to blur artifacts and loss of fine details. To overcome these limitations, we propose a novel frequency-aware multiscale vision transformer model for E2V reconstruction (MSViT-E2V). Our model employs wavelet-based decomposition to extract features at multiple scales, preserving fine-grained details through multilevel wavelet-based downsampling blocks, followed by transformer blocks for multiscale feature aggregation and long-range dependency modeling. Extensive experiments on various event datasets demonstrate that our model not only minimizes artifacts and preserves fine details but also reduces computational costs by up to 50% compared to the transformer-based model ET-Net.
Agradecimentos/Acknowledgements
--
Palavras-chave
Event-based vision,Frequency-domain analysis,Video reconstruction,Vision transformer
Registos de financiamentos
| Referência de financiamento | Entidade Financiadora |
|---|---|
| UID/50008:Instituto de Telecomunicações | FCT/MECI |
English