Comunicação em evento científico
A generalized parallelization algorithm for Particle-In-Cell Simulations
Ricardo Fonseca (Fonseca, R. A.);
Título Evento
Laser-Plasma Accelerators Workshop 2025
Ano (publicação definitiva)
2025
Língua
Inglês
País
Itália
Mais Informação
Web of Science®

Esta publicação não está indexada na Web of Science®

Scopus

Esta publicação não está indexada na Scopus

Google Scholar

N.º de citações: 0

(Última verificação: 2026-08-26 22:09)

Ver o registo no Google Scholar

Esta publicação não está indexada no Overton

Abstract/Resumo
Particle-in-cell (PIC) codes have been a cornerstone of plasma-based accelerator development. These work at the most fundamental, microscopic level, making few physics approximations, and are ideally suited to this problem. However, this makes them some of the most computationally expensive models in plasma physics. The current ecosystem of scientific computing systems relies on many different hardware approaches and vendors, each with specific programming models, memory architectures, and processor types, and efficiently deploying PIC codes on these architectures is paramount. In this paper, we present a generalized parallelization algorithm for PIC simulations that is shown to work across all of the main architectures available today, including both CPUs (x86 / Arm) and GPUs (NVIDIA, AMD, Intel). The algorithm is based on a micro-spatial domain decomposition, with a high-performance particle manager to move particles between domains. Each domain is then assigned to a different thread (CPU) or thread block (GPU), achieving good parallel load balancing even for realistic simulation scenarios. The implementation is done using different programming models for different architectures, namely OpenMP (CPU), CUDA, ROCm (GPU), and SYCL (CPU/GPU/FPGA). While the implementations are effectively different code bases, given that the overall algorithm is the same, there are great similarities between all the implementations, making porting between them relatively straightforward. We present a performance comparison between different architectures/programming models for a test 2D problem, demonstrating very high performance for the architectures explored.
Agradecimentos/Acknowledgements
--
Palavras-chave