Exportar Publicação

A publicação pode ser exportada nos seguintes formatos: referência da APA (American Psychological Association), referência do IEEE (Institute of Electrical and Electronics Engineers), BibTeX e RIS.

Exportar Referência (APA)
Fonseca, R. A. (2025). A generalized parallelization algorithm for Particle-In-Cell Simulations. Laser-Plasma Accelerators Workshop 2025,.
Exportar Referência (IEEE)
R. P. Fonseca,  "A generalized parallelization algorithm for Particle-In-Cell Simulations", in Laser-Plasma Accelerators Workshop 2025, Ischia, 2025
Exportar BibTeX
@misc{fonseca2025_1791219218625,
	author = "Fonseca, R. A.",
	title = "A generalized parallelization algorithm for Particle-In-Cell Simulations",
	year = "2025",
	url = "https://agenda.infn.it/event/42311/"
}
Exportar RIS
TY  - CPAPER
TI  - A generalized parallelization algorithm for Particle-In-Cell Simulations
T2  - Laser-Plasma Accelerators Workshop 2025
AU  - Fonseca, R. A.
PY  - 2025
CY  - Ischia
UR  - https://agenda.infn.it/event/42311/
AB  - Particle-in-cell (PIC) codes have been a cornerstone of plasma-based accelerator development. These work at the most fundamental, microscopic level, making few physics approximations, and are ideally suited to this problem. However, this makes them some of the most computationally expensive models in plasma physics. The current ecosystem of scientific computing systems relies on many different hardware approaches and vendors, each with specific programming models, memory architectures, and processor types, and efficiently deploying PIC codes on these architectures is paramount.

In this paper, we present a generalized parallelization algorithm for PIC simulations that is shown to work across all of the main architectures available today, including both CPUs (x86 / Arm) and GPUs (NVIDIA, AMD, Intel). The algorithm is based on a micro-spatial domain decomposition, with a high-performance particle manager to move particles between domains. Each domain is then assigned to a different thread (CPU) or thread block (GPU), achieving good parallel load balancing even for realistic simulation scenarios. The implementation is done using different programming models for different architectures, namely OpenMP (CPU), CUDA, ROCm (GPU), and SYCL (CPU/GPU/FPGA). While the implementations are effectively different code bases, given that the overall algorithm is the same, there are great similarities between all the implementations, making porting between them relatively straightforward. We present a performance comparison between different architectures/programming models for a test 2D problem, demonstrating very high performance for the architectures explored.
ER  -