TY - GEN
T1 - LegoIndex
T2 - 34th International Symposium on High-Performance Parallel and Distributed Computing
AU - Guo, Chang
AU - Yan, Ning
AU - Wan, Lipeng
AU - Cao, Zhichao
N1 - Publisher Copyright:
© 2025 Copyright is held by the owner/author(s). Publication rights licensed to ACM.
PY - 2025/9/9
Y1 - 2025/9/9
N2 - Particle-in-Cell (PIC) simulations play a critical role in various scientific domains, including plasma physics, astrophysics, and fusion energy research, by enabling the modeling of complex interactions between charged particles and electromagnetic fields. As PIC simulations scale up in size and complexity, they generate massive volumes of particle data at enormous speeds (TBs/hour). This enormous amount of data presents significant challenges for post-simulation analysis, as existing analysis tools (typically designed for smaller datasets) struggle with low query performance and high resource utilization. While incorporating indexes for PIC data can alleviate some of these inefficiencies, current indexing solutions often fall short of addressing the diverse analysis needs of scientists, substantial index construction overhead during simulation runs, and inefficient small I/O operations.To address these challenges, we present LegoIndex, a scalable, modular, and elastic post-simulation indexing framework designed to improve data access efficiency while minimizing unnecessary data I/Os and memory usage. LegoIndex offers users the flexibility to customize indexing components, structures, and statistic metrics based on the data scale and specific analysis needs. LegoIndex parallelizes the index construction process, enabling efficient processing of datasets of varying sizes. To further enhance query performance, LegoIndex intelligently clusters scattered index results that are spatially or temporally related and optimizes computation logic to effectively reduce data I/Os and memory usage. We conducted comprehensive evaluations of LegoIndex based on large-scale real-world PIC datasets. Integrating LegoIndex with the existing analysis tool can achieve up to a 2276× improvement in overall performance, a 3068× reduction in memory usage, and a 3001× decrease in I/O numbers for large datasets.
AB - Particle-in-Cell (PIC) simulations play a critical role in various scientific domains, including plasma physics, astrophysics, and fusion energy research, by enabling the modeling of complex interactions between charged particles and electromagnetic fields. As PIC simulations scale up in size and complexity, they generate massive volumes of particle data at enormous speeds (TBs/hour). This enormous amount of data presents significant challenges for post-simulation analysis, as existing analysis tools (typically designed for smaller datasets) struggle with low query performance and high resource utilization. While incorporating indexes for PIC data can alleviate some of these inefficiencies, current indexing solutions often fall short of addressing the diverse analysis needs of scientists, substantial index construction overhead during simulation runs, and inefficient small I/O operations.To address these challenges, we present LegoIndex, a scalable, modular, and elastic post-simulation indexing framework designed to improve data access efficiency while minimizing unnecessary data I/Os and memory usage. LegoIndex offers users the flexibility to customize indexing components, structures, and statistic metrics based on the data scale and specific analysis needs. LegoIndex parallelizes the index construction process, enabling efficient processing of datasets of varying sizes. To further enhance query performance, LegoIndex intelligently clusters scattered index results that are spatially or temporally related and optimizes computation logic to effectively reduce data I/Os and memory usage. We conducted comprehensive evaluations of LegoIndex based on large-scale real-world PIC datasets. Integrating LegoIndex with the existing analysis tool can achieve up to a 2276× improvement in overall performance, a 3068× reduction in memory usage, and a 3001× decrease in I/O numbers for large datasets.
KW - high-performance computing
KW - indexing
KW - particle data
KW - scientific data
UR - https://www.scopus.com/pages/publications/105017969820
UR - https://www.scopus.com/pages/publications/105017969820#tab=citedBy
U2 - 10.1145/3731545.3731591
DO - 10.1145/3731545.3731591
M3 - Conference contribution
AN - SCOPUS:105017969820
T3 - HPDC 2025 - Proceedings of the 34th International Symposium on High-Performance Parallel and Distributed Computing
BT - HPDC 2025 - Proceedings of the 34th International Symposium on High-Performance Parallel and Distributed Computing
PB - Association for Computing Machinery, Inc
Y2 - 20 July 2025 through 23 July 2025
ER -