Modern HPC systems combine increasingly heterogeneous processors, accelerators, networks, and storage technologies. Achieving good application performance therefore requires more than optimizing individual code sections: computation, communication, data movement, I/O, and energy consumption must be considered across the complete software and system stack.
The Method Lab Performance Engineering supports the NHR community in systematically analyzing, modeling, and optimizing scientific applications on modern HPC systems. Its expertise ranges from code optimization, parallelization, and scalability analysis to GPU porting, GPU performance analysis, parallel I/O, storage, and energy-efficient computing.
A particular focus lies on data-intensive simulations and AI/ML workloads. The Method Lab helps identify I/O bottlenecks, performance variability, inefficient data movement, and resource limitations. It develops and applies suitable tools, metrics, and tuning strategies to make system behavior observable, explainable, and reproducible.
Through individual consulting, collaborative optimization projects, hands-on training, and community workshops, the Method Lab enables researchers to use current and emerging HPC infrastructures more effectively.
Competencies
Performance Analysis
Analysis, modeling, benchmarking, and optimization of scientific applications.
Parallelism & Scalability
Parallelization, communication analysis, and scalable execution.
GPU Performance
GPU porting, profiling, optimization, and performance portability.
Data Movement
Parallel I/O, storage systems, and data-intensive workflows.
Sustainable Performance
Reproducible, explainable, and energy-aware performance engineering.
Services
- Performance and scalability analysis of scientific applications
- Identification of computational, communication, I/O, and storage bottlenecks
- Support for code optimization, parallelization, and GPU porting
- Development of application-specific profiling and tuning workflows
- Consulting and training on performance-analysis tools and methods
Research
Tools
DataCrumbs
Cross-layer data capture for explainable I/O performance analysis.
Mango-IO
Analysis and comparison of I/O performance metrics across tools.
FBench
Exploration of HPC I/O patterns through transparent tuning and replay.
ML-supported Workflows
Interactive, automated, and tool-agnostic performance analysis.
Training Activities
- HPC I/O Performance Analysis Hands-on performance analysis using DFTracer and Darshan · November 2026
- All-Female Trainers Performance Tuning Workshop In cooperation with VI-HPS, POP, and the MAR WHPC Chapter · March 2027
- Optimization of Deep-Learning Applications Application optimization with CUDA and C++ · Q1 2027
- Multi-Vendor GPU Portability Portable GPU workflows using CUDA, ROCm, and PyTorch
- Introduction to GPU Profilers Performance analysis using ROCprofiler and NVIDIA Nsight
- Use of Ad Hoc File Systems Application-oriented use of temporary file systems for data-intensive workloads
Community Activities
REX-IO
Re-envisioning Extreme-Scale I/O for Emerging Hybrid HPC Workloads.
Visit the workshop website →PERMAVOST
Performance Engineering, Modelling, Analysis, and Visualization Strategy.
Visit the workshop website →DREAM
Data Reduction and Energy-Aware Data Movement.
Visit the workshop website →Talks & Community Sessions
Tutorials, minisymposia, invited talks, and NHR Sofa Talks on performance engineering, I/O, storage, and sustainable HPC.
Team
Lead PIs
Team Members
Publications
2026
- “Grammar-Based Workload Synthesis for What-If Exploration of HPC I/O”Accepted for the 33rd IEEE International Conference on High Performance Computing (HiPC 2026).
- “Compositional Energy Modeling of Layer-Resolved GPU Training”Accepted for the Sustainable Supercomputing Workshop at SC26.
- “DataCrumbs: Efficient Cross-Layer Data Capture for Explainable I/O Performance in HPC Storage Stacks”Accepted for the 8th Workshop on Programming and Performance Visualization Tools (ProTools 2026) at SC26.
- “Beyond Throughput: Cross-Vendor Kernel-Level Characterization of DNN on Modern GPUs”Accepted for the 38th IEEE/SBC International Symposium on Computer Architecture and High Performance Computing (SBAC-PAD 2026).
- “How Caching Distorts MPI Point-to-Point Performance”Accepted for the 33rd European MPI Users’ Group Meeting (EuroMPI 2026).
- “DPA-Store: An Ordered Network Data Path Key-Value Store”20th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2026, pp. 1513–1530. Publication
- “Towards Scalable Storage Architectures for GPU Clusters Running Large Language Models”Proceedings of the ACM on Measurement and Analysis of Computing Systems, vol. 10, no. 1, article 24, pp. 1–26, 2026. DOI
- “Aligning Storage Benchmark Metrics with Application-Level Performance”38th International Conference on Scalable Scientific Data Management, SSDBM 2026. DOI
- “Robust I/O Characterization of Machine Learning Workloads Across Performance Analysis Tools”REX-IO at the 35th International Symposium on High-Performance Parallel and Distributed Computing, HPDC 2026. DOI
- “Metis: Agentic Knowledge Synthesis for Explainable I/O Performance in HPC Systems”REX-IO at the 35th International Symposium on High-Performance Parallel and Distributed Computing, HPDC 2026. DOI
- “ETP4HPC SRA 6 White Paper — Energy Efficiency and Sustainability”European Technology Platform for High Performance Computing, Strategic Research Agenda 6, January 2026. DOI
- “Quantifying the Energy Cost of Performance Inefficiency in HPC Applications”Proceedings of the Supercomputing Asia and International Conference on High Performance Computing in Asia Pacific Region Workshops, SCA/HPCAsiaWS 2026, pp. 418–427. DOI
- “Exploring I/O Performance and Power Consumption Trade-offs in Production Environments”Proceedings of the Supercomputing Asia and International Conference on High Performance Computing in Asia Pacific Region Workshops, SCA/HPCAsiaWS 2026, pp. 399–406. DOI
2025
- “Quantifying AWS S3 I/O Performance Boundaries Using the Roofline Model”Proceedings of the SC ’25 Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis, pp. 1415–1423. DOI
- “Comprehensive Performance Analysis of Portals4 Communication Primitives on BXI Hardware”33rd International Symposium on Modeling, Analysis and Simulation of Computer and Telecommunication Systems, MASCOTS 2025, pp. 1–8. DOI
- “A Comparative Study of OpenMP Scheduling Algorithm Selection Strategies”IEEE Access, vol. 13, pp. 151216–151234, 2025. DOI
- “A Comparative Study of Ad-Hoc File Systems for Extreme Scale Computing”Future Generation Computer Systems, vol. 170, article 107815, 2025. DOI
- “Contenders: Predicting Cache Contention of Co-Scheduled Applications”25th IEEE International Symposium on Cluster, Cloud and Internet Computing, CCGrid 2025, pp. 343–352. DOI
- “No Time to Halt: In-Situ Analysis for Large-Scale Data Processing via Virtual Snapshotting”28th International Conference on Extending Database Technology, EDBT 2025, pp. 438–450. DOI
- “Enhancing mmap Scalability by Saving TLB Shootdowns During Page Recycling”33rd Euromicro International Conference on Parallel, Distributed, and Network-Based Processing, PDP 2025, pp. 112–120. DOI
- “Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation”IEEE International Conference on Cluster Computing Workshops, CLUSTER Workshops 2025, pp. 1–8. DOI
- “Benchmarking Darshan and Recorder for HPC I/O Profiling and Tracing”IEEE International Conference on Cluster Computing Workshops, CLUSTER Workshops 2025, pp. 1–6. DOI
- “XIO: Toward eXplainable I/O for HPC Systems”37th International Conference on Scalable Scientific Data Management, SSDBM 2025, article 23, pp. 1–6. DOI
- “Advancing HPC Performance Modeling with an Interactive, Automated and Tool-Agnostic ML-Driven Workflow”37th International Conference on Scalable Scientific Data Management, SSDBM 2025, article 16, pp. 1–6. DOI
- “Streamlining HDF5’s AI Workloads Benchmarking”IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2025, pp. 722–730. DOI
- “ETP4HPC SRA 6 White Paper — I/O and Storage”European Technology Platform for High Performance Computing, Strategic Research Agenda 6, January 2025. DOI
2024
- “Challenges in Understanding Metadata Performance: A Case of Metadata Analysis Using Score-P”IEEE International Conference on Cluster Computing Workshops, CLUSTER Workshops 2024, pp. 113–117. DOI
- “Comparability and Reproducibility in HPC Applications’ Energy Consumption Characterization”15th ACM International Conference on Future and Sustainable Energy Systems, e-Energy 2024, pp. 560–568. DOI
- “Automated Network Performance Characterization for HPC Systems”International Journal of Networking and Computing, vol. 14, no. 1, pp. 2–25. DOI
- “Malleability in Modern HPC Systems: Current Experiences, Challenges, and Future Opportunities”IEEE Transactions on Parallel and Distributed Systems, vol. 35, no. 9, pp. 1551–1564, 2024. DOI
- “A Hierarchical Modeling Approach for Assessing the Reliability and Performability of Burst Buffers”Architecture of Computing Systems, ARCS 2024, Lecture Notes in Computer Science, vol. 14842, pp. 266–281. DOI
- “IO-SEA: Storage I/O and Data Management for Exascale Architectures”21st ACM International Conference on Computing Frontiers, Workshops and Special Sessions, CF 2024. DOI
- “Combining Buffered I/O and Direct I/O in Distributed File Systems”22nd USENIX Conference on File and Storage Technologies, FAST 2024, pp. 17–33. Publication
2023
- “Characterization of Large-Scale HPC Workloads with Non-Naïve I/O Roofline Modeling and Scoring”29th IEEE International Conference on Parallel and Distributed Systems, ICPADS 2023, pp. 737–744. DOI
- “Toward Reproducible Benchmarking of PGAS and MPI Communication Schemes”29th IEEE International Conference on Parallel and Distributed Systems, ICPADS 2023, pp. 1959–1967. DOI
- “How Do OS and Application Schedulers Interact? An Investigation with Multithreaded Applications”Euro-Par 2023: Parallel Processing, pp. 214–228. DOI
- “Automated Scheduling Algorithm Selection in OpenMP”22nd International Symposium on Parallel and Distributed Computing, ISPDC 2023, pp. 106–109. DOI
- “Mango-IO: I/O Metrics Consistency Analysis”IEEE International Conference on Cluster Computing Workshops, CLUSTER Workshops 2023, pp. 18–24. DOI
- “Code Modernization Strategies for Short-Range Non-Bonded Molecular Dynamics Simulations”Computer Physics Communications, vol. 290, article 108760, 2023. DOI
- “Adaptive Multi-Tier Intelligent Data Manager for Exascale”20th ACM International Conference on Computing Frontiers, CF 2023, pp. 285–290. DOI
- “Xfast: Extreme File Attribute Stat Acceleration for Lustre”International Conference for High Performance Computing, Networking, Storage and Analysis, SC 2023, article 96, pp. 1–12. DOI
- “The I/O Trace Initiative: Building a Collaborative I/O Archive to Advance HPC”Proceedings of the SC ’23 Workshops of the International Conference on High Performance Computing, Network, Storage, and Analysis, pp. 1216–1222. DOI
- “From Static to Malleable: Improving Flexibility and Compatibility in Burst Buffer File Systems”ISC High Performance 2023 International Workshops, pp. 3–15. DOI