Performance Engineering

NHR@SW Method Lab

To the overview of the Method Labs →

Modern HPC systems combine increasingly heterogeneous processors, accelerators, networks, and storage technologies. Achieving good application performance therefore requires more than optimizing individual code sections: computation, communication, data movement, I/O, and energy consumption must be considered across the complete software and system stack.

The Method Lab Performance Engineering supports the NHR community in systematically analyzing, modeling, and optimizing scientific applications on modern HPC systems. Its expertise ranges from code optimization, parallelization, and scalability analysis to GPU porting, GPU performance analysis, parallel I/O, storage, and energy-efficient computing.

A particular focus lies on data-intensive simulations and AI/ML workloads. The Method Lab helps identify I/O bottlenecks, performance variability, inefficient data movement, and resource limitations. It develops and applies suitable tools, metrics, and tuning strategies to make system behavior observable, explainable, and reproducible.

Through individual consulting, collaborative optimization projects, hands-on training, and community workshops, the Method Lab enables researchers to use current and emerging HPC infrastructures more effectively.

Compute Communication GPU I/O & Storage Energy
01

Competencies

Performance Analysis

Analysis, modeling, benchmarking, and optimization of scientific applications.

Parallelism & Scalability

Parallelization, communication analysis, and scalable execution.

GPU Performance

GPU porting, profiling, optimization, and performance portability.

Data Movement

Parallel I/O, storage systems, and data-intensive workflows.

Sustainable Performance

Reproducible, explainable, and energy-aware performance engineering.

02

Services

  • Performance and scalability analysis of scientific applications
  • Identification of computational, communication, I/O, and storage bottlenecks
  • Support for code optimization, parallelization, and GPU porting
  • Development of application-specific profiling and tuning workflows
  • Consulting and training on performance-analysis tools and methods
03

Research

  • Holistic performance engineering
  • Performance modeling
  • Reproducible benchmarking
  • Parallel I/O and storage
  • GPU performance
  • Performance portability
  • Explainable performance analysis
  • Energy and resource efficiency
  • AI/ML workloads
04

Tools

I/O observability

DataCrumbs

Cross-layer data capture for explainable I/O performance analysis.

Metric consistency

Mango-IO

Analysis and comparison of I/O performance metrics across tools.

What-if analysis

FBench

Exploration of HPC I/O patterns through transparent tuning and replay.

Performance modeling

ML-supported Workflows

Interactive, automated, and tool-agnostic performance analysis.

Additional tools: Score-P · Darshan · Recorder · DFTracer · NVIDIA Nsight · ROCm profiling tools
05

Training Activities

  • HPC I/O Performance Analysis Hands-on performance analysis using DFTracer and Darshan · November 2026
  • All-Female Trainers Performance Tuning Workshop In cooperation with VI-HPS, POP, and the MAR WHPC Chapter · March 2027
  • Optimization of Deep-Learning Applications Application optimization with CUDA and C++ · Q1 2027
  • Multi-Vendor GPU Portability Portable GPU workflows using CUDA, ROCm, and PyTorch
  • Introduction to GPU Profilers Performance analysis using ROCprofiler and NVIDIA Nsight
  • Use of Ad Hoc File Systems Application-oriented use of temporary file systems for data-intensive workloads
06

Community Activities

07

Team

Lead PIs

Prof. Dr. André Brinkmann Saarland University
Prof. Dr. Sarah Neuwirth JGU Mainz · Speaker

Team Members

Julius Athenstaedt JGU Mainz
Niklas Bartelheimer JGU Mainz
Julius Dörhöfer JGU Mainz
Dr. Jonas Korndörfer Saarland University
Radita Liem JGU Mainz
Martin Machajewski JGU Mainz
Christian Sendlinger Goethe University Frankfurt
Amritanshu Verma JGU Mainz
Zhaobin Zhu JGU Mainz
08

Publications

2026

  • Z. Zhu, C. Wang, K. Mohror, and S. Neuwirth“Grammar-Based Workload Synthesis for What-If Exploration of HPC I/O”Accepted for the 33rd IEEE International Conference on High Performance Computing (HiPC 2026).
  • A. Verma and S. Neuwirth“Compositional Energy Modeling of Layer-Resolved GPU Training”Accepted for the Sustainable Supercomputing Workshop at SC26.
  • H. Devarajan, S. Neuwirth, C. Wang, R. Sinurat, N. Rajesh, J. Lofstead, V. Sochat, D. Milroy, and T. Scogland“DataCrumbs: Efficient Cross-Layer Data Capture for Explainable I/O Performance in HPC Storage Stacks”Accepted for the 8th Workshop on Programming and Performance Visualization Tools (ProTools 2026) at SC26.
  • A. Verma and S. Neuwirth“Beyond Throughput: Cross-Vendor Kernel-Level Characterization of DNN on Modern GPUs”Accepted for the 38th IEEE/SBC International Symposium on Computer Architecture and High Performance Computing (SBAC-PAD 2026).
  • N. Bartelheimer, J. Domke, and S. Neuwirth“How Caching Distorts MPI Point-to-Point Performance”Accepted for the 33rd European MPI Users’ Group Meeting (EuroMPI 2026).
  • F. Schimmelpfennig, J. Sass, R. Salkhordeh, M. Kröning, S. Lankes, and A. Brinkmann“DPA-Store: An Ordered Network Data Path Key-Value Store”20th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2026, pp. 1513–1530. Publication
  • A. Sedaghatgoo, R. Salkhordeh, A. Brinkmann, and H. Asadi“Towards Scalable Storage Architectures for GPU Clusters Running Large Language Models”Proceedings of the ACM on Measurement and Analysis of Computing Systems, vol. 10, no. 1, article 24, pp. 1–26, 2026. DOI
  • R. Liem, K. Nguyen, J. Kunkel, J. Lofstead, and S. Neuwirth“Aligning Storage Benchmark Metrics with Application-Level Performance”38th International Conference on Scalable Scientific Data Management, SSDBM 2026. DOI
  • Z. Masih, R. Liem, and J. Kunkel“Robust I/O Characterization of Machine Learning Workloads Across Performance Analysis Tools”REX-IO at the 35th International Symposium on High-Performance Parallel and Distributed Computing, HPDC 2026. DOI
  • K. Youssef, S. Neuwirth, N. Rajesh, and H. Devarajan“Metis: Agentic Knowledge Synthesis for Explainable I/O Performance in HPC Systems”REX-IO at the 35th International Symposium on High-Performance Parallel and Distributed Computing, HPDC 2026. DOI
  • A. Wierse, P. Seroul, O. Vysocký, M. Barth, M. Chinnici, G. Da Costa, E. Heiskanen, S. Neuwirth, F. Pariente, G. Svensson, and C. Vercellino“ETP4HPC SRA 6 White Paper — Energy Efficiency and Sustainability”European Technology Platform for High Performance Computing, Strategic Research Agenda 6, January 2026. DOI
  • R. Liem and D. Djebarov“Quantifying the Energy Cost of Performance Inefficiency in HPC Applications”Proceedings of the Supercomputing Asia and International Conference on High Performance Computing in Asia Pacific Region Workshops, SCA/HPCAsiaWS 2026, pp. 418–427. DOI
  • Z. Zhu, A. Henkel, and S. Neuwirth“Exploring I/O Performance and Power Consumption Trade-offs in Production Environments”Proceedings of the Supercomputing Asia and International Conference on High Performance Computing in Asia Pacific Region Workshops, SCA/HPCAsiaWS 2026, pp. 399–406. DOI

2025

  • M. Tang, Z. Zhu, L. Guo, J. G. Bandy, T. Carlson, S. Neuwirth, A. Kougkas, X.-H. Sun, and N. R. Tallent“Quantifying AWS S3 I/O Performance Boundaries Using the Roofline Model”Proceedings of the SC ’25 Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis, pp. 1415–1423. DOI
  • N. J. Bartelheimer and S. M. Neuwirth“Comprehensive Performance Analysis of Portals4 Communication Primitives on BXI Hardware”33rd International Symposium on Modeling, Analysis and Simulation of Computer and Telecommunication Systems, MASCOTS 2025, pp. 1–8. DOI
  • J. H. Müller Korndörfer, A. Mohammed, A. Eleliemy, Q. Guilloteau, R. Krummenacher, and F. M. Ciorba“A Comparative Study of OpenMP Scheduling Algorithm Selection Strategies”IEEE Access, vol. 13, pp. 151216–151234, 2025. DOI
  • N. O. Almaaitah, F. J. García Blas, G. Sanchez-Gallegos, J. Carretero, M.-A. Vef, and A. Brinkmann“A Comparative Study of Ad-Hoc File Systems for Extreme Scale Computing”Future Generation Computer Systems, vol. 170, article 107815, 2025. DOI
  • Z. Wang, T. Süß, A. Brinkmann, and L. Nagel“Contenders: Predicting Cache Contention of Co-Scheduled Applications”25th IEEE International Symposium on Cluster, Cloud and Internet Computing, CCGrid 2025, pp. 343–352. DOI
  • R. Salkhordeh, F. M. Schuhknecht, H. Asadi, S. Eiden, and A. Brinkmann“No Time to Halt: In-Situ Analysis for Large-Scale Data Processing via Virtual Snapshotting”28th International Conference on Extending Database Technology, EDBT 2025, pp. 438–450. DOI
  • F. Schimmelpfennig, A. Brinkmann, H. Asadi, and R. Salkhordeh“Enhancing mmap Scalability by Saving TLB Shootdowns During Page Recycling”33rd Euromicro International Conference on Parallel, Distributed, and Network-Based Processing, PDP 2025, pp. 112–120. DOI
  • H. Ahmad, R. Liem, and J. Lofstead“Towards an Optimal IO500 Configuration: Literature Meets Empirical Evaluation”IEEE International Conference on Cluster Computing Workshops, CLUSTER Workshops 2025, pp. 1–8. DOI
  • Z. Zhu, L. Derstroff, and S. Neuwirth“Benchmarking Darshan and Recorder for HPC I/O Profiling and Tracing”IEEE International Conference on Cluster Computing Workshops, CLUSTER Workshops 2025, pp. 1–6. DOI
  • S. Neuwirth, H. Devarajan, C. Wang, and J. Lofstead“XIO: Toward eXplainable I/O for HPC Systems”37th International Conference on Scalable Scientific Data Management, SSDBM 2025, article 23, pp. 1–6. DOI
  • Z. Zhu, C. Wang, and S. Neuwirth“Advancing HPC Performance Modeling with an Interactive, Automated and Tool-Agnostic ML-Driven Workflow”37th International Conference on Scalable Scientific Data Management, SSDBM 2025, article 16, pp. 1–6. DOI
  • D. Djebarov, R. Liem, S. Neuwirth, J. L. Bez, and S. Byna“Streamlining HDF5’s AI Workloads Benchmarking”IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2025, pp. 722–730. DOI
  • S. Neuwirth, P. Deniel, J.-T. Acquaviva, M. Golasowski, M. Hennecke, A. Jackson, T. Leibovici, J. Luettgau, and R. Nou“ETP4HPC SRA 6 White Paper — I/O and Storage”European Technology Platform for High Performance Computing, Strategic Research Agenda 6, January 2025. DOI

2024

  • B. Kosmynin and R. Liem“Challenges in Understanding Metadata Performance: A Case of Metadata Analysis Using Score-P”IEEE International Conference on Cluster Computing Workshops, CLUSTER Workshops 2024, pp. 113–117. DOI
  • T. Hilgers and R. Liem“Comparability and Reproducibility in HPC Applications’ Energy Consumption Characterization”15th ACM International Conference on Future and Sustainable Energy Systems, e-Energy 2024, pp. 560–568. DOI
  • N. Bartelheimer, Z. Zhu, and S. Neuwirth“Automated Network Performance Characterization for HPC Systems”International Journal of Networking and Computing, vol. 14, no. 1, pp. 2–25. DOI
  • A. Tarraf, M. Schreiber, A. Cascajo, J.-B. Besnard, M.-A. Vef, D. Huber, S. Happ, A. Brinkmann, D. E. Singh, H.-C. Hoppe, A. Miranda, A. J. Peña, R. Machado, M. Garcia-Gasulla, M. Schulz, P. M. Carpenter, S. Pickartz, T. Rotaru, S. Iserte, V. López, J. Ejarque, H. Sirwani, J. Carretero, and F. Wolf“Malleability in Modern HPC Systems: Current Experiences, Challenges, and Future Opportunities”IEEE Transactions on Parallel and Distributed Systems, vol. 35, no. 9, pp. 1551–1564, 2024. DOI
  • E. Borba, R. Salkhordeh, S. Mimouni, E. Tavares, P. Maciel, H. Asadi, and A. Brinkmann“A Hierarchical Modeling Approach for Assessing the Reliability and Performability of Burst Buffers”Architecture of Computing Systems, ARCS 2024, Lecture Notes in Computer Science, vol. 14842, pp. 266–281. DOI
  • D. Medeiros, E. B. Gregory, P. Couvée, J. N. Hawkes, S. Gougeaud, M. Gilliot, O. Bressand, Y. Valeri, J. Jaeger, D. Chapon, F. Bournaud, L. Strafella, D. Caviedes-Voullième, G. Tashakor, J. Zjupa, M. Holicki, T. Ridley, Y. Müller, F. S. M. Guimarães, W. Frings, J.-O. Mirus, I. Zhukov, E. R. Borba, N. Moti, R. Salkhordeh, N. Derbey, S. Mimouni, S. Derr, B. B. Gursoy, J. A. Grogan, R. Furmánek, M. Golasowski, K. Slaninová, J. Martinovic, J. Faltýnek, J. Wong, M. Cakircali, T. Quintino, S. D. Smart, O. Iffrig, S. Narasimhamurthy, S. Happ, M. Rauh, S. Krempel, M. Wiggins, J. Nováček, A. Brinkmann, S. Markidis, and P. Deniel“IO-SEA: Storage I/O and Data Management for Exascale Architectures”21st ACM International Conference on Computing Frontiers, Workshops and Special Sessions, CF 2024. DOI
  • Y. Qian, M.-A. Vef, P. Farrell, A. Dilger, X. Li, S. Ihara, Y. Fu, W. Xue, and A. Brinkmann“Combining Buffered I/O and Direct I/O in Distributed File Systems”22nd USENIX Conference on File and Storage Technologies, FAST 2024, pp. 17–33. Publication

2023

  • Z. Zhu and S. Neuwirth“Characterization of Large-Scale HPC Workloads with Non-Naïve I/O Roofline Modeling and Scoring”29th IEEE International Conference on Parallel and Distributed Systems, ICPADS 2023, pp. 737–744. DOI
  • N. Bartelheimer and S. Neuwirth“Toward Reproducible Benchmarking of PGAS and MPI Communication Schemes”29th IEEE International Conference on Parallel and Distributed Systems, ICPADS 2023, pp. 1959–1967. DOI
  • J. H. Müller Korndörfer, A. Eleliemy, O. S. Simsek, T. Ilsche, R. Schöne, and F. M. Ciorba“How Do OS and Application Schedulers Interact? An Investigation with Multithreaded Applications”Euro-Par 2023: Parallel Processing, pp. 214–228. DOI
  • F. M. Ciorba, A. Mohammed, J. H. Müller Korndörfer, and A. Eleliemy“Automated Scheduling Algorithm Selection in OpenMP”22nd International Symposium on Parallel and Distributed Computing, ISPDC 2023, pp. 106–109. DOI
  • R. Liem, S. Oeste, J. Lofstead, and J. Kunkel“Mango-IO: I/O Metrics Consistency Analysis”IEEE International Conference on Cluster Computing Workshops, CLUSTER Workshops 2023, pp. 18–24. DOI
  • J. Vance, Z.-H. Xu, N. Tretyakov, T. Stuehn, M. Rampp, S. Eibl, C. Junghans, and A. Brinkmann“Code Modernization Strategies for Short-Range Non-Bonded Molecular Dynamics Simulations”Computer Physics Communications, vol. 290, article 108760, 2023. DOI
  • J. Carretero, J. García-Blas, M. Aldinucci, J.-B. Besnard, J.-T. Acquaviva, A. Brinkmann, M.-A. Vef, E. Jeannot, A. Miranda, R. Nou, M. Riedel, M. Torquati, and F. Wolf“Adaptive Multi-Tier Intelligent Data Manager for Exascale”20th ACM International Conference on Computing Frontiers, CF 2023, pp. 285–290. DOI
  • Y. Qian, W. Cheng, L. Zeng, X. Li, M.-A. Vef, A. Dilger, S. Lai, S. Ihara, Y. Fan, and A. Brinkmann“Xfast: Extreme File Attribute Stat Acceleration for Lustre”International Conference for High Performance Computing, Networking, Storage and Analysis, SC 2023, article 96, pp. 1–12. DOI
  • N. Moti, A. Brinkmann, M.-A. Vef, P. Deniel, J. Carretero, P. H. Carns, J.-T. Acquaviva, and R. Salkhordeh“The I/O Trace Initiative: Building a Collaborative I/O Archive to Advance HPC”Proceedings of the SC ’23 Workshops of the International Conference on High Performance Computing, Network, Storage, and Analysis, pp. 1216–1222. DOI
  • M.-A. Vef, A. Miranda, R. Nou, and A. Brinkmann“From Static to Malleable: Improving Flexibility and Compatibility in Burst Buffer File Systems”ISC High Performance 2023 International Workshops, pp. 3–15. DOI