AI and ML

NHR@SW Method Lab

To the overview of the Method Labs →

Artificial intelligence and machine learning are becoming essential tools across scientific disciplines. Running these methods effectively on HPC systems, however, requires more than selecting a model: training strategies, parallelism, software environments, reproducibility, energy efficiency, and the underlying system architecture all need to work together.

The Method Lab AI and ML supports the NHR community in analyzing, optimizing, and adapting AI/ML methods for HPC-scale workloads. Its expertise ranges from data- and model-parallel training to framework deployment, reproducible workflows, containerization, and stable operation on NHR systems.

A particular focus lies on explainable and trustworthy AI as well as hybrid and neuro-symbolic methods. The Method Lab helps researchers visualize and debug complex models, assess robustness and scientific interpretability, and combine domain-specific models, physical knowledge, simulations, and data-driven learning.

Through individual consulting, collaborative projects, practical training, and hands-on workshops, the Method Lab enables researchers to select and use modern AI technologies efficiently, sustainably, and responsibly.

Scalable AIScientific MLTrustworthy AIHybrid AIReproducible Workflows
01

Competencies

Scaling and Efficiency

Adapting AI/ML workloads to HPC scaling models, including data-parallel and model-parallel training and tightly coupled processes.

Frameworks and MLOps

Deploying and operating established frameworks through reproducible, containerized, and versioned training workflows.

Explainable and Trustworthy AI

Visualizing, debugging, and explaining AI systems to improve transparency, robustness, and scientific interpretability.

Hybrid and Neuro-symbolic AI

Combining explicit domain models, physical knowledge, and simulations with data-driven neural methods.

Scientific Machine Learning

Applying machine learning, differentiable programming, and algorithmic differentiation to scientific simulations and experiments.

Responsible Knowledge Transfer

Helping interdisciplinary research communities evaluate and use modern AI/ML methods independently and sustainably.

02

Services

  • Individual consulting on suitable AI/ML methods, frameworks, and HPC resources
  • Support for scalable data- and model-parallel training
  • Framework deployment, containerization, and reproducible training pipelines
  • Model visualization, debugging, explainability, and robustness analysis
  • Support for hybrid, neuro-symbolic, and physics-informed approaches
  • Hands-on training and collaborative AI/ML projects
03

Research

  • Distributed and Scalable Training
  • Energy-Efficient AI on HPC
  • Scientific Machine Learning
  • Differentiable Programming
  • Algorithmic Differentiation
  • Explainable and Trustworthy AI
  • Hybrid and Neuro-symbolic AI
  • Large Language Models
  • Multilingual and Low-Resource AI
  • Fairness and Robust Refusal
  • AI Workload I/O
  • Transfer Learning for Performance Prediction
04

Supported Tools

Deep Learning

PyTorch

Deep-learning development and training on HPC systems.

Deep Learning

TensorFlow

Framework support for scalable machine-learning workflows.

Differentiable Computing

JAX

High-performance numerical computing and differentiable programming.

Algorithmic Differentiation

CoDiPack

Algorithmic differentiation for scientific and engineering codes.

Workflow support: Containers · Versioned training pipelines · Weights & Biases
05

Training Activities

  • Training Large Language Models on HPC SystemsHands-on development and scaling of LLM training workflows on HPC infrastructure
  • Reproducible LLM EvaluationReliable evaluation across models, datasets, prompts, frameworks, and hardware configurations
  • Explainability in PracticePractical methods for visualizing, interpreting, and debugging machine-learning models
  • Introduction to In-Context LearningUsing demonstrations and prompts to adapt and evaluate large language models
  • Introduction to PyTorchCore concepts and practical building blocks for scientific deep-learning applications
  • Algorithmic DifferentiationHands-on introduction to differentiable scientific computing and tools such as CoDiPack
  • Fairness and Responsible AIAssessing fairness, robustness, and responsible behavior in AI systems
06

Community Activities

2nd International Conference on AI for Science

Co-organization of the AI for Science Week 2026 in Mainz.

MODE Workshop

Contributions on differentiable particle tracking, electromagnetic shower simulations, and differentiable silicon-pixel detector modeling.

Next-Generation Particle Detectors

Hands-on machine-learning training for Heidelberg University’s graduate school in Bergen.

CERN DRD6 Collaboration

Contribution to the Collaboration on Calorimeter Design; Nicolas R. Gauger serves on its Collaboration Board.

07

Team

Lead PIs

NGProf. Dr. Nicolas GaugerRPTU Kaiserslautern-Landau
DKProf. Dr. Dietrich KlakowSaarland University
SNProf. Dr. Sarah NeuwirthJGU Mainz
PSProf. Dr. Philipp SlusallekSaarland University / DFKI
SKProf. Dr. Michael WandJGU Mainz

Team Members

IAIsrael Abebe AzimeSaarland University
TKDr. Alexander BlattSaarland University
TKTobias KortusRPTU Kaiserslautern-Landau
RLRadita LiemJGU Mainz
EÖEmre ÖzkayaRPTU Kaiserslautern-Landau
MSMax SagebaumRPTU Kaiserslautern-Landau
ASAlexander SchillingRPTU Kaiserslautern-Landau
AVAmritanshu VermaJGU Mainz
08

Selected Publications

2026

  • T. Kortus, R. Keidel, N. R. Gauger, J. Kieseler, and the Bergen pCT CollaborationConstrained Collaborative Optimization of Charged Particle Tracking with Multi-Agent Reinforcement LearningMachine Learning: Science and Technology, 2026
  • Z. Masih, R. Liem, and J. KunkelRobust I/O Characterization of Machine Learning Workloads Across Performance Analysis ToolsProceedings of the 35th International Symposium on High-Performance Parallel and Distributed Computing, REX-IO session, 2026 · DOI
  • A. Verma and S. NeuwirthBeyond Throughput: Cross-Vendor Kernel-Level Characterization of DNN on Modern GPUs38th IEEE/SBC International Symposium on Computer Architecture and High Performance Computing, 2026

2025

  • T. Kortus, R. Keidel, N. R. Gauger, and the Bergen pCT CollaborationExploring End-to-End Differentiable Neural Charged Particle Tracking: A Loss Landscape PerspectiveTransactions on Machine Learning Research, 2025
  • D. Djebarov, R. Liem, S. Neuwirth, J. L. Bez, and S. BynaStreamlining HDF5’s AI Workloads Benchmarking2025 IEEE International Parallel and Distributed Processing Symposium Workshops, pp. 722–730 · DOI
  • A. Schilling, M. Aehle, J. Alme, G. G. Barnaföldi, G. Bíró, T. Bodova, V. Borshchov, A. van den Brink, V. Eikeland, G. Feofilov, C. Garth, N. R. Gauger, O. Grøttvik, H. Helstrup, S. Igolkin, J. G. Johansen, R. Keidel, C. Kobdaj, T. Kortus, et al.Modeling Charge Collection in Silicon Pixel Detectors for Proton Therapy ApplicationsBiomedical Physics & Engineering Express, 2025
  • A. Voß, R. Liem, J. Kunkel, J. Lofstead, P. Carns, and M. MüllerMaximizing Insights, Minimizing Data: I/O Time Prediction Using Transfer Learning2025 IEEE 32nd International Conference on High Performance Computing, Data, and Analytics · DOI

2024

  • M. Aehle, M. Novák, V. Vassilev, N. R. Gauger, L. Heinrich, M. Kagan, and D. LangeOptimization Using Pathwise Algorithmic Derivatives of Electromagnetic Shower SimulationsarXiv:2405.07944, 2024
  • M. Aehle, X. T. Nguyen, M. Novák, T. Dorigo, N. R. Gauger, J. Kieseler, M. Klute, and V. VassilevEfficient Forward-Mode Algorithmic Derivatives of Geant4arXiv:2407.02966, 2024