Artificial intelligence and machine learning are becoming essential tools across scientific disciplines. Running these methods effectively on HPC systems, however, requires more than selecting a model: training strategies, parallelism, software environments, reproducibility, energy efficiency, and the underlying system architecture all need to work together.
The Method Lab AI and ML supports the NHR community in analyzing, optimizing, and adapting AI/ML methods for HPC-scale workloads. Its expertise ranges from data- and model-parallel training to framework deployment, reproducible workflows, containerization, and stable operation on NHR systems.
A particular focus lies on explainable and trustworthy AI as well as hybrid and neuro-symbolic methods. The Method Lab helps researchers visualize and debug complex models, assess robustness and scientific interpretability, and combine domain-specific models, physical knowledge, simulations, and data-driven learning.
Through individual consulting, collaborative projects, practical training, and hands-on workshops, the Method Lab enables researchers to select and use modern AI technologies efficiently, sustainably, and responsibly.
Competencies
Scaling and Efficiency
Adapting AI/ML workloads to HPC scaling models, including data-parallel and model-parallel training and tightly coupled processes.
Frameworks and MLOps
Deploying and operating established frameworks through reproducible, containerized, and versioned training workflows.
Explainable and Trustworthy AI
Visualizing, debugging, and explaining AI systems to improve transparency, robustness, and scientific interpretability.
Hybrid and Neuro-symbolic AI
Combining explicit domain models, physical knowledge, and simulations with data-driven neural methods.
Scientific Machine Learning
Applying machine learning, differentiable programming, and algorithmic differentiation to scientific simulations and experiments.
Responsible Knowledge Transfer
Helping interdisciplinary research communities evaluate and use modern AI/ML methods independently and sustainably.
Services
- Individual consulting on suitable AI/ML methods, frameworks, and HPC resources
- Support for scalable data- and model-parallel training
- Framework deployment, containerization, and reproducible training pipelines
- Model visualization, debugging, explainability, and robustness analysis
- Support for hybrid, neuro-symbolic, and physics-informed approaches
- Hands-on training and collaborative AI/ML projects
Research
Supported Tools
PyTorch
Deep-learning development and training on HPC systems.
TensorFlow
Framework support for scalable machine-learning workflows.
JAX
High-performance numerical computing and differentiable programming.
CoDiPack
Algorithmic differentiation for scientific and engineering codes.
Training Activities
- Training Large Language Models on HPC SystemsHands-on development and scaling of LLM training workflows on HPC infrastructure
- Reproducible LLM EvaluationReliable evaluation across models, datasets, prompts, frameworks, and hardware configurations
- Explainability in PracticePractical methods for visualizing, interpreting, and debugging machine-learning models
- Introduction to In-Context LearningUsing demonstrations and prompts to adapt and evaluate large language models
- Introduction to PyTorchCore concepts and practical building blocks for scientific deep-learning applications
- Algorithmic DifferentiationHands-on introduction to differentiable scientific computing and tools such as CoDiPack
- Fairness and Responsible AIAssessing fairness, robustness, and responsible behavior in AI systems
Community Activities
2nd International Conference on AI for Science
Co-organization of the AI for Science Week 2026 in Mainz.
MODE Workshop
Contributions on differentiable particle tracking, electromagnetic shower simulations, and differentiable silicon-pixel detector modeling.
Next-Generation Particle Detectors
Hands-on machine-learning training for Heidelberg University’s graduate school in Bergen.
CERN DRD6 Collaboration
Contribution to the Collaboration on Calorimeter Design; Nicolas R. Gauger serves on its Collaboration Board.
Team
Lead PIs
Team Members
Selected Publications
2026
- Constrained Collaborative Optimization of Charged Particle Tracking with Multi-Agent Reinforcement LearningMachine Learning: Science and Technology, 2026
- Robust I/O Characterization of Machine Learning Workloads Across Performance Analysis ToolsProceedings of the 35th International Symposium on High-Performance Parallel and Distributed Computing, REX-IO session, 2026 · DOI
- Beyond Throughput: Cross-Vendor Kernel-Level Characterization of DNN on Modern GPUs38th IEEE/SBC International Symposium on Computer Architecture and High Performance Computing, 2026
2025
- Exploring End-to-End Differentiable Neural Charged Particle Tracking: A Loss Landscape PerspectiveTransactions on Machine Learning Research, 2025
- Streamlining HDF5’s AI Workloads Benchmarking2025 IEEE International Parallel and Distributed Processing Symposium Workshops, pp. 722–730 · DOI
- Modeling Charge Collection in Silicon Pixel Detectors for Proton Therapy ApplicationsBiomedical Physics & Engineering Express, 2025
- Maximizing Insights, Minimizing Data: I/O Time Prediction Using Transfer Learning2025 IEEE 32nd International Conference on High Performance Computing, Data, and Analytics · DOI
2024
- Optimization Using Pathwise Algorithmic Derivatives of Electromagnetic Shower SimulationsarXiv:2405.07944, 2024
- Efficient Forward-Mode Algorithmic Derivatives of Geant4arXiv:2407.02966, 2024