Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
|
Daily Overview |
| Session | ||
Session 4b: Infrastructure & Tools
Session Topics: Agentic AI, Foundation Models, Generative Models, Graph Neural Networks, Physics-informed Machine Learning, Reinforcement Learning, Probabilistic Methods, Uncertainty Quantification, Audio, Other, Graphs, Image, Multimodal Data, Simulation Data, Tabular Data, Text, Time Series, Video, Other, Core Machine Learning, Aeronautics, Space & Transport, Energy, Earth & Environment, Health, Information, Matter
| ||
| Session Abstract | ||
|
Choose from expert-led talks running simultaneously to explore AI topics that match your interests. | ||
| Presentations | ||
9:00am - 9:12am
ID: 1161 / Thu | GERN 9h Parallel S 4b: 001 Modalities: Image Methods: Other Application Domain: Health AI-assisted Labeling and its Pitfalls: A Case Study in Electron Microscopy Segmentation 1: Helmholtz AI, Helmholtz Center Munich, Germany; 2: Institute of Toxicology and Environmental Hygiene, TUM School of Medicine and Health, Technical University of Munich, Germany; 3: Institute of Molecular Toxicology and Pharmacology, Helmholtz Center Munich, Germany High-quality labels are of utmost importance in biomedicine and instrumental for tasks such as disease understanding, drug discovery, and medical diagnosis. Human-in-the-loop approaches and AI-assisted annotations are currently widely accepted as the safest approach for data curation, offering both speed and human supervision, while lever-aging the power of AI. Indeed, foundation models are often used to ob-tain initial labels which are then manually corrected by a human expert and subsequently used for downstream tasks. In this work, we uncover a rarely discussed risk of such semi-automated approaches and present a case study to demonstrate this. Focusing on microscopy imaging seg-mentation, where annotation quality is critical and costly, we collected and annotated a Transmission Electron Microscopy dataset in a semi-automated approach, using BATS, a software we built for AI-assisted segmentation labeling in microscopy which encourages data centric prac-tices. We show how batch effects can be derived by a mix of human and modelannotationsaffectingthedownstreamanalysis.Finally, wedemon-strate a solution which mitigates these batch effects and discuss how future researchers can adopt responsible and accurate data annotation pipelines.
9:12am - 9:24am
ID: 149 / Thu | GERN 9h Parallel S 4b: 002 Modalities: Image Methods: Other Application Domain: Health A decentralized Swarm Learning framework for 90-Day outcome prediction for acute ischaemic stroke 1: DZNE, Germany; 2: CISPA, Germany Acute ischaemic stroke remains the predominant cause of disability and a major contributor to mortality globally. The damage caused by ischemia, triggered by vascular occlusion, progresses rapidly, with brain tissues beginning necrosis within minutes. This necessitates urgent clinical decision-making. This is particularly important for reperfusion therapies such as intravenous thrombolysis and mechanical thrombectomy. The efficacy of these treatments is time-sensitive and associated with risk of intracranial hemorrhage. In addition, treatment efficacy varies considerably across patients and depends upon several factors. This underscores the need for individualized risk stratification. In this project, we aim to develop a deep learning framework for predicting 90-day functional outcomes using multimodal data acquired at hospital admission. We are integrating heterogeneous data sources across 25 different hospitals within Germany (German Stroke Registry data), including clinical scores, patient history, and MRI images. Traditional centralized learning approaches face limitations related to data privacy, small sample sizes, and institutional barriers. Swarm Learning provides a privacy-preserving, decentralized alternative. However, a decentralized pipeline for multimodal MRI-based model development is currently lacking. To ensure robustness and generalizability, we will implement model training in a decentralized swarm learning manner. Our objective is to validate a multimodal deep learning model for individualized 90-day outcome prediction in a decentralized AI infrastructure. We aim to advance precision medicine for acute stroke care and establish a foundation for globally collaborative, privacy-preserving AI development. 9:24am - 9:36am
ID: 228 / Thu | GERN 9h Parallel S 4b: 003 Modalities: Graphs, Image, Time Series Methods: Graph Neural Networks Application Domain: Matter Microsecond Latency Graph Neural Network Inference on Point Clouds Karlsruhe Institute of Technology, Germany Graph Neural Networks are powerful machine learning techniques for processing sparse data with irregular geometries, as encountered in high-energy physics detectors. However, deploying such models within hardware triggers remains challenging due to stringent real-time constraints in terms of both latency and throughput. State-of-the-art hardware triggers in collider experiments impose hard latency deadlines on the order of 1 to 10 microseconds, necessitating the development of custom machine learning accelerators based on Field Programmable Gate Arrays. This work presents a deployment methodology for mapping Graph Neural Networks onto such platforms. By implementing commonly used neural network operators as reusable architecture templates, and leveraging model quantization and pruning, our approach achieves microsecond inference latencies. We demonstrate the methodology by deploying a Graph Neural Network based clustering algorithm for the Electromagnetic Calorimeter of the Belle II experiment. Through hardware-algorithm co-design, we achieve an end-to-end system latency of 1.050 microseconds, while preserving clustering quality, and meeting the real-time constraints required for hardware triggers. We validate our approach through cycle-accurate simulation and direct deployment on hardware, achieving complete agreement between simulation and measured results. Furthermore, we investigate the use of heterogeneous System-on-Chip architectures, such as AMD Versal platforms, as a path to deploy even larger neural network models. To conclude, this work establishes a deployment methodology for graph-based machine learning inference under extreme real-time constraints.
9:36am - 9:48am
ID: 146 / Thu | GERN 9h Parallel S 4b: 004 Modalities: Graphs, Simulation Data Methods: Agentic AI, Foundation Models, Generative Models, Graph Neural Networks, Physics-informed Machine Learning, Reinforcement Learning Application Domain: Core Machine Learning, Aeronautics, Space & Transport, Energy, Matter GENIUS: An Agentic AI Framework for Autonomous Design and Execution of Simulation Protocols 1: Karlsruhe Institute of Technology, Germany; 2: Helmholtz-Zentrum Hereon Atomistic simulations are at the forefront of materials discovery, yet their complex setup and debugging often require specialized expertise, limiting the widespread adoption of Integrated Computational Materials Engineering (ICME). To bridge this critical know-do gap, we introduce GENIUS, a novel AI-driven framework that autonomously designs and executes simulation protocols. GENIUS seamlessly integrates a smart knowledge graph tailored for Density Functional Theory (DFT) calculations with a hierarchical architecture of advanced Large Language Models (LLMs), supervised by a robust finite-state error-recovery machine. Focusing initially on DFT, GENIUS translates human-generated prompts into validated input files, achieving successful execution on a diverse set of 295 benchmarks. A key strength of GENIUS is its autonomous error handling, which repairs errors in 76% of failed runs, significantly boosting reliability. Compared to LLM-only baselines, GENIUS halves inference costs and virtually eliminates the 'hallucinations' that can lead to incorrect results. By intelligently automating protocol generation, validation, and repair, GENIUS democratizes access to electronic-structure simulations, enabling researchers to focus on scientific discovery rather than computational complexities. This framework empowers large-scale materials screening, accelerates ICME design loops, and promotes innovation across academia and industry by bridging the gap between experimental work and simulations, democratizing advanced simulation methods for users lacking extensive computational experience. 9:48am - 10:00am
ID: 137 / Thu | GERN 9h Parallel S 4b: 005 Modalities: Image Methods: Foundation Models Application Domain: Core Machine Learning, Information The Road to Exascale: Lessons Learned from Scaling a Scientific AI Workflow to 16,384 GPUs 1: Institute of Neuroscience and Medicine (INM-1), Forschungszentrum Jülich (FZJ), Germany; 2: Helmholtz AI, Forschungszentrum Jülich (FZJ), Germany; 3: Jülich Supercomputing Centre (JSC), Forschungszentrum Jülich, Germany; 4: German BioImaging, Gesellschaft für Mikroskopie und Bildanalyse e.V, Konstanz, Germany; 5: Cécile & Oskar Vogt Institute for Brain Research, University Hospital Düsseldorf, Germany; 6: Computer Vision, Institute for Computational Visualistics, University of Koblenz, Germany Foundation models have progressed by scaling parameters and training data, driving rapidly increasing computational demands. This trend is especially pronounced in scientific imaging, where datasets can span terabytes to petabytes. Exascale systems provide the compute to train models at this scale. However, using these machines efficiently is non-trivial. At extreme scale, bottlenecks shift from GPU throughput to end-to-end workflow behavior, including startup overheads, storage access, communication, and synchronization. We study these effects and derive practical scaling lessons from a real training pipeline. This work was carried out on the JUPITER system at Jülich Supercomputing Centre within the JUPITER Research and Early Access Program (JUREAP) and the GCS Exascale Pioneer project brainfm. We adapt the execution environment and I/O path to reduce indirect I/O and metadata pressure caused by containers, runtime-generated artifacts, and logging. In parallel, we evaluate model- and loss-level choices that reduce synchronization and collective communication. To quantify data access performance, we compare HDF5 and Zarr for highly concurrent random access across file layouts and backends. We demonstrate the approach by training a neuroscience vision foundation model with contrastive learning on terabyte-scale microscopic images of histological human brain sections (CytoNet, https://arxiv.org/abs/2511.01870). We scale the workflow up to 16,384 NVIDIA GH200 superchips across 4,096 compute nodes. Across large runs, indirect I/O emerges as a primary scalability limiter, driven by container image access, startup scripts, bytecode generation, temporary-directory traffic, and uncontrolled logging. Staging container images into node-local memory and redirecting runtime-generated files away from shared storage reduces filesystem metadata storms and improves startup robustness. On the algorithmic side, synchronization-heavy components constrain scaling, motivating architecture choices that avoid batch-level collectives (e.g., batch normalization) and a contrastive-loss implementation that reduces redundant per-rank compute while limiting collective communication. For highly concurrent data access, we find that Zarr with the TensorStore backend provides the lowest and most stable access times. We distill these findings into practical guidelines that link workflow engineering, model design, and storage choices for training scientific foundation models at extreme scale. | ||