Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
Please note that all times are shown in the time zone of the conference. The current conference time is: 4th Aug 2026, 10:37:03am CEST
|
Daily Overview |
| Session | |||||||||||||||||||||||||||||||||
Poster Session II
| |||||||||||||||||||||||||||||||||
| Session Abstract | |||||||||||||||||||||||||||||||||
|
Discover cutting-edge AI research and innovations at the Poster Session. Connect with authors, ask questions, and engage in lively discussions that spark collaboration and new ideas. | |||||||||||||||||||||||||||||||||
| Presentations | |||||||||||||||||||||||||||||||||
ID: 243
/ Poster No. # 10: 001
Modalities: Image Methods: Other Application Domain: Health HemAutomaton: A lightweight NCA-based pipeline for large-scale extraction of white blood cells from whole slide images 1: Helmholtz Munich, Germany; 2: Universitätsklinikum Erlangen, Germany Hematological smears contain rich morphological information at single-cell resolution and are central to routine diagnostics. Increasingly, such smears are digitized as whole slide images (WSIs) for computational analysis. Major challenges for automated full-slide analysis arise from the gigapixel resolution of WSIs and from variability in smear preparation across institutions, as well as heterogeneous regions within individual slides. As a result, many existing solutions focus on selected regions or small subsets of the slide and rely on computationally heavy architectures that require substantial training effort. Here, we propose HemAutomaton, a lightweight and generalizable pipeline for large-scale single-cell extraction from whole slide images. The method uses neural cellular automata (NCA) to localize white blood cells on downsampled WSIs, enabling efficient extraction under heterogeneous smear conditions without manual region selection or extensive labeling. The resulting pipeline enables large-scale unsupervised training of feature extractors, automated cell counting, and statistical analysis of cellular distributions across smear regions. To promote reproducible research, we release the complete pipeline as an open-source package and evaluate it on publicly available datasets. Together, this work provides an accessible and scalable foundation for whole-slide single-cell analysis in computational hematopathology. ID: 290
/ Poster No. # 10: 002
Modalities: Other Methods: Foundation Models Application Domain: Health DeepRVAT2: Unified Modeling of Coding and Regulatory Rare Variation at Genome Scale for Enhanced Gene Discovery and Diagnostics 1: Computational Health Center, Helmholtz Center Munich, Neuherberg, Germany; 2: chool ofComputation, Information and Technology, Technical University of Munich, Garching, Germany; 3: Institute of Human Genetics, School of Medicine, Technical University of Munich, Munich, Germany; 4: German Cancer Research Center (DKFZ), Heidelberg, Germany; 5: European Molecular Biology Laboratory (EMBL), Heidelberg, Germany; 6: Heidelberg University, Germany Rare genetic variants can strongly influence human health and disease, providing a direct link to causal genes and therapeutic targets. However, the vast majority of high-impact variants are ultra-rare or unique to a single individual, which makes their prioritization infeasible without AI models capable of strong generalization across variant classes and genes. While our previous DeepRVAT framework used exome data to increase power in rare variant association testing (RVAT), it omitted non-coding variation captured by whole-genome sequencing (WGS). We present DeepRVAT2, a major extension that integrates coding and non-coding signals into a unified, interpretable gene impairment score. DeepRVAT2 utilizes a resource-efficient pipeline that compresses terabytes of raw genetic data into a gigabyte-scale, ML-ready parquet format, enabling efficient training across 379,645 individuals and incorporating over 100 functional annotations to capture coding and non-coding variant effects for genome-wide deployment. We demonstrate that DeepRVAT2 successfully captures non-coding regulatory signals without compromising its strong performance on coding variants. To enhance biological interpretability, the model scale is anchored such that a score of one corresponds to the average effect of a loss-of-function variant. Compared to DeepRVAT1, DeepRVAT2 identifies 12% more gene-trait associations across 190 UK Biobank traits. Notably, while previous RVAT studies reported little difference between WES and WGS data, ablation studies of DeepRVAT2 reveal that its performance gains stem from both expanding coverage to the entire gene body (5% increase over exonic-only) and incorporating comprehensive non-coding features (5.8% increase), demonstrating that holistic integration is essential to unlock the full potential of WGS data. Beyond association testing, DeepRVAT2 provides critical utility for clinical interpretation. In rare disease diagnostics (SolveRD cohort, n=1023), it more than doubles Recall@1 for true causal genes compared to state-of-the-art alternatives. Furthermore, a novel gene constraint metric derived from DeepRVAT2 scores outperforms established benchmarks in identifying genes known to cause genetic disorders (AUC 0.27 vs 0.24). Finally, by approximating gene scores with pre-computed, shareable variant-level impairment scores, DeepRVAT2 seamlessly integrates into standard pipelines to unlock the full potential of WGS for drug discovery and precision diagnostics. ID: 1215
/ Poster No. # 10: 003
Modalities: Image Methods: Foundation Models, Other Application Domain: Health The Mean is the Mirage: Entropy-Adaptive Model Mergingunder Heterogeneous Domain Shifts in Medical Imaging 1: School of Computation, Information and Technology, Technical University of Munich, Germany; 2: Institute of Machine Learning in Biomedical Imaging, Helmholtz Munich, Germany; 3: relAI – Konrad Zuse School of Excellence in Reliable AI; 4: Munich Center for Machine Learning (MCML); 5: Institute of Pathology, Technical University of Munich, Germany; 6: School of Biomedical Engineering and Imaging Sciences, King’s College London, UK. Model merging under unseen test-time distribution shifts often renders naive strategies, such as mean averaging unreliable. This challenge is especially acute in medical imaging, where models are fine-tuned locally at clinics on private data, producing domain-specific models that differ by scanner, protocol, and population. When deployed at an unseen clinical site, test cases arrive in unlabeled, non-i.i.d. batches, and the model must adapt immediately without labels. In this work, we introduce an entropy-adaptive, fully online model-merging method that yields a batch-specific merged model via only forward passes, effectively leveraging target information. We further demonstrate why mean merging is prone to failure and misaligned under heterogeneous domain shifts. Next, we mitigate encoder classifier mismatch by decoupling the encoder and classification head, merging with separate merging coefficients. We extensively evaluate our method with state-of-the-art baselines using two backbones across nine medical and natural-domain generalization image classification datasets, showing consistent gains across standard evaluation and challenging scenarios. These performance gains are achieved while retaining single-model inference at test-time, thereby demonstrating the effectiveness of our method.
ID: 1145
/ Poster No. # 10: 004
Modalities: Tabular Data Methods: Other Application Domain: Health ConvexGating infers gating strategies from clusters in single cell cytometry data 1: University of Leipzig, Institute for Medical Informatics, Statistics, and Epidemiology, Leipzig, Germany; 2: Center for Scalable Data Analytics and Artificial Intelligence (ScaDS.AI), Leipzig, Germany; 3: Systems Medicine, Deutsches Zentrum für Neurodegenerative Erkrankungen (DZNE), Bonn, Germany; 4: Modular High Performance Computing and Artificial Intelligence, German Center for Neurodegenerative Diseases (DZNE), Bonn, Germany; 5: Research Group Tissue Control of Immunocytes, Helmholtz Center Munich, Munich, Germany; 6: Life and Medical Sciences (LIMES) Institute, University of Bonn, Bonn, Germany; 7: PRECISE Platform for Single Cell Genomics and Epigenomics, DZNE and University of Bonn and West German Genome Center (WGGC), Bonn, Germany; 8: Immunogenomics & Neurodegeneration, Deutsches Zentrum für Neurodegenerative Erkrankungen (DZNE), Bonn, Germany; 9: Department of Microbiology and Immunology, The Peter Doherty Institute for Infection and Immunity, University of Melbourne, Melbourne, VIC, Australia; 10: Molecular Immunology in Neurodegeneration, German Center for Neurodegenerative Diseases and the University of Bonn, Germany; 11: Medical Microbiology and Immunology Department, Faculty of Medicine, Mansoura University, Egypt; 12: Institute of Innate Immunity, Biophysical Imaging, Medical Faculty, University of Bonn, Bonn, Germany; 13: Max Planck Institute for Metabolism Research, Center for Endocrinology, Diabetes and Preventive Medicine (CEDP), Cologne, Germany; 14: University of Leipzig, Faculty of Mathematics and Computer Science, Leipzig, Germany; 15: Institute of Computational Biology, Helmholtz Center Munich, Germany; 16: Department of Mathematics, Technical University of Munich, Germany; 17: TUM School of Life Sciences Weihenstephan, Technical University of Munich, Germany Manual expert gating remains the standard approach for defining specific cell populations in flow cytometry. However, as the number of measured parameters per cell continues to increase, manual gating becomes increasingly difficult to scale. In addition, high inter-rater variability limits reproducibility, particularly in multicentre studies, making the process both labour-intensive and inconsistent. Here, we introduce ConvexGating, an AI tool that automatically learns gating strategies for cell sorting in an unbiased, fully data-driven, and interpretable manner. A central advantage of ConvexGating over existing computational gating approaches is that the inferred strategies are expressed as polygon gates on conventional two-dimensional scatter plots. These gates can be directly implemented on standard cell sorters without modification. By preserving compatibility with established workflows, ConvexGating bridges the gap between advanced computational analysis and practical cell sorting. ConvexGating scales efficiently to high-dimensional parameter spaces and generates robust strategies with low contamination of the target population, applicable to both well-characterized and previously undefined or poorly resolved cell types. Importantly, the inferred gating strategies are independent of predefined parent populations. For example, plasmacytoid dendritic cells (pDCs) can be fully identified as CD57- CD13- CD45RA+ CD123+ cells using a single, self-contained gating hierarchy. We validated ConvexGating-derived strategies for CD8+ naive and CD8+ Temra cells, as well as progenitor populations from mouse white adipose tissue, using 384-well–based single-cell RNA sequencing (scRNA-seq). In all cases, ConvexGating yielded more homogeneous sorted populations compared to conventional manual strategies. Beyond flow cytometry, ConvexGating derives transferable gating strategies for CyTOF (Cytometry by Time of Flight) and CITE-seq (Cellular Indexing of Transcriptomes and Epitopes by Sequencing) data, and it supports optimal marker panel design for targeted cell sorting applications.
ID: 1278
/ Poster No. # 10: 005
Modalities: Multimodal Data Methods: Foundation Models Application Domain: Earth & Environment Cross Modalities Pretraining of Sparse Lidar and Dense Image Foundation Model for Global Carbon Stock Mapping 1: Helmholtz-Zentrum Dresden-Rossendorf, Dresden, Germany; 2: Chair of Data Science in Earth Observation, Technical University of Munich, Munich, Germany Foundation Models (FMs) are built by pretraining on extensive datasets, followed by fine-tuning for specific downstream applications. By leveraging large-scale unlabeled data, FMs learn generalized task-agnostic feature representations that enables the integration of diverse Earth Observation (EO) data, for applications such as global carbon stock mapping, forest canopy height estimation, and long-term monitoring. Most EO modality, such as optical imagery, are represented as dense 2D grids, making them well suited for image-based pretraining. However, not all EO data follow this structure. Some sensors, such as the Global Ecosystem Dynamics Investigation (GEDI), produce sparse and irregular observations, creating a fundamental challenge for multimodal learning. GEDI is a spaceborne LiDAR sensor mounted on the ISS that emits laser pulses toward the Earth's surface and records the full waveform of the reflectance. GEDI produces precise measurements of forest canopy height, vertical structure and surface elevation. Despite its powerful capability for measuring vegetation height, GEDI shots are very sparse compared to gridded EO imagery. The spacing between GEDI shots is approximately 600 m across-track and 60 m along-track. In contrast, products such as the 30 m Harmonized Landsat Sentinel-2 (HLS) dataset provide dense spatial coverage. This mismatch makes direct integration of GEDI with grid-based 2D modalities difficult during training. To address this challenge, we treat the GEDI full waveform as a 1D continuous signal and apply a masking strategy in which portions of the signal are randomly masked for each GEDI shot and then reconstructed in a masked autoencoder framework, enabling self-supervised learning of general waveform representations from large collections of unlabeled GEDI data. The pretrained waveform representation are then tokenized into embeddings and fused with other modality to capture cross-modal dependencies. We then refine the map with the pretrained GEDI and finally estimate the vegetation height for each pixel of the 2D dense images. ID: 1151
/ Poster No. # 10: 006
Modalities: Tabular Data, Text Methods: Agentic AI, Uncertainty Quantification Application Domain: Health ICD-Code Extraction from Clinical Notes using Large Language Models in a RAG pipeline 1: Hybrid Methods in Artificial Intelligence and Machine Learning, University of Rostock, Germany; 2: German Center for Neurodegenerative Diseases, Rostock, Germany Datasets created in clinical settings, such as imaging, biosignals, or genetic data, often lack the structured metadata (age, sex, medication, medical diagnoses, etc.) required for research. Restrospective creation of such structured metadata for research is complicated and expensive. Electronic health records typically contain rich free-text clinical notes describing medical conditions, patient’s history, symptoms, and others, which represent a valuable but unstructured source for generating such metadata. Retrieval-augmented generation (RAG) has been shown to enhance LLMs on knowledge-intensive tasks, but has not yet been systematically applied to diagnosis coding on large open-access clinical datasets. This work investigates whether large language models (LLMs), grounded on reliable diagnosis ontologies, can be used to automatically and reliably extract diagnosis codes from free-text clinical notes, in a tracable and explainable manner, to produce research-ready metadata. A three-step pipeline was devised. First, Named Entity Recognition (NER) is applied to identify spans in clinical text indicating disorders and symptoms. Second, semantically similar diagnosis codes are retrieved from a vector database of embedded International Classification of Diseases (ICD) descriptions and complementary ontologies with equivalent terminologies. Third, 'gpt-oss' (120B) is prompted for self-verification and output a final set of ICD codes. This pipeline was evaluated on 200 clinical notes of the MIMIC IV 3.1 benchmark. On the level of a three-digit ICD code, i.e., a concrete diagnosis that groups its subtypes, it attained a recall of 0.56 and precision of 0.16 compared to single-shot LLM prompt recall of 0.42 and precision 0.55. On the level of two-digit ICD codes, it attained a recall of 0.67. We propose a metric of Jiang-Conrath (JC) "closeness” between diagnoses, based on PheCodeX or ICD-9, enabling relaxed evaluations that can be tuned on case-by-case needs. The long-tail distribution of the MIMIC dataset, with a few very frequent codes and many codes not so frequently used, revealed a clear trade-off that most of the proposed systems apply toward prioritizing the accurate identification of the most common diagnoses, at the cost of additional false positives. Overall, this work proposes a framework for creating structured and machine-readable metadata tables from routine clinical documentation for research use, grounding LLM capabilities to reliable ontologies.
ID: 1225
/ Poster No. # 10: 007
Modalities: Tabular Data Methods: Other Application Domain: Health Causal Machine Learning for Predictive Biomarker Discovery and Subgroup Refinement in Metastatic Colorectal Cancer 1: Department of Biochemistry and Pharmacology, Bio21 Molecular Science and Biotechnology Institute, The University of Melbourne, Australia; 2: Department of Medicine III and Comprehensive Cancer Center Munich, University Hospital, Ludwig-Maximilians University Munich, Germany; 3: Comprehensive Cancer Center Munich, Germany; 4: German Cancer Consortium (DKTK), partner site Munich, German Cancer Research Center (DKFZ), Germany; 5: LMU Munich School of Management, LMU Munich, Germany; 6: Munich Center for Machine Learning, Germany; 7: Computational Health Center, Institute of Computational Biology, Helmholtz Munich, Germany Treatment responses in oncology vary strongly across patients, and identifying actionable biomarkers of heterogeneous treatment benefit remains a central challenge in precision medicine. While traditional machine learning (ML) approaches provide limited insight into how therapeutic benefit differs between individuals, causal ML offers a principled approach to address this problem by estimating individual and subgroup-specific treatment effects from high-dimensional molecular and clinical data, enabling optimization of personalized treatment strategies. However, robust computational frameworks that combine systematic characterization of treatment effect heterogeneity with biologically grounded predictive biomarker discovery remain scarce. Here we present a multimodal causal ML framework for robust predictive biomarker discovery and subgroup refinement in randomized controlled trials (RCTs) with deep molecular characterization. Using state-of-the-art causal inference methods, treatment effects are estimated at the individual level and systematically analyzed to characterize treatment heterogeneity and identify clinically relevant predictive biomarkers and subgroups. To enable biologically informed discovery and clinically actionable results, the approach combines two complementary strategies: (1) bottom-up data-driven discovery of novel predictive biomarkers, and (2) top-down validation and refinement of clinically established biomarkers and domain-informed hypotheses. We demonstrate our framework using the randomized phase III FIRE-3 trial (AIO KRK-0306) in metastatic colorectal cancer (mCRC), integrating genomics, transcriptomics, and clinical data to identify molecular and clinical features associated with sensitivity or resistance to anti-EGFR (cetuximab) versus anti-VEGF (bevacizumab) therapy. The analysis enables biologically informed characterization of predictive biomarkers and systematic evaluation of clinically established and candidate subgroups in mCRC, including RAS mutation status and primary tumor sidedness, while uncovering additional sources of variation in treatment benefit to refine current patient stratification strategies. This work highlights the promise of causal ML to move toward data-driven, biologically informed personalized decision-making, providing a scalable computational strategy for systematic predictive biomarker discovery and subgroup refinement in RCTs, ultimately improving individual patient outcomes. ID: 1187
/ Poster No. # 10: 008
Modalities: Simulation Data, Tabular Data, Time Series Methods: Physics-informed Machine Learning Application Domain: Matter Neural Operator-Based Surrogate Modeling for Efficient Prediction of Temperature and Residual Stresses in Tempered Glass Universität Augsburg, Germany Numerical simulation of temperature and stress evolution in multi-physics problems is typically performed using physics-based methods such as the finite element method (FEM). While these methods provide accurate solutions, they are computationally expensive, especially for fine discretizations and large domains or repeated evaluations required in process optimization and design studies. In this work, we investigate neural operator-based surrogate models as efficient alternatives for predicting thermo-mechanical fields in tempered glass. The proposed approach learns a parameter-to-field mapping, where a vector of process parameters is directly transformed into spatio-temporal temperature and residual stress distributions. Two neural operator architectures are analyzed: Multi-Input Fourier Neural Operators (MIFNO) and Multi-Input Operator Networks (MIONet). Both models incorporate multiple physical parameters as inputs and predict full 1D temperature and stress fields over space and time. The models are trained using data generated from finite element simulations of the glass tempering process and evaluated on unseen parameter combinations. The results show that MIFNO achieves prediction errors of approximately 2.1% for temperature and 3.2% for stress on unseen or untrained cases, while providing a computational speed-up of about 22× compared to FEM. The MIONet model demonstrates even greater efficiency, achieving speed-ups of over 600× while maintaining prediction errors of about 1.1% for temperature and 3.5% for stress. These results demonstrate that neural operator-based surrogate models provide a promising framework for efficient prediction of thermo-mechanical fields in glass tempering processes. Such models enable rapid parametric analysis and support real-time decision making and optimization in glass manufacturing and other industrial purposes. ID: 1368
/ Poster No. # 10: 009
Modalities: Graphs, Other Methods: Graph Neural Networks, Physics-informed Machine Learning Application Domain: Earth & Environment, Health GRIP: Physics-Informed Neural Network for Gradient Retention Time Prediction in Liquid Chromatography 1: Helmholtz Institute for Pharmaceutical Research Saarland (HIPS), Helmholtz Centre for Infection Research (HZI), Germany; 2: German Research Center for Artificial Intelligence (DFKI) Kaiserslautern Gradient high-performance liquid chromatography (HPLC), often coupled with mass spectrometry, is widely used to separate and identify small molecules in complex samples. During chromatographic separation, each compound travels through the column at a different rate and elutes at a characteristic retention time (RT), which provides complementary information for compound identification. Predicting RT from molecular structure enables the integration of chromatographic information into computational analysis pipelines. However, RT depends not only on molecular properties but also on system-specific parameters such as column type, solvent gradient, and temperature. As a result, existing machine-learning approaches are typically restricted to a single experimental setup and require retraining or transfer learning to adapt to new chromatographic systems. In this work, we employ the physical principles of liquid chromatography to create GRIP, a physics-informed deep learning model for retention time prediction across different chromatographic setups. The model uses a message-passing graph neural network to encode molecular structure and a feed-forward network to represent chromatographic conditions. Together they predict linear solvent strength (LSS) parameters describing the interaction of a compound with the mobile phase, which are then used in the fundamental equation of gradient elution to compute retention times. We trained GRIP on 65 reverse-phase HPLC datasets spanning diverse column, gradient, and temperature conditions and evaluated its ability to generalize across chromatographic setups. The model demonstrates zero-shot prediction on previously unseen systems while matching or outperforming transfer-learning-based baselines fine-tuned with varying amounts of system-specific data. We further assessed generalization across chemical space using a similarity-based split that separates structurally related compounds between training and test sets. These results suggest that GRIP enables reliable retention time prediction across diverse chromatographic setups. This can support in-silico optimization of chromatographic methods by predicting RT under previously untested conditions to improve compound separation. ID: 1252
/ Poster No. # 10: 010
Modalities: Text Methods: Foundation Models, Other Application Domain: Core Machine Learning, Information Who Owns Human Experience? Ethical Implications of Transforming Tacit Knowledge into Neural Models 1: University Augsburg, Germany; 2: ergonoi GbR Recent advances in artificial intelligence (AI) increasingly rely on large-scale interaction data generated through the everyday use of digital tools. Beyond explicit data, such interactions may also capture patterns of tacit knowledge, including decision strategies, problem-solving heuristics, and domain-specific expertise. Tacit knowledge refers to forms of knowing that are not easily articulated or formalised (Polanyi, 1966; Collins, 2010). As AI systems learn from these interaction traces, elements of human experiential reasoning may become functionally represented within neural models. ID: 1325
/ Poster No. # 10: 011
Modalities: Tabular Data, Other Methods: Generative Models, Probabilistic Methods Application Domain: Health Towards Useful and Private Synthetic Omics: Community Benchmarking of Generative Models for Transcriptomics Data 1: European Molecular Biology Laboratory (EMBL), Genome Biology Unit, Heidelberg, Germany; 2: Division of Computational Genomics and Systems Genetics, German Cancer Research Center (DKFZ), Heidelberg, Germany; 3: CISPA Helmholtz Center for Information Security, Saarbrücken, Germany; 4: University of Helsinki, Finland; 5: Heidelberg University, Germany; 6: Helmholtz Munich, Germany; 7: Division of Tumorigenesis and Molecular Cancer Prevention, German Cancer Research Center (DKFZ), Heidelberg, Germany; 8: DKFZ Hector Cancer Institute at the University Medical Center Mannheim, Germany; 9: Eberhard Karls Universität Tübingen, Germany; 10: University of Washington Tacoma, USA; 11: Sage Bionetworks, Seattle, USA; 12: Ghent University, Ghent, Belgium; 13: European Bioinformatics Institute (EMBL-EBI), UK Background: The synthesis of anonymized data derived from real-world cohorts offers a promising strategy for regulatory-compliant and privacy-preserving biological data sharing, potentially facilitating model development that can improve predictive performance. However, the extent to which generative models can preserve biological signals while remaining resilient to adversarial privacy attacks in high-dimensional omics contexts remains underexplored. To address this gap, the CAMDA 2025 Health Privacy Challenge launched a community-driven effort to systematically benchmark synthetic and privacy-preserving data generation for bulk RNA-seq cohorts (https://benchmarks.elsa-ai.eu/?ch=4). Results: Building on this initiative, we systematically benchmarked 11 generative methods across two cancer cohorts (~1,000 and ~5,000 patients) over 978 landmark genes. Methods were evaluated across complementary axes of distributional fidelity, downstream utility, biological plausibility and empirical privacy risk, with emphasis on trade-offs between vulnerability to membership inference attacks (MIA) and other evaluation dimensions. Expressive deep generative models achieved strong predictive utility and differential expression recovery, but were often more vulnerable to membership inference risk. Differentially private methods improved resistance to attacks at the cost of reduced utility, while simpler statistical approaches offered competitive utility with moderate privacy risk and fast training. Conclusions: Synthetic bulk RNA-seq quality is inherently multi-dimensional and shaped by trade-offs between utility, biological preservation and privacy. Our results indicate that differences in model architecture drive distinct trade-offs across these axes, suggesting that model choice should align with dataset characteristics, intended downstream use and privacy requirements. Privacy risk should also be assessed using multiple complementary attack methods and, where possible, formal differential privacy protection. Keywords: Synthetic data generation, private data generation, RNA-seq, generative models, membership inference attack, reproducibility, evaluation metrics and trade-offs ID: 372
/ Poster No. # 10: 012
Modalities: Graphs, Multimodal Data, Tabular Data, Text Methods: Uncertainty Quantification, Other Application Domain: Core Machine Learning, Health Multimodal Representation Learning for Pan-Cancer Type Classification Using Attention-Based Set Encoding and Contrastive Pre-Training Max Delbrück Center for Molecular Medicine, Germany Precision oncology demands reliable cancer type identification from molecular data to inform therapeutic decisions, particularly for rare diagnoses. Most current approaches depend on engineered features, single data modalities, or data types such as whole-genome sequencing or RNA-seq not routinely available in clinical practice. We present a representation learning approach for pan-cancer classification across 69 cancer types, built on clinically feasible panel sequencing data from the AACR GENIE registry, a multi-institutional dataset reflecting what hospitals can realistically collect under current regulatory standards. We use a Perceiver IO architecture processing multimodal molecular and clinical features as an unordered set of tokens. The model integrates six modalities: protein language model, protein interaction network, and text embeddings, somatic mutations, copy number variations, and clinical variables. Each patient's altered genes form a variable-length token sequence of 450 dimensions, compressed into a fixed-size latent representation through cross-attention, supporting variable numbers of alterations without fixed-size feature engineering. We propose a multi-phase contrastive training schedule. The first phase applies symmetric InfoNCE loss as self-supervised warmup, generating views through token-level dropout, feature-level dropout, and additive Gaussian noise. The second phase transitions to a weighted combination of InfoNCE and supervised contrastive loss, with a linearly decaying weight shifting emphasis from instance-level alignment toward class-discriminative structure. Classification uses an MLP on mean-pooled encoder latents trained with cross-entropy loss. On a development cohort of 42,405 training and 5,300 validation patients, the model achieves a weighted F1 of 0.73, with contrastive pretraining producing embedding spaces showing clear class separation. Current state-of-the-art on this 69-type task reports 0.78–0.79 weighted F1 using comparable modalities. We attribute the gap primarily to dataset scale: ongoing work increases training to 135,000–176,000 patients. We additionally investigate the learned embeddings for use cases beyond cancer type classification and MC Dropout for uncertainty estimation to flag ambiguous cases. ID: 175
/ Poster No. # 10: 013
Modalities: Simulation Data, Time Series Methods: Probabilistic Methods, Uncertainty Quantification Application Domain: Core Machine Learning Higher-Order Hit-&-Run Samplers for Linearly Constrained Densities 1: Forschungszentrum Jülich, Germany; 2: Department of Statistics, LMU Munich; 3: Munich Center for Machine Learning (MCML) Markov chain Monte Carlo (MCMC) sampling of densities restricted to linearly constrained domains is an important task arising in Bayesian treatment of inverse problems in the natural sciences. While efficient algorithms for uniform polytope sampling exist, much less work has dealt with more complex constrained densities. In particular, gradient information as used in unconstrained MCMC is not necessarily helpful in the constrained case, where the gradient may push the proposal’s density out of the polytope. In this work, we propose novel constrained sampling algorithms, which combine strengths of higher-order information, like the target log-density’s gradients and curvature, with the Hit-&-Run (HR) proposal, a simple mechanism which guarantees the generation of feasible proposals, fulfilling the linear constraints. To this end, we combine the HR proposal mechanism with existing first- and second-order samplers like Metropolis-adjusted Langevin algorithm (MALA) and simplified manifold MALA (Girolami and Calderhead, 2011). As a result, we develop three new samplers Langevin Hit-&-Run (LHR), simplified manifold Hit-&-Run (smHR), and simplified manifold Langevin Hit-&-Run (smLHR). Our approach is mainly driven by the observation that a Gaussian random variable can be sampled in a HR-like fashion by decomposing it into direction and magnitude. We provide theoretical analysis to prove convergence of our samplers to the desired target densities. We numerically evaluate our algorithms on a systematically constructed benchmark consisting of 2240 problems, which combine parametrizable polytopes and probability densities of varying geometry or ill condition, as well as on real-world examples from the field of Bayesian 13C metabolic flux analysis. Our results show that our combination of both the incorporation of higher-order information and the constraining of the proposal density by the HR proposal mechanism improve sampling efficiency.
ID: 251
/ Poster No. # 10: 014
Modalities: Multimodal Data, Tabular Data Methods: Probabilistic Methods, Other Application Domain: Core Machine Learning, Health Simple Yet Effective: Basic Population and Evolutionary Statistics Contain Richest Information to Infer Systematic Variant Effects 1: 1Institute of Translational Genomics, Helmholtz Zentrum München, German Research Center for Environmental Health, Neuherberg, Germany; 2: School of Computation, Information and Technology, Technical University of Munich, Garching, Germany; 3: Helmholtz Association—Munich School for Data Science (MUDS), Munich, Germany; 4: Helmholtz Pioneer Campus, Helmholtz Zentrum München, German Research Center for Environmental Health, Neuherberg, Germany; 5: Department of Pediatrics, Dr. von Hauner Children's Hospital, University Hospital, Ludwig Maximilian University Munich, Munich, Germany Prioritization of mutations that critically modulate protein function or organism fitness remains a cornerstone of clinical diagnostics, therapeutic development, and protein engineering. With unprecedented volumes of labelled data, from clinically annotated variants (e.g., ClinVar) to experimentally profiled protein fitness landscapes (e.g., ProteinGym), generated, the computational prediction of variant effects has seen a parallel surge in methodological innovation. This includes the rapid proliferation of machine learning-based variant effect predictors (VEPs), ranging from traditional annotation based statistical models to modern deep neural network architectures with millions of parameters. While recent methods increasingly aim to incorporate features from diverse modalities and scale up to models with billions of parameters, a systematic understanding of the key features that drive predictive performance remains lacking. Here, we present a comparative analysis that evaluates the predictive power of diverse feature classes for the prediction of missense variant effects in a systematic framework. We demonstrate that a parsimonious model (XGBoost) that combines only 30 population and evolution informed features outperforms state-of-the-art deep learning models like AlphaMissense and ESM1b in discriminating pathogenic from benign mutants on ClinVar. We further used different annotations for rare variant burden tests on UKBioBank data, further explaining which group of features would give the best statistical power to infer variant-trait interactions. Finally, the minimalist approach achieves close-to-top accuracy in predicting single-mutation fitness effects from in vitro assays (DMS score), and provides top performance on inferring effects from double mutants. Our findings challenge the prevailing assumption that model complexity inherently improves performance, instead highlighting the sufficiency of information contained in populational and evolutional statistics. By prioritizing interpretability and biological relevance over architectural complexity, we outline a reference lens for evaluation and development of robust, generalizable future predictors that balance computational efficiency with real-world applicability. This approach emphasizes the pressing need to integrate our understanding of sequence, function, individual health, and evolutionary context—an area where current deep learning models still face significant limitations. ID: 351
/ Poster No. # 10: 015
Modalities: Text Methods: Probabilistic Methods, Uncertainty Quantification, Other Application Domain: Earth & Environment Bringing Gaps between Farmers and AI with XAI Principles 1: German Aerospace Center (DLR), Germany; 2: German Aerospace Center (DLR), Germany; 3: German Aerospace Center (DLR), Germany Explainable AI (XAI) aims to bridge the gap between intelligent algorithms and technology and the end-users. Considering applications in precision agriculture, much is to be done to help agricultural stakeholders understand and adapt to AI-based solutions. A black-box solution may perform well in specific tasks; however, its lack of interpretability hinders its acceptance in practice. In this research. We are interviewing farmers to understand the issue of limited technology adaptability and to develop a mechanism using XAI tools to address it. The activities focus on an early pest-detection technology using satellite data. The pest-detection model works with openly available satellite data and in situ data and is launched with a simple user interface. Two parallel models are discussed and compared in terms of efficiency, explainability, and their ability to be adopted by non-expert users. The study is divided into four major parts: 1) Understanding the level of technology penetration in two countries, Germany and India, and the role of XAI. 2) Analyzing the situation for a new technology and opportunities in XAI 3) The driving factors in improving technology adoption and connection with model explainability. The three sections will be discussed separately for each case study. The case studies here are Coconut farming in India and soybean farming in Germany. ID: 322
/ Poster No. # 10: 016
Modalities: Text Methods: Agentic AI, Generative Models Application Domain: Information Abstracts Explorer: LLM-Powered Semantic Exploration of Scientific Conference Proceedings Helmholtz Zentrum Dresden Rossendorf (HZDR), Germany Keeping pace with the rapidly growing machine learning literature is a major challenge. Conferences such as NeurIPS, ICLR, and ICML now accept thousands of papers annually, making it difficult for researchers to identify relevant work and track emerging trends. We present Abstracts Explorer, an open-source Python toolkit that combines LLM-based semantic search, unsupervised clustering, and retrieval-augmented generation (RAG) to enable intelligent exploration of conference proceedings at scale. ID: 332
/ Poster No. # 10: 017
Modalities: Multimodal Data Methods: Uncertainty Quantification, Other Application Domain: Core Machine Learning Systematic benchmarking of imaging-based spatial transcriptomics preprocessing 1: Institute of Computational Biology, Helmholtz Munich, Germany; 2: Institute of Lung Health and Immunity, Helmholtz Munich, Germany; 3: Indiana University School of Medicine Spatial and dissociated single cell transcriptomics have revolutionized our understanding of tissue-specific gene expression. Imaging-based spatial transcriptomics (iST) methods capture spatial gene expression on single cell- and subcellular resolution for a predefined set of genes (between 400 and 5000). However, although these technologies measure the same mRNA molecules, the resulting data distributions and structure generally differ, affecting data interpretation and research outcomes. The underlying causes of those differences are poorly understood and often ignored in downstream analysis. We hypothesize that one of the contributing factors to the difference is lack of established best practices for iST preprocessing. Processing of image-based spatial transcriptomics comes with new challenges compared to scRNA-seq including, cell segmentation from images, a reduced gene panel, and an inherently different sampling process for measuring transcript counts. To address this, we developed an iST preprocessing pipeline benchmark that systematically compares available methods, ranging from statistical approaches to deep-learning based strategies for each step. We introduce novel metrics that assess the similarity to matched scRNA-seq data, going beyond classical evaluation schemes for cell segmentation which are based on expert annotated cells. Our metrics address multiple data characteristics that are relevant for downstream analysis of single cell transcriptomics. Integrated into the Open Problems living benchmarking framework, our dynamic benchmark ensures frequent updates with the latest computational methods. We envision that the insights gained will be instrumental in refining computational methods and models for spatial transcriptomics, establishing best practices for data processing, and minimizing false positive results. Thereby we expect this benchmark to significantly enhance the accuracy and reliability of imaging-based spatial transcriptomics analyses and serve as a recommendation guidelines for pre-processing strategies for novel analysis in context of the tissue and the methodology. ID: 191
/ Poster No. # 10: 018
Modalities: Image, Multimodal Data, Text Methods: Agentic AI, Uncertainty Quantification Application Domain: Health Dynamic Decision Learning: Test-Time Evolution for Abnormality Grounding in Rare Diseases 1: Technical University of Munich, Germany; 2: University of Trento, Italy; 3: Imperial College London, United Kingdom; 4: Helmholtz Munich, Germany; 5: King's College London, United Kingdom Large vision-language models (LVLMs) show promise in medical imaging but struggle with clinical abnormality grounding, particularly for rare diseases and long-tailed distributions where data scarcity and domain shifts make traditional fine-tuning impractical. We find that standard single-pass inference is highly sensitive to minor variations in instructions and visual inputs, leading to hallucinated detections that lack spatial consistency across different perspectives (Fig. 1). In contrast, clinicians typically cross-reference multiple views and expert priors to characterize findings, a process currently missing in static LVLM pipelines. We propose Dynamic Decision Learning (DDL), a training-free test-time framework that transforms static grounding into an adaptive process of hypothesis testing (Fig. 2). DDL synergizes two evolutionary pathways: (1) a Language Path that introduces Distribution-Aware Prompt Evolution (DAPE) to iteratively search for optimal task-specific instructional priors, effectively minimizing linguistic sensitivity; and (2) a Visual Path that employs Visual-Perception Uncertainty Probing (V-PUP) and Referenced Hungarian Consolidation (RHC) to verify and consolidate spatial hypotheses across multiple stochastic augmented views. By treating consensus as a structured bipartite matching problem relative to a reference anchor, DDL isolates stable clinical signals from transient noise without requiring parameter updates. Extensive evaluation on brain MRI benchmarks, including the NOVA dataset with 281 rare pathologies and heterogeneous imaging protocols, demonstrates that DDL consistently outperforms supervised fine-tuning and state-of-the-art adaptation baselines. Across LVLMs ranging from 3B to 72B parameters, DDL achieves up to a 105% relative improvement in localization precision (mAP). Furthermore, we reveal an emergent spatial calibration scaling law: while small models exhibit decoupled confidence, increasing model size unlocks a strong alignment where DDL’s reliability scores become significantly more predictive of actual grounding success, providing a trustworthy self-verifying signal for high-stakes clinical decision support.
ID: 177
/ Poster No. # 10: 019
Modalities: Graphs Methods: Graph Neural Networks Application Domain: Core Machine Learning Geometry-Aware Edge Pooling for Graph Neural Networks 1: Helmholtz Munich, Germany; 2: Technical University of Munich, Germany; 3: University of Montreal, Canada; 4: Mila - Quebec Artificial Intelligence Institute, Canada; 5: University of Fribourg, Switzerland Graph Neural Networks (GNNs) have shown significant success for graph-based tasks. Motivated by the prevalence of large datasets in real-world applications, pooling layers are crucial components of GNNs. By reducing the size of input graphs, pooling enables faster training and potentially better generalisation. However, existing pooling operations often optimise for the learning task at the expense of discarding fundamental graph structures, thus reducing interpretability. This leads to unreliable performance across dataset types, downstream tasks and pooling ratios. Addressing these concerns, we propose novel graph pooling layers for structure-aware pooling via edge collapses. Our methods leverage diffusion geometry and iteratively reduce a graph's size while preserving both its metric structure and its structural diversity. We guide pooling using magnitude, an isometry-invariant diversity measure, which permits us to control the fidelity of the pooling process. Further, we use the spread of a metric space as a faster and more stable alternative ensuring computational efficiency. Empirical results demonstrate that our methods (i) achieve top performance compared to alternative pooling layers across a range of diverse graph classification tasks, (ii) preserve key spectral properties of the input graphs, and (iii) retain high accuracy across varying pooling ratios. ID: 157
/ Poster No. # 10: 020
Modalities: Multimodal Data Methods: Foundation Models Application Domain: Health Phenotype-driven virtual screening accelerates therapeutic discovery for idiopathic pulmonary fibrosis 1: Helmholtz Munich; 2: Faculty of Mathematics, Informatics and Mechanics, University of Warsaw Idiopathic pulmonary fibrosis (IPF) is a chronic, progressive lung disease characterized by activation of pathogenic fibroblasts and excessive extracellular matrix (ECM) deposition. Current antifibrotic therapies only slow down disease progression, with a median survival of 3-5 years. To accelerate antifibrotics discovery, we develop an AI-augmented virtual screening framework, trained on high-content phenotypic data from primary patient-derived lung fibroblasts, that integrates chemical language models, graph representation learning and structure-based descriptors with ensemble learning. By framing virtual screening as out-of-distribution prediction, we show that ensemble learning consistently outperforms individual models and identifies complementary candidate hits. The resulting pipeline enabled screening of 1.5 million molecules and prioritized ~2,000 chemically diverse, drug-like candidates for downstream validation, yielding a fivefold increase in hit rate. Our results suggest that disease-specific phenotypic virtual screening provides a scalable strategy for antifibrotic discovery in IPF. ID: 347
/ Poster No. # 10: 021
Modalities: Tabular Data, Other Methods: Generative Models, Other Application Domain: Health Machine Learning-based RNA Design for the Development of Efficient mRNA Therapeutics Helmholtz Munich, Computational Health Center, Computational Biology (ICB), Germany The rapid development and high effectiveness of mRNA vaccines during the COVID-19 pandemic highlighted the potential of mRNA-based therapeutics, driven by the adaptability and efficiency of synthetic mRNA design. While substantial efforts have been dedicated to optimizing the coding sequences of therapeutic mRNAs, the untranslated regions (UTRs), key elements in mRNA stability, translation, have been less explored. Recent advances in artificial intelligence (AI) and machine learning (ML) tools have opened new avenues for optimizing UTR design by uncovering complex regulatory relationships within large-scale sequence datasets. In this study, we present a comprehensive experimental and computational framework to optimize mRNA sequences by leveraging Massive Parallel Reporter Assays (MPRAs) for systematic identification of UTRs that influence mRNA stability and translation. The dataset comprises approximately 35,000 synthetic mRNAs, along with corresponding measurements from four distinct cell lines, allowing for a detailed analysis of in-cell survivability. We developed Cell Line-Specific convolutional neural network (CNN) models to predict survivability based on RNA sequence features, achieving high predictive accuracy for two cell lines and uncovering critical sequence composition variants. To exploit the richness of the full dataset and uncover universal patterns, we advanced our approach by implementing a multiclass classification model capable of effectively distinguishing between different experimental pools. Additionally, we explored transfer learning, which notably outperformed both single-task and multi-task models in predictive accuracy. Finally, we embedded these CNN models into a sequence design pipeline utilizing genetic algorithms and in silico mutagenesis to generate novel mRNA sequences optimized for cellular survivability. The highest-ranking candidate sequences were selected for experimental validation, and their biological performance is currently being evaluated. This work highlights the power of combining AI-driven models with experimental methods to optimize mRNA design, providing a robust framework for developing more effective and tailored mRNA-based therapeutics. ID: 328
/ Poster No. # 10: 022
Modalities: Text Methods: Other Application Domain: Core Machine Learning Reinterpreting Protein Sequences as Structured Signals: A CNN-Residual Approach to Neuropeptide Classification 1: Federal University of Technology-Paraná, Brazil; 2: Helmholtz Centre for Environmental Research, Deutschland Traditional machine learning approaches to protein sequence classification commonly treat amino acid sequences as textual data. Standard models such as Decision Trees, Logistic Regression, Random Forests, Support Vector Machines, K-Nearest Neighbors, and Naïve Bayes are often combined with handcrafted feature extraction and dimensionality reduction approaches, such as PCA and ICA. Although effective, these approaches depend heavily on manual feature engineering and remain constrained by text-based modeling paradigms, which may limit their ability to identify hierarchical and spatial sequence patterns. In this work, we reinterpret protein sequences as structured signals and introduce an image-processing-inspired deep learning framework for neuropeptide classification. Our supervised model uses convolutional residual blocks inspired by computer vision architectures to learn directly from numerically encoded, padded amino acid sequences, eliminating the need for handcrafted descriptors. The architecture consists of an embedding layer followed by dilated Conv1D residual blocks that capture multi-scale local dependencies. To model long-range interactions across sequences, a multi-head self-attention module is incorporated into the network. The resulting convolutional representations are aggregated using global max pooling and passed to fully connected layers with dropout regularization, producing a sigmoid-activated output for binary classification. Experiments were carried out using a UniProt/Swiss-Prot dataset containing validated neuropeptides and negative protein samples. Hyperparameters, including learning rate, batch size, and number of training epochs, were optimized using Optuna. The proposed framework was evaluated against traditional machine learning baselines, including XGBoost, SVM, Random Forest, KNN, and Naïve Bayes. Our model achieved an accuracy of 0.899, approaching the benchmark performance of the NeuroPpred-Fuse model (0.906). These results suggest that image-inspired convolutional architectures applied directly to primary protein sequences offer a competitive, conceptually distinct alternative to conventional text-based sequence classification methods. Our study reveals the potential of computer vision–inspired inductive biases for modeling biological sequences. ID: 333
/ Poster No. # 10: 023
Modalities: Image, Multimodal Data, Other Methods: Other Application Domain: Aeronautics, Space & Transport Extracting 1D JWST Spectra from 2D Spectrograms with Deep Learning 1: Centre de Physique des Particules de Marseille, France; 2: Laboratoire d'Informatique et des Systèmes, Marseille, France With the launch of the Euclid Telescope in 2023, in 2026, with DR1, we are looking at an unprecedented amount of spectral data. Euclid is designed for slitless spectroscopy, which allows the telescope to produce spectra for all objects in its field of view at once. Although this produces a huge amount of data, the spectra from two objects can overlap, making it very difficult to ascertain the spectral properties of the objects we are observing. As a first step in creating a fully automated deep-learning-based pipeline to extract spectral information from Euclid data, we need to extract the 1D spectra from clean 2D spectrograms. For this purpose, we use 2D spectrgorams (images) from JWST and their corresponding 1D sepctra (sequences) to train a deep-learning model our model has two stages trained simultaneously, the first stage predicts the spectra for the input spectrogram, and the second stage predicts the residual between the prediction from the first stage and the expected spectra, conditioned on the initial 2D spectrogram and the prediction from the first stage. The JWST spectrometer observes between 0.5 − 5.5μm, with an approximate resolution of 0.01μm per pixel, leading to a full spectra size of around 450 pixels. We divide the full spectra into X equal parts and train the network on these sub-samples. We then stitch the predictions together to form a single full-length spectrum. We find that the two-stage model, when trained on sub-samples, is able to predict the 1D spectra, and the model finds it very difficult to accurately predict when fed the full spectrogram. The mode cannot be trained one stage at a time to reproduce the quoted level of performance. We find that the network only works accurately in some regimes of SNR and spectra types. This is the first step towards creating a fully automated deep-learning-based spectral extraction pipeline for Euclid data. ID: 204
/ Poster No. # 10: 024
Modalities: Image Methods: Uncertainty Quantification Application Domain: Health Segmentation of glands in big cohort data by deep learning Universitätsmedizin Greifswald, Germany Thyroid hormones are essential regulatory molecules with critical roles in growth, development and metabolism. There is raising evidence that even small variation in thyroid function, even within the reference range, is associated with adverse clinical outcomes. The synthesis and secretion of thyroid hormones is tightly regulated by the hypothalamus – pituitary – thyroid axis (HPT axis) through a negative feedback loop. The pituitary gland releases thyroid stimulating hormone (TSH) to stimulate the thyroid gland to produce T3 and T4 hormones, which in turn signal the pituitary to reduce TSH production at high levels. The morphology parameters including the volume of these glands may affect the variation in thyroid hormones. The aim of our project is the automatic segmentation of the pituitary and thyroid glands from MRI images in big cohort data. In our study, U-Net 3D model was successfully trained on the segmentation of these glands from MRI images in two population-based cohorts of the Study of Health in Pomerania (SHIP), and will be applied on the data of the German National Cohort (NAKO). As quality control, we applied Monte-Carlo dropouts to set up a pipeline for uncertainty quantification. So far, the model reached a dice score of 0.87 ± 0.06 for pituitary glands and 0.80 ± 0.06 for thyroid glands, compared with the dice score of 0.76 for thyroid segmentation in VIBESegmentor project by Robert Graf et al, 2024. The model prediction on unseen data was verified against volume data provided by experienced radiologists on SHIP data, showing high correlation with Pearson correlation coefficient (PCC) scores. PCC scores of pituitary volumes were 0.75 in SHIP-TREND-0 (N=1830) and 0.81 in SHIP-START-2 (N=1005). For thyroid volumes, PCC scores were 0.85 in SHIP-TREND-0 (N=1597) and 0.84 in SHIP-START-2 (N=638). This model showed robustness in a big cohort dataset, establishing the basis for future analysis in NAKO data. Besides, our work successfully overcame the difficulty of small glands in segmentation tasks as well as the high uncertainty in big dataset. ID: 198
/ Poster No. # 10: 025
Modalities: Image Methods: Generative Models Application Domain: Health Subclass-Conditioned Flow Matching for Long-Tailed Chest X-Ray Augmentation 1: Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU), Germany; 2: Imperial College London, UK Long-tailed class imbalance in chest X-ray datasets limits classifier performance on rare but clinically important findings. Generative augmentation can help, yet coarse disease labels often mix multiple latent patterns, causing conditional generators to overproduce dominant submodes. [1] Holste, Gregory et al. “Long-Tailed Classification of Thorax Diseases on Chest X-Ray: A New Benchmark Study.”Data augmentation, labelling, and imperfections : second MICCAI workshop, DALI 2022, Singapore, September 22, 2022, proceedings. DALI (Workshop) (2nd : 2022 : Singapore) vol. 13567 (2022): 22-32. doi:10.1007/978-3-031-17027-0_3 ID: 376
/ Poster No. # 10: 026
Modalities: Image Methods: Foundation Models, Other Application Domain: Health MoA: Mixture of Aggregators Improves Slide-Level Diagnosis in Computational Pathology 1: Helmholtz Munich, Germany; 2: TUM; 3: LMU Multiple instance learning (MIL) is the dominant paradigm for deriving patient-level predictions from whole slide images and blood smear cytology. Standard pipelines encode patches or cells into embeddings and compress them through a single attention-based pooling module into one slide-level vector. This assumes one aggregation strategy suffices for all diagnostic contexts. However, pathology and hematology specimens exhibit rich morphological heterogeneity: different diseases show varying cell compositions and distributional signatures that a single pooling mechanism cannot sufficiently capture. To address this, we introduce Mixture of Aggregators (MoA), training several aggregation modules in parallel within a unified MIL pipeline. Each aggregator is provided with shared instance features but each of them aggregates the instances in a unique way to create patient representation. A lightweight gating network selects the two most informative aggregators per slide, whose weighted outputs fuse into a patient-level representation. Load-balancing regularization, stochastic gating perturbation, and temperature annealing maintain training stability and encourage diverse specialization. We evaluate on 19 clinical tasks from 16 public datasets covering histopathology and hematologic cytology across 13 anatomical regions, including genetic subtype classification, biomarker prediction, immune categorization, and tumor grading. Pathology slides are encoded with UNI, cytology with DinoBloom-B. A single configuration selected on one dataset is applied unchanged to all benchmarks. With attention-based experts, MoA achieves 4.5% average relative improvement; with Transformer experts, gains reach 12.6%, exceeding 50% on challenging cohorts like endometrial immune classification. Attention analyses on leukemia cytology reveal that aggregators highlight distinct disease-relevant cell populations. Divergence measurements confirm meaningful separation (mean JSD = 0.42), showing that each aggregator captures complementary morphological evidence contributing to the final diagnosis. MoA provides an architecture-agnostic strategy boosting diagnostic accuracy while producing complementary, clinically interpretable attention patterns with minimal computational overhead. Our extensive evaluation across diverse organs, modalities demonstrates that MoA generalizes robustly making it a practical drop-in upgrade for existing MIL pipelines.
ID: 231
/ Poster No. # 10: 027
Modalities: Text Methods: Other Application Domain: Aeronautics, Space & Transport Topic Modeling and Opinion Analysis of Hyperloop Discussions on Twitter Using Latent Dirichlet Allocation 1: University of Stuttgart , Germany; 2: University of Southern California, USA; 3: Boston University, USA; 4: ARENA2036 e.V.; 5: University of Mumbai, India; 6: Hochschule Bremen, Germany; 7: Swinburne University of Technology, Melbourne, Australia The Hyperloop concept, proposed by Elon Musk in 2013, has attracted global attention as a potential next-generation transportation system. Understanding public perception of such emerging technologies is important for evaluating societal acceptance and long-term feasibility. This study applies Latent Dirichlet Allocation (LDA) topic modeling to analyze discussions about Hyperloop on Twitter and identify underlying themes and shifts in public sentiment. The research follows three stages: data collection, text preprocessing, and topic modeling. Twitter data related to Hyperloop was collected using Apify and processed through a natural language processing pipeline including tokenization, stop-word removal, and text normalization. LDA was then applied to extract latent topics from the dataset and evaluate topic coherence. To validate the robustness of the approach, the methodology was additionally applied to a Vande Bharat train dataset, enabling comparison of topic coherence and preprocessing effects. The final analysis focused on the Hyperloop dataset, where LDA revealed several prominent themes including technological innovation, transportation infrastructure, public perception, and references to Elon Musk. In addition, Google Trends data was incorporated to examine temporal changes in public attention and opinion dynamics related to Hyperloop discussions. The results indicate a gradual shift in public discourse from early enthusiasm toward increased skepticism and uncertainty regarding the feasibility and implementation of Hyperloop technology. These findings demonstrate the potential of topic modeling techniques to analyze large-scale social media discussions and provide insights into public attitudes toward emerging transportation innovations. ID: 229
/ Poster No. # 10: 028
Modalities: Text Methods: Other Application Domain: Aeronautics, Space & Transport Sentiment Analysis of Indian Hyperloop Tweets for Opinion Inversion Prediction 1: University of Stuttgart, Germany; 2: University of Southern California, USA; 3: Boston University, USA; 4: ARENA2036, Germany; 5: Deakin University; 6: Swinburne University of Technology, Australia; 7: University of Mumbai, India Social media platforms have become influential spaces for public discourse on emerging technologies. This study analyzes sentiments expressed by Indian users toward the Hyperloop transportation system using a large dataset of tweets. By examining public discussions, the work aims to identify prevailing attitudes and assess the potential acceptance and market viability of Hyperloop technology in India. From a dataset of approximately 18,557 relevant tweets, 1,667 source–quote associations were extracted, representing 18.84% of the overall score–quote pairs. These quote pairs were used to investigate the phenomenon of opinion inversion (OI), where the meaning or sentiment of quoted text differs from the original context. Detecting such inversions is important for understanding how opinions evolve and propagate in online discussions. Five machine learning models were evaluated for the task of predicting opinion inversion in the dataset. Among them, the Random Forest model achieved the best performance with the highest mean ROC–AUC score. The selected model was further optimized using RandomizedSearchCV for hyperparameter tuning, resulting in a test accuracy of approximately 80%. The results demonstrate that opinion inversion within complex social media discussions can be effectively detected using relatively simple machine learning approaches. These findings highlight the importance of identifying shifts in sentiment and meaning within online discourse surrounding emerging transportation technologies. Such predictive insights can support better analysis of public perception and contribute to understanding the societal acceptance of innovations such as Hyperloop. ID: 268
/ Poster No. # 10: 029
Modalities: Image, Simulation Data Methods: Physics-informed Machine Learning, Other Application Domain: Health Test-time augmentation with synthetic data addresses distribution shifts in spectral imaging 1: Division of Intelligent Medical Systems (IMSY), German Cancer Research Center (DKFZ), Heidelberg, Germany; 2: Helmholtz Information and Data Science School for Health, Karlsruhe/Heidelberg, Germany; 3: Faculty of Mathematics and Computer Science, Heidelberg University, Heidelberg, Germany; 4: Department of General, Visceral, and Transplantation Surgery, Heidelberg University Hospital, Heidelberg, Germany; 5: National Center for Tumor Diseases (NCT), A Partnership between DKFZ and University Medical Center Heidelberg, Heidelberg, Germany; 6: Department of Urology, University Medical Center Mannheim, Heidelberg University, Mannheim, Germany; 7: Medical Faculty, Heidelberg University, Heidelberg, Germany Purpose: Methods: Results: Conclusion: ID: 241
/ Poster No. # 10: 030
Modalities: Image Methods: Generative Models, Physics-informed Machine Learning Application Domain: Core Machine Learning HoloZip: A Machine Learning-Based Approach for Extreme X-ray Holographic Data Compression 1: Deutsches Elektronen-Synchrotron DES, Germany; 2: Department Physik, Universität Hamburg UHH, Germany; 3: HELMHOLTZ IMAGING Modern X-ray free-electron laser (XFEL) and synchrotron facilities generate data at an unprecedented and rapidly accelerating scale. Advances in detector technology, high-repetition-rate sources, and large-area pixel detectors have dramatically increased the data produced in X-ray imaging, diffraction, and holographic experiments. In addition, the development of next-generation synchrotron light sources, such as diffraction-limited storage rings (e.g., PETRA IV), further increases experimental throughput and data acquisition rates. As a result, large experimental campaigns routinely generate massive datasets, creating significant challenges for data storage, transfer, and downstream analysis. This rapid growth makes efficient data compression an essential component of modern X-ray science infrastructures. ID: 183
/ Poster No. # 10: 032
Modalities: Audio, Graphs, Image, Multimodal Data, Simulation Data, Tabular Data, Text Methods: Probabilistic Methods Application Domain: Core Machine Learning Optimal conversion from Rényi Differential Privacy to $f$-Differential Privacy 1: Helmholtz Munich, Germany; 2: Technical University of Munich, Germany; 3: Harvard University, USA; 4: Hasso-Plattner-Institut, Germany Differential Privacy (DP) (Dwork et al., 2014) has become the rigorous standard for privacy-preserving AI. In this work, we solve the problem of optimally converting an RDP profile $\rho(\cdot)$ into $f$-DP, formally proving a key standing conjecture stated in Appendix F.3 of Zhu et al. (2022). ID: 260
/ Poster No. # 10: 033
Modalities: Image, Multimodal Data, Text Methods: Foundation Models Application Domain: Core Machine Learning SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport 1: Helmholtz Munich, Germany; 2: Technical University of Munich, Germany; 3: Munich Center for Machine Learning, Germany; 4: Munich Data Science Institute, Germany; 5: Télécom Paris, France; 6: École Polytechnique, France The Platonic Representation Hypothesis posits that neural networks trained on different modalities converge toward a shared statistical model of the world. Recent work exploits this convergence by aligning frozen pretrained vision and language models with lightweight alignment layers, but typically relies on contrastive losses and millions of paired samples. In this work, we ask whether meaningful alignment can be achieved with substantially less supervision. We introduce a semi-supervised setting in which pretrained unimodal encoders are aligned using a small number of image-text pairs together with large amounts of unpaired data. To address this challenge, we propose SOTAlign, a two-stage framework that first recovers a coarse shared geometry from limited paired data using a linear teacher, then refines the alignment on unpaired samples via an optimal-transport-based divergence that transfers relational structure without overconstraining the target space. Unlike existing semi-supervised methods, SOTAlign effectively leverages unpaired images and text, learning robust joint embeddings across datasets and encoder pairs, and significantly outperforming supervised and semi-supervised baselines. ID: 330
/ Poster No. # 10: 034
Modalities: Other Methods: Probabilistic Methods, Uncertainty Quantification Application Domain: Health AtlasAlign: PointTransformer-based alignment of spatial omics to mouse brain common coordinate framework 1: Technical University of Munich (TUM), Germany; 2: Institute for Stroke and Dementia Research (ISD), Germany; 3: Institute of Neuroscience and Medicine (INM-1) Forschungszentrum Jülich, Germany; 4: Institute of Computational Biology, Helmholtz Munich, Germany Spatial transcriptomics (ST) datasets capture gene expression with cellular resolution, but integrating them into a shared anatomical reference remains challenging due to diverse gene panels, batch effects, and the need for section-wide global context in existing methods. We present AtlasAlign, a point transformer-based approach that directly maps local ST patches (1 mm diameter) to coordinates in the Allen Mouse Brain Common Coordinate Framework (CCF), enabling fine-grained 3D anatomical registration without requiring full tissue sections or manual region labels (Fig 1). To achieve cross-dataset generalizability, we unify gene panels to 6,000 genes via Tangram and the ABC Atlas [1], and remove batch effects using Harmony/Symphony. Cell patches are downsampled to 128 cells, normalizing for varying cell densities and enabling uncertainty quantification via repeated stochastic sampling. The model is pretrained on 2 million patches from two whole-brain MERFISH datasets (Zhuang [2] and Zeng [1]) using a joint coordinate regression and domain-adversarial (DANN) objective, learning dataset-invariant spatial representations from gene expression and local cell neighborhoods. New datasets can be incorporated via lightweight DANN fine-tuning with a SpatialNCE loss, requiring only a single tissue section. Here, the pretrained model's zero-shot predictions identify a subsample of source patches from the same brain region as implicit spatial anchors. On held-out MERFISH data, AtlasAlign achieves R² scores of 0.96–0.99 for CCF coordinate prediction (Fig 2). The stochastic patch sampling further enables calibration analysis, with empirical calibration errors of 0.10–0.17 (Dorsal-Ventral) and 0.04–0.13 (Left-Right). Comparison against FuseMap demonstrates competitive anatomical annotation performance (Fig 3). AtlasAlign's direct coordinate mapping also enables fine-grained annotation: fine-tuned on an in-house coronal section, it correctly resolves isocortex layers. Unlike global registration methods such as STAlign, AtlasAlign handles fragmented sections, as its patch-wise predictions rely only on local neighborhoods. In future work, we will extend AtlasAlign to sections with arbitrary cutting angles, both hemispheres, and other ST technologies. [1] Yao et al. A high-resolution transcriptomic and spatial atlas of cell types in the whole mouse brain. Nature 2023 [2] Zhang et al. A Molecularly defined and spatially resolved cell atlas of the whole mouse brain. Nature 2023
ID: 158
/ Poster No. # 10: 035
Modalities: Other Methods: Foundation Models, Physics-informed Machine Learning Application Domain: Health UdonPred: Untangling Protein Intrinsic DisorderPrediction 1: School of Computation, Information and Technology (CIT), Faculty of Informatics, Chair of Bioinformatics & Computational Biology - i12, Technical University of Munich (TUM), Boltzmannstr. 3, 85748 Garching/Munich, Germany; 2: Institute of Computational Biology (ICB), Helmholtz München, Ingolstädter Landstraße 1, 85764 Neuherberg, Germany; 3: Translational Microbiome Data Integration, School of Life Sciences, Technical University of Munich (TUM), Alte Akademie 8, 85354 Freising, Germany; 4: Institute for Advanced Study (TUM-IAS), Technical University of Munich (TUM), Lichtenbergstr. 2a, 85748 Garching/Munich, Germany; 5: School of Life Sciences Weihenstephan (TUM-WZW), Technical University of Munich (TUM), Alte Akademie 8, 85354 Freising, Germany Motivation: Regions in intrinsic disordered proteins (IDPs) constitute important continuous aspects of Results: Building on recently released datasets of continuous protein disorder and flexibility, we introduce ID: 254
/ Poster No. # 10: 036
Modalities: Text Methods: Foundation Models Application Domain: Core Machine Learning Fused Transformer Blocks Improve Pretraining Efficiency and Contextual Language Understanding factorize.bio, Germany Transformers conventionally compose self-attention and feed-forward networks serially through the residual stream. We revisit this design and study a fused transformer block in which the self-attention output is fused directly into the feed-forward network. Across parameter- and FLOP-matched decoder-only language models spanning 50M to 400M parameters, the fused design consistently improves validation loss relative to a standard serial baseline. These pretraining gains translate into disproportionally larger improvements on context-dependent downstream evaluations, including Lambada, SQuAD, and CoQA. The strongest gains appear in settings where task-relevant information is explicitly present in the prompt, such as known-answer few-shot CoQA variants, suggesting that the fused design improves the model’s use of provided context information. Our findings indicate that transformer block wiring remains an underexplored design dimension, and that tighter attention–FFN coupling can improve both pretraining efficiency and context-sensitive language understanding. ID: 396
/ Poster No. # 10: 037
Persona Vectors: Steering Towards Faithful Verbalized Uncertainty Expression in LLMs 1: Helmholtz AI, Munich; 2: Technical University of Munich; 3: University of Toronto; 4: Technical University of Nuremberg Reliable uncertainty communication is essential for trustworthy LLM use. Beyond uncertainty estimation itself, we focus on whether models’ utterances faithfully express their intrinsic uncertainty in text for users, which is critical in safety-sensitive settings. Recently, an interest has grown in assessing and improving this uncertainty verbalization and its faithfulness. This work extends the human metacognition-inspired prompt methodology MetaFaith by Liu et al. [2025] and tackles the problem through mechanistic interpretability. Our goal is to identify model components that drive faithful and accurate uncertainty descriptions, then steer LLM behavior at test time using those components, emulating the effect of metacognitive prompts. Our approach builds on methods for discovering persona vectors [Chen et al., 2025], demonstrated to be useful for identifying and controlling the effects of harmful or unhelpful LLM behaviors. We trace the persona vectors that reproduce system-prompt effects within the model and apply contrastive steering [Panickssery et al., 2024] to modulate activations at test time, yielding more faithful utterances. To uncover these directions, we construct paired sets of faithful and deliberately unfaithful prompts and derive a contrastive direction in representation space. Unlike rigid MetaFaith prompts, our method supports graded control over steering strength and simultaneously tracks which internal components contribute most. References Gabrielle Kaili-May Liu, Gal Yona, Avi Caciularu, Idan Szpektor, Tim G. J. Rudner, and Arman Cohan. Metafaith: Faithful natural language uncertainty expression in llms, 2025. URL https://arxiv.org/abs/2505.24858. Runjin Chen, Andy Arditi, Henry Sleight, Owain Evans, and Jack Lindsey. Persona vectors: Monitoring and controlling character traits in language models, 2025. URL https://arxiv.org/abs/2507.21509. Nina Panickssery, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner. Steering llama 2 via contrastive activation addition, 2024. URL https://arxiv.org/abs/2312.06681 ID: 131
/ Poster No. # 10: 038
Modalities: Image, Multimodal Data, Text Methods: Other Application Domain: Core Machine Learning Scalable and interpretable representation alignment with ordinal similarity 1: Helmholtz Munich, AIH, Germany; 2: Technical University of Munich, Germany; 3: University of Warsaw, Poland Evaluating representation similarity is fundamental to representation learning. However, existing metrics suffer from significant limitations: they are difficult to interpret due to shifting baselines, lack robustness to outliers, and are frequently computationally intractable for large datasets, forcing a reliance on heuristic approximations. To address these shortcomings, we develop an ordinal-similarity framework, instantiated by the Triplet (TSI) and Quadruplet (QSI) Similarity Indices, which measure alignment by quantifying the consistency of ordinal relationships. We provide a theoretical analysis demonstrating that this formulation is inherently interpretable, robust to outliers, and computationally efficient. Finally, we establish a formal equivalence between TSI alignment and the alignment of local neighborhood structures, as measured by Mutual Nearest Neighbors. Through empirical analysis, we validate these properties and show that ordinal similarity offers a scalable, practical approach to measuring alignment, enabling practitioners to better understand and design representations. ID: 244
/ Poster No. # 10: 039
Modalities: Multimodal Data Methods: Reinforcement Learning Application Domain: Earth & Environment Multimodal Data-driven Transfer Learning Impelled Framework for Spatiotemporal Forecasting of Urban Pollution (PM2.5) Research Institute for Sustainability, Germany More and more people are moving to cities, and air pollution in cities is now a big hazard for the environment and public health. For city planning and long-term environmental protection, it is very important to accurately measure fine particulate matter (PM2.5), which is a major air pollutant. The complicated way that weather, air pollution, and human activity are related makes it very hard to utilize standard statistical forecasting methods to capture the highly nonlinear patterns in space and time. Recent advancements in AI have enabled the analysis of extensive environmental information, facilitating the elucidation of complex linkages. This proposed method uses a multimodal deep learning framework to predict urban air pollution (PM2.5) over time and space. The suggested method uses a lot of different kinds of environmental data from urban air quality monitoring stations, such as meteorological data. These data sources offer many modalities that, when combined, show how the urban air system is changing. The proposed approach has the potential to encapsulate both temporal dependencies and inter-variable interactions influencing pollutant dynamics by integrating data from several dimensions into a unified deep learning framework. Deep neural networks are used by the multimodal architecture to discover hidden delays in each item of environmental input. Then, it uses a deep learning layer to integrate both representations by moving shared attributes from one to the other. This way of combining helps the model learn about the complicated relationships between pollution levels and weather conditions at different times and places. Deep learning-based feature learning will help with a method for weighting features based on attention that dynamically prioritizes changes in the environment based on different atmospheric conditions. The model accurately represents nonlinear pollutant interactions and improves forecasting stability during high pollution alerts in urban settings. The learnt multimodal variational lags not only make forecasts more accurate, but they also explain how diverse atmospheric variables are connected to each other. This helps make better decisions on policies and monitoring the environment. The suggested method shows how multidimensional data-driven AI may make urban air quality intelligence systems better by combining different environmental data sources into a single prediction framework.
ID: 166
/ Poster No. # 10: 040
Modalities: Other Methods: Probabilistic Methods Application Domain: Health A Hierarchical Bayesian Multiple Instance Learning Framework for Joint Localization and Testing of Spatial Differential Expression 1: Helmholtz Munich, Germany; 2: Technical University of Munich, Germany; 3: Regensburg University, Germany; 4: ETH Zurich, Switzerland Recent multi-subject spatial transcriptomics datasets enable differential expression (DE) analysis across conditions in tissue context, yet most comparative approaches aggregate expression within predefined regions, fixing the spatial scale and location of analysis. Here we introduce SpaceDX, a hierarchical Bayesian framework for subject-level spatial DE that replaces hard regional aggregation with gene-specific attention weights that adaptively aggregate spot-level expression in the context of differential expression between subject groups. Using attention-based multiple instance learning, SpaceDX learns these weights jointly with the differential expression model, enabling unified spatial localization and statistical testing within a single probabilistic framework. Across two mouse brain datasets spanning multiple subjects, SpaceDX detected 45.8% more DE genes than region-based pseudobulk approaches in a chronic stress model and increased discovery by 19.1% when combined with DESeq2 in an Alzheimer’s model, while producing coherent spatial localization maps. These results show that integrating localization and testing improves sensitivity and spatial resolution in comparative spatial transcriptomics.
ID: 355
/ Poster No. # 10: 043
Modalities: Image, Simulation Data Methods: Probabilistic Methods, Uncertainty Quantification Application Domain: Information Trustworthy Deep Surrogates for Inverse Materials Design with Calibrated Uncertainty 1: Institute for Advanced Simulations - Materials Data Science and Informatics (IAS-9), Forschungszentrum Jülich GmbH; 2: Chair of Materials Data Science and Materials Informatics, Faculty 5 – Georesources and Materials Engineering, RWTH Aachen University The integration of Artificial Intelligence (AI) in the field of materials science has been key to modelling process-structure-property (PSP) linkages in accelerating the development of sustainable materials. Deep learning (DL) models are able to provide accurate point predictions for regression tasks, but they are known to be overconfident and lacking in information regarding the quantity, coverage and type of uncertainties. This prevents the application of DL models for inverse material design and combination with active learning for the exploration of the material design space. Therefore, uncertainty quantification (UQ) methods are incorporated to bridge this gap and provide information regarding model accuracy and uncertainty. UQ frameworks such as Bayesian Neural Network (BNN), Gaussian Process (GP), Conformal Predictor (CP) and Conformalized Quantile Regression (CQR) are studied and compared to a ResNet18 base architecture (He et al., 2016). BNNs incorporate Bayesian inference into neural networks (Jospin et al., 2020) while GPs are non-parametric Bayesian models that describe a distribution over functions (Rasmussen and Williams, 2006). CPs generate statistically guaranteed, distribution-free intervals (Angelopoulos and Bates, 2023), and when coupled with quantile regression in CQRs deliver calibrated prediction intervals with statistically guaranteed coverage (Romano et al., 2019). Case studies with datasets based on the Ising model and Cahn-Hilliard equations, which have known structure-property relations show that BNN, CQR and the pairing of GP with CP capture model and data uncertainties with well-calibrated prediction intervals. The combination of CP with ResNet18 delivers similar point-prediction accuracy with prediction intervals of comparable coverage, albeit being able to quantify only the model uncertainty. These results highlight the importance of combining DL models with UQ methods in instilling trustworthiness into black-box surrogate models. For data-driven materials design, trustworthy models equip forward PSP linkage modelling with prediction intervals that integrates risk-awareness in the design cycle. The combination of uncertainty-aware models and active learning algorithms with Bayesian optimisation also results in a closed-loop material discovery pipeline for inverse design. While benchmarked on simulated microstructural datasets, this approach is data-agnostic and applicable on experimental data. ID: 307
/ Poster No. # 10: 044
Modalities: Image Methods: Reinforcement Learning, Other Application Domain: Health In-Model Bias Mitigation Methods for Age Confounding in Dementia Detection 1: Deutsches Zentrum für Neurodegenerative Erkrankungen (DZNE), Rostock, Germany; 2: Klinik für Psychosomatische Medizin und Psychotherapie (KPM), Universitätsmedizin Rostock, Germany Background: Deep learning models for MRI-based dementia detection are vulnerable to confounding by age, which is strongly associated with both brain structure and disease prevalence. This can lead convolutional neural networks (CNNs) to exploit age-related signals instead of disease-specific biomarkers. We therefore investigated two approaches for age-bias mitigation: an adversarial confounder-free network (CF-Net) and Penalty-based Metadata Normalization (PMDN).
ID: 242
/ Poster No. # 10: 045
Modalities: Audio, Multimodal Data, Time Series, Video, Other Methods: Other Application Domain: Health Modeling Neuro–Behavioral State Dynamics from Longitudinal Multimodal Recordings 1: Deutsches Zentrum für Neurodegenerative Erkrankungen e. V. (DZNE), Germany; 2: Helmholtz Munich, Neuherberg, Germany; 3: University of Bonn, Medical Faculty, Bonn, Germany Continuous advances in neurotechnology allow for increasingly precise, large-scale read-outs of neuronal activity. These approaches have been complemented with markerless pose estimation via multi-camera tracking and miniaturized motion sensors or wearable devices for vital measures. Together, novel multimodal read-outs allow for synchronous data streams of neuronal and bodily parameters, which yield unprecedentedly rich datasets. Machine learning algorithms were recently utilized to extract multimodal neuronal (Schneider et al. 2023, Nature) and behavioral structures (Weinreb et al. 2024, Nature Methods) at scale. However, despite these advances, the field is still lacking models that enable joint prediction of behavioral and neuronal states and capture the transitions between these states that shape neural dynamics during behavior. To this end, we collected a longitudinal, multimodal dataset of mice (> 400 recording hours across several weeks) that includes neuronal recordings (calcium imaging, multi-site EEG), multi-angle video, ambient sound as well as head-movement sensor data. Mice moved freely in a homecage-like environment and expressed a wide range of self-paced behaviors including exploration and sleep. In addition, we extracted fine-grained pose trajectories, behavioral motifs, and sleep states to generate a synchronized, ML-ready dataset linking neuronal, behavioral, and physiological signals. Using this dataset, we established a scalable framework for multimodal neural state modeling that enables the prediction of state transitions and delivers interpretable insights into the dynamics of state encoding. Our final goal is the deployment of a real-time neuro-behavioral state inference system to facilitate closed-loop state-dependent perturbations. This framework will facilitate the systematic study of neuronal state dynamics across learning, aging, and disease states and serve as a foundational step toward real-time, multimodal brain–machine interfaces in preclinical research. ID: 360
/ Poster No. # 10: 046
Modalities: Audio, Multimodal Data, Time Series, Video, Other Methods: Generative Models, Other Application Domain: Core Machine Learning, Health Consistent Representation Learning for Modeling Neuro-Behavioral State Dynamics 1: Helmholtz Zentrum München Deutsches Forschungszentrum für Gesundheit und Umwelt ( GmbH ) Germany; 2: German Center for Neurodegenerative Diseases (DZNE), Bonn, Germany; 3: University of Bonn, Medical Faculty, Bonn, Germany; 4: Munich Center for Machine Learning Complex biological systems are often observed through high-dimensional measurements. In these cases, it is common to assume that observations are generated by lower-dimensional latent states whose dynamics capture the system's relevant structure. Recovering these latent representations and their dynamics is a central challenge in systems neuroscience. ID: 288
/ Poster No. # 10: 047
Modalities: Image Methods: Other Application Domain: Matter Self-supervised Denoising Of Raw Tomography Detector Data For Improved Image Reconstruction Helmholtz-Zentrum Dresden - Rossendorf e. V., Germany Ultrafast electron beam X-ray computed tomography produces noisy data due to short acquisition times, resulting in reconstruction artifacts and reduced image quality. To address this challenge, we investigated two self-supervised deep-learning methods for denoising raw detector data and compared them with a conventional non-learning-based denoising approach. The methods were evaluated at both the sinogram and reconstruction levels using quantitative image quality measures, including PSNR and SSIM, as well as line-based analyses such as CTF and MTF. We found that the deep-learning-based methods enhanced the signal-to-noise ratio and led to overall improvements in the reconstructed images, outperforming the conventional non-learning-based method. Although the improvements in CTF and MTF were marginal, the overall findings indicate that self-supervised denoising of raw detector data can improve reconstruction quality. Future work will focus on further optimizing denoising performance and on investigating downstream tasks such as segmentation for additional validation.
ID: 222
/ Poster No. # 10: 048
Modalities: Image, Multimodal Data, Text Methods: Agentic AI Application Domain: Health Semantic Harmonization of Noisy Medical Imaging Cohorts via Vision-Language Model 1: Friedrich-Alexander-Universität Erlangen-Nürnberg, Germany; 2: Imperial College London, UK; 3: University of Zurich, Switzerland; 4: ETH Zurich, Switzerland; 5: Istanbul Medipol University, Turkey Introduction With the growth of population-scale CT datasets, the main challenge for medical foundation models has shifted from data scarcity to semantic reliability. Anatomical labels are often inferred from unreliable DICOM metadata [1], introducing semantic noise that degrades downstream robustness and generalization. Manual curation is infeasible at scale, while single-source automated methods remain vulnerable to corrupted signals. Method We propose a vote-aware semantic aggregation framework that formulates anatomical label inference as a multi-source consensus problem. Imaging content, DICOM metadata, and report-derived cues act as independent voters generating modality-specific hypotheses. Each modality is first processed by a large language model (LLM) to generate candidate labels conditioned solely on its own evidence. The LLM then performs structured cross-source arbitration by jointly evaluating candidate hypotheses and contextual signals. By reasoning over cross-source consistency and uncertainty, the system resolves conflicts and produces a coherent decision, enabling robust consensus under noisy or partially corrupted inputs. Evaluation and Conclusion The framework is applied to a dataset conceptually aligned with the scale and heterogeneity of CT-RATE [2], spanning head–neck to pelvis; evaluation uses 362 expert-annotated studies. Our method achieves a macro-F1 of 0.892 ± 0.023, outperforming DICOM baselines (0.662 ± 0.032) and state-of-the-art body-part regression (BPR) models [3] (0.835 ± 0.021). Single-source methods show region-specific instability: DICOM labeling drops to 0.577 ± 0.157 in the neck, while BPR degrades to 0.541 ± 0.110. In contrast, our framework maintains stable performance (0.790 ± 0.108). Gains in macro-precision (0.911 ± 0.012) and macro-recall (0.895 ± 0.030) further indicate improved cross-class consistency and robustness to modality failures. These results demonstrate that vote-aware multi-source reasoning improves the semantic reliability of large-scale CT datasets and offers a robust framework for reliable anatomical labeling in clinical deployment. [1] Gueld, M.O. et al.: Quality of DICOM header information for image categorization. Proc. SPIE 4685 (2002). ID: 346
/ Poster No. # 10: 049
Modalities: Tabular Data Methods: Other Application Domain: Health Enhancing Clinical Immunology and Disease Diagnostics with Global Swarm Learning 1: DZNE, Germany; 2: Calico Life Sciences; 3: Hewlett Packard Enterprise, Houston, Texas, USA; 4: The Global Swarm Learning Group High-dimensional cytometry (HDC) has transformed human immunology, enabling the study of health and disease at single-cell granularity. Especially during the COVID-19 pandemic, HDC with specialized panels drove new disease understanding. However, cytometry data’s heterogeneity limits its potential for clinical machine learning (ML). The large variety of platforms, instruments, and experimental settings introduces batch effects; low marker panel overlap renders many datasets incompatible; and data privacy regulations and technical barriers to large data transfer make data pooling difficult. Earlier, we proposed Swarm Learning as a democratic ML framework that overcomes limitations of local ML training while providing full data confidentiality. Here, we show how Swarm Learning can be applied to heterogeneous HDC data. We established a Global Swarm Learning Network for COVID-19 prediction, connecting 14 leading immunology institutions with data totaling 6182 samples from 1223 donors, combining different marker panels, technologies and COVID-19 disease states. We mapped all data to a common data model, developing a standardized pipeline for analysis, annotation and quality assessment. We replicated major biological findings across independent sites, such as reduced T-cell frequencies in patients with severe COVID-19. In the disease prediction setting, local models can achieve high performance when trained and tested on samples from the same site, but do not generalize across datasets from other sites; only models trained on diverse and sufficiently standardized data show broad generalization. Analysis of the common markers driving performance revealed the contribution of the T cell and monocyte compartments to disease prediction. In summary, this study highlights the potential of Swarm Learning to enable global collaboration, accelerate data harmonization and standardization efforts, and support a path towards ML-based clinical disease diagnostics. ID: 159
/ Poster No. # 10: 050
Modalities: Image, Video Methods: Generative Models Application Domain: Health Flow-aware Latent Diffusion for Synthetic Temporal Angiography 1: LMU Munich, Germany; 2: Heidelberg University, Germany Digital Subtraction Angiography (DSA) is a key imaging modality for the diagnosis and treatment of cerebrovascular diseases, providing high-temporal-resolution visualization of contrast agent dynamics in cerebral vessels. However, large-scale temporal DSA datasets are difficult to obtain due to acquisition costs, radiation exposure, and strict clinical data privacy constraints that limit data sharing. These challenges motivate the development of generative models capable of synthesizing realistic temporal angiography sequences. In this work, we present a framework for fully synthetic temporal cerebral DSA generation using a flow-aware latent diffusion model. Our approach introduces a content–motion factorization that explicitly disentangles static vascular anatomy from the dynamic propagation of contrast agent. The motion component is explicitly modeled via a flow-based mechanism, providing an inductive bias for temporally coherent contrast evolution while preserving anatomical structure. We evaluate the proposed method on real clinical DSA data using quantitative metrics, including Fréchet Video Distance, and a blinded clinical expert study with Likert-scale ratings. The results demonstrate that the generated sequences exhibit realistic vascular appearance and coherent temporal dynamics, highlighting the potential of generative models for scientific data synthesis and simulation in medical imaging.
ID: 306
/ Poster No. # 10: 051
Modalities: Graphs, Multimodal Data, Simulation Data, Text, Time Series Methods: Foundation Models, Graph Neural Networks Application Domain: Matter MatBind: Probing the multimodality of materials science with contrastive learning 1: Institute for Advanced Simulations (IAS-9), Forschungszentrum Jülich GmbH, Germany; 2: Helmholtz Institute for Polymer in Energy Applications Jena, Germany; 3: Institute of Nanotechnology, Karlsruhe Institute of Technology, Germany; 4: University of Cologne, Germany; 5: RWTH Aachen University, Germany; 6: Friedrich Schiller University Jena, Germany; 7: Jena Center for Soft Matter, Germany; 8: Center for Energy and Environmental Chemistry Jena, Germany; 9: Jülich Supercomputing Centre, Forschungszentrum Jülich GmbH, Germany Crystalline materials characterization is traditionally fragmented across isolated modalities—atomic structures, X-ray diffraction patterns, electronic density of states, and natural language—which are rarely integrated despite describing the same physical entity. We present MatBind, a contrastive learning framework that aligns these four disparate modalities into a unified embedding space by using the crystal structure as a central physical anchor. By training pairwise alignments between the structure and other data formats, the model induces a shared representation that enables strong cross-modal retrieval and allows modalities never explicitly paired during training, such as text and density of states, to develop meaningful mutual retrieval. Evaluations demonstrate that our model captures the properties of materials without directly related information and that combining modalities at query time improves retrieval beyond what either modality achieves alone. From these joint observations, MatBind learns a common language in which a diffraction pattern, an electronic structure, and a descriptive sentence can be compared, combined, and used to find one another, providing a scalable path for multimodal materials discovery. ID: 250
/ Poster No. # 10: 052
Modalities: Image Methods: Other Application Domain: Earth & Environment Creating a Benchmark Image Dataset for Reliable Quantification of Arctic Phytoplankton using AI 1: Alfred-Wegener-Institut, Germany; 2: University of Bremen, Germany Monitoring phytoplankton and other microscopic organisms is essential for assessing the status of marine ecosystems and carbon cycles and is directly linked to several Sustainable Development Goals (SDGs) and the European Water Framework Directive. Yet most of the microscopic environmental monitoring still depends on time intensive and subjective manual microscopy. Multispectral imaging flow cytometry combined with automated image recognition using Convolutional Neural Networks (CNNs) offers a high throughput alternative. However, the robustness and transferability of such approaches are currently still limited by e.g. batch effects and sensitivity to data shifts and image input from different modalities. The project AIMBIS (“Artificial Intelligence for Microscopic Biodiversity Screening”) addresses this issue by developing benchmark datasets for label-efficient machine learning to build reliable and precise recognition algorithms for taxonomic classification of phytoplankton and other microscopic organisms. We are currently creating an image dataset for Arctic phytoplankton, including images from cultivated species and natural samples collected in the Arctic Ocean. Samples were analysed using the Amnis Imaging Flow Cytometer and are currently being annotated. The annotation process highlights several challenges, including high morphological variability within species (differences in cell orientation, life stage, focus, health of the cell), taxonomic uncertainty in natural assemblages, images containing multiple cells as well as variability and consistency within different annotators. To date, 20 cultured phytoplankton samples have been annotated, with image numbers ranging from 50 to 8,000 per sample. The finalized Arctic benchmark dataset will provide a standardized foundation for improving the development and reliability of AI- based phytoplankton monitoring in the rapidly changing Arctic environment. ID: 235
/ Poster No. # 10: 053
Modalities: Graphs, Multimodal Data, Text Methods: Foundation Models Application Domain: Core Machine Learning, Health PromptGate: Federated Context Optimizationfor Dynamic VLM Gating in Open-Set MedicalImaging University of Bonn, University Hospital Bonn, Clinic for Diagnostic and Interventional Radiology, Bonn, Germany Background Purpose Materials and Methods Results Conclusion ID: 284
/ Poster No. # 10: 054
Modalities: Simulation Data, Time Series Methods: Physics-informed Machine Learning, Probabilistic Methods, Uncertainty Quantification Application Domain: Energy Neural Posterior Estimation for Empirical Power System Time Series Institute for Automation and Applied Informatics, Karlsruhe Institute of Technology, Germany Methods for state and parameter estimation are widely used to analyze complex systems, such as power systems. Estimating parameters is often necessary for follow-up simulations or effective control. Bayesian approaches allow us to go beyond point estimates and include domain knowledge when identifying parameters. However, such Bayesian approaches are often limited to simulated settings and not well-tuned for noisy, empirical data. We propose to investigate power systems, specifically the power grid frequency dynamics, as an important empirical use case with non-trivial data. ID: 398
/ Poster No. # 10: 055
Underconfident Predictions in MFVI Helmholtz Munich, Germany While conventional wisdom suggests that Mean Field Variational Inference (MFVI) consistently underestimates posterior variance, recent work in conjugate regression suggests that predictive uncertainty can be overestimated by MFVI for i.i.d. test data. This work extends that analysis to the broader exponential family, demonstrating that variance underestimation in the natural parameter space does not monotonically translate to the predictive manifold. We propose methods for reweighting the prior and likelihood terms in the ELBO to counteract this variance overestimation and validate these theoretical insights across several regression benchmarks. ID: 150
/ Poster No. # 10: 056
Modalities: Graphs, Simulation Data, Time Series Methods: Other Application Domain: Energy Differentiable Power-Flow Optimization 1: Karlsruhe Institute of Technology, Germany; 2: Helmholz AI, Karlsruhe With the rise of renewable energy sources and their high variability in generation, the management of power grids becomes increasingly complex. Traditional AC-power-flow simulations, which use the Newton-Raphson (NR) method, suffer from poor scalability, necessitating methods that can handle larger grids more effectively. We propose Differentiable Power-Flow (DPF) as a method to integrate first order optimizations into power-flow calculations. With DPF, the parameters are learned using reverse-mode automatic differentation. ID: 220
/ Poster No. # 10: 057
Modalities: Simulation Data, Tabular Data, Other Methods: Generative Models, Probabilistic Methods Application Domain: Core Machine Learning, Health A generative model for dimensionality reduction with millions of features and few samples 1: Helmholtz AI, Helmholtz Zentrum Munchen, Ingolstadter Landstraße 1,Neuherberg, Germany; 2: Center for Health Data Science, Department of Public Health, University of Copenhagen, Copenhagen, Denmark; 3: Department of Medical Sciences, University of Torino, Italy Training deep generative models for dimensionality reduction in extremely high-dimensional settings remains a major challenge, particularly when the number of samples is limited. In this work, we demonstrate that it is feasible to train a deep generative model with millions of features using only a few thousand samples. We hypothesize that, for decoder-only architectures, the number of training samples required is largely independent of the input feature dimensionality across a broad class of network architectures. We first validate this hypothesis through an extensive set of experiments on synthetic non-linear data, systematically varying both feature dimensionality and sample size. To further assess robustness in realistic genomic settings, we conduct preliminary experiments on the 1000 Genomes Project (1KGP) dataset, evaluating performance across different sample sizes prior to applying the method to cancer data. Finally, we train a Deep Generative Decoder (DGD) on a curated dataset from the International Cancer Genome Consortium (ICGC) comprising 4.4 million genomic features. The model is trained on approximately 4,000 samples and evaluated on 1,000 held-out samples. The learned latent representations exhibit clear biological structure, with distinct clustering by tumor type. When constrained to the same latent dimensionality, DGD outperforms both PCA and VAE in downstream tumor type classification. Notably, the proposed approach is computationally efficient and can be trained on a single 16GB GPU. ID: 152
/ Poster No. # 10: 058
Modalities: Time Series Methods: Other Application Domain: Earth & Environment "Hello World": How to Have a Conversation with the Global Weather 1: Karlsruhe Institute of Technology, Germany; 2: Helmholtz AI Data-driven models for weather and climate prediction are rapidly outperforming traditional, physics-based methods. However, the foundation of this revolution, the ERA5 global reanalysis dataset at over 950TB, remains practically inaccessible to the broader engineering and research communities. The ERA5 dataset contains dozens of highly specialized variables, such as potential vorticity, geopotential height, and specific cloud ice water content, spread across 37 distinct atmospheric pressure levels. To a software engineer, data scientist, or researcher without formal meteorological training, these variables are fundamentally difficult to interpret. Additionally, the sheer volume of this multi-dimensional data means that simply parsing it is challenging and requires parallel processing clusters, whilst use of the data is only possible with distributed computing and petabyte-scale storage. As we approach the even larger ERA6 generation, this gatekeeping by complexity and compute will only grow. We overcome this bottleneck by exposing ERA5 data through a Model Context Protocol (MCP) server. MCP provides a standardized, open-source architecture that allows Large Language Models (LLMs) to securely and dynamically connect to external tools and data stores. By successfully turning a targeted subset of ERA5 into an MCP-accessible resource, we demonstrate a functional proof-of-concept for natively querying decades of global climate data. While scaling to the full ERA5 data set is the next milestone, this initial implementation already completely bypasses traditional data-analysis, replacing complex code with intuitive natural language queries. Researchers, policymakers, and the general public can now rapidly extract insights, test hypotheses, and perform exploratory analytics without a single line of code. For the first time, instead of writing complex distributed computing scripts, you can finally just type "Hello World" and start a conversation with the global weather. ID: 335
/ Poster No. # 10: 059
Modalities: Image Methods: Other Application Domain: Health Segmentation of the intracranial cavity in T1-weighted MR images of the human head by a location-specific 3D U-Net 1: Institute for Neuroscience and Medicine (INM-1), Forschungszentrum Jülich, Germany; 2: C. and O. Vogt Institute for Brain Research, Medical Faculty, University Hospital, Heinrich-Heine-University of Düsseldorf, Germany; 3: Institute for Anatomy I, Medical Faculty, University Hospital, Heinrich-Heine-University of Düsseldorf, Germany The intracranial cavity (ICC) is a space in the head which contains the brain, cerebrospinal fluid (CSF), meninges and blood vessels. Its volume is often used as a reference variable in studies of brain structure in order to account for the inter-individual variability in brain size. However, the segmentation of the ICC in T1-weighted MR images is challenging because the visual appearance of its boundary is heterogenous depending on the location in the head, but also between images: E.g., the dorsal skull consists of several thin layers, which can differ in thickness and contrast, whereas the neck below the brain shows larger portions of fat tissue and muscles, and the brain can directly touch the surrounding meninges, but there might be also larger CSF filled spaces. Here we introduce a method for the ICC segmentation based on a convolutional neural network (CNN) with U-Net architecture to overcome these difficulties. All MR images were affinely registered to a reference brain image to reduce variability in global size and orientation, so that a data augmentation by geometrical transformations could be avoided. For the CNN training we used 98 T1 MR images from previous studies with manual ICC segmentation in every tenth sagittal section. Firstly, we trained a 2D U-Net CNN with the segmented sections to predict the ICC in the interleaving sections. This yielded 3D ICC masks, which were visually checked for correctness. Next, we used these 3D masks for the training of a 3D U-Net CNN. The images were partitioned in 27 overlapping tiles of length 96 voxels, and one model for each tile was trained, so that these were more specific for head parts. The binary predictions of models in overlapping parts of the tiles were averaged with weights depending on the distance to the tile centers. To validate this method, manual segmentations of 192 MR images (different from training images) where compared with their predicted segmentations. This yielded a mean Dice coefficient of 0.962 ± 0.017, and a mean relative mask size difference of 2.0 % per image. Next, ICC masks were predicted for 433 longitudinal image pairs of the 1000BRAINS study1 (mean time difference 3.7 ± 0.8 years), which yielded a mean relative ICC volume difference of 0.16 ± 0.28 %, showing the robustness of this method. The CNN trainings used an NVIDIA V100 GPU with 16 GB RAM, whereas the predictions were run within 11 sec per image on a QUADRO P2200 GPU (5 GB RAM). [1] Caspers S, et al (2014) Front Aging Neurosci 6 ID: 168
/ Poster No. # 10: 060
Modalities: Image Methods: Foundation Models Application Domain: Earth & Environment BioCUDA: An Accessible AI Pipeline for Fast, Accurate 3D MRI Reconstruction Across Species 1: Data Science Group, Alfred Wegener Institute, Helmholtz Centre for Polar and Marine Research, Bremerhaven; 2: Integrative Ecophysiology, Alfred Wegener Institute, Helmholtz Centre for Polar and Marine Research, Bremerhaven; 3: Piraud Team, Helmholtz AI, Helmholtz Munich Quantifying volumetric changes in biological tissues using imaging techniques such as magnetic resonance imaging (MRI) is essential for understanding how organisms respond to environmental stressors. This need is especially pronounced in studies involving multiple species and long‑term imaging series, where accurate 3D reconstructions are required to detect subtle structural changes under varying environmental conditions. BioCUDA is a desktop application designed to streamline the analysis of MRI data across diverse species and experimental setups. It addresses common challenges in reconstructing 3D volumes from 2D slice data, including motion artifacts, inconsistent image contrast, and variability in imaging protocols. Traditional reconstruction workflows are slow and labor‑intensive, while many AI‑based alternatives require specialized expertise and lack accessible interfaces, limiting their usability in environmental research. Built on the Napari ecosystem, BioCUDA supports DICOM formats, corrects temporal misalignment, and integrates MedSAM2‑3D for fast few‑shot segmentation, enabling accurate 3D reconstructions with minimal user annotations across unseen species. Additional tools for volume calculation and intuitive user interaction further simplify the workflow. Across heterogeneous datasets and multiple species, BioCUDA substantially reduces analysis time while maintaining high accuracy in volumetric measurements. By combining AI‑driven segmentation, automated image registration, and an accessible interface, it enables fast, reliable, and reproducible analysis of complex biological tissues. Its broad applicability lowers technical barriers for non‑experts, fostering interdisciplinary research and advancing the study of ecophysiological responses to environmental change. ID: 299
/ Poster No. # 10: 061
Modalities: Graphs, Simulation Data Methods: Foundation Models, Graph Neural Networks, Physics-informed Machine Learning Application Domain: Matter Machine-Learned Interatomic Potentials for Fe–Ni Alloys: Benchmarking MACE Models for Magnetism and Phase Stability 1: Helmholtz-Zentrum Dresden-Rossendorf (HZDR); 2: Center for Advanced Systems Understanding (CASUS) Iron–nickel (Fe–Ni) alloys are central to industrial metallurgy and planetary science, exhibiting rich composition-dependent phase stability, itinerant magnetism, and magneto-elastic effects including the Invar anomaly. Accurate and efficient modeling of these properties across compositions and pressures remains a key challenge in computational materials science. We develop a system-specific MACE potential (MACE-sqs) trained on spin-polarized DFT calculations using special quasirandom structures (SQS) to represent chemical disorder in Fe–Ni alloys. The training dataset spans bcc and fcc phases, isotropic and shear strains, and the full composition range from pure Fe to pure Ni. We benchmark MACE-sqs against four MACE foundation models across experimental and DFT reference data for equations of state, elastic constants, lattice volumes, pressure-induced phase transitions, and finite-temperature properties. MACE-sqs outperforms all foundation models — including those employing Hubbard U corrections — for structural and elastic properties in both crystal phases. We further demonstrate that DFT+U is fundamentally ill-suited for itinerant Fe–Ni alloys, systematically distorting equilibrium volumes and phase energetics. A key open challenge remains: all models invert the experimentally observed decrease in bcc–hcp transition pressure with increasing Ni content, revealing the need for training data that explicitly encodes high-pressure magnetic collapse. This work illustrates how targeted, domain-specific machine-learned interatomic potentials (MLIPs) trained on carefully curated DFT datasets can surpass large-scale foundation models for complex magnetic alloys, enabling near-DFT-accurate simulations at a fraction of the DFT computational cost — directly relevant to alloy design and geophysical modeling. ID: 203
/ Poster No. # 10: 062
Modalities: Image, Multimodal Data, Tabular Data Methods: Other Application Domain: Health Hybrid Multimodal Late Fusion Frameworks for Robust bvFTD Classification in Imbalanced Dementia Cohorts 1: German Center for Neurodegenerative Diseases (DZNE), Germany; 2: Institute and Policlinic of Radiology, Pediatric Radiology and Neuroradiology, University Medical Center Rostock, Rostock, Germany; 3: Department of Psychosomatic Medicine, Rostock University Medical Center, Rostock, Germany Background Dementia is a leading cause of disability among older adults, and its increasing prevalence demands accurate diagnostic methods. The behavioral variant of frontotemporal dementia (bvFTD) is an early-onset dementia subtype characterized by progressive neurodegeneration. While Alzheimer’s disease (AD) has been extensively studied, bvFTD remains underrepresented in computational frameworks. MRI is widely used to detect brain alterations, and deep learning has shown promise in dementia classification. However, the clinical presentation of bvFTD partially overlaps with AD, making differentiation challenging. Moreover, because bvFTD is less prevalent, its inclusion as a minority class in machine learning models often leads to reduced diagnostic sensitivity. To address this challenge, we propose two hybrid multimodal late-fusion frameworks integrating CNN-based image features with volumetric measures to improve bvFTD classification under class imbalance. Method T1-weighted structural MRI data from 5,928 participants were retrospectively analyzed, including individuals diagnosed with bvFTD, AD, and cognitively normal (CN) controls. To address the minority status of bvFTD, augmentation techniques were applied. Twelve 3D DenseNet models with different hyperparameters were trained. FastSurfer (v2.0.4) estimated regional volumes and cortical thickness. These features, together with age and sex, were used as tabular input to a multilayer perceptron (MLP) for classification. Two hybrid multimodal late-fusion frameworks were developed, as shown in Figure 1. Result Using the MLP hybrid fusion framework with augmented data, the average F1-score improved by approximately 3% for AD, 260% for bvFTD, and 4% for CN. The macro-average AUC and F1-score increased by about 5% and 21%. The overall mean AUC and F1-score reached 85.6±1.1% and 70.5±2.1%. The meta-classifier hybrid fusion framework achieved an average accuracy of 87±11.4%, AUC of 87.5±6.3%, and F1-score of 74.7±17.2%. Conclusion By combining complementary multi-scale information, the proposed hybrid fusion frameworks address limitations of single-representation models while incorporating covariates such as age and sex. The ensemble-based multimodal strategy improves classification performance, particularly enhancing bvFTD sensitivity despite its minority representation. Improved differentiation between bvFTD and AD may support reliable MRI-based diagnostic workflows and reduce misclassification in overlapping cases.
ID: 386
/ Poster No. # 10: 063
Amortising Inference and Meta-Learning Priors in Neural Networks 1: Helmholtz AI, Germany; 2: MCML, Germany; 3: TUM, Germany; 4: UTN, Germany One of the core facets of Bayesianism is in the updating of prior beliefs in light of new evidence—so how can we maintain a Bayesian approach if we have no prior beliefs in the first place? This is one of the central challenges in the field of Bayesian deep learning, where it is not clear how to represent beliefs about a prediction task by prior distributions over model parameters. Bridging the fields of Bayesian deep learning and probabilistic meta-learning, we introduce a way to learn a weights prior from a collection of datasets by introducing a way to perform per-dataset amortised variational inference. The model we develop can be viewed as a neural process whose latent variable is the set of weights of a BNN and whose decoder is the neural network parameterised by a sample of the latent variable itself. This unique model allows us to study the behaviour of Bayesian neural networks under well-specified priors, use Bayesian neural networks as flexible generative models, and perform desirable but previously elusive feats in neural processes such as within-task minibatching or meta-learning under extreme data-starvation. ID: 337
/ Poster No. # 10: 064
Modalities: Image, Multimodal Data, Other Methods: Other Application Domain: Health SCHEMA: profiling Spatial Cancer HEterogeneity across modalities to benchmark Metastasis risk prediction 1: Institute of Computational Biology, Helmholtz Munich, Germany; 2: Institute of Lung Health and Immunity, Helmholtz Munich, Germany; 3: University Hospital Tübingen, Germany; 4: Karlsruhe Institute of Technology, Germany; 5: TUM School of Life Sciences Weihenstephan,Technical University of Munich, Germany; 6: TUM School of Computing, Information and Technology, Technical University of Munich, Munich, Germany Cancer remains one of the leading causes of death and the incident rates remain stable despite constant scientific advances. The main determinant of poor outcomes is metastatic disease to distant organs, which can arise very early in cancer progression, yet often remains undetected at diagnosis. Understanding differences between primary and metastatic tumors, as well as mechanisms of tumor progression that lead to metastasis, is essential for the early detection and further therapy development, though we still lack robust predictive biomarkers for such early metastasis prediction. Machine learning niche detection methods can be a powerful tool to identify such biomarkers, however proper training and benchmarking datasets are lacking. We address this by introducing SCHEMA - a large-scale resource designed for machine learning research on metastasis prediction. Here we focus on the multimodal definition of cellular niches - functional units of neighboring cells with distinct molecular profiles as biomarkers of cancer progression. We hypothesize that spatial multimodal data contains the necessary signal for building predictive models of metastasis occurrence. SCHEMA enables systematic evaluation and comparison of machine learning approaches for spatial representation learning, niche detection, and outcome prediction by framing tasks within a unified dataset: (1) predicting the occurrence of distant metastasis within 6–12 months after sample collection from spatial molecular profiles of primary tumors; (2) distinguishing primary from metastatic tumor samples; (3) identifying spatial cellular niches associated with patient outcomes. As the backbone of this benchmarking resource, we curated publicly available spatial omics datasets across lung, breast, and colorectal cancers. The collection spans multiple spatial modalities, including spatial transcriptomics and spatial proteomics, with associated clinical metadata regarding metastatic status (if available). We harmonize and standardize these datasets into a unified resource hosted with the support of Lamin.ai. We expect SCHEMA to become a benchmarking dataset for spatial machine learning in oncology. To our knowledge, it represents the largest curated resource on single-cell resolution linking tumor heterogeneity with metastatic outcomes. This dataset is an important step towards enabling reproducible evaluation of machine learning methods and supporting future advances in spatially informed cancer prognosis. ID: 282
/ Poster No. # 10: 065
Modalities: Graphs, Image, Simulation Data Methods: Physics-informed Machine Learning Application Domain: Information Machine Learning-Accelerated Prediction of Hydrogen Embrittlement fromTEM Microstructures 1: Forschungszentrum Juüich GmbH; 2: Helmholtz Zentrum Hereon, Germany Hydrogen embrittlement (HE) poses a critical threat to high-strength metallic components across the automotive, energy, aerospace, and infrastructure sectors, often leading to sudden and catastrophic failures with significant safety and economic consequences. Predicting HE-induced crack initiation and propagation using conventional physics-based approaches such as phase-field (PF) models is computationally demanding due to the need for very fine spatial resolution and small time steps. ID: 133
/ Poster No. # 10: 066
Modalities: Time Series Methods: Foundation Models, Graph Neural Networks, Other Application Domain: Earth & Environment Multimodal Data-Driven Forecasting of Coastal Environmental Health in the Baltic Sea 1: Christian-Albrechts-Universität zu Kiel (CAU), Germany; 2: GEOMAR Helmholtz Centre for Ocean Research Kiel, Germany; 3: Technische Universität Hamburg, Germany Forecasting environmental time series poses a significant challenge due to complex nonlinear dependencies, heterogeneous sampling frequencies, and the integration of heterogeneous data sources. Traditional physics-based models often struggle to capture these complexities at high spatiotemporal resolution, motivating the need for robust data-driven forecasting models for these systems. We address this challenge in the context of Baltic Sea coastal health monitoring, where accurate short-term predictions of key parameters including dissolved oxygen, salinity, temperature, and nutrients are critical for detecting environmental change and enabling early warning systems. We present a multivariate time series forecasting framework based on the state-of-the-art transformer architecture SageFormer (Series-aware Graph-enhanced Transformer). SageFormer's ability to jointly model intra-series temporal dependencies and inter-series relational correlations makes it particularly suitable for capturing the complex and coupled dynamics among heterogeneous water quality indicators. We design a multimodal forecasting pipeline that fuses meteorological variables (atmospheric temperature and weather drivers) with high-frequency sensor observations from monitoring stations, enabling the model to learn interactions between heterogeneous data streams that drive changes in environmental health of Baltic Sea. We evaluate our approach against established forecasting baselines including persistence (i.e., predicting the next value as a copy of the last observed value) and exponentially weighted moving averages (EWMA). For dissolved oxygen (DO), our pipeline outperforms persistence across multiple stations, reducing mean squared error (MSE) by up to 40% and mean absolute error (MAE) by up to 29%. Furthermore, DO prediction errors remain below the instrumental accuracy threshold of ±10 µmol/L across all stations, demonstrating both statistical and practical significance. Our results show that combining high-frequency marine observations with meteorological data and modern transformer architectures provides support for DO forecasting. The approach strengthens AI-driven environmental monitoring and provides a foundation for real-time early warning systems. ID: 170
/ Poster No. # 10: 067
Modalities: Audio, Image Methods: Uncertainty Quantification, Other Application Domain: Earth & Environment Improvement of Whale Vocalization Detection with Temporal Context and Soft Labels 1: Data Science Group, Alfred Wegener Institute, Helmholtz Centre for Polar and Marine Research, Germany; 2: Ocean Acoustics Group, Alfred Wegener Institute, Helmholtz Centre for Polar and Marine Research Automated detection and classification of whale vocalizations in passive acoustic monitoring (PAM) data is an important method in marine ecology. A common strategy is to convert audio into spectrograms (time-frequency-amplitude representation) and to apply convolutional neural network–based object detectors such as YOLO (You Only Look Once). Although effective, this approach faces several difficulties: detections often appear in isolation without sufficient temporal context, the data are highly imbalanced with many empty or acoustically complex segments such as ice noise, shipping, or seismic activity. Further, manual annotations can be inaccurate or inconsistent, especially for lowsignaltonoiseratio calls. To overcome these limitations, we implement three key enhancements. First, we add a postprocessing stage that rescores detections by incorporating the temporal context of neighboring predictions. This suppresses isolated false alarms and increases confidence in detections that form part of a coherent call sequence. Second, we employ hard negative mining to concentrate training on challenging nonwhale acoustic events. Third, we adopt softlabel training to explicitly represent uncertainty in the annotations. For evaluation, we use the BioDCASE 2025 benchmark dataset together with its baseline YOLO model. We introduce each technique incrementally and quantify the corresponding performance gains. To better characterize the model’s sensitivity to noise, we additionally incorporate signaltonoise ratio (SNR) information, which can be derived from the dataset. This includes filtering training vocalizations by their SNR values and assessing whether supplying SNRbased soft labels helps the model interpret uncertain ground truth. We plan to extend these analyses once reliable SNR thresholds become available. ID: 240
/ Poster No. # 10: 068
Modalities: Simulation Data Methods: Graph Neural Networks, Physics-informed Machine Learning Application Domain: Earth & Environment, Health Learning Coarse-Grained Potentials for Intrinsically Disordered Proteins Forschungszentrum Jülich, Germany Proteins regulate a wide range of biological functions. Machine-learning approaches such as AlphaFold [1] have revolutionized structural biology by enabling accurate prediction of globular protein structures from sequence alone. However, globular proteins represent only ~30% of the human proteome. Intrinsically disordered proteins (IDPs) form a distinct class that lacks a stable three-dimensional structure and instead explore a highly dynamic conformational landscape. Many IDPs are implicated in neurodegenerative diseases such as Alzheimer’s, where aggregation into amyloid fibrils contributes to disease progression [2]. Their pronounced structural heterogeneity poses major challenges for both experimental and computational characterization. ID: 135
/ Poster No. # 10: 069
Modalities: Text Methods: Other Application Domain: Information A modular Retrieval Augmented Generation (RAG) Toolbox for Building Transparent AI-Search Engines FZJ, Germany RAG is widely used to enable LLM Solutions to ingest new or private data. However, many existing RAG Solutions limit transparency of crucial processes, like data ingestion and retrieval strategies. This limitation undermines understanding, and in consequence impairs trust in the LLM-generated results. Especially in the context of known LLM issues like hallucinations and inaccuracies in information-retrieval. To counter this trend, we introduce the RAG Search Engine Toolbox ( RAG-SET ). Focused on transparency, the RAG-SET allows end-users to view raw retrieval data. Furthermore, Admins have control over all internal steps from data parsing to LLM-Answer generation. Together these transparency and accessibility features create an open playground, enabling developers to quickly assemble understandable RAG solutions that produce trusted results. This poster presents the design principles of RAG-SET, illustrating how the transparency focus leads to expanded analytic capabilities, simplified debugging and accessible systematic experimentation. ID: 387
/ Poster No. # 10: 070
Better Uncertainties Don’t Guarantee Better Decisions 1: HAI, Germany; 2: TUM; 3: relAI; 4: MCML We aim to raise awareness of the disconnect between uncertainty quantification metrics and real-world decision-making proficiency in machine learning. For three classes of decision-making problems, we demonstrate through a comprehensive empirical study that the commonly used uncertainty quantification metrics frequently misalign with the quality of downstream decisions. Our experimental setup and codebase serve as a starting point for a richer evaluation protocol that better reflects the utility of uncertainty estimates in real-world decision-making. ID: 205
/ Poster No. # 10: 071
Modalities: Graphs Methods: Other Application Domain: Energy Attributing the effects of multiple line outages via Shapley values 1: Institute of Climate and Energy Systems (ICE-1), Forschungszentrum Jülich, Germany; 2: Amprion GmbH, Transmission System Operation, Germany The impact of multiple line outages on power flows is essential for power systems operation and planning. Commonly used Line Outage Distribution Factors (LODFs) quantify the effect of single line outages on power flows, but they do not capture important interaction effects in case of multiple line outages. These interactions can even lead to sign reversals of power flows. Shapley values from game theory have been proposed to attribute the effect of multiple line outages to individual lines while accounting for interaction effects. However, they are rarely applied due to their high computational complexity. We propose an approximation framework for Shapley values that identifies negligible interactions leading to faster computation times. Furthermore, we introduce Shapley-Taylor interaction indices (STIs), which combine standard LODFs and Shapley-based attribution, enabling intuitive interpretation. STIs can identify adverse interaction effects between simultaneous outages, which is particularly useful for outage planning. Overall, our approach improves both the computational feasibility and the interpretability of Shapley-based analysis for multiple line outages. ID: 234
/ Poster No. # 10: 072
Modalities: Image Methods: Probabilistic Methods, Uncertainty Quantification Application Domain: Aeronautics, Space & Transport, Earth & Environment Can Uncertainty Quantification Benefit From Label Embeddings? A Case Study on Local Climate Zone Classification 1: German Aerospace Center, Germany; 2: Interhyp; 3: LMU Munich; 4: TUM Modern deep learning models have achieved superior ID: 294
/ Poster No. # 10: 073
Modalities: Multimodal Data, Tabular Data Methods: Probabilistic Methods Application Domain: Health Benchmarking Taxonomic Classifiers: Bridging 16S rRNA metabarcoding and Shotgun Metagenomics for AI-Driven Microbiome Decoding 1: Department of Tumor Immunology and Tumor Immunotherapy, Helmholtz Center for Translational Oncology (HI-TRON), Mainz, Germany; 2: German Cancer Research Center (DKFZ), Heidelberg, Germany; 3: Department of Hematology, Medical Oncology and Pneumology, University Medical Center, Mainz, Germany; 4: Division of Microbiome and Cancer, German Cancer Research Center (DKFZ), Heidelberg, Germany; 5: Epithelium Microbiome lnteractions, German Cancer Research Center (DKFZ), Heidelberg, Germany; 6: Department of Systems Immunology, Weizmann Institute of Science, Rehovot, Israel Accurate and reproducible taxonomic classification of microbial communities is fundamental to microbiome research, yet methodological differences between sequencing platforms introduce divergent error profiles, genomic targets, and database biases. These inconsistencies pose a significant challenge for building robust AI-driven predictive models from multi-platform microbiome datasets. Despite growing adoption of both full-length amplicon and shotgun metagenomic sequencing, comprehensive benchmarking frameworks that systematically reconcile taxonomic classifiers across both platforms remain lacking. This study addresses this gap by establishing a rigorous computational benchmarking framework to evaluate and standardize taxonomic outputs across both sequencing strategies. Clinical samples stratified by distinct clinical phenotypes were profiled using two parallel sequencing workflows. Full-length 16S rRNA sequencing (~1500 bp) was performed using Oxford Nanopore chemistry, with taxonomic classification conducted using Emu and Kraken2, optimized for long-read amplicon data. Concurrently, Illumina shotgun metagenomics was performed with taxonomic profiling via MetaPhlAn4 and Kraken2, and functional annotation using HUMAnN3. Unified reference databases (GTDB/SILVA) were applied across both pipelines to quantify cross-platform concordance and isolate platform-specific biases. Beta-diversity analysis (PCoA/Bray-Curtis) demonstrated that biological signals strongly outweighed platform-specific variance, with samples clustering primarily by clinical phenotype. A cross-platform comparison revealed significant concordance for dominant core taxa (ρ = 0.567, p < 0.001). Taxonomic accumulation curves confirmed sufficient sequencing depth across both platforms; however, distinct performance divergence emerged at lower taxonomic ranks. Shotgun metagenomics demonstrated a steeper discovery rate for low-abundance rare taxa, whereas full-length 16S rRNA sequencing provided superior species-level resolution for phylogenetically similar clades that were frequently collapsed in short-read profiles. By defining these translational boundaries, this framework establishes essential quality-control parameters for taxonomic sensitivity, specificity, and cross-platform concordance. This benchmarking framework provides a reproducible and scalable foundation for integrating multi-platform microbiome data into machine learning pipelines, advancing the development of AI-driven precision diagnostics. ID: 289
/ Poster No. # 10: 074
Modalities: Graphs Methods: Generative Models, Probabilistic Methods Application Domain: Matter Toward Condition-Aware MOF Topology Prediction via Latent Generative Modeling 1: Forschungszentrum Jülich GmbH, Germany; 2: Helmholtz Institute for Polymers in Energy Applications Jena (HIPOLE Jena), Jena, Germany Metal-organic frameworks (MOFs) are crystalline porous materials constructed from metal clusters and organic ligands through reticular chemistry. The modularity of these building blocks creates a vast combinatorial design space and a wide variety of network topologies, which strongly influence properties relevant to applications such as energy storage and catalysis. Predicting which topology will form from a given set of building blocks remains challenging, as experimental outcomes depend not only on geometric feasibility but also on coordination preferences and synthesis conditions. In this exploratory work, we formulate topology prediction as the modeling of the conditional distribution p(T|B), where T denotes the MOF topology and B denotes the set of building blocks. Using several MOF structure databases (totaling about 400,000 structures), we build a model that links building blocks to candidate topology assignments. Building blocks are represented as molecular or cluster graphs, encoded with crystal graph convolutional neural networks, and aggregated using permutation-invariant set representations. Topology reconstruction is performed through diffusion in a learned latent space. We then refine this topology model using MOF synthesis datasets that pair structures with experimental synthesis conditions C extracted from the literature, resulting in a p(T|B,C) model. A key challenge is the strong heterogeneity of MOF building blocks: both metal clusters and organic linkers vary widely in size, composition, and connectivity, making it difficult to define a single representation that is both expressive and stable across the full design space. A second challenge is data quality in synthesis metadata mined from the literature, where extraction errors, missing fields, inconsistent units, and ambiguous terminology can introduce substantial noise into condition-aware modeling and complicate reliable topology attribution. The proposed framework supports uncertainty-aware topology prediction and candidate net ranking for a given set of building blocks. This submission presents an exploratory methodological study and establishes a foundation for probabilistic modeling of topology formation in reticular materials, with future extensions toward condition-aware generative design of MOFs. ID: 361
/ Poster No. # 10: 075
Modalities: Image Methods: Foundation Models Application Domain: Health Blur-aware Single-cell Segmentation and Feature Extraction for Drug-activity Measurement in Brightfield Organoid Imaging 1: Helmholtz Zentrum München GmbH, Germany; 2: Department for Translational Medical Oncology, NCT/UCC Dresden and DKFZ; 3: Translational Medical Oncology, Carl Gustav Carus Faculty of Medicine, Dresden Brightfield microscopy provides a scalable, label-free modality for studying cellular responses in organoid models, but reliable single-cell analysis remains challenging due to low contrast and variable focus. Pretrained segmentation models such as Cellpose can segment cells across imaging modalities, yet their performance degrades in blurred regions commonly found in brightfield microscopy. To address this limitation, we build upon a blur-aware segmentation pipeline that applies a Laplacian-variance blur map to filter unreliable instance predictions from pretrained Cellpose models. The proposed approach improves segmentation quality and enables more reliable benchmarking on brightfield organoid datasets. Experiments on approximately 10,000 organoid slices demonstrate that blur-aware instance filtering substantially improves segmentation consistency and instance-level metrics without requiring model retraining. Following segmentation, each cell instance is used for downstream feature extraction and phenotype analysis to measure drug activity. Three complementary feature extraction strategies are evaluated: classical morphology and texture features, radiomics features extracted using the PyRadiomics framework, and deep feature embeddings obtained from domain specific pretrained foundation models. These features are compared in terms of predictive efficacy and computational efficiency and are used to train supervised classifiers that assign each cell to biologically relevant phenotypes, including healthy, apoptotic, necrotic, ferroptotic, and other drug-induced states. Aggregating cell-level predictions across treatments provides a quantitative measure of drug response and enables characterization of drug-induced cellular effects. ID: 128
/ Poster No. # 10: 076
Modalities: Text Methods: Foundation Models, Generative Models Application Domain: Core Machine Learning, Health, Information Decoding Values: Measuring Culture Embedded in Large Language Models 1: Perelyn GmbH, Germany; 2: Technical University Munich, Germany Large Language Models learn biases and values implicitly from their training data. With the models’ increasing performance and reasoning capabilities, research has started to investigate the cultural values of LLMs. A misalignment in values between the user and the model can lead to miscommunication or the imposition of values from one culture onto another. Recent work has found LLMs to align best with the values of English-speaking and Western countries. The primary technique of assessing an LLM’s cultural alignment is prompting the models to answer questionnaires designed to evaluate human culture. However, most methods rely solely on the generated outputs of the models without considering their internal states. We bridge the gap between cultural research in LLMs and model interpretability by developing an algorithm that measures the models’ values in their embedding spaces by observing their associations of cultural concepts to selected antonym pairs. Our adaptation of the POLAR algorithm with additional robustness and normalization measures aims at enabling a cross-model comparison of values - independent of the LLMs’ architecture and the dimensionality of their embedding spaces. Our comparison of candidate models on Hofstede’s cultural dimensions suggests that models from the same country of origin do not necessarily achieve the same position on each dimension. In a case study, we evaluate the LLMs’ socioeconomic values and find the overall trend of the models being work-driven, assessing economic growth and freedom as more important than fair wages and equality, as well as favoring capitalism and socialism over communism. Moreover, our statistical analysis of POLAR dimensions provides insights into the generalization of POLAR embeddings to Decoder-Only LLMs and the limitations we face. Our work contributes to the research of cultural values in LLMs and provides the first white-box algorithm based on POLAR embeddings, measuring cultural values in the embedding spaces of models while enabling cross-model comparability. ID: 274
/ Poster No. # 10: 077
Modalities: Image, Text Methods: Foundation Models, Generative Models Application Domain: Health Cytoarchitecture in Words: Weakly Supervised Vision–Language Modeling for Human Brain Microscopy 1: Institute of Neuroscience and Medicine (INM-1), Research Centre Jülich, Germany; 2: Helmholtz AI, Research Centre Jülich, Germany; 3: Cécile & Oskar Vogt Institute for Brain Research, University Hospital Düsseldorf, Germamny; 4: Computer Vision, Institute for Computational Visualistics, University of Koblenz, Germany Foundation models increasingly offer potential to support interactive, agentic workflows that assist researchers during analysis and interpretation of image data. Such workflows often require coupling vision to language to provide a natural-language interface. However, paired image–text data needed to learn this coupling are scarce and difficult to obtain in many research and clinical settings. One such setting is microscopic analysis of cell-body–stained histological human brain sections, which enables the study of cytoarchitecture: cell density and morphology and their laminar and areal organization. Here, we propose a label-mediated method that generates meaningful captions from images by linking images and text only through a label, without requiring curated paired image–text data. Given the label, we automatically mine area descriptions from related literature and use them as synthetic captions reflecting canonical cytoarchitectonic attributes. An existing cytoarchitectonic vision foundation model (CytoNet) is then coupled to a large language model via an image-to-text training objective, enabling microscopy regions to be described in natural language. Across 57 brain areas, the resulting method produces plausible area-level descriptions and supports open-set use through explicit rejection of unseen areas. It matches the cytoarchitectonic reference label for in-scope patches with 90.6% accuracy and, with the area label masked, its descriptions remain discriminative enough to recover the area in an 8-way test with 68.6% accuracy. These results suggest that weak, label-mediated pairing can suffice to connect existing biomedical vision foundation models to language, providing a practical recipe for integrating natural-language in domains where fine-grained paired annotations are scarce. ID: 366
/ Poster No. # 10: 078
Modalities: Multimodal Data, Text Methods: Agentic AI Application Domain: Core Machine Learning The (AI)^2 Consultant - A Co-consultant for AI Consultants Helmholtz Zentrum Munich, Germany AI consultants routinely face incomplete project specifications, fragmented stakeholder input, and insufficient time to survey the technical literature that underpins sound recommendations. We present The (AI)^2 Consultant, a multi-stage agentic system designed to augment early-phase AI consulting. A conversational agent first elicits structured project context through a guided machine learning canvas, capturing goals, data characteristics, constraints, and stakeholder expectations. A seven-phase analytical workflow then synthesizes this context into actionable guidance: it distills the project framing, formulates targeted research questions, retrieves and indexes relevant scientific literature, performs retrieval-augmented question answering over acquired papers, ranks findings by project relevance, proposes implementation directions subject to scientific, ethical, and feasibility critique, and delivers a final evidence-grounded report. Our results show that this approach produces traceable, literature-backed consulting recommendations directly from unstructured project intake. ID: 312
/ Poster No. # 10: 079
Modalities: Image, Text Methods: Generative Models, Probabilistic Methods Application Domain: Health Accessible pipeline for open set dermatological captioning with scarce data 1: Helmholtz AI, Helmholtz Zentrum München, Germany; 2: Technical University of Munich, Klinikum rechts der Isar, Germany Dermatological documentation is time-consuming, yet accurate descriptions of skin lesions are essential for diagnosis and follow-up. Building automated captioning systems for this setting is challenging because dermatological datasets are often small, privacy constraints are strict, fairness must be considered across skin tones, and deployment should remain feasible on limited hardware. In this work, we present an accessible pipeline for open set dermatological captioning under scarce-data conditions. We focus on a GDPR-compliant training and evaluation setup, knowledge distillation from a large model into a smaller deployable model, and open-set handling that allows marking concepts by specifying their name. To support practical use, the model architecture is chosen to run on a single consumer GPU. We also propose an evaluation strategy that includes skin-tone–based subgroup analysis to assess potential performance degradation, as well as Hungarian matching to compare predicted and reference concepts in a robust way. Together, these contributions provide a practical framework for developing and evaluating image-to-text systems that are resource-efficient and suited for real-world clinical constraints. ID: 297
/ Poster No. # 10: 080
Modalities: Tabular Data, Time Series Methods: Other Application Domain: Earth & Environment Machine Learning for Rogue Wave Prediction in the Southern North Sea 1: Helmholtz-Zentrum Hereon, Germany; 2: Federal Waterways Engineering and Research Institute (BAW); 3: Helmholtz AI Rogue waves are unexpectedly large and highly unpredictable ocean waves that can cause major disruptions to offshore structures and maritime operations. Although short‑term prediction is now a pressing need, the underlying physical mechanisms remain poorly understood, making it difficult to develop reliable forecasting tools. For maritime safety, there is growing international interest in establishing robust early‑warning systems capable of identifying potential rogue wave events in advance. Machine learning (ML) offers a promising direction, yet systematic comparisons across models and the use of explainable AI to uncover the conditions that lead to rogue waves remain limited. In this work, we benchmark five ML algorithms using a metocean dataset collected at the K14 radar station in the southern North Sea. Our key objective is to develop better predictive models for the abnormality index, defined as the relative height of the largest wave, in the next 10‑minute window using 17 meteorological and oceanic features describing the preceding 30‑minute period. The models evaluated are Elastic-Net regression (ENR), support vector regression (SVR), random forests (RF), XGBoost (XGB), and a feed forward neural network (FFNN). Our results show that XGB and RF achieve the strongest predictive performance (each with rs = 0.98), though RF achieves a slightly lower recall (79%) compared with XGB (85%), despite both offering low inference cost. SVR performs well in detecting rogue wave events (rs = 0.97, recall 89%), but at the expense of more false alarms, while both SVR and FFNN require substantially more computational resources. The FFNN reaches only moderate skill (rs = 0.91, recall 77%), and ENR shows the weakest performance overall. Feature‑importance analysis consistently identifies dimensionless water depth as the most influential predictor across all models, with atmospheric pressure and wind gustiness also contributing meaningfully. The varied influence of the spectral parameters suggests that rogue waves may arise through several overlapping physical processes rather than a single universal mechanism. As next steps, we aim to develop a real‑time short‑term rogue‑wave forecasting system for the southern North Sea. To further advance physical understanding, we also plan to explore additional explainable AI approaches, including forest‑guided clustering and causal machine learning. ID: 279
/ Poster No. # 10: 081
Modalities: Time Series Methods: Other Application Domain: Health Nanopore- and AI-Empowered Microbial Viability Inference Helmholtz Munich, Germany The ability to differentiate between viable and dead microorganisms in metagenomic data is crucial for various microbial inferences, ranging from assessing ecosystem functions of environmental microbiomes to inferring the virulence of potential pathogens from metagenomic analysis. Established viability-resolved genomic approaches are labor-intensive as well as biased and lacking in sensitivity. We here introduce a new fully computational framework that leverages nanopore sequencing technology to assess microbial viability directly from freely available nanopore signal data. Our approach utilizes deep neural networks to learn features from such raw nanopore signal data that can distinguish DNA from viable and dead microorganisms in a controlled experimental setting of UV-induced Escherichia cell death. The application of explainable artificial intelligence (AI) tools then allows us to pinpoint the signal patterns in the nanopore raw data that allow the model to make viability predictions at high accuracy. Using the model predictions as well as explainable AI, we show that our framework can be leveraged in a real-world application to estimate the viability of obligate intracellular Chlamydia, where traditional culture-based methods suffer from inherently high false-negative rates. This application shows that our viability model captures predictive patterns in the nanopore signal that can be utilized to predict viability across taxonomic boundaries. We finally show the limits of our model’s generalizability through antibiotic exposure of a simple mock microbial community, where a new model specific to the killing method had to be trained to obtain accurate viability predictions. While the potential of our computational framework’s generalizability and applicability to metagenomic studies needs to be assessed in more detail, we here demonstrate for the first time the analysis of freely available nanopore signal data to infer the viability of microorganisms, with many potential applications in environmental, veterinary, and clinical settings.ID: 301
/ Poster No. # 10: 082
Modalities: Image, Multimodal Data, Time Series Methods: Other Application Domain: Aeronautics, Space & Transport, Earth & Environment Counterfactual Ship Attribution in Cloud Fields Using Deep Learning and Satellite Observations German Aerospace Center (DLR), Germany Anthropogenic emissions from maritime transport influence local cloud microphysics, affecting albedo and atmospheric radiative forcing. However, quantifying the contribution of individual vessels to these modifications remains a complex problem due to high natural variability and confounding meteorological factors. This work presents a novel framework for attributing localized cloud changes to specific ships using deep learning and explainable artificial intelligence (xAI). We approach this as a counterfactual change-detection problem. A convolutional neural network is trained to predict per-pixel cloud features, specifically brightness and thickness, utilizing Earth observation data, climate reanalysis data, and ship trajectory data. To isolate ship influence, the model generates two outputs for each scene: one using observed input features including ship positions, and another where all ship-dependent inputs are artificially set to zero to simulate a no-shipping baseline. The difference between these outputs reveals the perturbed regions attributable to shipping activity. Sentinel-5P satellite data serves as ground truth for cloud characteristics, while ERA5 reanalysis data and Automatic Identification System (AIS) provides climate variables and high-resolution vessel location and identity data to construct feature vectors for the model inputs. Following change detection, we extract regions of significant deviation. To link these changes to individual sources, we apply xAI methods, specifically Gradient-weighted Class Activation Mapping (Grad-CAM). This technique highlights the input features most responsible for the model's prediction, allowing us to spatially associate the cloud alterations with specific ship locations in the AIS dataset. This research is currently ongoing. By decomposing total cloud change into individual contributions, we aim to provide a more granular assessment of vessel-specific environmental impacts than current emission inventories allow. This approach has potential applications in verifying compliance with maritime regulations and improving climate feedback parameters in global circulation models. ID: 208
/ Poster No. # 10: 083
Modalities: Time Series Methods: Other Application Domain: Energy, Earth & Environment RenewBench: Real Energy Data You Can Actually Use 1: Karlsruhe Institute of Technology (KIT), Germany; 2: Helmholtz-Center Hereon, Germany; 3: Helmholtz AI, Germany Transitioning to renewable energy is essential for mitigating climate change, but the variable and decentralised nature of such generation systems presents major challenges when maintaining grid stability for reliable operation. AI-driven solutions have the potential to address these challenges, particularly in the form of more powerful and robust forecasting. However, progress at scale is hampered by the lack of standardised, high-quality renewable energy datasets. Existing models are therefore often limited in geographic scope, restricted to a specific generation or data type, and evaluated on proprietary datasets that prevent broad comparison. Additionally, these models disregard the spatio-temporal couplings influencing long-term grid stability, as they consider only local weather inputs or ignore weather dependencies altogether. RenewBench addresses these limitations by fusing renewable energy generation and weather data to create a global, open-source energy benchmark. We consolidate openly available generation datasets with high temporal and spatial resolution into a standardised Zarr-based structure. These are combined with meteorological reanalysis data and enriched with comprehensive SpatioTemporal Asset Catalog (STAC) metadata to create an AI-ready findable, accessible, interoperable, and reusable (FAIR) dataset. By leveraging a STAC FastAPI and PgSTAC backend, researchers can perform nearly instantaneous metadata queries and automated retrieval via a dedicated Python package to facilitate many benchmarking tasks. In this poster we present the current status of RenewBench, including incorporated geographic regions, database setup, and initial benchmarking ideas. By building this open and unified global benchmark, we contribute to democratising data access for the next generation of data-driven energy-meteorology solutions. ID: 315
/ Poster No. # 10: 084
Modalities: Text Methods: Agentic AI, Generative Models Application Domain: Health, Information An AI System for Robust Plain Language Translation of Clinical Trials 1: Institute of Computational Biology, Computational Health Center, Helmholtz Munich, Germany; 2: German Center for Diabetes Research, Germany Clinical trial registry entries often contain complex medical terminology, creating barriers for patients who search online for relevant studies. Limited accessibility of trial information contributes to challenges in patient recruitment and retention, with up to 80% of trials failing to meet enrollment timelines. Large Language Models (LLMs) are increasingly used to generate plain language summaries (PLS) to improve accessibility. In the high-stakes domain of health information, where registries provide patients access to novel therapeutics and contribute to scientific advancement, summaries must preserve accuracy and clarity. Although current approaches aim to improve readability, they rarely incorporate structured validation to detect distortions or omissions of crucial details. We investigate whether multi-agent critique systems can reduce distortions and improve alignment between AI-generated PLS and their source clinical trial registry entries. Specifically, we assess whether iterative revision preserves clinically relevant details. We compare two pipelines: (1) a baseline single-pass LLM summarization system and (2) an agentic architecture with a generator agent followed by critique agents targeting eligibility preservation and explicit uncertainty representation. Critique agents provide structured feedback that triggers iterative revision until alignment criteria are met. We construct a dataset of ClinicalTrials.gov registry entries and implement a multi-dimensional evaluation framework. (1) Content preservation is measured with embedding-based similarity scores, such as BERTScore. (2) Consistency checks ensure key clinical details remain aligned with the source, combining rule-based extraction (e.g., regex for numeric values or age ranges) with a ground-truth–driven LLM-as-a-judge. (3) Readability is assessed using standard metrics, such as the Flesch Reading Ease Score. This study quantifies whether agentic validation layers reduce clinically relevant distortions while maintaining accessibility and proposes a reproducible benchmarking methodology for evaluating generative models in high-stakes health communication, where fidelity to source material is essential. ID: 309
/ Poster No. # 10: 085
Modalities: Image, Text Methods: Foundation Models Application Domain: Health K-MaT: Knowledge-Anchored Manifold Transport For Cross-Modal Prompt Learning In Medical Imaging University of Bonn, University Hospital Bonn, Clinic for Diagnostic and Interventional Radiology, 53127 Bonn, Germany Background: Deep learning models for medical imaging often degrade during cross-modal transfer because distinct acquisition physics encourage modality-specific shortcuts. Large-scale biomedical vision-language models adapted on high-end imaging (e.g., CT) frequently fail to generalize to frontline low-end modalities (e.g., radiography), suffering from catastrophic knowledge forgetting. Purpose: We propose K-MaT (Knowledge-Anchored Manifold Transport), a prompt-learning framework designed to reliably transfer diagnostic semantics from high-end visual data to low-end modalities in a strict zero-shot regime. The goal is to eliminate the need for low-end visual training data while preventing the decision boundary from collapsing into modality-specific statistics. Materials and Methods: K-MaT builds on the BiomedCLIP backbone and introduces a factorized prompt parameterization using Class-Specific Context (CSC) and Modality-Specific Context (MSC). The framework utilizes LLM-generated clinical descriptions as semantic anchors to prevent deviation from meaningful clinical knowledge. It further employs Fused Gromov-Wasserstein (FGW) optimal transport to align the low-end prompt manifold with the visually-grounded high-end space. Results: Evaluated on four cross-modal benchmarks (including Dermoscopy to Clinical and CT to X-ray), K-MaT achieved state-of-the-art results with an average harmonic mean accuracy of 44.1% and macro-F1 of 36.2%. On challenging breast imaging tasks, it improved low-end accuracy to 38.4%, whereas standard methods like CoOp dropped to 27.0%. Conclusion: Aligning prompt manifolds via optimal transport provides a highly effective route for the zero-shot cross-modal deployment of medical VLMs. K-MaT successfully mitigates catastrophic forgetting and preserves robust diagnostic performance across diverse medical imaging modalities without requiring target-domain visual training. ID: 154
/ Poster No. # 10: 086
Modalities: Multimodal Data, Simulation Data, Time Series Methods: Other Application Domain: Health Analysis of immune responses in renal cell carcinoma with synergistic ex vivo and in silico simulation tumor models 1: German Cancer Research Center (DKFZ), Helmholtz Institute for Translational Oncology (HI-TRON), Tumor Immunology and Tumor Immunotherapy, Mainz, Germany; 2: University Medical Center Mainz, Department of Hematology and Medical Oncology, Mainz, Germany; 3: Faculty of Biosciences, Ruprecht Karls University, Heidelberg, Germany; 4: National Center for Tumor Diseases, Department of Medical Oncology, Heidelberg, Germany; 5: University Medical Center Mainz, Institute for Pathology, Mainz, Germany; 6: University Medical Center Mainz, University Cancer Center (UCT), Mainz, Germany Background Immune checkpoint-inhibitors (ICI) and anti-angiogenic tyrosine kinase inhibitors (TKI) have improved treatment outcomes for clear cell renal cell carcinoma (ccRCC). However, heterogeneous treatment responses underline the need for further therapeutic options. A better understanding of functional mechanisms governing immune responses in ccRCC is therefore needed. We have established a synergistic tumor model platform that combines ex vivo human tumor explant models with patient-specific in silico modeling, enabling detailed longitudinal analyses of tumor-immune interactions. Methods Our fully human tissue-explant model system preserves the cellular and soluble components of the tumor microenvironment, allowing short-term culture and direct assessment of treatment effects. Explants of ccRCC primary tumor tissues were treated with ICI and TKI as well as immune receptor antagonists. Spatial and functional parameters from ex vivo experiments were incorporated in an agent-based simulation model (PhysiCell, Ghaffarizadeh et al., PLoS Comput. Biol., 2018) for unlimited exploration of functional cellular dynamics in the tumor. Results Immunohistochemical analyses and cytokine profiling of ccRCC-tissue explants reflected heterogeneous immune responses following treatment among different patients. Spatial analysis revealed clusters of T cells and CD163+ macrophages localized closely to tumor blood vessels as well as high CCR5 expression in the tumor, particularly on the tumor blood vessels. Of note, treatment of the tissue explants with an anti-CCR5 inhibitor led to an increase of CD8+ T cells and cytotoxic cytokines in comparison to ICI monotherapy. In silico simulation demonstrated an increase of T cells and cytotoxic cytokines upon blockade of the interaction between T cells with endothelial cells or macrophages. Conclusion Our combined ex vivo and in silico analyses provide evidence for immunosuppression in the ccRCC tumor microenvironment mediated by macrophages and endothelial cells proposing an immunosuppressive perivascular area as potential target to improve immune responses in ccRCC. ID: 120
/ Poster No. # 10: 087
Modalities: Time Series Methods: Probabilistic Methods, Uncertainty Quantification Application Domain: Energy Bad Data Detection and Signal Recovery in Large-Scale ICT-Platforms 1: ICE-1, Institute of Climate and Energy Systems, Forschungszentrum Jülich, 52428 Jülich, Germany; 2: RWTH Aachen University, Aachen 52056, Germany; 3: JARA-Energy, Jülich 52425, Germany High-quality measurement data is essential for reliably monitoring and controlling modern power systems. However, large-scale sensing infrastructures often experience corrupted or missing data due to device failures, communication issues or cyber-related events. This work presents a digital twin for managing data quality in power system ICT platforms, which is designed to detect bad data automatically and recover signals. The digital twin is based probabilistic machine learning. In particular, correlated Gaussian processes are applied to local subsystems, enabling multi-channel signal reconstruction without the need for full system models or large training datasets. Unlike conventional single-channel interpolation and data-driven approaches, this method uses locally correlated measurements to consistently and accurately recover long missing data intervals. The digital twin provides uncertainty-aware reconstructions and incorporates self-aware failure indications, enabling operators to assess the reliability of the recovered data directly. By delivering validated, AI-ready measurement signals, the approach also supports advanced, machine learning–based monitoring, anomaly detection and decision support applications in future power system operations. This approach has been validated using a real-world power system operated by Forschungszentrum Jülich. ID: 119
/ Poster No. # 10: 088
Modalities: Time Series Methods: Physics-informed Machine Learning, Probabilistic Methods Application Domain: Energy System Characterization by Affine Operators in Physics-Informed Gaussian Processes 1: ICE-1, Institute of Climate and Energy Systems, Forschungszentrum Jülich, 52428 Jülich, Germany; 2: RWTH Aachen University, Aachen 52056, Germany; 3: JARA-Energy, Jülich 52425, Germany Non-invasive inference of system parameters from observational data provides a framework for understanding complex systems under real operating conditions. In this work, we formulate the problem of system characterization in terms of the inverse solutions of differential equations via Physics-Informed Gaussian Processes using an augmented framework that incorporates affine structures, yielding operators beyond the linear regime. Besides explicitly accounting for the application-specific noise characteristics, the proposed framework enables the estimation of arbitrary parameters of interest depending on the operating conditions. We show that the method yields system estimates whose accuracy is comparable to that obtained through experimental characterization. In particular, it is shown that parameters can be estimated up to experimentally observed error bounds across different operational regimes. From a modeling perspective, we introduce a dedicated noise-handling strategy that improves estimation performance relative to alternative inference approaches like Physics-Informed Neural Networks when applied to real systems operated in the field. We further analyze the numerical stability of the proposed method and present a self-consistent normalization scheme that ensures robust and stable model training and inference. ID: 223
/ Poster No. # 10: 089
Modalities: Simulation Data Methods: Physics-informed Machine Learning, Other Application Domain: Information, Matter NEOPIC: A Neural Operator Framework for Particle-based Kinetic Plasma Simulations 1: Forschungszentrum Jülich GmbH, Germany; 2: Helmholtz Zentrum Dresden-Rossendorf, Germany Kinetic plasma simulations play a critical role in applications of societal relevance such as nuclear fusion and building the next-generation of compact particle accelerators. They are also widely used in studying astrophysical phenomena and industrial plasma processes. Particle methods, specifically particle-in-cell (PIC), have been the method of choice for such simulations for many decades. PIC simulations demand high computational costs imposed by the stability requirements of the solvers making them prohibitively expensive for parameter studies, uncertainty quantification and long simulation times. In this work we combine particle schemes with the recently introduced neural operators towards fast, accurate kinetic plasma simulations. The key idea of our approach is to obtain the electric and magnetic fields at each time step based on the input particle quantities from a neural operator instead of via conventional mesh-based solvers or tree-based approaches. The resulting particle-in-neural operator (PINOP) scheme is mesh-free, naturally adaptive, applicable to any geometry and not constrained by spatial and time stepping stability requirements. We choose a version of Fourier neural operator which can handle nonuniform particle inputs and train with simulation data from different particle-based schemes on different test cases in plasma simulations. Our results show that the PINOP scheme generalizes well beyond the training regime and maintains acceptable conservation and accuracy in quantities of interest. We also obtain speedups compared to reference GPU implementation of PIC schemes. For the first time, we will demonstrate the applicability of neural operators in speeding up a complex, nonlinear, highdimensional, multi-scale, coupled PDE system. ID: 384
/ Poster No. # 10: 090
Learning PDE Dynamics from Small Data: Out-of-Distribution Generalization of Deep Learning Surrogates 1: Institute for Advanced Simulations -- Materials Data Science and Informatics (IAS-9), Forschungszentrum Juelich GmbH, Juelich 52425, Germany; 2: Chair of Materials Data Science and Materials Informatics, Faculty 5 -- Georesources and Materials Engineering, RWTH Aachen University, Aachen 52056, Germany Partial differential equations (PDEs) are essential for modeling complex physical, engineering, and materials systems. However, high-fidelity simulations are still computationally expensive. Recent advances in scientific machine learning aim to create surrogate models that can directly approximate PDE dynamics from data. However, many existing approaches rely on large training datasets and increasingly complex architectures. This raises a key question: How effective can simple, data-efficient architectures be for learning PDE dynamics? ID: 163
/ Poster No. # 10: 091
Modalities: Graphs, Image, Multimodal Data, Simulation Data, Tabular Data, Time Series, Other Methods: Foundation Models, Physics-informed Machine Learning, Probabilistic Methods, Other Application Domain: Energy, Information, Matter Human-Explainable, Compact, Clustering-based Latents for Fast Proton Energy Spectra Estimation 1: Helmholtz-Zentrum Dresden-Rossendorf e.V., Germany; 2: Center for Advance Systems Understanding (CASUS), Görlitz, Saxony, Germany; 3: TUD Dresden University of Technology, 01062 Dresden, Germany; 4: Technische Universität Chemnitz, Institute of Physics, Chemnitz, Saxony, Germany A bottleneck in gaining a deeper understanding of the complex laser plasma interaction that generates laser-accelerated protons is the lack of robust and near real-time information extraction from high frequency shot data, due to human intervention required in the process. Here, we present an approach to employ deep learning methods to reduce the need for human input into the analysis of Thomson Parabola Spectrometer (TPS) measurements of proton energy spectra given relatively limited labelled data. Our approach builds on deep feature extraction using general pre-trained autoencoders, and self organizing map-based clustering of global image features and the spectra that are available as labels, to effectively reduce the dimension of input and output modalities. Lower-dimensional representations then enable a small model to be trained on limited data to estimate proton spectra and to help with spectrometer re-calibrations.
ID: 263
/ Poster No. # 10: 092
Modalities: Audio, Image, Multimodal Data, Tabular Data, Text Methods: Generative Models, Other Application Domain: Core Machine Learning, Information Collaborative Research Data Management for AI Hof University of Applied Sciences, Germany SARA (System for AI Research and Assessment) is a centralized system that allows managing multimodal datasets, collaborative annotation and evaluation, as well as cross-group sharing around different research projects. AI-research and the involved dataset management typically end up with many different incompatible approaches. Fragmentation is driven both by the specialization for individual modalities (with their own tooling) as well as the lack of centralized systems that incentivize consolidation through cross-group sharing and reuse. SARA supports all stages of a typical evaluation process (dataset creation → annotation → model selection → evaluation). It handles diverse modalities and allows fine-grained operations including creation, viewing, editing, and extension of datasets in common formats (.csv, .jsonl, .xlsx, .parquet) via an intuitive Web-UI. Collaboration features include multi-user editing, versioning, and changelogs. The system also features reusable components, such as AI models, which can be added to a central library by the user before linking them to individual projects. This model library additionally includes structured documentation and fosters peer exchange about the models by allowing users to share code snippets, comments and experiences. If a model is linked to a finished evaluation, the results are also shown on a leaderboard amongst other models tested on the same dataset. Datasets and media files can be stored on any mountable file system allowing integration with already existing storage infrastructure. SARA was piloted in a university AI course, where it was used to manage and grade student submissions spanning text, image, and audio modalities across multiple encoding formats. It is also used in ongoing research projects with up to ~5 GB of data per dataset. The data-management system addresses a practical gap that affects most academic AI projects: the challenges of managing datasets across modalities, teams and user types (console-focused like developers vs. UI-driven like data annotators and project managers). By combining data exploration, annotation, model evaluation and documentation (e.g. licenses) in a single collaborative system, SARA lowers the barrier to reproducible workflows while aiding research. Future plans include embedding AI-services for enhancing data like speaker diarization for audio, semantic segmentation for images and named entity recognition for text (AI for AI).
ID: 280
/ Poster No. # 10: 093
Modalities: Multimodal Data Methods: Other Application Domain: Health From FAIR Multimodal Data Infrastructure to AI-Driven Clinical Interpretation: The MINDset Platform for Precision Psychiatry 1: Institute of Neuroscience and Medicine 4 (INM-4), Forschungszentrum Jülich GmbH, Jülich, Germany; 2: Department of Psychiatry, Psychotherapy and Psychosomatics, University Hospital Aachen,RWTH Aachen University, Aachen, Germany; 3: Department of Computer and Information Science and Engineering, University of Florida, USA; 4: Faculty of Medical Engineering and Technomathematics, FH Aachen University Applied Sciences, Aachen, Germany; 5: Department of Neurology, RWTH Aachen University, Aachen, Germany; 6: Department of Psychiatry II, Ulm University and BKH Günzburg, Germany. Advances in neuroimaging and digital phenotyping are producing increasingly rich multimodal datasets, yet translating these heterogeneous data into clinically interpretable insights remains a major challenge. Existing neuroimaging platforms primarily focus on data storage and sharing, often lacking mechanisms for semantic integration and direct use of AI methods for patient-level interpretation. Fragmented storage, inconsistent metadata, and limited interoperability further restrict the effective reuse of multimodal data in clinical research. We present MINDset, a FAIR-aligned research data management infrastructure designed to transform multimodal psychiatric datasets into an AI-ready analytical environment. The platform integrates heterogeneous multimodal data types—including neuroimaging data (MRI, PET, EEG, fNIRS), tabular data (cognitive and clinical variables), and unstructured textual reports— within a unified ecosystem based on structured metadata and interoperable data models. Its cloud-native, containerized architecture enables scalable multimodal data processing, interactive dataset exploration, and reproducible AI workflows across large neuroimaging cohorts (Figure.1). By combining standardized metadata with flexible querying and visualization, MINDset enables systematic multimodal analysis beyond conventional neuroimaging repositories. Building on this infrastructure, we have develop an AI-driven pipeline for automated patient-level interpretation using LLMs. A retrieval-augmented generation (RAG) framework integrates multimodal imaging-derived features, clinical metadata, and curated scientific literature to produce structured, clinician-readable reports. In a proof-of-concept study involving patients with major depressive disorder scanned before and after treatment, resting-state fMRI metrics were combined with clinical information and literature evidence to generate individualized summaries linking patient-specific neural changes to prior findings (Figure. 2). By integrating semantic metadata and generative AI within the research data management layer, MINDset transforms multimodal datasets into AI-ready resources for interpretable analysis. The framework enables scalable, reproducible AI workflows for psychiatric research and provides a foundation for future developments, including multimodal foundation models, automated hypothesis generation, and clinical decision-support tools for precision mental health.
ID: 238
/ Poster No. # 10: 095
Modalities: Graphs, Time Series Methods: Generative Models, Graph Neural Networks Application Domain: Aeronautics, Space & Transport Measuring Diversity in Multi-Agent Motion Generation 1: Karlsruhe Institute of Technology, Germany; 2: Aumovio SE Recent advances in motion generation models have substantially improved the ability to simulate the behavior of multiple interacting agents in complex autonomous driving scenarios. These motion generation models are commonly evaluated using plausibility measures, that asses how closely the generated motion resembles real-world observations. However, the evaluation of motion diversity is typically neglected, despite its importance in ensuring mode coverage and sampling efficiency. We address this gap by introducing a similarity-based diversity metric, that complements state-of-the-art evaluation protocols. The proposed metric evaluates the variability of generated motions at the level of individual agents or multi-agent scenarios, providing a direct assessment of how many distinct behaviors the model captures in a finite number of samples. We further demonstrate, how this metric enables a clearer differentiation between modeling strategies, such as diffusion or autoregressive generation, and sampling approaches, such as top-k and top-p samplers. ID: 140
/ Poster No. # 10: 096
Modalities: Graphs, Image, Simulation Data, Time Series, Video Methods: Generative Models, Graph Neural Networks, Physics-informed Machine Learning Application Domain: Earth & Environment GEOMAGFOR - Geomagnetic Core Field Forecasting: Excursions and Polarity Reversals 1: Institute of Fluid Dynamics, Helmholtz-Zentrum Dresden-Rossendorf (HZDR), Germany; 2: Geomagnetism, GFZ Helmholtz Centre for Geosciences, Potsdam, Germany The GEOMAGFOR project aims to explore the potential of AI-based methods to improve geomagnetic core field forecasting on both short and long timescales. For the short term (a decade), a particular focus lies on enhancing the accuracy of the International Geomagnetic Reference Field (IGRF) and its predictive capabilities (GFZ Potsdam team). In parallel, the project applies physics-informed artificial intelligence techniques to develop more realistic models of geomagnetic reversals and excursions (HZDR team). At GFZ, we implemented the Machine Learning tools to test the predictability of geomagnetic core field spherical harmonic coefficients for short-term forecasts. Based on the frameworks of scikit-learn, PyTorch, and TensorFlow, we implemented classical time series forecasts as well as multi-input and multi-window methods. Utilising linear regression, we implemented Neural Networks, Convolutional Neural Networks, Long-Short-Term-Memory, Recurrent Neural Networks and Gated Recurrent Units. Accompanying efforts used the existing ‘sktime’-framework, dedicated to time series forecasting, on a paleomagnetic model pfm9k2 covering a longer period than the IGRF, covering about the last 9000 years. This mean model is given with an ensemble of alternate solutions, which increases the available amount of time series data. Our results indicate that even short-term geomagnetic field forecasts might not be feasible without physical information. At HZDR, a 2D α²-dynamo model implemented in the Dedalus framework was developed and parameterised by the radial and meridional distribution of the α-effect. A region of parameter space capable of reproducing geomagnetic excursions and polarity reversals was identified. Building on this model, physics-informed neural networks (PINNs) are applied to solve the inverse problem of reconstructing the dynamical evolution of excursions and reversals from paleomagnetic data. Numerical modelling is performed using our own PINN-Sandmännchen code and NVIDIA PhysicsNeMo PINNs, an open-source Python framework for building, training, and scaling physics-informed AI models. Training is performed using spherical harmonics models of the Laschamps excursion and the Brunhes–Matuyama reversal, provided by the GFZ team. We present results on the forecasting of geomagnetic excursions and polarity reversals, demonstrating the feasibility of AI-assisted mid-term geomagnetic prediction when physical information about the geodynamo process is included. ID: 226
/ Poster No. # 10: 097
Modalities: Simulation Data Methods: Probabilistic Methods, Uncertainty Quantification Application Domain: Health Clinical Evidence to Individualized Care: AI Powered Clinical Decision Support Based on Predicted Individual Treatment Effect (PITE) 1: Faculty for Informatics and Data Science, Regensburg University, Germany; 2: MRC Biostatistics Unit, University of Cambridge, UK Advances in computational systems have enabled machines to process large amounts of data, identify complex patterns, and support decision-making across diverse domains. These developments are contributing to a fundamental shift in healthcare systems, enabling clinicians to move toward more data-driven and individualized approaches to medical decision-making. Medical decisions—ranging from diagnosis and prognosis to treatment selection—are inherently subject to uncertainty. In practice, clinicians often rely on complex, experience-driven, and heuristic decision processes that integrate heterogeneous sources of information. In this context, Predicted Individual Treatment Effects (PITE) provide a principled statistical framework to quantify how much a specific patient is expected to benefit from one treatment compared with another. This work explores how PITE can be used to support clinical decision-making across different diseases and clinical contexts. In particular, we address practical questions such as: which AI methods are most appropriate for specific clinical datasets; how PITE should be evaluated and validated; and how different outcome structures and levels of complexity affect the development of reliable decision support tools. We illustrate how PITE-based models can be implemented and evaluated in diverse disease settings, each presenting its own methodological and practical challenges. Our results show that even under real-world conditions—such as missing data or complex outcomes—predictive models can maintain interpretability and generate individualized treatment effect estimates that support clinical decision-making. At the same time, our findings highlight that internal validation alone is insufficient, and that external validation is essential to ensure robust and reliable predictions. Overall, our work suggests that effective clinical decision support based on PITE requires more than predictive modeling alone. It requires an adaptive and continuously learning system integrating data management, modeling strategies, regulatory-grade explainability, and ongoing validation. Such systems can help translate clinical evidence into individualized treatment decisions while continuously improving as new data become available. ID: 275
/ Poster No. # 10: 098
Modalities: Other Methods: Generative Models Application Domain: Earth & Environment, Information Integrating an LLM Chatbot into an Existing Web Frontend via DASF and MCP Helmholtz-Zentrum Hereon, Germany Large-language-model (LLM) chat interfaces are ubiquitous but often live beside domain applications rather than within them. We present a chatbot embedded into an existing web framework by combining the Data Analytics Software Framework (DASF)—a secure, message‑broker–based RPC system—with the Model Context Protocol (MCP) for lightweight tool exposure. DASF exposes Python classes and functions as remotely callable procedures without opening internet‑facing ports, enabling the chatbot backend to run near high‑value infrastructure (e.g., HPC), minimizing data movement and aligning with institutional security policies. Building on DASF’s generated Python client stubs, we add an MCP server that orchestrates requests and is used by an OpenAI‑compatible API (the Blablador service at Forschungszentrum Jülich) for conversational processing. On the frontend, we reuse generic, composable components from prior work (doi:10.5194/egusphere-egu25-3120) and exploit DASF’s asynchronous execution to stream results into the interface. Unlike detached chat UIs, our approach embeds conversations within operational frontend so responses appear as domain‑native views — maps, plots, tables, dashboards — rather than plain text only. The chatbot acts as a copilot: it triggers analyses, parameterizes workflows, and visualizes outcomes in context while keeping compute close to data. This tight coupling yields a secure, scalable, and maintainable architecture that augments user workflows, lowers the barrier to advanced analytics, and improves accessibility without re‑platforming or duplicating interfaces. In effect, conversation becomes an interaction modality for the host application. The system orchestrates computations near HPC resources without inbound ports and returns model outputs rendered natively by the application. Thanks to DASF’s generic design, the framework generalizes across scientific knowledge‑transfer scenarios and stakeholder engagements with minimal need for supplemental web development. ID: 381
/ Poster No. # 10: 099
CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training 1: Thomson Reuters Foundational Research; 2: Tübingen AI Center, University of Tübingen; 3: Helmholtz Munich; 4: MCML, Technical University of Munich; 5: Imperial College London Large language model (LLM) post-training enhances latent skills, unlocks value alignment, improves performance, and enables domain adaptation. Unfortunately, post-training is known to induce forgetting, especially in the ubiquitous use-case of leveraging third-party pre-trained models, which is typically understood as a loss of parametric or factual knowledge. We argue that this accuracy-centric view is insufficient for modern foundation models and instead define forgetting as systematic model drift that degrades behavior and user experience. In this context, we introduce CapTrack, a capability-centric framework for analyzing forgetting in LLMs that combines a behavioral taxonomy with an evaluation suite built on established benchmarks and targeted adaptations. Using CapTrack, we conduct a large-scale empirical study across post-training algorithms, domains, and model families, including models up to 80B parameters. We find that forgetting extends beyond parametric knowledge, with pronounced drift in robustness and default behaviors. Instruction fine-tuning induces the strongest relative drift, while preference optimization is more conservative and can partially recover lost capabilities. Differences across model families persist, and no universal mitigation emerges. ID: 214
/ Poster No. # 10: 100
Modalities: Time Series Methods: Other Application Domain: Matter Patch-MLP-Based Predictive Control: Simulation of Upstream Pointing Stabilization for PHELIX Laser System 1: Helmholtz-Zentrum Dresden-Rossendorf, Dresden, Germany; 2: GSI Helmholtzzentrum für Schwerionenforschung, Darmstadt, Germany; 3: Amplitude laser group—Dresden operations, Dresden, Germany; 4: Extreme Light Infrastructure—Nuclear Physics, National Institute for Physics and Nuclear Engineering, Ilfov, Romania; 5: University of Bucharest, Ilfov, Romania; 6: Engineering and Applications of Lasers and Accelerators Doctoral School (SDIALA), National University of Science and Technology Politehnica of Bucharest, Bucharest RO-060042, Romania; 7: Technische Universität Dresden, Dresden, Germany; 8: Chemnitz University of Technology, Chemnitz, Germany High-energy laser facilities such as PHELIX at GSI require excellent beam-pointing stability to ensure reproducibility and reliable operation. Conventional PID control mitigates slow drift but is fundamentally limited by diagnostic latency and mirror inertia. We introduce a predictive control scheme in which beam-pointing errors are forecast using a patch-based multilayer perceptron, and the predicted errors are converted into correction signals via a PID controller. This feed back strategy compensates for system delay and is trained directly on diagnostic time-series data. Simulations with an upstream correction mirror at the PHELIX pre-amplifier bridge show reduced residual jitter compared with conventional PID control. Across a 10-hour dataset, the predictive controller remained drift-free and improved pointing metrics by approximately 10%-20%. This research was published in Machine Learning Science and Technology (DOI: 10.1088/2632-2153/ae393d). ID: 1398
/ Poster No. # 10: 101
Modalities: Graphs Methods: Foundation Models, Graph Neural Networks Application Domain: Health Evaluating 3D Molecular Representation Learning with Multipolar Atom Types 1: University of Warsaw; 2: Helmholtz Munich Pretrained molecular representations are central to modern drug discovery pipelines, yet their intrinsic quality is rarely evaluated independently of downstream finetuning. Existing benchmarks assess finetuned models on molecular property prediction and operate exclusively at the level of whole molecules. No benchmark probes how well learned representations capture fine-grained, atom-level structural and electronic features, nor how this capacity depends on input modality. We address this gap by introducing the Multipolar Atom Types Dataset (MATD) and the Molecular Atom Types Evaluation (MATE) framework, a combination of supervised linear and nearest-neighbor probing with unsupervised clustering, grounded in quantum-mechanically derived multipolar atom types from the Multipolar Atom Types from Theory and Statistical clustering (MATTS) data bank. MATD annotates 74,283 QM9 molecules with 366 physically grounded atom-type labels at atomic resolution. Using MATE we evaluate a suite of graph-based and 3D-based pretrained molecular models under controlled variations in conformer source and hydrogen inclusion. We find that (i) explicit 3D structural input consistently improves atom-level representation quality; (ii) models pretrained via 3D structure reconstruction retain a geometric advantage even when only 2D-derived coordinates are available at inference, a result not previously documented; and (iii) including explicit hydrogens at inference time yields consistent gains across all models and evaluation regimes, challenging the common practice of omitting them. Finally, systematic confusion between atom types in learned representations identifies several pairs with statistically indistinguishable multipole parameter distributions, enabling principled improvement of the MATTS data bank. HELMHOLTZ AI PROJECT CALL AWARDEES 2025
ID: 1403 / Poster No. # 10: 102 Modalities: Image, Multimodal Data Methods: Foundation Models, Other Application Domain: Core Machine Learning, Health, Information, Matter LIVR: Learnable Implicit Volumetric Representations for High-resolution 3D Images 1: Division of Medical Image Computing, German Cancer Research Center (DKFZ), Heidelberg, Germany; 2: Institute of Materials Physics, Helmholtz-Zentrum Hereon, Geesthacht, Germany; 3: Institute of Metallic Biomaterials, Helmholtz-Zentrum Hereon, Geesthacht, Germany Modern 3D imaging in materials science and biomedicine produces volumetric data at micrometer and even nanometer scale, often reaching billions of voxels per sample. This overwhelms current deep-learning-based data analysis pipelines. Patching, slicing, and strong downsampling remain common but break spatial continuity, remove volumetric context, or blur subtle geometric cues that define cracks, vessels, and tissue microstructure. These artifacts undermine the reliability of downstream tasks and limit the scientific value of high-resolution data. HELMHOLTZ AI PROJECT CALL AWARDEES 2025
ID: 1404 / Poster No. # 10: 103 Modalities: Multimodal Data, Simulation Data, Tabular Data, Time Series Methods: Physics-informed Machine Learning, Probabilistic Methods, Uncertainty Quantification Application Domain: Earth & Environment The AEON-UP Helmholtz AI Project - Adaptive Environmental Prediction System using Neural Processes for Urban Air Quality Prediction 1: Helmholtz-Zentrum Hereon, Germany; 2: RIFS Forschungsinstitut für Nachhaltigkeit, Germany Air pollution remains a major environmental health challenge in urban areas, where strong spatial and temporal variability complicates exposure assessment and policy development. While regulated pollutants such as nitrogen dioxide (NO2) and particulate matter (PM2.5) are routinely monitored, pollutants of emerging concern, and especially ultrafine particles (UFPs), remain sparsely measured despite growing evidence of health impacts. While existing numerical chemistry transport models (CTMs) can provide spatio-temporal "complete" estimates, they are computationally demanding at urban scales. Also, current machine-learning approaches often require extensive retraining for new cities and typically lack robust uncertainty quantification. The Helmholtz AI project (2026) AEON-UP (Adaptive Environmental Prediction System using Neural Processes for Urban Air Quality Prediction) addresses these challenges by developing a transferable, uncertainty-aware AI framework for high-resolution urban air quality prediction. AEON-UP combines physics-based CTM simulations with heterogeneous and sparse observational data within a unified probabilistic framework based on neural processes (NPs). The approach integrates gridded model outputs with irregularly distributed measurements to generate hourly two-dimensional concentration fields at approximately 100 m resolution. The project builds on complementary expertise across the Helmholtz Association. Helmholtz-Zentrum Hereon contributes operational urban air-quality modelling and machine-learning downscaling experience, while RIFS provides unique observational datasets, including UFP measurements from the Net4Cities network. Helmholtz Munich and Helmholtz AI consultants will give guidance for methodological development in Bayesian deep learning and scalable NP architectures. AEON-UP aims to deliver the first transferable AI system capable of predicting both regulated pollutants and UFPs at urban scale while providing uncertainty estimates to improve interpretability and support evidence-based decision-making. Beyond reducing computational requirements by orders of magnitude compared with high-resolution CTMs, the project advances trustworthy AI for Earth-system applications through probabilistic multimodal learning. All datasets, model components, and documentation will be openly released to support reproducibility and future applications in environmental research, health studies, and urban policy.
HELMHOLTZ AI PROJECT CALL AWARDEES 2025
ID: 1411 / Poster No. # 10: 104 Modalities: Image Methods: Foundation Models Application Domain: Health SCALE.FM: Multi-Scale Foundation Models for Precision Neuroimaging 1: Deutsches Zentrum für Neurodegenerative Erkrankungen (DZNE), Bonn, Germany; 2: Helmholtz Zentrum München - Deutsches Forschungszentrum für Gesundheit und Umwelt (HMGU), Munich, Germany The SCALE.FM project develops novel AI methods for resolution- and modality-adaptive neuroimaging by innovating Voxel Size Independent Neural Networks (VINNs) and extending them into foundation models (VINN4FM). Self-supervised learning and foundation models (FMs) have shown promise in reducing annotation needs and enhancing domain generalization, but are not natively resolution-adaptive and lack dedicated mechanisms to preserve fine anatomical detail across modalities and voxel sizes. To address these issues, we will augment foundation models with learned interpolation and transformer-based image fusion for anatomically precise, robust segmentation across resolutions and modalities. Our core application is high-resolution neuroimage segmentation of hypothalamic nuclei and choroid plexus, enabling new insights into insulin resistance and associated metabolic and cognitive dysfunction. The primary innovations in this project will be 1) to integrate a resolution- and modality-adaptive transformer architecture with our previously developed VINN architecture that significantly improves generalization by internal resolution normalization, modality fusion, and feature-space augmentation, and 2) to supplement this advanced architecture with foundation model training paradigms, resulting in VINN for Foundation Models (VINN4FM). VINN4FM will be trained, benchmarked, and fine-tuned for application in metabolism research, enabling detailed segmentation and morphometric analysis of the hypothalamus and choroid plexus, two key brain regions involved in insulin signaling and homeostatic regulation. The resulting neuroimaging measures will be integrated with comprehensive metabolic phenotyping in a large case-control study to characterize how metabolic dysfunction impacts brain structure and function in individuals with low metabolic health. Altogether, the novel VINN4FM architecture will, for the first time, enable detailed analysis of critical neuroanatomical structures in the human brain implicated in cognitive and metabolic diseases. We will make available the software via the popular and award-winning open-source framework FastSurfer. VINN4FM thus closes the loop between AI method development and biomedical translation, providing a technically robust and openly available foundation for high-resolution, generalizable brain segmentation that directly supports future advances in neuroscience, metabolism, and medical research. HELMHOLTZ AI PROJECT CALL AWARDEES 2025
ID: 1416 / Poster No. # 10: 105 Modalities: Simulation Data, Time Series Methods: Foundation Models, Reinforcement Learning, Probabilistic Methods, Uncertainty Quantification Application Domain: Matter Machine Learning Green Functions for Magnetic Materials and Spintronics (MLGREEN) 1: Helmholtz-Zentrum Dresden-Rossendorf, Germany; 2: Forschungszentrum Jülich, Germany
MLGREEN develops a machine-learning approach to accelerate Korringa–Kohn–Rostoker Green-function calculations for magnetic materials and spintronic systems. By learning the energy-resolved Green function from local atomic environments, the project aims to replace computationally expensive global calculations with a scalable, locality-based model. This could reduce the cost of KKR simulations from cubic to linear scaling, enabling quantum-accurate modeling of realistic nanoscale systems with defects, interfaces, disorder, and complex magnetic textures. Target applications include magnetic skyrmions, spintronic device geometries, transport properties, and rare-earth-free magnetic alloys.
HELMHOLTZ AI PROJECT CALL AWARDEES 2025
ID: 1418 / Poster No. # 10: 106 Modalities: Graphs, Multimodal Data, Simulation Data, Text Methods: Generative Models, Physics-informed Machine Learning Application Domain: Health, Matter Generative AI Framework for Steered Protein Structure Prediction via NMR Chemical Shift Inputs 1: Helmholtz Munich, Germany; 2: Technical University of Munich, Germany State-of-the-art protein 3D structure prediction tools, such as AlphaFold2, predominantly model low-energy, ground-state protein conformations due to biases in their training data. However, higher-energy conformational states are frequently the biologically active and pharmacologically relevant forms of proteins, making them critical targets in drug discovery. While there are many AI-based approaches to sample hypothetical alternative protein conformations, there is currently no computational method to identify biologically relevant states or guarantee their inclusion in the predictions. One uniquely sensitive method to detect these hidden states is NMR spectroscopy, which provides conformation-dependent observables such as chemical shifts. However, directly integrating NMR chemical shifts into AI-based protein structure inference remains underexplored. We propose a generative AI framework that proceeds in three stages. First, a transformer-based chemical shift prediction model is trained that takes amino acid sequence and protein structural representations as input. Second, using this model, a large-scale dataset of ~170,000 protein structures paired with simulated dynamics (MD) gets annotated with predicted chemical shifts. Third, a generative diffusion-based model is trained and conditioned on measured and predicted (previous step) chemical shifts to obtain protein conformations that are consistent with the input chemical shifts. As a proof of concept, we trained a chemical shift prediction model on a small publicly available dataset of 250 chemical shift measurements paired with PDB structures. Our model achieves state-of-the-art performance, and we show that the predicted shifts preserve relevant structural information. Next, we will expand the training dataset to the entire Biological Magnetic Resonance Bank (BMRB) and include Protein Language Model embeddings in our model. Because NMR chemical shifts reflect dynamic conformational ensemble averages, we collaborate with partners from Jülich Forschungszentrum (FZJ) to generate additional training conformations with MD simulations. Ultimately, when provided with protein chemical shifts, our model is expected to reveal plausible alternative high-energy conformational states of proteins that are presently inaccessible to the mainstream protein structure prediction models. The predictions will be validated with NMR experiments performed in a collaborative effort across Helmholtz centers (BNMRZ and FZJ).
ID: 1149
/ Poster No. # 10: 107
Modalities: Image Methods: Other Application Domain: Health A decentralized Swarm Learning framework for 90-Day outcome prediction for acute ischaemic stroke 1: DZNE, Germany; 2: CISPA, Germany Acute ischaemic stroke remains the predominant cause of disability and a major contributor to mortality globally. The damage caused by ischemia, triggered by vascular occlusion, progresses rapidly, with brain tissues beginning necrosis within minutes. This necessitates urgent clinical decision-making. This is particularly important for reperfusion therapies such as intravenous thrombolysis and mechanical thrombectomy. The efficacy of these treatments is time-sensitive and associated with risk of intracranial hemorrhage. In addition, treatment efficacy varies considerably across patients and depends upon several factors. This underscores the need for individualized risk stratification. In this project, we aim to develop a deep learning framework for predicting 90-day functional outcomes using multimodal data acquired at hospital admission. We are integrating heterogeneous data sources across 25 different hospitals within Germany (German Stroke Registry data), including clinical scores, patient history, and MRI images. Traditional centralized learning approaches face limitations related to data privacy, small sample sizes, and institutional barriers. Swarm Learning provides a privacy-preserving, decentralized alternative. However, a decentralized pipeline for multimodal MRI-based model development is currently lacking. To ensure robustness and generalizability, we will implement model training in a decentralized swarm learning manner. Our objective is to validate a multimodal deep learning model for individualized 90-day outcome prediction in a decentralized AI infrastructure. We aim to advance precision medicine for acute stroke care and establish a foundation for globally collaborative, privacy-preserving AI development. | |||||||||||||||||||||||||||||||||
