Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
Please note that all times are shown in the time zone of the conference. The current conference time is: 4th Aug 2026, 10:33:42am CEST
|
Daily Overview |
| Session | |||||||||||||||||||||||||||||||
Poster Session I
| |||||||||||||||||||||||||||||||
| Session Abstract | |||||||||||||||||||||||||||||||
|
Discover cutting-edge AI research and innovations at the Poster Session. Connect with authors, ask questions, and engage in lively discussions that spark collaboration and new ideas. | |||||||||||||||||||||||||||||||
| Presentations | |||||||||||||||||||||||||||||||
ID: 358
/ Poster No. # 9: 001
Modalities: Image Methods: Foundation Models, Other Application Domain: Health Interpretable Representations for Hematology 1: Helmholtz Munich, Germany; 2: Medical Department for Hematology and Oncology, Technical University Munich Sparse autoencoders (SAEs) emerged as a promising tool for mechanistic interpretability of transformer-based foundation models. Very recently, SAEs were also adopted for the visual domain, enabling the discovery of visual concepts and their patch-wise attribution to tokens in the transformer model. While a growing number of foundation models emerged for medical imaging, tools for explaining their inferences are still lacking. In this work, we show the applicability of SAEs for hematology. We propose CytoSAE, a sparse autoencoder which is trained on over 40,000 peripheral blood single-cell images. CytoSAE generalizes to diverse and out-of-domain datasets, including bone marrow cytology, where it identifies morphologically relevant concepts which we validated with medical experts. Furthermore, we demonstrate scenarios in which CytoSAE can generate patient-specific and disease-specific concepts, enabling the detection of pathognomonic cells and localized cellular abnormalities at the patch level. We quantified the effect of concepts on a patient-level AML subtype classification and eight hematologic disease classification tasks and show that CytoSAE concepts reach performance comparable to the state-of-the-art, while offering explainability on the sub-cellular level. Moreover, we built ConceptHub, a shared concept library to systematically store and annotate concepts extracted from CytoSAE. ID: 255
/ Poster No. # 9: 002
Modalities: Graphs, Multimodal Data Methods: Graph Neural Networks Application Domain: Health Analysis of multi-omics data using graph neural networks identifies novel schizophrenia-associated genes suitable as drug targets 1: Helmholtz Munich, Germany; 2: Boehringer Ingelheim Pharma, Germany; 3: Technical University of Munich, Germany Identifying novel therapeutic targets for psychiatric disorders such as schizophrenia requires a comprehensive understanding of altered molecular neurobiology in disease states. Traditional approaches typically analyse single data modalities in isolation, e.g., genetic or transcriptomic data. By contrast, graph neural networks can integrate diverse multimodal molecular data with prior knowledge, such as protein-protein interactions and known disease-associated genes. Here, we employed a graph attention network (GAT) model that combines individual-level RNA-seq, DNA methylation, and genotyping data from the PsychEncode consortium with protein-protein and functional interaction networks to identify and prioritise novel schizophrenia-associated genes. First, we assessed how different interaction networks and omics modalities contribute to model performance. Using literature-derived schizophrenia-associated genes as ground-truth labels, we benchmarked six tissue-agnostic and four brain-specific networks in the GAT model against a multi-layer perceptron (MLP) baseline that uses only molecular data without network information. While models trained on almost all of the networks outperformed the MLP baseline, the tissue-agnostic STRING network showed superior performance compared to other networks. We also demonstrated that the majority of network models performed best when data from all three omics modalities were used. Using this optimised GAT model, we identified 103 novel schizophrenia-associated genes not previously reported in the literature. We characterised these genes through gene set enrichment analyses and examined data modality and patient-level contributions for the most promising candidate genes. For example, we found an interesting cluster of microtubule-associated genes, suggesting a role for impaired axonal transport in schizophrenia pathophysiology. Finally, based on their predicted tractability and expression patterns, we evaluated shortlisted genes regarding their potential as drug targets for treating schizophrenia. Here, several potassium channel subunits emerged as attractive targets, with KCNC4 (Kv3.4) as the most interesting candidate. This work demonstrates that integrating multimodal molecular data with interaction networks through graph neural networks can advance our understanding of psychiatric disorder neurobiology and systematically prioritise biologically plausible and therapeutically tractable targets for psychiatric disorders. ID: 1218
/ Poster No. # 9: 003
Modalities: Audio, Image, Tabular Data, Text, Time Series, Video Methods: Other Application Domain: Core Machine Learning MLflow pilot service for Helmholtz researchers Karlsruhe Institute of Technology (KIT), Germany As the field of Artificial Intelligence / Machine Learning (AI/ML) advances, managing and monitoring intelligent models during their lifecycle, also known as machine learning operations (MLOps), has become essential [1]. MLflow is an open-source platform [2] to assist AI/ML practitioners and teams in handling the AI/ML lifecycle, ensuring that each stage is manageable, traceable, and reproducible. MLflow v3 has many comprehensive features, including Experiment Tracking for parameter logging, metrics visualisation, and artifact storage; Model Registry for model versioning and lineage tracking; Datasets module for dataset management. It also features GenAI capabilities like LLM Observability (Tracing), Evaluation and Monitoring of GenAI applications, Prompt Management, and AI Gateway. Our team provides MLflow instances within several EU projects, including AI4EOSC and iMagine, and upcoming FLUID-AI and EOSC-ARENA, covering use cases from various scientific domains. Recently, we also brought a new pilot instance for Helmholtz researchers [3]. The Helmholtz instance is coupled with Helmholtz ID/AAI and LSDF large storage and allows experiment and model sharing among MLflow users for collaborative development. A major upgrade that also includes group-sharing capabilities is foreseen before June 2026. In this contribution, we are going to present the advantages of the MLflow platform for AI/ML researchers, share our experience in providing MLflow for various use-cases and supporting them with their workflows, and demonstrate the main features of the available Helmholtz pilot instance. [1] Berberi, L., Kozlov, V., Nguyen, G. et al. Machine learning operations landscape: platforms and tools. Artif Intell Rev 58, 167 (2025). https://doi.org/10.1007/s10462-025-11164-3 [2] https://mlflow.org/ and https://mlflow.org/docs [3] https://mlflow.scc.kit.edu and https://mlops.data.kit.edu ID: 1273
/ Poster No. # 9: 004
Modalities: Image Methods: Foundation Models Application Domain: Health Towards Robust Foundation Models for Digital Pathology 1: Berlin Institute for the Foundations of Learning and Data (BIFOLD), Berlin, Germany; 2: Machine Learning Group, Technische Universität Berlin, Berlin, Germany; 3: Aignostics, Berlin, Germany; 4: The Netherlands Cancer Institute Amsterdam (NKI), Antoni van Leeuwenhoek Hospital (AvL), Amsterdam, Netherlands; 5: Institute of Pathology, Ludwig Maximilian University, Munich, Germany; 6: German Cancer Research Center, Heidelberg, and German Cancer Consortium, Munich, Germany; 7: Institute of Pathology, Charité Universitätsmedizin, Berlin, Germany; 8: Department of Artificial Intelligence, Korea University, Seoul, Korea; 9: Max-Planck Institute for Informatics, Saarbrücken, Germany Biomedical foundation models (FMs) pre-trained on large-scale histopathology datasets are rapidly advancing AI-enabled diagnostics and tissue analysis. However, their self-supervised training objectives capture any variation in data, including non-biological confounders such as differences in staining protocols or scanners across medical centers. While recent benchmarks focused on performance, a systematic evaluation of FM robustness to such artifacts has been lacking. We introduce PathoROB, a comprehensive robustness benchmark for pathology FMs, comprising 4 multi-center datasets (~99k patches, 28 biological classes, 34 medical centers). We propose 3 complementary metrics: (i) a robustness index quantifying the local dominance of biological over confounding features in FM embedding space; (ii) average performance drop, measuring downstream model vulnerability to shortcut learning under increasingly spurious training data; and (iii) a clustering score assessing global embedding space organization. We evaluated 20 FMs spanning diverse architectures, pre-training objectives, and dataset scales. We also studied post-hoc robustification strategies — stain normalization, ComBat batch correction, and domain-adversarial learning — not requiring FM retraining. All 20 FMs exhibited robustness deficits, with robustness indices ranging from 0.45 to 0.86. Larger pre-training datasets and vision-language objectives yielded higher robustness, with Virchow2 and Atlas achieving the best performance–robustness tradeoff. Under spurious correlations, downstream models suffered accuracy drops up to 47 pp with less robust FMs failing to detect tumor regions entirely. These failures extended to slide-level MIL models and unsupervised clustering and retrieval. The robustness index strongly correlated with downstream robustness (Spearman ρ up to 0.90). Combined stain normalization and ComBat improved the robustness index by up to 70%, and stain normalization with domain-adversarial training reduced performance drops, though no method fully eliminated them. Our findings demonstrate that robustness evaluation is essential before clinical deployment of pathology FMs, as non-robust representations can cause diagnostic failures even in state-of-the-art models. Future FM development should integrate robustness as a key criterion, improving robustness potentially via post-training alignment. PathoROB provides a blueprint for systematic robustness assessment across biomedical domains. ID: 1253
/ Poster No. # 9: 005
Modalities: Image, Simulation Data, Text, Video Methods: Generative Models Application Domain: Information, Matter Sailing Past Syntax: A Human-in-the-Loop Framework for Safe Generative AI in Science Centre de Physique des Particules de Marseille, France Integrating Large Language Models (LLMs) into scientific workflows presents a critical challenge across both academia and industry: AI agents excel at functional syntax but lack specific scientific reasoning and intuition. They frequently hallucinate domain-specific logic or silently discard governing scientific laws and principles to optimize performance. This poses a severe risk in scientific visualizations and software development, as unguided generative AI can produce convincing yet fundamentally invalid tools. To safely harness LLM code generation, we advocate for the Scientist-AI-Loop (SAIL), a human-in-the-loop framework designed to structurally decouple scientific logic from coding syntax. In SAIL, the researcher acts as the conceptual architect enforcing theoretical boundaries and phenomenological constraints, while the AI exclusively handles code implementation and rendering. Originally designed to overcome bottlenecks in building public outreach tools, this domain-agnostic framework provides a generalized blueprint for broader scientist-AI workflows. We validate SAIL via two astrophysical visualization tools: a real-time gravitational lensing application (nicosmo.github.io/lensing_visualization/) and a dynamic cosmic structure formation simulation (nicosmo.github.io/cosmic_web_explorer/). During development, SAIL exposed a series of critical, often invisible AI failures, instances where agents confidently fabricated physics or silently discarded governing laws simply to satisfy code compilation. By establishing a structured progression from rapid prototyping to agentic IDE integration, we demonstrate how SAIL not only safeguards scientific integrity against these probabilistic failures, but compresses development timelines from months to under 80 hours. Ultimately, SAIL establishes the necessary protocols to stop AI from breaking physics, ensuring generative tools can be safely leveraged for professional modeling, theoretical sandboxing, and interactive communication. ID: 1331
/ Poster No. # 9: 006
Modalities: Audio, Graphs, Image, Multimodal Data, Simulation Data, Tabular Data, Text, Time Series, Video Methods: Other Application Domain: Core Machine Learning, Information The Helmholtz Model Zoo: Enabling AI Model Sharing and Inference in the Helmholtz Cloud Deutsches Elektronen-Synchroton DESY, Germany The Helmholtz Model Zoo (HMZ) is a cloud-based platform that provides remote access to deep learning models within the Helmholtz Association. It enables seamless inference execution via both a web interface and a REST API, lowering the barrier for scientists to integrate state-of-the-art AI models into their research. Scientists from all 18 Helmholtz centers can contribute their models to HMZ through a streamlined, well-documented submission process on GitLab. This process minimizes effort for model providers while ensuring flexibility for diverse scientific use cases. Based on the information provided about the model, HMZ automatically generates the web interface and API, tests the model, and deploys it. The REST API further allows for easy integration of HMZ models into other computational pipelines. With the launch of HMZ, researchers can now run AI models directly within the Helmholtz Cloud, ensuring that all data remain within the association and that our data sovereignty is preserved. The platform imposes no strict limits on the number of inferences or the volume of uploaded data, while Helmholtz Virtual Organizations (VOs) enable fine-grained access control for specialized models. External researchers can also access HMZ through Helmholtz VOs upon invitation by a Helmholtz representative, facilitating collaborative research beyond the association's boundaries. Data uploaded for inference is stored within HIFIS dCache InfiniteSpace and remains under the ownership of the uploading user. HMZ is powered by GPU nodes, hosted as part of the DESY Hamburg HPC cluster. Model inference is managed through the NVIDIA Triton Inference Server, ensuring efficient GPU utilization. The development and maintenance of HMZ are led by the Helmholtz Imaging Support Team at DESY, with support from Helmholtz Federated IT Services (HIFIS) and the Helmholtz AI platform. Hardware and implementation have been supported by funds from the Haicore initiative. Our presentation will provide an overview of HMZ architecture and its integration into a professional HPC environment. It will also address the scientific foundations of selected models and emphasise the benefits of operating them entirely within the Helmholtz infrastructure. ID: 1230
/ Poster No. # 9: 007
Modalities: Simulation Data, Text Methods: Foundation Models, Generative Models Application Domain: Core Machine Learning Improving Reliability of LLM-Based Robotic Task Planning Through Domain Adaptation and Benchmarking ARENA2036, Germany Recent advances in Large Language Models (LLMs) have enabled natural language interfaces for robotic systems, allowing robots to interpret human instructions and generate executable task plans. However, general-purpose LLMs often lack grounding in robot capabilities, which can lead to hallucinated actions, incomplete plans, and incorrect task ordering when generating robotic task sequences. In this work, we investigate the reliability of LLMs for robotic task planning under constrained action spaces. We introduce a benchmark based on a fixed robot skill library that represents the available capabilities of a robotic system. Using this environment, we evaluate several planning approaches, including prompting-based baselines, general-purpose LLMs such as Mistral, and robotics-oriented planning models. To address common failure modes observed in baseline systems, we construct an extended instruction-to-plan dataset derived from publicly available and synthetically generated data. The dataset focuses on structured action sequences and realistic robotic task constraints. Using this dataset, we apply parameter-efficient Low-Rank Adaptation (LoRA) fine-tuning to adapt language models for robotic planning tasks. We evaluate the models across several reliability metrics, including plan validity, step completeness, hallucination rate, and action ordering correctness. Experimental results demonstrate that lightweight domain-specific fine-tuning significantly improves planning reliability compared to zero-shot prompting approaches. In addition, we analyze training and inference performance across CPU and high-performance computing environments to assess practical deployment considerations. Overall, this work provides a reproducible framework for benchmarking LLM-based robotic task planners and highlights the importance of dataset design and domain adaptation for improving the reliability of language-driven robotic systems. ID: 1184
/ Poster No. # 9: 009
Modalities: Image, Multimodal Data Methods: Generative Models, Uncertainty Quantification Application Domain: Health Uncertainty-Guided Generation of Dark-Field Radiographs 1: School of Computation, Information and Technology, Technical University of Munich, 85748 Garching, Germany; 2: Munich Center for Machine Learning (MCML), Munich, Germany; 3: Institute of Machine Learning in Biomedical Imaging, Helmholtz Munich, 85764 Neuherberg, Germany; 4: Department of Physics, School of Natural Sciences, Technical University of Munich, 85748 Garching, Germany; 5: Munich Institute of Biomedical Engineering, Technical University of Munich, 85748 Garching, Germany; 6: Institute for Diagnostic and Interventional Radiology, School of Medicine and Health, TUM University Hospital Klinikum rechts der Isar, Technical University of Munich, 81675 Munich, Germany; 7: Institute for Advanced Study, Technical University of Munich, 85748 Garching, Germany; 8: School of Biomedical Engineering & Imaging Sciences, King’s College, London, UK X-ray dark-field radiography provides complementary diagnostic information to conventional attenuation imaging by visualizing microstructural tissue changes through small-angle scattering. Early clinical studies suggest that dark-field chest radiography offers unique diagnostic value for quantifying pulmonary emphysema in COPD patients and improving COVID-19 diagnosis. While dark-field scanners provide paired attenuation and dark-field radiographs, the limited availability of such data poses challenges for developing robust deep learning models. In our recent work [1], we present the first framework for generating dark-field images directly from conventional 2D chest X-ray radiographs using an Uncertainty-Guided Progressive Generative Adversarial Network. This approach enables the large-scale generation of virtual dark-field data from widely available chest X-rays. The proposed framework follows a progressive learning scheme in which aleatoric uncertainty estimates are used as attention maps to guide model refinement across stages. We explicitly incorporate both aleatoric and epistemic uncertainty to improve interpretability and reliability, as they reflect different aspects of model confidence. High aleatoric uncertainty marks regions with inherently ambiguous signals, while elevated epistemic uncertainty highlights areas where the model may not generalize well. Together, these uncertainty estimates provide a more comprehensive view of model reliability and data quality. Our results demonstrate that the proposed progressive model can accurately reconstruct dark-field images from attenuation data, achieving high image fidelity and structural consistency. Quantitatively, all metrics (MSE, PSNR, and SSIM) show consistent improvement across the model stages, confirming that progressive refinement effectively enhances fine structural detail and reduces artifacts. The final stage achieves the best overall performance, indicating that the model learns to recover increasingly realistic and anatomically consistent dark-field representations. Furthermore, out-of-distribution evaluation demonstrates that the proposed model generalizes well and provides a promising foundation for future clinical applications. [1] Lina Felsner, Henriette Bast, Tina Dorosti, Florian Schaff, Franz Pfeiffer, Daniela Pfeiffer, Julia Schnabel, ‘Uncertainty-guided Generation of Dark-field Radiographs’, IEEE International Symposium on Biomedical Imaging (ISBI) 2026
ID: 1167
/ Poster No. # 9: 010
Modalities: Image, Video Methods: Generative Models Application Domain: Health, Information SoraCT: Unconditioned 3D CT Synthesis via Video Diffusion Transformers 1: Friedrich-Alexander University Erlangen-Nürnberg, Germany; 2: Department of Computing, Imperial College London, UK The advent of Diffusion Transformers (DiT), exemplified by Sora, has revolutionized video generation by capturing complex spatiotemporal dependencies. However, their potential in 3D medical imaging remains largely underexplored. In this paper, we present a novel approach to unconditioned CT volume generation by adapting the Open-Sora architecture. Treating axial CT slices as temporal frames, we leverage the spatiotemporal attention mechanisms of video diffusion models to learn the high-dimensional distribution of anatomical structures without relying on explicit conditions (e.g., text prompts or segmentation maps). Our method addresses the challenges of volumetric consistency and data scarcity in medical domains. Extensive experiments demonstrate that our model generates high-fidelity, diverse 3D CT volumes that preserve structural integrity and tissue texture. Quantitative metrics and qualitative assessments by radiologists confirm the realism of the synthetic data. This work highlights the feasibility of repurposing state-of-the-art video generation models for 3D medical synthesis, paving the way for privacy-preserving data augmentation and foundational anatomy modeling. Furthermore, this model serves as a foundational prior that can be adapted for text-conditioned or image-conditioned generation tasks.
ID: 1394
/ Poster No. # 9: 011
Modalities: Image Methods: Foundation Models Application Domain: Health A 3D Foundation Model for Generalizable Biological Structure Segmentation in Tissue Clearing Images (DISCO-CAR: whole-body mapping of engineered cells and diseases at single-cell resolution) 1: Institute for Stroke and Dementia Research, Klinikum der Universität München, Ludwig-Maximilians University Munich, Munich, Germany; 2: Institute for Intelligent Biotechnologies (iBIO), Helmholtz Center Munich, Neuherberg, Germany; 3: Faculty of Medicine, Ludwig-Maximilians University Munich, Munich, Germany; 4: Munich Cluster for Systems Neurology (SyNergy), Munich, Germany; 5: Munich Medical Research School (MMRS), Munich, Germany; 6: Deep Piction GmbH, Munich, Germany; 7: School of Medicine, Koç University, İstanbul, Turkey; 8: TUM School of Computation, Information and Technology, Technical University of Munich, Munich, Germany; 9: German Center for Infection Research (DZIF), Partner Site Munich, Munich, Germany; 10: Institute for Medical Microbiology, Immunology and Hygiene, Technische Universität München (TUM), Munich, Germany; 11: Focus Group 'Clinical Cell Processing and Purification', Institute for Advanced Study, TUM, Munich, Germany Tissue clearing combined with light-sheet microscopy (LSM) enables the 3D visualization and analysis of intricate cellular and subcellular structures across tissues and organisms, characterized by high contrast and super-resolution capabilities. However, segmenting diverse biological structures in 3D LSM images remains a major challenge due to significant variations in morphology, artifacts, signal-to-noise ratio, and surrounding tissue context. Conventional supervised learning-based segmentation models typically require extensive voxel-wise annotations and the training structure-specific models, thereby limiting their generalizability and scalability across datasets and applications. Foundation models (FMs) are expected to mitigate this limitation. FMs trained on large-scale data are anticipated to achieve zero-shot or few-shot generalization. Despite their success in other domains, their application to LSM data remains underexplored. In this study, we propose a foundation model tailored for comprehensive segmentation of diverse biological structures in tissue-cleared mouse LSM dataset and evaluate its domain generalization capability. We construct a large-scale dataset containing more than 9 biological structures and 50,000 patches of size 300³ and perform self-supervised learning (SSL) to learn robust representations from diverse LSM data. The model is evaluated across extensive 3D segmentation tasks, including out-of-distribution datasets. Our findings demonstrate that a self-supervised pretrained foundation model enables effective cross-structure transfer for 3D image segmentation and indicate its strong ability to generalize to previously unseen biological structures. In several downstream tasks, it outperforms state-of-the-art task-specific 3D segmentation models. Overall, this study underscores the potential of FMs in LSM image domain and demonstrates their capability as a unified approach for segmentation across diverse biological domains. ID: 295
/ Poster No. # 9: 012
Modalities: Graphs, Tabular Data, Text Methods: Agentic AI, Foundation Models, Uncertainty Quantification Application Domain: Aeronautics, Space & Transport, Information M²S³OM-graph: A Hybrid Deterministic-LLM Pipeline for Automated Metadata Crosswalk Extraction and Graph-Grounded Interoperability Deutsches Zentrum für Luft- und Raumfahrt e. V. (DLR), Germany Scientific metadata interoperability depends on crosswalk specifications, field mappings between standards, currently trapped in heterogeneous web artifacts: HTML tables, XSL stylesheets, PDFs, and prose that no existing tool systematically harvests. We present M²S³OM-graph, a hybrid AI pipeline that extracts, structures, and applies these rules at scale. The core design principle is deterministic-first: format-specific parsers handle structured artifacts reproducibly without model calls, while ambiguous cases are delegated to Blablador (Helmholtz-AI LLM infrastructure) as a targeted fallback. Notably, the parsers themselves were developed with LLM-assisted coding, illustrating a dual role for AI: accelerating development, while runtime inference is reserved for formats that genuinely resist deterministic parsing. A key finding: deterministic logic alone solved 51% of the 37 RDAMSC crosswalks (19/37), while 8 out of 20 LLM extraction attempts returned nothing, suggesting LLMs are most effective as selective recovery tools, not default parsers. The pipeline has extracted 599 mapping rules, exported as SSSOM TSV files (a community standard for sharing ontology mappings) in a version-controlled repository. A DataCite to Dublin Core benchmark yields 85.7% field coverage and 79.2% value overlap. All rules are stored in a SurrealDB knowledge graph (standards as nodes, mapping rules as edges), currently being extended toward transitive path discovery (A to B to C) and Graph-RAG-assisted candidate generation for uncharted standard pairs. This work demonstrates that hybrid deterministic-LLM pipelines can turn fragmented, human-readable documentation into machine-executable infrastructure, a pattern applicable well beyond metadata to any domain where specification documents have resisted automation.
ID: 370
/ Poster No. # 9: 013
Modalities: Image Methods: Foundation Models, Uncertainty Quantification Application Domain: Earth & Environment EO Foundation Models for Pre-Training Dataset Uncertainty Analysis German Aerospace Center, Remote Sensing Technology Institute, Germany Reliable real-world performance of supervised machine learning in Earth observation (EO) depends not only on model architecture, but critically on the data. This includes the adequate representation of the data distribution as well as the task definition based on input-target pairs. In EO it is important that the training data cover the real-world data distribution, including various acquisition conditions and edge cases. While a lot of data is available, the labeling process is often difficult and error-prone, and the tasks themselves are often affected by high levels of noise, which in combination lead to ambiguity and label noise. We evaluate a model-agnostic framework for dataset auditing prior to training a model on the task. The approach utilizes pre-trained EO foundation-model embeddings to quantify two major sources of failure risk: epistemic risk caused by insufficient distributional coverage, and aleatoric risk caused by ambiguity or label noise. Building on embeddings of Sentinel-2 imagery, our approach measures global coverage by assigning samples to semantically meaningful clusters derived from a large reference corpus of the foundation models’ training and validation datasets. The approach further estimates task clarity with simple one-vs-rest probing in the embedding space and by the analyzing label homogeneity in embedding neighbourhoods. The neighbourhood evaluation is additionally utilized as an analysis tool for label noise. The procedure is presented on a configurable benchmark setting of land cover classification and on commonly used standard datasets for land cover classification. The dummy dataset controlled benchmark enables us to adjust and evaluate different levels of aleatoric and epistemic risks. Across a range of foundation models and real datasets, the method recovers known differences in class separability and dataset diversity, and highlights uncovered but interpretable regions of the embedding space, including cloud-contaminated, artifact-heavy, and snowy patterns that can confuse downstream classifiers when such conditions are not present in training data. In addition, various examples of label noise can be detected. Overall, the resulting framework supports data-centric AI for science by turning foundation-model representations into actionable diagnostics for dataset curation, task design, and targeted data acquisition before costly training cycles begin. ID: 160
/ Poster No. # 9: 014
Modalities: Simulation Data, Other Methods: Physics-informed Machine Learning Application Domain: Health, Information FrustrAI-Seq: Scaling Local Energetic Frustration to the Protein Sequence Space 1: Helmholtz Munich, Germany; 2: Technical University of Munich, Germany; 3: Barcelona Supercomputing Center, Spain; 4: University College London, United Kingdom Proteins fold into their native three-dimensional (3D) structures by navigating complex energy landscapes shaped by the biophysical and biochemical properties of their sequence. Once folded, some sequence positions (dubbed residues) remain locally frustrated, reflecting functional constraints incompatible with optimal packing. This local energetic frustration provides important insights into protein function and dynamics, but its analysis typically relies on structure-based energy calculations and remains energetically costly at scale. Here, we introduce an ultra-fast sequence-based prediction of local energetic frustration directly from protein sequences using embeddings from protein language models (pLMs). Our method, coined FrustrAI-Seq, enables proteome-wide frustration profiling in minutes (∼17 minutes for the entire human proteome on a single NVIDIA H100 GPU) while retaining biologically relevant performance as shown for the α-globin and β-lactamase family. By eliminating the need for explicit structural or evolutionary information, this approach expands frustration analysis to protein regions and classes that were previously inaccessible, including intrinsically disordered regions and high-throughput de novo designed protein datasets. To support reproducibility and large-scale applications, we provide the largest freely available resource of precomputed local frustration scores to date (∼106 proteins), along with model weights and complete training and inference code at: github.com/leuschjanphilipp/FrustrAI-Seq. ID: 199
/ Poster No. # 9: 015
Modalities: Image Methods: Generative Models Application Domain: Information muBRAND: Multi-Branch Diffusion for Flexible Generation and Inpainting of Image Series 1: Forschungszentrum Jülich, Germany; 2: Helmholtz AI, Jülich, Germany; 3: C. und O. Vogt Insitut, Universitätsklinikum Düsseldorf, Düsseldorf, Germany; 4: Institute of Computational Visualistics, Koblenz University, Koblenz, Germany We introduce a unified diffusion model framework, muBRAND, for performing slice interpolation, cross-domain style transfer and damage repair in ordered image sequences such as serial sections, z-slices and temporal frames. ID: 367
/ Poster No. # 9: 016
Modalities: Graphs, Image, Multimodal Data, Time Series, Other Methods: Reinforcement Learning, Probabilistic Methods, Uncertainty Quantification, Other Application Domain: Health, Information Multimodal Analysis of Organoid Network Dynamics Using Multielectrode Arrays and Calcium Imaging Helmholtz Munich, Germany Brain organoids provide a promising platform for studying the emergence of neural network activity, yet interpreting their dynamics across developmental stages, recording modalities, and perturbation paradigms remains an open challenge. Here, we present a multimodal computational framework that integrates multi-electrode array (MEA) recordings and calcium imaging to characterize spontaneous and stimulus-evoked activity in developing brain organoids. By combining dimensionality reduction, representation learning, synchrony analysis, and stimulus-aligned response characterization, we aim to identify interpretable patterns of spontaneous and evoked network activity across temporal and spatial scales. Across MEA recordings, we observe developmental changes in population activity and network organization. Low-dimensional analyses separate activity patterns by recording day, suggesting that global population dynamics evolve systematically over time. In addition, spatial synchrony analyses indicate a shift from broadly synchronized, network-wide activity in early organoid recordings toward more spatially differentiated and localized communication patterns at later stages of development. Together, these findings suggest that organoid network organization evolves measurably over development and can be quantified using data-driven analysis methods. To complement the population-level temporal resolution of MEA recordings, we analyzed calcium imaging data from 2D human brain organoids to characterize stimulus-evoked responses at single-cell resolution. Our current pipeline extracts stimulus-aligned responses, performs baseline correction, and compares spontaneous and evoked activity across regions of interest. This framework enables comparison of organoid responses across spontaneous recording, multiple optogenetic stimulation patterns and durations, and pharmacological perturbation with ketamine. These analyses reveal heterogeneous response profiles across cells and provide a spatially resolved view of perturbation-driven dynamics that is not accessible through MEA alone. Together, these results demonstrate the value of AI-guided multimodal analysis for quantifying the development of neural dynamics in brain organoids. This framework provides a foundation for future studies of perturbation response, functional state discovery, and closed-loop modeling in organoid systems. ID: 206
/ Poster No. # 9: 017
Modalities: Multimodal Data, Simulation Data Methods: Physics-informed Machine Learning Application Domain: Energy, Matter Towards Multimodal Inference of Electron Bunch Structures using Gradient-Based Reconstruction with Differentiable Detector Models 1: Helmholtz-Zentrum Dresden-Rossendorf, Germany; 2: Ludwig-Maximilians-Universität München, Germany; 3: Shanghai Jiao Tong University, China; 4: Center for Advanced Systems Understanding, Germany; 5: Technische Universität Dresden, Germany; 6: Technische Universität Chemitz, Germany Reconstructing electron bunch structure from measured diagnostics is often ill-posed and non-unique, motivating methods that can incorporate physical priors and, when available, additional constraints from complementary measurements. Coherent transition radiation (CTR) spectroscopy provides single-shot sensitivity to longitudinal microstructure, yet retrieving the bunch profile from a measured spectrum remains a challenging phase-retrieval problem influenced by detector response and transfer functions. We present a gradient-based reconstruction framework built around differentiable detector models, designed to serve as a modular foundation for multimodal inference. As a first modality, we formulate CTR phase retrieval using a fully differentiable forward model of the measurement chain and introduce a phase-only gradient descent (GD) approach: the measured spectral amplitude is enforced as a hard constraint while the unknown Fourier phase is optimized under real-space physical priors. This re-parametrization improves optimization conditioning relative to amplitude-parametrized updates and allows experimental effects such as system response and bandwidth limits to be integrated seamlessly into the reconstruction loop. Using synthetic CTR spectra spanning strongly modulated and multi-peaked profiles, we benchmark against classical Gerchberg–Saxton-type projection methods and real-space GD baselines. The differentiable formulation achieves comparable reconstruction fidelity while providing a direct pathway to (i) principled inclusion of additional diagnostics as extra forward-model terms, (ii) uncertainty-aware reconstructions via ensembles or perturbation-based approximations, and (iii) future incorporation of physics-informed neural priors without changing the underlying detector model. This establishes a scalable basis for multimodal, uncertainty-aware electron bunch inference in realistic accelerator experiments. ID: 262
/ Poster No. # 9: 018
Modalities: Image Methods: Other Application Domain: Health NURBSeg: bridging the gap between medical imaging and CAD via end-to-end parametric segmentation 1: German Aerospace Center (DLR), Germany; 2: German Cancer Research Center (DKFZ), Germany Traditional semantic segmentation in medical imaging is fundamentally constrained by the discrete voxel grid. This limitation not only restricts geometric fidelity to the underlying image resolution but also necessitates extensive post-processing for integration into Computer-Aided Design (CAD) workflows. To overcome these barriers, we introduce NURBSeg, an emerging deep learning framework that shifts the segmentation paradigm from discrete label maps to continuous parametric surface representations. As a joint Helmholtz initiative, NURBSeg synergizes the expertise in geometric deep learning from the DLR with the medical imaging experience of the DKFZ. By predicting Non-Uniform Rational B-Spline (NURBS) parameters directly from 3D volumetric data, the framework enables the reconstruction of smooth, analytically defined anatomical models with inherent subpixel precision. Our approach utilizes a 3D convolutional backbone to deform a topological template, ensuring that the resulting geometry is manifold and natively compatible with digital engineering pipelines. While still in its early stages, NURBSeg is being evaluated across diverse clinical domains, including cardiac MRI (ACDC dataset) for high-precision stroke volume estimation. Beyond diagnostics, the project emphasizes interdisciplinary downstream applications, providing a direct bridge between raw medical imaging and precision engineering tasks like CAD-ready modeling. ID: 216
/ Poster No. # 9: 019
Modalities: Other Methods: Generative Models Application Domain: Health From prediction to generation of antigen-specific, full-length T-cell receptors 1: Helmholtz Munich, Germany; 2: Technical University of Munich, Germany Engineered T-cell therapies represent a promising approach for treating previously incurable diseases. However, the identification or generation of suitable TCRs remains a bottleneck. Computational methods hold the potential to accelerate the development of TCRs binding to target antigens. While the computational investigation of the TCR-epitope landscape has mainly focused on binding prediction, synthetic TCR design has recently emerged as the next frontier. Here, we present a proof-of-concept study on in silico generation of paired full-length TCR sequences reactive to a fixed epitope. To explore the feasibility and challenges of successful design, we utilized a unique dataset comprising thousands of TCRs experimentally validated as reactive towards four murine epitopes to train an autoregressive transformer model, called TCRGenesis. The model generated realistic TCR sequences, as assessed through biophysical and sequence-based properties, and achieved high binding scores according to a dedicated TCR-specificity predictor. Because the simultaneous generation of both TCR chains introduces the additional challenge of identifying compatible chain pairs, we developed a chain-pairing predictor to guide the selection of reactive from dysfunctional chain combinations. Crucially, generated sequences were experimentally validated using re-expressed Jurkat reporter T cells. Notably, 15 out of 27 (56%) tested synthetic TCRs showed reactivity towards the model epitope SIINFEKL, demonstrating the feasibility of full-chain TCR design. To guide future efforts, we systematically analyzed the relationship between the number of reactive TCRs in the training set and the model’s predictive and generative performances. Our results indicate that the tested autoregressive generative model requires approximately 500 reactive TCRs before diversity and generative capabilities begin to plateau, whereas approximately 100 examples are sufficient to train categorical prediction models. Notably, generative and predictive performance strongly depends on the natural diversity of the TCR repertoire of a nominative epitope-specificity. This work represents a first step towards paired full-sequence de novo design of epitope-specific TCRs in silico. By combining autoregressive sequence generation, chain-pairing prediction, and experimental validation, our results establish a practical framework for data-driven TCR design and guide future efforts toward scalable receptor engineering. ID: 261
/ Poster No. # 9: 020
Modalities: Graphs, Tabular Data, Text, Time Series Methods: Agentic AI, Probabilistic Methods, Other Application Domain: Energy, Matter Provenance-Aware Knowledge Graphs for Reusable PEFC and PEWE Experiments 1: Institute of Energy Technologies (IET-4), Electrochemical Process Engineering, Forschungszentrum Jülich GmbH, 52425 Jülich, Germany; 2: Theory and Computation of Energy Materials (IET-3), Institute of Energy Technologies, Forschungszentrum Jülich GmbH, 52425 Jülich, Germany; 3: Chair for Theory and Computation of Energy Materials, Faculty of Georesources and Materials Engineering, RWTH Aachen University, 52062 Aachen, Germany; 4: Data Analytics, Access and Applications, Scientific Computing Center (SCC), Karlsruher Institut für Technologie, 76344 Eggenstein-Leopoldshafen, Germany; 5: Helmholtz AI, Karlsruher Institut für Technologie, 76344 Eggenstein-Leopoldshafen, Germany Reaching the EU's climate goal of net zero emissions until 2050 depends on the rapid commercialization of advanced energy technologies. Hydrogen energy systems - including polymer electrolyte fuel cells (PEFCs) and electrolyzers (PEWEs) - are critical for decarbonization, yet traditional materials development remains slow. Meanwhile, advances in AI and cloud computing enable networks of laboratories, opening the possibility of accelerated, data-driven discovery cycles. However, most ML models require vast validation data stored per FAIR guidelines (Findable, Accessible, Interoperable, Reusable). In practice, experimental data from individual laboratories are still stored in diverse formats, and publications often lack the complete experimental context needed for reproducibility, reuse, and cross-institutional comparison. Zimmer’s meta-analysis of 127 PEWE papers reveals that essential parameters are missing in up to 70% of publications [1], rendering benchmarking and data-driven modeling unreliable. We address this challenge by developing an ontology-driven data pipeline that converts raw experimental outputs into provenance-aware knowledge graphs (KGs) for PEFCs and PEWEs. An AI-assisted workflow handles data extraction and semantic mapping aligned with existing H2 ontologies and metadata schemas. The KG links each derived quantity to a physical device (MEA and hardware), measurement protocol, and post-processing steps, transforming heterogeneous inputs into structured graph representations. This enables semantic search and traceable analysis of provenance, e.g. identifying comparable polarization curves under identical conditions or determining degradation rates associated with specific accelerated stress test histories. As a proof of concept, we mapped a dataset on PEFC cathode catalyst degradation from Schneider et al. [2] into the KG, producing a machine-readable performance history linking material properties, operating conditions, and degradation behavior. Our approach enables statistical correlation of performance metrics such as Tafel slopes and impedance-derived resistances while preserving the measurement and analysis context. The framework will be extended to PEWE datasets and scaled within the EU project DECODE to enable cross-lab benchmarking and AI-orchestrated collaborative R&D. [1] Zimmer, S. et al., J. Electrochem. Soc. (2026) doi: 10.1149/1945-7111/ae335e ID: 327
/ Poster No. # 9: 021
Modalities: Multimodal Data Methods: Foundation Models, Generative Models, Graph Neural Networks Application Domain: Core Machine Learning, Health Deep Learning Modeling of RNA Structure Perturbations Induced by Small Molecules 1: Computational Health Center, Helmholtz Munich, Munich, Germany; 2: School of Computation, Information and Technology, Technical University of Munich, Munich, Germany; 3: Department of Molecular Genetics, Groningen Biomolecular Sciences and Biotechnology Institute (GBB), University of Groningen, Groningen, The Netherlands; 4: Equal contributions Ribonucleic acid (RNA) plays a crucial role in regulating cellular processes, far beyond its status as an intermediate molecule between DNA and proteins. It has recently reemerged as a therapeutic target of choice for small molecules (SMs), with potential to address diseases associated with "undruggable" proteins, and opening the field to all non-coding RNAs (75% of the human transcriptome). Predicting modulations of RNA structures by SMs is limited by the scarcity of experimentally solved RNA 3D structures and the difficulties in modeling these structures computationally. Moreover, there is a lack of systematic approaches to identify functional RNA structures at large scale, that could alter molecular phenotypes when targeted. SHAPE-based chemical probing is a high-throughput protocol that measures nucleotide accessibility transcriptome-wide, linking RNA structures to measured mutation rates. By probing structures in a control condition and in a drug-treated condition, we can derive a functional readout of transcriptomic changes in response to drug exposure. In this work, we present an AI framework to predict the effect of eleven SMs on RNA structures of the human transcriptome. To this end, we leverage the measured probing profiles to model the effect of the SMs on the RNA using Generative AI. Specifically, we develop a diffusion model to simulate the effects of the drug treatment by generating probing profiles for given SMs and RNA sequences. Our model combines the multimodal input using specific encoders, notably RNA-based large language models (LLMs) and graph neural networks (GNNs) to obtain SM representations, to condition its generated output. We compare our model to various baselines, ranging from simple machine learning classification models to advanced classifiers and regression models. Additionally, we evaluate the performance of our model against state-of-the-art baseline methods, which were trained on data from RNA 3D structures. Overall, our AI framework confirms the effectiveness of the SHAPE protocol to measure small-molecule impact on the transcriptome, and reveals the binding preferences of the studied SMs. Moreover, by learning to predict / generate probing profiles from both RNA sequence and SM representations, we provide an in silico DRUG-seq screening pipeline that enables the interrogation of candidate SMs. This pipeline can be used to guide new experiments as well as to prioritize candidate drugs of interest. ID: 189
/ Poster No. # 9: 022
Modalities: Multimodal Data, Simulation Data, Other Methods: Probabilistic Methods, Other Application Domain: Health T-time: Multimodal Manifold Learning for Clonally Constrained Trajectory Inference 1: Helmholtz Munich, Germany; 2: Mila – Quebec AI Institute, Montréal, Canada; 3: Department of Mathematics and Statistics, Université de Montréal A central goal of single-cell transcriptomics is to reconstruct dynamic cellular trajectories from static snapshots of a tissue. Most trajectory inference (TI) methods rely on transcriptomic similarity as a proxy for developmental relatedness, an assumption that breaks when distinct lineages converge to similar transcriptional states or when intermediate states are sparsely sampled. Prospective lineage tracing can resolve these issues, but requires genetic perturbations and longitudinal experiments that are rarely available. In T cells, the T Cell Receptor (TCR) encodes clonal ancestry and is routinely co-sequenced with the transcriptome, yet current TI methods cannot exploit this information. We introduce T-time, a multimodal framework that integrates lineage information with transcriptomic profiles to constrain TI. T-time uses optimal transport-based clonal similarities to impose lineage constraints on diffusion maps, enabling lineage-consistent TI without longitudinal sampling or genetic perturbation. Lineage information guides the construction of a clonally constrained diffusion operator supporting representation learning, terminal fate prediction, and pseudotime inference while maintaining lineage consistency. To prevent clonal collapse and preserve local transcriptomic neighborhoods, we introduce a step-size-controlled alternating diffusion scheme that balances the modalities. Across biologically motivated simulations, T-time consistently outperformed RNA-only baselines (PHATE, diffusion maps, Palantir, CellRank) and multimodal alternatives (Alternating Diffusion, Integrated Diffusion, RF-PHATE). In cross-converging trajectories, T-time improved global topology preservation (Mantel 0.653 vs. 0.582; p=2×10⁻⁵), reduced fate total variation (TV 0.170 vs. 0.343; p=3×10⁻¹⁶), and increased per-cell fate accuracy (0.825 vs. 0.639; p=2×10⁻⁶), while preserving local neighborhoods. In cyclic topologies, pseudotime ordering improved substantially (r²=0.854 vs. 0.693; p=0.002), where RNA-only methods failed to resolve directionality. Gains were largest when transcriptomic and lineage signals were complementary, though injecting transcriptomic information into clonal distances reduced robustness in high-noise settings. Overall, T-Time provides a principled lineage-constrained TI framework. While motivated by T-cell differentiation, it generalizes to other endogenous lineage barcodes and offers a generalstrategy for lineage-constrained trajectory inference. ID: 248
/ Poster No. # 9: 023
Modalities: Text, Other Methods: Foundation Models, Generative Models, Reinforcement Learning Application Domain: Health End-To-End Protein Engineering With Protein Language Models 1: Institute of Computational Biology, Computational Health Center, Helmholtz Munich, Germany; 2: Department of Bioinformatics and Computational Biology , Technical University of Munich, Germany Protein language models (PLMs), large language models (LLMs) pre-trained on millions of protein sequences, have proven highly effective for predicting diverse structural and functional protein properties. The goal of protein engineering is to optimize properties such as thermostability or enzymatic activity for an existing wild-type protein. We propose PLM-based methods that together serve as an end-to-end framework for protein engineering. Alignment is an essential post-training step for LLMs and natural language applications which ensures that, e.g., users receive helpful rather than toxic responses from chatbots. To this end, we investigate how different Preference Optimization (PO) methods compare for the task of generating new protein variants, when aligning PLMs with preferred, more functional proteins. We show that unaligned PLMs are biased toward generating protein variants with few mutations, and demonstrate that selecting the right PO method is crucial to generate diverse variants with improved functionality at a high rate. In addition, we find that fine-tuning PLMs with ranking-based objectives yields a superior in silico screening of generated variants, enabling the selection of the most promising variants for experimental validation. In this context, we further explore how evolutionary data in the form of a multiple sequence alignment (MSA) of the wild type against a large sequence database can be leveraged to reduce the need for experimental protein fitness data during fine-tuning. Furthermore, we observe a performance scaling law for MSA Pairformer, a relatively small PLM pre-trained on MSAs, when ranking many protein variants with cross-sequence attention in a single forward pass. This hints at limitations in most current PLMs, which typically accept only a single sequence as input. ID: 161
/ Poster No. # 9: 024
Modalities: Image Methods: Other Application Domain: Health AI-assisted Labeling and its Pitfalls: A Case Study in Electron Microscopy Segmentation 1: Helmholtz AI, Helmholtz Center Munich, Germany; 2: Institute of Toxicology and Environmental Hygiene, TUM School of Medicine and Health, Technical University of Munich, Germany; 3: Institute of Molecular Toxicology and Pharmacology, Helmholtz Center Munich, Germany High-quality labels are of utmost importance in biomedicine and instrumental for tasks such as disease understanding, drug discovery, and medical diagnosis. Human-in-the-loop approaches and AI-assisted annotations are currently widely accepted as the safest approach for data curation, offering both speed and human supervision, while lever-aging the power of AI. Indeed, foundation models are often used to ob-tain initial labels which are then manually corrected by a human expert and subsequently used for downstream tasks. In this work, we uncover a rarely discussed risk of such semi-automated approaches and present a case study to demonstrate this. Focusing on microscopy imaging seg-mentation, where annotation quality is critical and costly, we collected and annotated a Transmission Electron Microscopy dataset in a semi-automated approach, using BATS, a software we built for AI-assisted segmentation labeling in microscopy which encourages data centric prac-tices. We show how batch effects can be derived by a mix of human and modelannotationsaffectingthedownstreamanalysis.Finally, wedemon-strate a solution which mitigates these batch effects and discuss how future researchers can adopt responsible and accurate data annotation pipelines. ID: 323
/ Poster No. # 9: 025
Modalities: Image, Multimodal Data, Time Series Methods: Foundation Models Application Domain: Core Machine Learning, Earth & Environment Spatial and Temporal Evaluation of Extreme Precipitation in ML-Based Global Weather Forecasts Forschungszentrum Jülich, Germany Machine learning (ML) weather models have achieved remarkable skill for bulk meteorological variables, yet consistently struggle to capture extreme precipitation in both magnitude and location. We present an evaluation framework for diagnosing how ML forecasts fail on the extreme tail of the precipitation distribution, applied to the WeatherGenerator, a global ML forecasting system. ID: 311
/ Poster No. # 9: 026
Modalities: Audio, Text, Time Series, Video Methods: Other Application Domain: Core Machine Learning, Information Recurrent Neural Networks and Tensor Ring Decomposition KU Leuven, Belgium Managing over-parameterization in deep learning, particularly for sequence models, has driven extensive research into model compression. Tensor Ring Decomposition lets us represent massive data arrays as a closed loop of smaller, interconnected tensors. This allows systems to store and process enormous datasets using just a fraction of the memory and computing power. ID: 326
/ Poster No. # 9: 027
Modalities: Graphs, Tabular Data, Text Methods: Foundation Models, Graph Neural Networks, Other Application Domain: Health AI and Hierarchical Clustering Techniques for Accurate Patient Stratification 1: Hochschule München, Campus for Health and Engineering, Munich, Germany; 2: PerMediQ, Stuttgart, Germany; 3: University of Stuttgart, HLRS - High Performance Computing Center, Stuttgart, Germany; 4: Klinikum Stuttgart, Stuttgart Cancer Center - Tumorzentrum Eva Mayr-Stihl DE, Stuttgart, Germany; 5: University Hospital Regensburg, Department of Internal Medicine I, Regensburg, Germany Graph-based methods for data representation and analysis are well suited for encoding both data points and their interrelationships. This approach integrates data and topology, enabling the representation of interrelated information. In this study, we represent patient cohorts as cohort graphs (CG) and discuss their application for real-world patient data (RWD) extracted using natural language processing (NLP) methods, comparing fine-tuned transformer models with large language models (LLMs). We focus on developing methods to cluster patients with similar symptoms and examine how demographic stratification parameters (sex and age group) influence interlinking within CGs, thereby refining hierarchical clustering for more accurate patient stratification and supporting personalized decision-making in a clinical context. We illustrate how incorporating sex and age group information improves symptom-based clustering in a population of patients with lung and gastrointestinal cancer presenting to the emergency department. Finally, we outline how high-performance computing (HPC) can enable the upscaling of graph-based analytical methods to larger, multicentric cohorts, where pairwise similarity computation scales quadratically with cohort size.
ID: 362
/ Poster No. # 9: 028
Modalities: Text, Other Methods: Agentic AI, Graph Neural Networks, Other Application Domain: Core Machine Learning, Aeronautics, Space & Transport, Energy, Earth & Environment, Information, Matter SciTracer - the reliable and logically provable autonomous scientific discovery. Unaffiliated, Ukraine AI scientists are shaping and revolutionizing the future of scientific discovery process and laboratory workflow. However, solutions of highly complex tasks that are defined beyond standard school exercises or exam problems require much more resources and decision-making steps than a simple heuristic-based solution. Moreover, the discovery trajectory of an AI scientist can be knotted and entangled because of the lack of consistency, agentic consensus, and huge number of redundant LLM calls, and entails human intervention for control of the AI researcher. These drawbacks makes autonomous discovery ureliable in case of complex and compute demanding task. In this paper, we present and analyze an AI agent tailored for automating the scientific discovery process supported by a verification protocol that mitigates the mentioned problem and guarantees the safe discovery plan execution. The Sci-Tracer is grounded on the Hoare Triplet Logic text decomposition and Abstract Syntax Tree-based planning and verification. The integration of the software verification tools and logic provers leads to greater truthfulness and sustainability of autonomous research and science. ID: 300
/ Poster No. # 9: 029
Modalities: Image Methods: Other Application Domain: Earth & Environment Benchmarking for Microscopic Biodiversity Screening (AIMBIS-project) 1: Helmholtz-Centre for Environmental Research GmbH (UFZ), Germany; 2: German Centre for integrative Biodiversity Research e.V. Halle-Jena-Leipzig (iDiv), Germany; 3: Helmholtz-Centre for Polar and Marine Research – AWI, Department/Section Polar Terrestrial Environmental Systems; 4: Helmholtz-Centre for Polar and Marine Research – AWI, Department/Section Polar Biological Oceanography Many microscopic biodiversity monitoring tasks are currently still dependent on manual microscopy. These tasks would greatly benefit from higher throughput but also enhanced objectivity and data preservation. Multispectral imaging flow cytometry (MIFC) enables automated analysis with a capacity of 100 samples per day with up to 5000 particles/s while also facilitating image archiving, leading to scalable monitoring efforts. Innovative AI solutions are needed to overcome limitations of automated image recognition and to further develop label-efficient machine learning approaches using benchmark datasets from various domains. The AIMBIS project has already jointly collected images from relevant domains (recent/ fossil pollen and phytoplankton of polar/ temperate environments). These data are processed to gain appropriate benchmark datasets for different applications in microscopic ecological monitoring tasks.
ID: 245
/ Poster No. # 9: 030
Modalities: Simulation Data, Tabular Data, Text, Other Methods: Generative Models, Probabilistic Methods, Other Application Domain: Core Machine Learning, Matter MolecuLA: Learning Molecules As A Language For Generative And Interpretable Chemistry 1: CASUS/Helmholtz-Zentrum Dresden-Rossendorf, Germany; 2: University of Wrocław, Poland We introduce MolecuLa; a framework for learning generative and chemically interpretable molecular latent spaces from sequence-based molecular representations. Our approach is built on a Transformer variational autoencoder with efficient linear-attention components, enabling modeling of large molecular corpora while preserving a continuous latent space suitable for analysis and control. Molecules are represented as "SELFIES", a robust string representation in which every valid token sequence maps to a valid molecular graph, making it particularly attractive for generative modeling. Our model achieves strong reconstruction fidelity, reaching 99% sequence-level accuracy, and provides a stable basis for latent-space interrogation. On top of this generative backbone, we identify global latent directions associated with chemically relevant properties, including molecular weight, lipophilicity, polarity, hydrogen-bonding capacity, saturation, and ring content, using linear-probe evaluation with R² based analysis. To distinguish genuine chemical structure from artifacts of the underlying string representation, we further introduce a confound-aware evaluation based on SELFIES length, branch and ring token counts, and token entropy. This reveals that while some highly predictable directions are driven by representation-level shortcuts, several important chemical properties remain strongly encoded after confound removal, yielding stable and traversal-ready directions. Together, these results position scalable Transformer VAEs as a promising foundation for interpretable molecular generation and controlled property-aware molecule editing. ID: 342
/ Poster No. # 9: 031
Modalities: Simulation Data Methods: Generative Models, Other Application Domain: Health Construct-valid and adaptive benchmarks for single-cell RNA-seq via a hierarchical, streaming simulator 1: Chair for Clinical Bioinformatics, Saarland University, Germany; 2: Helmholtz Institute for Pharmaceutical Research Saarland We study how to design evaluation benchmarks for single-cell RNA-seq methods with construct validity, reliability under adaptivity and reuse, and scaling. We use a hierarchical simulator spanning study, batch, plate, donor, and cell-type levels, with modular effects for library-size shifts, ambient RNA, doublets, logistic dropout, and negative-binomial dispersion. Data are emitted in streaming micro-batches into fixed-size buffers with logging of outputs and ground truth, enabling large runs while keeping memory constant per chunk. Construct validity is assessed by calibrating a parameter set to public resources including open CellxGene and the ROSMAP/COMPASS AD atlas (~16 million cells). Calibration fits the mean and standard deviation of a log-normal library-size model, a negative-binomial dispersion parameter from per-gene means and variances, and logistic dropout parameters from the relationship between zero fraction and mean expression; when stable proxies exist, an ambient RNA parameter is also included. We quantify divergence between real and simulated data using summary distributions of library size, gene-wise mean-variance relationships, and dropout as a function of mean expression. Reliability under adaptivity is studied by regenerating multiple distribution-matched test sets from the same calibrated configuration and measuring variability in downstream conclusions across repeated draws. To link simulator parameters to benchmark behavior, we inject known confounders and signals. Batch-level perturbations such as library-size offsets or ambient contamination are evaluated with integration procedures, with recoverability quantified by iLISI and cLISI before and after correction. In separate experiments, predefined differential-expression effects are injected into subsets of genes, and recovery is quantified by AUROC against logged ground truth after a DE method. Scaling is summarized by wall-clock time, peak resident memory, and output size as a function of cell number. For fixed gene dimension and cell counts from 10,000 to 1,000,000, we fit a linear runtime model and report its slope with hardware specifications. The contribution is an evaluation protocol rather than a single benchmark: minimal calibration with quantitative construct-validity checks, controlled stress tests, adaptivity analysis from regenerated test sets, and a reproducible scaling profile. All configurations and scripts will be released for reproduction on a single workstation. ID: 344
/ Poster No. # 9: 032
Modalities: Image, Multimodal Data Methods: Generative Models, Probabilistic Methods, Uncertainty Quantification Application Domain: Core Machine Learning, Health From T-FAKE to TFAN: Reusable Thermal Face Alignment for AI in Scientific Imaging 1: Chair for Clinical Bioinformatics, Saarland University, Germany; 2: Technical University of Berlin, Germany; 3: MPI for Informatics, SIC, Germany; 4: Helmholtz Institute for Pharmaceutical Research, Germany; 5: Equal author contribution Facial landmarking is a bottleneck for thermal-imaging workflows in health, physiology, and behavioral analysis because annotated thermal datasets are small, heterogeneous, and rarely support dense alignment. We present T-FAKE, the first large-scale synthetic thermal face dataset with 200k images from at least 2k subjects under warm and cold conditions, 70 sparse landmarks, 478 dense landmarks, and segmentation masks, together with TFAN, a released Python package providing pretrained sparse and dense models for thermal face alignment. T-FAKE is generated from synthetic RGB faces using our RGB2Thermal loss, which combines supervised reconstruction, patch-wise Wasserstein matching, and facial temperature priors to produce realistic thermal structure. Models trained with T-FAKE achieve state-of-the-art sparse thermal landmarking on the challenging CHARLOTTE benchmark while maintaining competitive RGB performance on 300W, and they enable, to our knowledge, the first dense thermal facial landmarker. Our probabilistic landmark predictor also supports uncertainty-aware integration of face detection into the alignment pipeline, and TFAN makes the models directly usable on temperature-valued or grayscale thermal images, including multi-face inference. Beyond the dataset and models, we define a benchmark protocol that harmonizes incompatible landmark conventions with a learned label-adaptation model and evaluates the practical combined task of face detection and landmarking. Together, the dataset, protocol, and released TFAN package are a reusable AI solution for thermal face analysis by making training, evaluation, and deployment comparable and reproducible across studies. This turns thermal facial landmarking from an ad hoc preprocessing preprocessing task into an accessible, standardized, and comparable step for registration, tracking, and region-based temperature analysis across health-related and behavioral imaging studies. The code is available from GitHub and as the thermal-face-alignment package on PyPI.
ID: 201
/ Poster No. # 9: 033
Modalities: Image, Other Methods: Probabilistic Methods, Other Application Domain: Core Machine Learning From Feature Learning to Neural Collapse: A Data-Dependent Kernel View 1: Helmholtz Munich, Germany; 2: Technical University of Munich, TUM School of Computation, Information and Technology; 3: Faculty of Mathematics, Informatics and Mechanics, University of Warsaw Neural collapse is an emergent phenomenon that occurs during the final phase of training, where the representations in the last layer of a neural network converge into a highly structured geometry. In here, we focus on the first characteristic: the vanishing of within-class variance. Previous theoretical approaches to understanding this phenomenon, like the unconstrained feature models, are data-independent. Here, we build on kernel methods from feature learning theories to capture neural collapse in deep networks after training. The training data explicitly enters into these feature learning kernels, allowing us to study its effect on neural collapse. Further, we investigate what drives neural collapse in terms of the network and training hyperparameters. Overall, we obtain a data-aware theory describing the emergence of neural collapse. ID: 354
/ Poster No. # 9: 035
Modalities: Multimodal Data, Time Series, Video Methods: Other Application Domain: Core Machine Learning Do world models model the world? Helmholtz Munich, Germany World models are widely seen as a promising route toward agents capable of long-horizon, goal-directed reasoning. However, it remains unclear whether current world models learn representations that are compact, physically grounded, and suitable for causal reasoning. ID: 173
/ Poster No. # 9: 036
Modalities: Image Methods: Foundation Models Application Domain: Core Machine Learning, Health Editing the Foundations: Task Arithmetic for 3D Medical Segmentation Foundation Models 1: Helmholtz Munich, Germany; 2: Technical University of Munich, Germany; 3: King's College London, UK 3D biomedical foundation models have enabled substantial progress in multi-organ segmentation. However, adapting these large-scale models to different imaging protocols, modalities, or institution-specific data distributions remains challenging. Conventional finetuning often leads to catastrophic forgetting, sacrificing generalizable knowledge for narrow domain specialization. We introduce a novel framework for enhancing 3D medical segmentation foundation models through weight-space task arithmetic. Each new domain is modeled as a task vector derived from a specialist model and composed with the base foundation model directly in weight space. This approach enables the composition of multiple domain-specific task vectors within a single parameter space, yielding a unified enhanced model. For experimental validation, we use AMOS dataset as the base-domain benchmark for multi-organ CT/MRI segmentation and we select organ-specific CT datasets from MSD (spleen, liver, pancreas) and LiTS (contrast-enhanced liver and tumor scans) for domain enhancement. Our results (Figure 1) demonstrate that domain-specific task vector scaling coefficient induces an implicit trade-off between generalization and specialization. The enhanced models consistently achieved substantial gains on target domains while preserving base-domain performance, with forgetting limited to negligible levels. Figure 2 shows that multiple domain-specific task vectors can be composed within a single parameter space to yield a unified, enhanced model. Although interactions between task vectors influence domain-specific outcomes, the proposed approach successfully balances preservation of base-domain performance with simultaneous improvement across multiple target domains. Moreover, our label-efficiency analysis (Table 1) reveals that near-optimal adaptation can be achieved with only a small fraction of annotated target-domain data. Achieving approximately 96% of peak performance with only 5% of the 102 total training cases underscores the approach's practical viability in annotation-scarce settings. Overall, our results establish weight-space composition as an efficient and scalable paradigm for enhancing 3D medical foundational models. Our approach enables controlled, non-destructive adaptation without requiring access to prior training data, architectural expansion, or task-specific routing mechanisms, making it particularly well aligned with the requirements of real-world clinical environments.
ID: 285
/ Poster No. # 9: 037
Modalities: Image, Multimodal Data Methods: Foundation Models, Other Application Domain: Health Spatial transcriptomics-informed inference of cell types and states from nuclear-stained whole slide images. 1: Institute of Computational Biology (ICB), Helmholtz Zentrum München, Munich, Germany; 2: Institute for Stroke and Dementia Research (ISD), Klinikum der Universität München, Munich, Germany; 3: German Center for Neurodegenerative Diseases (DZNE), Munich, Germany; 4: Munich Cluster of Systems Neurology (SyNergy), Munich, Germany; 5: Institute of Neuronal Cell Biology, Technical University Munich, Munich, Germany Identifying the spatial organization of cells of different types and in different functional states in tissues is fundamental to understanding their biological functions. In brain tissue, cell phenotyping relies on of cell-specific markers, such as antibody-based staining, or novel technologies such as Spatial Transcriptomics (ST). However, both approaches are costly and therefore not always accessible. Contrarily, DAPI nuclear staining is an accessible tool routinely used in laboratories. Trained microscopy experts are able to use nucleus features to distinguish between main brain cell types, and studies exist demonstrating that it is possible to automatize inference of cell cycle stage, type or functional state from DAPI-stained nuclei (Eulenberg, 2017; Narotamo, 2020; Narotamo, 2021; Li, 2025; Hua, Li, 2026). We are currently developing a supervised Deep Learning (DL) model to predict cell types and states from DAPI Whole Slide Images (WSIs) of mouse brain tissue, leveraging ST-derived labels for the classification (Fig.1). Our aim is to develop a fast and reliable tool for wet-lab scientists that can assist cell type and state annotation without expensive markers or technologies. Moreover, such a tool could provide insights into the role of different cell types in pathological mechanisms in brain tissue. Initially, we used one high resolution DAPI WSI from a MERFISH experiment performed on mouse brain (Fig.2). From this image we cropped 224x224 tiles around each cell nucleus, which (after oversampling of rare classes) resulted in a dataset of 33’124 cells in total. These were used to train a ResNet50 model to predict the corresponding ST-derived cell type labels (“Astroependymal”, “Immune”, “Neuronal”, “Oligodendrocytes”, “Vascular”), resulting in 73% balanced accuracy on training set, and 59% on the validation set (Fig.3). We plan to expand our training data to 20 additional MERFISH mouse brain DAPI WSIs. We also want to further develop our model with a hierarchical architecture, to account for additional local and global features that can inform cell type assignment. Overall, we expect these improvements to further increase the accuracy of our model; make it more robust to batch effects; increase its resolution in distinguishing cell sub-types; and make it capable of discriminating between not only cell types, but also functional and pathological states (which we will evaluate using DAPI WSIs from MERFISH data of an Alzheimer’s Disease (AD) mouse model).
ID: 277
/ Poster No. # 9: 038
Modalities: Image Methods: Uncertainty Quantification Application Domain: Health Uncertainty Quantification for Medical Image Segmentation HZDR, Germany Medical image segmentation is a critical task in medical image analysis, where the goal is to delineate anatomical structures or regions of interest within medical images. However, the performance of segmentation algorithms can be affected by various sources of uncertainty, such as noise in the input data, variability in the imaging process, and limitations of the segmentation model itself. Uncertainty quantification is an important aspect of medical image segmentation, as it provides insights into the reliability and confidence of the segmentation results and improves trustworthiness in clinical applications. In this work, we review and discuss two methods for uncertainty quantification: Mean-variance estimation and clustered inference ensembling. We compare both methods on a set of benchmark datasets and evaluate their performance in terms of uncertainty calibration, compute efficiency and implementation complexity. Our results show that both methods can provide valuable insights into the uncertainty of segmentation results, but they have different strengths and weaknesses. Mean-variance estimation is computationally efficient and well calibrated but it requires model changes and re-training, while clustered inference ensembling requires some effort for calibration is also computationally efficient but does not require model changes and re-training. We conclude that the choice of uncertainty quantification method should be based on the specific requirements of the application and the available computational resources. ID: 281
/ Poster No. # 9: 039
Modalities: Simulation Data, Text Methods: Agentic AI, Other Application Domain: Core Machine Learning Fast and Few-Shot: Equipping Scientific LLM Agents with Meta-Learned Surrogates 1: Institute for Advanced Simulations – Materials Data Science and Informatics (IAS-9), Forschungszentrum Jülich GmbH, 52425 Jülich, Germany; 2: Chair of Materials Data Science and Materials Informatics, Faculty 5 – Georesources and Materials Engineering, RWTH Aachen University, 52056 Aachen, Germany Autonomous Large Language Model (LLM) agents are reshaping scientific discovery, yet their integration with physical simulations remains inefficient. While machine learning-based surrogate models are frequently employed to accelerate high-throughput query scenarios such as design optimization and uncertainty propagation, they introduce an important bottleneck: static global surrogates often fail to generalize to diverse, out-of-distribution physical parameter regimes. Conversely, relying on iterative active learning or continuous ground-truth solver calls introduces computational latency that fundamentally disrupts the agent's autonomous reasoning loop. To address this, we propose an agentic framework integrating an adaptive, meta-learned surrogate model as a core tool. We formulate the surrogate as a task-specific operator, pre-trained across a broad distribution of physical parameter regimes using gradient-based meta-learning to capture the underlying governing dynamics. Upon encountering a novel physical regime, the agent orchestrates a targeted calibration phase: it queries the expensive numerical solver a minimal number of times to perform a rapid, few-shot gradient-based adaptation of the surrogate. This adaptation step instantiates a specialized, local surrogate, which enables the agent to execute computationally intensive downstream tasks in many-query applications without further overhead. In this work, we will demonstrate the efficacy of this paradigm through empirical validation against state-of-the-art baselines across parameterized physical systems. Specifically, we will highlight how our meta-surrogate recovers high-fidelity predictive accuracy via few-shot adaptation, and subsequently drives high-throughput downstream applications such as inverse design optimization with minimal computational latency and numerical solver calls. ID: 397
/ Poster No. # 9: 040
scooby: modeling multimodal genomic profiles from DNA sequence at single-cell resolution. 1: Technical University of Munich, School of Computation, Information and Technology, Munich, Germany; 2: Helmholtz Center Munich, Computational Health Center, Neuherberg, Germany; 3: Technical University of Munich, Munich Center for Machine Learning, Munich, Germany; 4: Technical University of Munich, Institute of Human Genetics, School of Medicine, Munich, Germany Understanding how regulatory sequences shape gene expression across individual cells is a fundamental challenge in genomics. Joint RNA sequencing and epigenomic profiling provides opportunities to build models capturing sequence determinants across steps of gene expression. However, current models, developed primarily for bulk omics data, fail to capture the cellular heterogeneity and dynamic processes revealed by single-cell multimodal technologies. Here, we introduce scooby, a framework to model genomic profiles of single-cell RNA-sequencing coverage and single-cell assay for transposase-accessible chromatin using sequencing insertions from sequence at single-cell resolution. For this, we leverage the pretrained multiomics profile predictor Borzoi and equip it with a cell-specific decoder. Scooby recapitulates cell-specific expression levels of held-out genes and identifies regulators and their putative target genes. Moreover, scooby allows resolving single-cell effects of bulk expression quantitative trait loci and delineating their impact on chromatin accessibility and gene expression. Finally, we present preliminary results when applying scooby to multi-species data. We anticipate scooby to aid unraveling the complexities of gene regulation at the resolution of individual cells. ID: 197
/ Poster No. # 9: 042
Modalities: Image, Multimodal Data, Time Series Methods: Physics-informed Machine Learning, Uncertainty Quantification, Other Application Domain: Earth & Environment Geospatial Causal Inference for Drought Driver Analysis DLR, Germany Satellite Earth Observation (EO) enables large-scale monitoring of vegetation dynamics and drought conditions, but most analyses rely on correlational relationships between vegetation indices and climate variables. While useful for detecting patterns, correlation alone cannot identify the underlying drivers of vegetation stress. Declines in vegetation greenness may coincide with drought indicators but can also result from confounding factors such as temperature extremes, radiation anomalies, land cover differences, or irrigation practices. Without causal analysis, EO-based assessments risk misattributing drought impacts. This study proposes a causal inference framework to estimate the impact of drought on vegetation using satellite time series. Drought severity is modeled as a treatment variable and vegetation greenness as the outcome within a structural causal model informed by hydrological and ecological processes. Key environmental variables—including temperature, precipitation, solar radiation, and land cover—are incorporated as confounders, while soil moisture is treated as a mediator representing the pathway through which atmospheric drought affects vegetation. Causal effects are estimated using linear causal models applied to observational EO datasets. Satellite vegetation indices derived from MODIS (NDVI/EVI) are combined with drought indicators such as SPEI or PDSI and climate reanalysis variables from ERA5. Additional datasets describing soil moisture, land cover, and irrigation are used to account for confounding influences and to explore how ecosystem characteristics modify drought responses. The framework enables estimation of vegetation sensitivity to drought while controlling for environmental variability. It also supports analyses of how drought impacts vary across land-use classes, whether irrigation reduces vegetation vulnerability to meteorological drought, and what fraction of drought effects are mediated by soil moisture deficits. By integrating causal reasoning with EO data, this approach provides more interpretable attribution of vegetation stress and contributes to improved drought impact assessment and environmental decision support.
ID: 283
/ Poster No. # 9: 043
Modalities: Graphs, Multimodal Data, Simulation Data, Text Methods: Generative Models, Graph Neural Networks, Other Application Domain: Core Machine Learning, Health When Protein Dynamics Matter: Integrating Molecular Dynamics into Protein Foundation Models jülich forschungszentrum, Germany Proteins are dynamic molecules whose function depends not only on sequence and structure but also on conformational changes over time. We investigate how Molecular Dynamics (MD) trajectories can be integrated as an additional modality in protein foundation models by extending OneProt with all-atom, time-resolved MD data curated from mdCATH, GPCRmd, and ATLAS databases. These trajectories encode protein flexibility, conformational variability, and thermodynamic sensitivity, complementing static sequence-, structure-, and text-based representations. Using a pre-trained transformer-based MDGen encoder, we perform systematic pre-training ablations and evaluate the resulting representations across diverse protein prediction tasks. We find that incorporating MD trajectories consistently improves performance on downstream tasks sensitive to dynamics or structural context, particularly when explicit structural information is limited. Our results demonstrate that transformer-based MD encoders capture biologically meaningful dynamic signals that enhance protein foundation models, highlighting the value of integrating protein dynamics for potential applications such as protein engineering and drug discovery.
ID: 314
/ Poster No. # 9: 044
Modalities: Image, Multimodal Data Methods: Probabilistic Methods, Uncertainty Quantification, Other Application Domain: Health Toward Automated MRI-Based Screening of Cortical Superficial Siderosis Using Multimodal 3D nnU-Net 1: Institute for Stroke and Dementia Research(ISD), University Hospital, LMU Munich,Munich, Bavaria, Germany; 2: Institute of Computational Biology (ICB), Helmholtz Munich, Neuherberg, Germany Cortical superficial siderosis (cSS) is a clinically relevant MRI manifestation of cerebral amyloid angiopathy (CAA), a severe type of cerebral small vessel disease. Manual delineation of cSS is time-consuming and particularly difficult in subtle cSS cases and longitudinal change assessment [1]. We developed a preliminary deep-learning workflow for cSS rim segmentation as a screening support tool. We trained a multimodal 3D residual-encoder nnU-Net (nnU-Net v2; R2* + QSM maps) on 10 expert-labeled cSS-positive cases using 5-fold cross-validation (8 training / 2 held-out per fold). All folds were completed and evaluated on fold-held-out subjects [2]. Voxel-level aggregation across all held-out predictions (10/10 cases; each evaluated out-of-fold) yielded Dice 0.664, precision 0.736, and recall 0.605. Fold-wise mean Dice values (from per-fold validation summaries) were 0.685, 0.681, 0.479, 0.663, and 0.542 (average 0.610), showing expected variability in this small cohort. Qualitative overlays indicate that the model captures cSS patterns in most held-out cases. Dominant errors are false positives in sulcal/vascular look-alike regions and misses in subtle cSS segments (with non-target deep punctate findings outside the current segmentation scope). These findings support feasibility for AI-assisted cSS pre-screening, but not autonomous cSS detection at this stage. Next steps are: (i) dataset expansion with additional labeled cSS-positive cases and cSS-negative scans, (ii) combine fold-model predictions and apply anatomy-aware post-processing to reduce false positives; and (iii) deploy an expert-in-the-loop active-learning workflow that prioritizes uncertain and heterogenous candidate cases to improve labeling efficiency and longitudinal progression tracking in large-cohort MRI pipelines [3]. Reference 1. van Harten, T. W., et al. "Quantitative measurement of cortical superficial siderosis in cerebral amyloid angiopathy." NeuroImage: Clinical 38 (2023): 103447. 2. Isensee, Fabian, et al. "nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation." Nature methods 18.2 (2021): 203-211. 3. Lüth, Carsten T., et al. "nnactive: A framework for evaluation of active learning in 3d biomedical segmentation." arXiv preprint arXiv:2511.19183 (2025).
ID: 321
/ Poster No. # 9: 045
Modalities: Multimodal Data, Tabular Data Methods: Generative Models, Probabilistic Methods, Uncertainty Quantification Application Domain: Core Machine Learning, Health Cross-View Latent Integration via Nonparametric Gamma Shrinkage Factor Analysis 1: Institute of AI for Health, Helmholtz Munich; 2: Faculty of Mathematics, Informatics and Mechanics, University of Warsaw; 3: Oncology Data Science, Merck Healthcare KGaA Factor analysis is a dominant paradigm for multi-omic heterogeneous data, but is challenged by partially redundant signals and noise across views and by an unknown true number of factors. We present CLING, an unsupervised multi-view factor model with hierarchical Bayesian sparsity priors: a product-of-Gammas prior inducing cumulative column-wise shrinkage (increasing with factor index) coupled with a Gamma–Gamma local-precision hierarchy on loadings yielding heavy-tailed marginals. This pairing enables automatic factor selection by adaptively deactivating unsupported factors while retaining active ones during inference, and induces selective sparsity that allows salient loadings to escape shrinkage while collapsing negligible ones. As a fully conjugate hierarchical model, CLING admits a scalable variational inference algorithm for multi-view data. Across synthetic benchmarks and multi-omics datasets, CLING recovers more accurate factors and more informative loadings while explaining at least as much variance as competitive multi-view baselines; on glioblastoma gene expression and DNA methylation data, CLING identifies pathways linked to tumor subtype and patient age. ID: 209
/ Poster No. # 9: 046
Modalities: Image Methods: Generative Models, Probabilistic Methods Application Domain: Health Physics-Guided Diffusion for Neonatal Ultra-Low-Field MRI enhancement 1: University of Bonn, University Hospital Bonn, Clinic for Diagnostic and Interventional Radiology, Germany; 2: University of Bonn, University Hospital Bonn, Department of Experimental Neonatology, Germany; 3: Technical University of Munich, School of Computation, Information and Technology, Germany; 4: Laboratoire Traitement du Signal et de l’Image, Inserm UMR 1099, Universit´ e de Rennes, France Clinical Motivation The neonatal brain is highly vulnerable during the perinatal period. Globally, ~59 million children under five experience developmental delays due to neonatal brain disorders, with ~2.3 million neonatal deaths annually (WHO). In NICUs, cranial ultrasound is fast and bedside-accessible but operator-dependent, artifact-prone, and offers limited views of deep structures. HF-MRI remains the gold standard yet requires sedation, long acquisition times (45-60 min), and costly infrastructure. Portable ultra-low-field MRI (uLF-MRI, 0.064 T) offers a safer bedside alternative, but its reduced signal-to-noise ratio and poor tissue contrast severely limit diagnostic reliability. Existing IQT methods - GAN- and CNN-based approaches such as LoHiResGAN, SFNet, LF-SynthSR, and GAMBAS - attempt to bridge this quality gap but suffer from poor structural fidelity through oversharpening, edge hallucination, or anatomical oversmoothing. Critically, they rely on simplistic degradation simulations that fail to capture true low-field physics, and none adequately address neonatal populations with diverse pathologies. Method A physics-guided pipeline simulates realistic uLF acquisitions for effective paired training. Starting from a partially noised uLF input, the model progressively refines it conditioned on the original scan, guided by a perceptual loss prioritizing anatomical fidelity and contrast in neonatal brains. Results MRIQT significantly outperforms all baselines on a neonatal dataset spanning diverse brain conditions: +1.8 dB PSNR (to 15.34 dB) and +5.9% Pearson correlation (to 0.413). Downstream tissue segmentation (CSF, gray matter, white matter) showed a +9.4% Dice improvement (to 0.474), reflecting clinically meaningful structural recovery. In a blinded reader study (3 clinicians, 34 scans), 85% of enhanced images were rated good quality, with 80% pathology detection accuracy and 82.3% sensitivity - and no notable hallucinations or distortions. Conclusion MRIQT demonstrates how physics-guided generative AI can enable reliable bedside neonatal neuroimaging without sedation or transport, paving the way for earlier pathology detection in resource-limited settings. ID: 286
/ Poster No. # 9: 047
Modalities: Image Methods: Foundation Models Application Domain: Health Atlas 2 - Foundation Models for Clinical Deployment 1: Aignostics, Germany; 2: Department of Laboratory Medicine and Pathology, Mayo Clinic, Rochester, MN, US; 3: Department of Radiology, Mayo Clinic, Rochester MN, US; 4: Department of Information Technology, Mayo Clinic, Rochester MN, US; 5: Mayo Clinic, Rochester MN, US; 6: Digital Pathology, Mayo Clinic, Rochester MN, US; 7: Machine Learning Group, Technische Universität Berlin, Germany; 8: BIFOLD – Berlin Institute for the Foundations of Learning and Data, Germany; 9: Department of Artificial Intelligence, Korea University, Republic of Korea; 10: Max-Planck Institute for Informatics, Germany; 11: German Cancer Research Center (DKFZ) & German Cancer Consortium (DKTK), Berlin & Munich Partner Sites, Germany; 12: Institute of Pathology, Ludwig-Maximilians-Universität München, Germany; 13: Institute of Pathology, Charité – Universitätsmedizin Berlin, Germany; 14: Bavarian Cancer Research Center (BZKF), Germany; 15: Helmholtz Munich, Germany; 16: Technical University Munich, Germany Pathology foundation models (FMs) underpin clinical-grade computational pathology applications, yet tradeoffs between prediction performance, robustness to confounding variations (scanner, staining, lab processing), and computational efficiency have limited clinical translation. Prior work has addressed these dimensions individually—through larger pretraining datasets, stain normalization, domain-adversarial training, or knowledge distillation—but no model family has jointly optimized all three. We present Atlas 2, Atlas 2-B, and Atlas 2-S, a family of vision FMs that bridge these shortcomings and set a new state of the art across 80 public benchmarks. Atlas 2 is a 2-billion-parameter Vision Transformer (ViT, patch size 8) trained on 5.5 million de-identified whole slide images from Charité Berlin, LMU Munich, and Mayo Clinic—the largest pathology FM dataset to date. Tiles are extracted at four resolutions (0.25–2.0 µm/px). Training builds on DINOv2/v3 with proprietary improvements. Two lightweight variants, Atlas 2-B (ViT-B, 86M params) and Atlas 2-S (ViT-S, 22M params), are obtained via knowledge distillation, with 3.4× and 9× faster inference. Evaluation uses five public frameworks—eva, HEST, PathoROB, Plismbench, and Patho-Bench—with frozen encoders and task-specific heads, covering morphology classification, mutation and gene expression prediction, robustness, and more. Atlas 2 achieves best performance in 22/27 main benchmark tasks, with average scores of 44.8% on HEST (+1.6 pp over the next best model), 82.9% on eva (+1.2 pp), and a robustness index of 85.7% (+9.7 pp). On Patho-Bench, Atlas 2 leads in 15/24 molecular and 14/19 morphology tasks. Distilled models retain strong performance: Atlas 2-B matches or exceeds much larger models (UNI2-H, Virchow2) while being 24× smaller, and Atlas 2-S rivals H-Optimus-0 and Virchow2 at 91× fewer parameters. Both distilled variants show considerably improved robustness over similarly sized competitors (+29 and +25 pp, respectively). Atlas 2 substantially minimizes the tradeoff between accuracy, robustness, and efficiency, establishing a new Pareto front across all three dimensions. Consistent gains across diverse tasks, tissue types, and evaluation protocols suggest multi-centric data diversity as a key driver. Efficient distillation further facilitates clinical adoption under real-world throughput and cost constraints. Atlas 2 represents a significant step toward FMs ready for routine pathology diagnostics. ID: 219
/ Poster No. # 9: 048
Modalities: Audio, Image, Multimodal Data, Simulation Data, Text, Time Series, Video Methods: Agentic AI, Foundation Models, Generative Models Application Domain: Core Machine Learning EOSC-ARENA: Towards an Open and Ethical GenAI Infrastructure for European Research 1: Scientific Computing Center (SCC), Karlsruhe Institute of Technology (KIT), Eggenstein-Leopoldshafen, 76344,Germany; 2: Research Data Alliance Association - RAD, Belgium; 3: Instituto de Física de Cantabria (IFCA), CSIC-UC, Avda. los Castros s/n, Santander, 39006, Cantabria, Spain; 4: Instituto de Instrumentación para Imagen Molecular (I3M), Centro Mixto CSIC — Universitat Politècnica de València, Camino de Vera s/n, 46022 Valencia, España; 5: Ethniko Kentro Erevanas Kai Technologikis Anaptyxis - CERTH, Greece; 6: nstitute of Informatics, Slovak Academy of Sciences (IISAS), Dúbravská cesta 9, Bratislava, 84507, Slovakia; 7: Instytut Chemii Bioorganicznej Polskiej Akademii Nauk - PSNC, Poland Generative Artificial Intelligence (GenAI) is transforming research by enabling new forms of discovery, automation, and collaboration. However, current AI ecosystems often lack transparency, reproducibility, and alignment with open science principles, raising concerns about data provenance, ethical use, and reliance on proprietary platforms. The recently approved European project, EOSC-ARENA (starting in June 2026), aims to create an independent, open, and agentic GenAI ecosystem integrated with the European Open Science Cloud (EOSC), supporting researchers across the entire research data lifecycle. Methodologically, EOSC-ARENA incorporates human-centered and responsible AI principles, with dedicated work packages focused on research integrity, ethics, and skills development. Building on the AI4EOSC platform, EOSC-ARENA will provide federated infrastructure for Federated GenAI model fine-tuning, scalable model serving, and trustworthy knowledge-augmentation services, including Retrieval-Augmented Generation and Model Context Protocols, across European e-infrastructures and EOSC Nodes. It will ensure model, data, and platform independence in line with the EU AI Act, FAIR principles, and EOSC policies. The project will validate its infrastructure through interdisciplinary use cases in life sciences, medical imaging and oncology, climate science, cosmology, materials science, food science, and research data management, aiming to improve data quality, FAIRness, and research productivity. A dedicated work package on research integrity, ethics, and responsible AI will develop cross-domain curricula, guidelines, and open educational resources, embedding human-centered, trustworthy, and reproducible practices into GenAI workflows. In this contribution, we present the EOSC-ARENA project, highlighting its vision, infrastructure, and planned agentic GenAI framework. ID: 121
/ Poster No. # 9: 049
Modalities: Graphs, Multimodal Data Methods: Graph Neural Networks Application Domain: Health Multimodal graph-based fusion learning for Alzheimer’s disease prediction 1: Modular High-Performance Computing and Artificial Intelligence, German Center for Neurodegenerative Diseases (DZNE), Bonn, North Rhine-Westphalia, Germany; 2: Systems Medicine, German Center for Neurodegenerative Diseases (DZNE), Bonn, North Rhine-Westphalia, Germany; 3: PRECISE Platform for Genomics and Epigenomics at DZNE and University of Bonn, Bonn, North Rhine-Westphalia, Germany; 4: Life and Medical Sciences (LIMES) Institute, Genomics & Immunoregulation, Bonn, North Rhine-Westphalia, Germany Alzheimer’s disease (AD) is a heterogeneous and complex neurodegenerative disease that is responsible for 60-70% of dementia cases worldwide. Early detection of individuals at risk remains a major challenge, particularly because disease progression begins years before clinical symptoms become severe. Peripheral blood mononuclear cells (PBMCs) provide a minimally invasive window into systemic immune alterations associated with AD and therefore represent a promising resource for early detection and disease monitoring. In this work, we propose a multimodal learning framework that integrates single-cell RNA sequencing, single-cell ATAC sequencing, flow cytometry, and clinical metadata to predict cognitive stages of Alzheimer’s disease. Instead of analysing each modality independently, we represent each patient as a graph that captures cellular heterogeneity and relationships between immune cells separately from each modality. We construct patient-specific graphs where nodes represent individual cells with edges defined by k-nearest neighbour (kNN) affinity in the respective latent manifold of each modality. We employ graph neural networks (GNNs) with hierarchical pooling to learn compact patient-level embeddings that summarize complex cellular interactions while preserving biologically meaningful structure. To effectively combine complementary information across modalities, we introduce a cross-attention-based fusion mechanism that adaptively weighs the contribution of each modality for patient-level prediction. This design enables patient-specific modality weighing and improves interpretability by revealing modality importance as well as cell-type and feature-level relevance scores derived from attention maps. Beyond prediction, the framework supports immune profiling by identifying cell populations and molecular signatures associated with disease progression. Our approach demonstrates how multimodal graph learning can bridge cellular, molecular, and clinical information to improve the prediction of Alzheimer’s disease stages. This work highlights the potential of immune-based biomarkers and graph-based deep learning as a step toward more precise and explainable risk assessment in neurodegenerative disease.
ID: 178
/ Poster No. # 9: 050
Modalities: Text Methods: Generative Models Application Domain: Information Can Large Language Models Grade Short Free-Text Exam Answers? An Evaluation on Real German Student Data Forschungszentrum Jülich, Germany Over the past few years, the capabilities of Large Language Models (LLMs) have rapidly improved, expanding their potential application domains. One widely discussed area is the use of LLMs for educational assessment. However, many existing studies rely on English-language datasets and often exhibit limitations such as synthetically generated answers or simplified evaluation schemes based on grade categories or classification rather than detailed point-based scoring. In this project, we investigate whether LLMs can support the grading of German-language short free-text exam answers using authentic educational data. Our dataset consists of more than 8,000 real student responses to 51 exam questions, each evaluated by human experts. In addition to the question and student answer, the dataset contains a reference solution, grading hints, and a point-based score assigned by human graders. To obtain a reliable evaluation basis, we extracted a validation and test subset that were manually re-evaluated to establish a high-quality ground truth. This allows us to analyze the alignment between LLM-generated scores and human grading while accounting for the inherent variability of human assessment. For systematic experimentation, we developed a highly reproducible evaluation framework based on Hydra that enables flexible configuration of models, prompts, and parameters. Using this framework, we evaluate several state-of-the-art LLMs and investigate how different prompting strategies influence grading performance. Our results provide insights into the potential and limitations of LLM-based grading for German-language assessments and highlight how model choice and prompting strategies affect alignment with expert evaluations. We also analyze the ethical implications of applying LLM-based grading in real-world educational settings. In particular, we discuss potential risks such as bias, transparency, and over-reliance on automated assessment, and outline mitigation strategies. Furthermore, we assess the regulatory context under the EU AI Act to evaluate compliance requirements for the responsible deployment of such systems. ID: 202
/ Poster No. # 9: 051
Modalities: Time Series Methods: Other Application Domain: Health Pseudo-Anomaly Augmented One-Class Autoencoder and Supervised Deep Learning for Murine ECG Anomaly Detection 1: Helmholtz AI, Helmholtz Center Munich, Germany; 2: Mary Lyon Centre at MRC Harwell, Harwell Campus, United Kingdom; 3: Integrative Physiology/Advanced Technology Cores, Baylor College of Medicine, Texas, United States; 4: Institute of Experimental Genetics, Helmholtz Center Munich, German Research Center for Environmental Health, Germany Reliable identification of abnormal patterns in murine electrocardiogram (ECG) recordings is essential for quantitative analysis in preclinical cardiovascular research. We investigate deep learning-based approaches for automated detection of abnormal ECG patterns in a murine dataset comprising ECG recordings from anesthetized mice generated at International Mouse Phenotyping Consortium (IMPC) testing centers, with beat-level ground-truth annotations provided by expert scorers. As preprocessing, ECG signals are segmented into doublets, i.e., beat-anchored windows of length twice the median inter-beat interval to account for inter-individual variability in inter-beat intervals. This representation provides sufficient temporal context around the beat of interest while limiting the inclusion of signal regions far from the beat of interest that may introduce irrelevant variability or noise. We compare two learning paradigms. First, we employ a supervised Conv1D-based classifier with an effective receptive field covering the entire doublet, allowing the learned features to jointly consider both beats. Second, we formulate the task as one-class anomaly detection using an autoencoder (AE) trained exclusively on normal doublets. To enhance the AE framework, we introduce a pseudo-anomaly augmentation strategy in which synthetic perturbations are generated from normal doublets during training. These pseudo anomalies encourage tighter modeling of physiological patterns and increase reconstruction sensitivity to abnormal deviations. Supervised learning achieves the strongest overall detection performance (test area under the precision–recall curve (PR-AUC): 98.23%). The baseline AE shows limited discrimination capability (test PR-AUC: 82.52%). However, incorporating pseudo-anomaly augmentation substantially improves performance to a test PR-AUC of 94.37%. Overall, our results highlight the effectiveness of supervised learning for murine ECG anomaly detection and demonstrate that pseudo-anomaly augmentation provides a simple yet impactful enhancement for one-class AE-based methods. Qualitative analysis of false positive and false negative predictions suggests that some disagreements with ground-truth labels may stem from borderline or ambiguous annotations, indicating potential benefits from further label refinement. ID: 349
/ Poster No. # 9: 052
Modalities: Text Methods: Other Application Domain: Energy, Earth & Environment A Physics-Based Machine Learning Workflow for Urban-Scale Residential Heat-Demand Estimation Chair of Methods for Model-based Development in Computational Engineering, RWTH Aachen University, Germany Physics-based building energy modeling and simulation provide a scientifically grounded basis for analyzing residential thermal energy demand, but their large-scale application is limited by high computational cost and, in some workflows, by complex software dependencies and fragmented toolchains. This work presents a physics-based machine learning pipeline that accelerates simulation-based heat-demand analysis through surrogate modeling. The approach builds on German residential TABULA archetypes and generates reduced-order thermal models with TEASER (Tool for Energy Analysis and Simulation for Efficient Retrofit) for a set of representative buildings in Germany. To efficiently explore the input space, parameters are sampled using Latin Hypercube Sampling, and annual simulations are carried out in OpenModelica. The resulting synthetic dataset spans diverse building and climate scenarios and is used to train machine learning surrogates that emulate the outputs of the underlying simulations. The surrogate serves as a computationally efficient emulator of a scientifically grounded modeling framework rather than replacing physics-based simulation with a fully data-driven approach. Across the evaluated model configurations, the trained surrogates closely reproduce the reference simulation results, achieving MAPE values between 0.88% and 3.89%, while reducing average evaluation times from approximately 23–24 s to about 1 ms, corresponding to a speed-up of more than four orders of magnitude. Such acceleration is critical for applications where the one-time cost of training is justified by many subsequent evaluations, such as large-scale scenario analysis and urban digital twins. In addition to the surrogate models, the work introduces a containerized end-to-end workflow that improves reproducibility and portability by encapsulating the required software stack and execution environment. This approach reduces practical barriers to reuse and supports more robust deployment across computing environments. The study demonstrates how machine learning can support urban-scale analysis of residential heat demand. The contribution lies in combining physics-based modeling, dynamic simulation, and surrogate model development within a reproducible framework. In this context, artificial intelligence serves not as a standalone predictor but as an enabler for accelerating and broadening access to scientific computation. ID: 305
/ Poster No. # 9: 053
Modalities: Multimodal Data, Tabular Data, Other Methods: Probabilistic Methods Application Domain: Health scDigital-Karyotypes: a genomic representation to learn shared variants 1: Max Delbrück Center for Molecular Medicine, Berlin, Germany; 2: Helmholtz AI, Helmholtz Munich, Germany; 3: Department of Nephrology and Hypertension, Hannover Medical School, Hannover, Germany; 4: Charité – Universitätsmedizin Berlin, Germany Precision diagnostics relies on identifying genomic signatures shared across patients, yet somatic mutations complicate these comparisons by generating extensive heterogeneity across individual cells. Single-cell sequencing technologies resolve genomic variation at cellular resolution, however scalable computational approaches that compare variation across cells and patients remain limited. Strand-seq is a single-cell DNA sequencing technique that enables genome-wide detection of structural variants (SVs) together with epigenetic signals derived from nucleosome occupancy (NO). Current analytical methods are limited to SV detection within individual samples and do not support systematic discovery of shared structural variant patterns across large patient cohorts. Here, we introduce a scalable machine-learning framework for multi-patient analysis of single-cell Strand-seq data. Our approach jointly segments genomes across all cells and patients, and employs Bayesian probabilistic modeling to estimate SV state probabilities for 70 disease-relevant SV classes. These probabilistic SV profiles are then used to compute cell–cell similarities by defining a similarity metric, enabling unsupervised clustering of cells based on shared genomic signatures. We then apply supervised dimensionality reduction to identify and visualise cluster-specific genomic markers. We subsequently integrate the epigenetic modality captured by NO to assess the potential functional impact of the genomic markers on cellular health. By learning SV-defined genomic signatures across heterogeneous cell populations, our framework enables data-driven identification of cellular subgroups and provides a scalable strategy for cross-patient genomic stratification. This work demonstrates how machine learning provides a disease-agnostic strategy for signature-associated patient risk stratification to better understand clonal dynamics and support precision diagnostics. ID: 139
/ Poster No. # 9: 054
Modalities: Image Methods: Foundation Models, Probabilistic Methods, Other Application Domain: Health Semi-Supervised Federated Learning for Medical Image Classification Ozyegin University, Turkey (Türkiye) Medical image classification is frequently hindered by two critical bottlenecks: the scarcity of high-quality labeled data and stringent privacy regulations that prevent data centralization. Federated Semi-Supervised Learning (FSSL) has emerged as a robust paradigm to address these challenges by enabling collaborative training across decentralized institutions. However, statistical heterogeneity among clients (Non-IID data) remains a significant hurdle, often degrading model performance and pseudo-label quality. In this work, we propose an enhanced FSSL framework that builds upon the SAGE (Semi-supervised Aggregation for Globally-Enhanced Ensemble) architecture to mitigate pseudo-label mismatches through confidence discrepancy. To further optimize global model stability, we integrate ShapFed, a contribution-aware aggregation strategy. Unlike conventional Federated Averaging (FedAvg), our framework weights client updates according to their Class-Specific Shapley Values (CSSV), accounting for both local data quality and its alignment with the global optimization path. Our methodology includes a comprehensive comparative analysis on the HAM10000 skin lesion dataset, benchmarking the proposed framework against the original SAGE and standard FedAvg schemes in realistic, heterogeneous medical scenarios. Initial validation on the CIFAR-10 dataset confirmed the functional integration of Shapley-driven weighting. While the computation of CSSV introduces a trade-off in convergence speed, we demonstrate that this additional complexity is a necessary investment for achieving superior accuracy and robustness in challenging Non-IID settings. This research offers a scalable path toward privacy-preserving, high-precision AI in decentralized healthcare ecosystems.
ID: 181
/ Poster No. # 9: 055
Modalities: Time Series Methods: Other Application Domain: Health ECG Template-Based Deep Learning for Amplified P-wave Duration Prediction 1: Karlsruhe Institute of Technology, Germany; 2: University Heart Center Freiburg - Bad Krozingen, Germany; 3: Lucerne Cantonal Hospital, Switzerland; 4: Justus Liebig University Giessen, Germany Atrial cardiomyopathy (AtCM) is associated with electrical and structural remodeling of the atria, frequently resulting in atrial fibrillation (AF), leading to increased mortality. Amplified P-wave duration (APWD), annotated in 12-lead ECG templates, i.e. signals averaged over multiple heart beats, enables non-invasive quantification of AtCM. In this study, we present a deep learning regression model for APWD prediction, detecting amplified P-wave (APW) onset and offset in ECG templates. A bidirectional LSTM architecture with 256 hidden units and two layers was utilized. The primary dataset comprised 1,044 ECG templates with expert APWD annotations from the University Heart Center Freiburg - Bad Krozingen, including 3 subgroups (individuals between 18 and 30 years of age, cardiovascular patients without AF, and AF patients. To address the limited size of the labeled dataset, transfer learning was applied. For pretraining, 129,302 ECG recordings from the publicly available MIMIC-IV database were processed to generate ECG templates. Labels were assigned using ECGdeli, an open-source ECG analysis toolbox. These labels were considered weak, as they reflected standard P-wave duration rather than APWD. Data were split into 70% training, 10% validation, and 20% testing sets. We systematically evaluated the effect of pretraining and performed subgroup analysis. Additionally, the model was tested on an independent dataset of 200 ECGs from patients with diagnosed AF. Performance was assessed using mean absolute error (MAE) relative to expert annotations and Bland-Altman analysis. The best-performing model achieved MAEs of 9.03±10.88, 11.81±10.03, and 12.09±10.76 ms on the test set for onset, offset, and APWD, respectively. Bland-Altman analysis revealed no systematic bias (bias=0.69 ms). Training without pretrained weights resulted in MAEs of 39.76±40.03, 38.51±41.00, and 20.00±16.18 ms. Pretraining therefore substantially improved model performance. Subgroup analysis showed similar MAEs for all subgroups (MAEs ranging from 09.73±10.50 to 13.29±11.76). On the independent AF cohort, MAEs for onset, offset, and APWD were 12.38±11.29, 13.53±11.94, 15.18±10.41 ms, respectively (bias=9.20 ms). In conclusion, the proposed algorithm demonstrates high generalizability and therefore a strong potential for automated APWD quantification, enabling objective and reproducible large-scale screening for AtCM. ID: 155
/ Poster No. # 9: 056
Modalities: Simulation Data Methods: Other Application Domain: Aeronautics, Space & Transport Reduction of Outflow Boundary Influence on Aerodynamic Performance Using Neural Networks 1: Jülich Supercomputing Centre (JSC), Forschungszentrum Jülich, Wilhelm-Johnen-Straße 52428 Jülich, Germany; 2: School of Mechanical and Mining Engineering, The University of Queensland, St Lucia 4072, Queensland, Australia; 3: Chair of Fluid Mechanics, University of Siegen, Paul-Bonatz-Straße 9-11, Siegen-Weidenau 57076, Germany; 4: Institute of Technology, Resource and Energy-efficient Engineering (TREE), Bonn-Rhein-Sieg University of Applied Sciences, Grantham-Allee 20, Sankt Augustin 53757, Germany; 5: Fraunhofer Institute for Algorithms and Scientific Computing (SCAI), Schloss Birlinghoven, Sankt Augustin 53754, Germany; 6: Institute of Aeronautics and Applied Mechanics, Warsaw University of Technology, Nowowiejska 24, Warszawa 00-665, Poland The accurate treatment of outflow boundary conditions remains a critical challenge in computational fluid dynamics when predicting aerodynamic forces and/or acoustic emissions. This is particularly evident when employing the lattice Boltzmann method (LBM) as the numerical solution technique, which often suffers from inaccuracies induced by artificial reflections from outflow boundaries. This paper investigates the use of neural networks to mitigate these adverse boundary effects and enable truncated domain requirements. Two distinct NN-based approaches are proposed: (1) direct reconstruction of unknown particle distribution functions at the outflow boundary; and (2) enhancement of established characteristic boundary conditions by dynamically tuning their parameters. The direct reconstruction model was trained on data generated from a two-dimensional (2D) flow over a cylindrical obstruction. The drag, lift, and Strouhal number were used to test the new boundary condition. We analyzed results for various Reynolds numbers and restricted domain sizes, where it demonstrated significantly improved predictions when compared with the traditional Zou and He boundary condition. To examine the robustness of the NN-based reconstruction, the same condition was applied to the simulation of a NACA0012 airfoil, again providing accurate aerodynamic performance predictions. The neural-enhanced characteristic boundary condition (CBC) was evaluated on a 2D convected vortex benchmark and showed superior performance in minimizing density errors compared to CBCs with fixed parameters. These findings highlight the potential of NN-integrated boundary conditions to improve accuracy and reduce computational expense of aerodynamic and acoustic emission simulations with the LBM. ID: 162
/ Poster No. # 9: 057
Modalities: Tabular Data, Text, Other Methods: Agentic AI Application Domain: Health Rolv.io - Multi-Agent AI For Healthcare And Research 1: Max Delbrück Center (MDC), Germany; 2: Hacettepe University, Turkey This work presents a multi-modal, multi-agent artificial intelligence platform designed specifically for healthcare and life sciences research. The system integrates domain-specialized reasoning, code execution, and scientific data analysis to support researchers across diverse biomedical workflows. Through its multi-modal capabilities, the platform interprets and generates text, data, and visual content, enabling seamless handling of complex biological datasets and analytical tasks. ID: 264
/ Poster No. # 9: 058
Modalities: Tabular Data Methods: Other Application Domain: Energy Understanding Green Technology Adoption in Germany: A Cluster-Based Analysis of Household-Scale Photovoltaics and Battery Electric Vehicles 1: Institute of Climate and Energy Systems: Energy Systems Engineering (ICE-1), Forschungszentrum Jülich, Germany; 2: Helmholtz AI, Helmholtz Munich, Germany; 3: Institute of Climate and Energy Systems: Jülich Systems Analysis (ICE-2), Forschungszentrum Jülich, Germany; 4: School of Business and Economics, RWTH Aachen University, Germany; 5: Institute for Theoretical Physics, University of Cologne, Germany Mitigating climate change requires a transformation away from fossil fuels towards renewable energy sources for energy generation and transportation. This transition necessitates the widespread adoption of green technologies not only by public entities and private companies, but also by individual private households. Since adoption rates vary greatly by region, identifying which factors drive high or low adoption rates and which regions follow similar adoption patterns is crucial for designing targeted policy interventions. In this work, we focus on the adoption of household-scale photovoltaic systems (PV) and battery electric vehicles (BEVs). We compiled a diverse set of demographic, geographic, political, and socioeconomic features from publicly available data sources at the level of municipal associations in Germany. We train Random Forest models to predict regional variations in the accumulated peak power of household-scale photovoltaics, and the share of battery electric vehicles. To analyze heterogeneity beyond standard feature importance measures, we apply Forest-Guided Clustering (FGC). FGC groups municipal associations into clusters based on similarities in their decision paths within the trained Random Forest models. In contrast to instance-level explanation methods, such as SHAP, which quantify feature contributions for individual observations, FGC identifies clusters of municipal associations that follow similar decision paths and estimates both overall and cluster-specific feature importances. This reveals not only that different features are important for the different technologies – adoption of PV is associated with large dwellings, and BEV is strongly correlated with income – but also how feature importances vary across clusters. Understanding the similarities and differences between the identified clusters — such as which features become more or less important across regions — provides actionable insights for designing more targeted studies of green technology adoption and for developing effective policy interventions. ID: 269
/ Poster No. # 9: 059
Modalities: Image Methods: Foundation Models Application Domain: Core Machine Learning From Learned Representation To Localization: Towards Semantic Integration Of Microscopic Data In The Human Brain 1: Institute of Neuroscience and Medicine (INM-1), Research Centre Jülich, Germany; 2: Helmholtz AI, Research Centre Jülich, Germany; 3: Computer Vision, Institute for Computational Visualistics, University of Koblenz, Germany; 4: C. & O. Vogt Institute of Brain Research, University Hospital Düsseldorf, Germany; 5: Department of Physics, University of Wuppertal, Germany Understanding the organization of the human brain requires analyzing data from different modalities and scales. Such multi-modal analysis is supported by integrating data into reference brain atlases. While microscopic imaging of histological sections provides high-resolution structural information, integration into atlases is challenging, as it typically requires 2D and 3D registration. Alignment becomes particularly challenging for small regions of interest—image patches or small tissue blocks—due to the lack of reliable anatomical landmarks in the limited field of view. ID: 1375
/ Poster No. # 9: 060
Modalities: Image Methods: Foundation Models, Other Application Domain: Earth & Environment AI-assisted Classification of Polar Phytoplankton’s Functional Biodiversity using Multispectral Imaging Flow Cytometry 1: Alfred Wegener Institute, Germany; 2: University of Bremen, Germany Ongoing environmental change in polar oceans is rapidly affecting phytoplankton communities living inside and below the sea ice. Assessing these changes in biodiversity requires determining species composition, as well as measuring physiological traits such as silicification and mixotrophy at a single-cell level. However, current methods are time-intensive, require expert training, and are subject to human bias. This project explores the use of AI in conjunction with multi-spectral imaging flow cytometry (MIFC) to develop a high-throughput methodology to assess biodiversity, silicification, and mixotrophy in polar phytoplankton. MIFC produces large datasets of multi-channel images capturing cell morphology, pigmentation, and fluorescence signals associated with physiological processes, providing a growing dataset of polar phytoplankton cells as a basis for automated taxonomic and functional classification. However, several challenges complicate the application of AI methods to plankton images. These include: limited expert-labelled training data, high morphological variability within species, and class imbalance across taxa. This project is working in conjunction with other plankton imaging initiatives such as AqQua and the Helmholtz UNLOCK project AIMBIS, to build a curated dataset of polar phytoplankton images collected during laboratory experiments and recent Arctic and Antarctic expeditions which will be used to investigate CNN approaches for classifying taxa and functional traits such as silicification and mixotrophic behaviour. By integrating imaging flow cytometry with machine learning approaches, this work aims to enable rapid identification of phytoplankton taxa and quantification of functional biodiversity traits such as silicification intensity and trophic modes whilst contributing to the development of automated and scalable tools for marine biodiversity monitoring. ID: 293
/ Poster No. # 9: 062
Modalities: Image Methods: Uncertainty Quantification Application Domain: Health LEVERAGING MULTI-RATER ANNOTATIONS TO CALIBRATE OBJECT DETECTORS IN MICROSOPY IMAGING Helmholtz AI, Helmholtz Zentrum Muenchen, Germany In microscopy image analysis, deep learning-based object detectors have achieved remarkable detection accuracies and are widely used to identify and quantify biological structures such as cells, nuclei, or organoids for various applications, including medical diagnostics and drug discovery. Despite their success, they often produce poorly calibrated confidence estimates, and relatively little attention has been paid to how confident the models are in their predictions. This is especially important for microscopy data, which often exhibit substantial variability due to diverse experimental conditions, imaging artefacts, and biological heterogeneity. Moreover, annotations are provided by human experts, who frequently disagree due to the subjective definitions of target structures and ambiguities from imaging artefacts or noise In this work, we introduce a new approach to improve model calibration by leveraging multi-rater annotations. We propose to train separate models on the annotations from single experts and aggregate their predictions to emulate consensus. ID: 276
/ Poster No. # 9: 063
Modalities: Image, Text Methods: Foundation Models, Generative Models Application Domain: Core Machine Learning, Health CytoDiff: AI-Driven Cytomorphology Image Synthesis for Medical Diagnostics Helmholtz Munich, Germany Biomedical datasets are often constrained by stringent privacy requirements and frequently suffer from severe class imbalance. These two aspects hinder the development of accurate machine learning models. While generative AI offers a promising solution, producing synthetic images of sufficient quality for training robust classifiers remains challenging. This work addresses the classification of individual white blood cells, a critical task in diagnosing hematological malignancies such as acute myeloid leukemia (AML). We introduce CytoDiff, a stable diffusion model fine-tuned with LoRA weights and guided by few-shot samples that generates high-fidelity synthetic white blood cell images. Our approach demonstrates substantial improvements in classifier performance when training data is limited. Using a small, highly imbalanced real dataset, the addition of 5,000 synthetic images per class improved ResNet classifier accuracy from 27% to 78% (+51%). Similarly, CLIP-based classification accuracy increased from 62% to 77% (+15%). These results establish synthetic image generation as a valuable tool for biomedical machine learning, enhancing data coverage and facilitating secure data sharing while preserving patient privacy. Paper code is publicly available at https://github.com/ JanCarreras24/CytoDiff.
ID: 143
/ Poster No. # 9: 064
Modalities: Time Series Methods: Probabilistic Methods Application Domain: Earth & Environment DeepElbe: Multi-Horizon Oxygen Forecasting and Hypoxia Early Warning Using Statistical and Deep Learning Methods Helmholtz-Zentrum Hereon, Germany Dissolved oxygen depletion poses a recurring threat to aquatic ecosystems in the tidal reach of the Elbe River within Hamburg, making reliable forecasting essential for timely risk management. We present DeepElbe, a reproducible framework for multi-horizon prediction of oxygen concentration and hypoxia events at monitoring stations along the tidal stretch. Building on established approaches, the framework compares a family of classical statistical baselines with data-driven neural networks under a unified train/validation/test protocol. Forecasts are formulated at exact lead times for both daily and hourly resolutions. The data-driven component includes dedicated artificial neural networks (ANNs) for both oxygen-level prediction and hypoxia-event classification. Evaluation uses task-specific metrics to analyze lead- and station-dependent trade-offs between modeling paradigms, and we show the complementarity of these approaches. DeepElbe is intended to provide the methodological foundation for an operational early warning system and is designed for extension to additional predictors, model classes, and other river systems. ID: 395
/ Poster No. # 9: 065
Cancer-label-free Tumor Segmentation in Whole-mouse Light Sheet Microscopy Images Using Vision-language Models 1: Institute for Stroke and Dementia Research, Klinikum der Universität München, Ludwig-Maximilians University Munich, Munich, Germany; 2: Institute for Intelligent Biotechnologies (iBIO), Helmholtz Center Munich, Neuherberg, Germany; 3: Faculty of Medicine, Ludwig-Maximilians University Munich, Munich, Germany; 4: Munich Cluster for Systems Neurology (SyNergy), Munich, Germany; 5: Munich Medical Research School (MMRS), Munich, Germany; 6: Graduate School of Neuroscience (GSN), Munich, Germany; 7: Deep Piction GmbH, Munich, Germany; 8: School of Medicine, Koç University, İstanbul, Turkey Detecting tumors in preclinical whole-body imaging is a critical step toward understanding both primary tumor growth and cancer metastasis. Light sheet microscopy (LSM) enables volumetric imaging of entire mouse bodies at high resolution, but current segmentation approaches often rely on cancer-specific fluorescent labels that cannot be translated to human clinical samples. Here, we present a framework for segmentation of both primary tumors and metastatic clusters using exclusively non-cancer imaging channels, designed for direct transferability to human tissue. We utilize whole-mouse LSM data from cancer mouse models exhibiting primary tumors and metastases predominantly in the lung, liver, and colon. Mice are imaged under two protocols: autofluorescence (AF) only; or AF combined with Propidium Iodide (PI). A cancer channel is used solely for generating ground truth annotations and is excluded from model input. Tumor regions remain identifiable in non-cancer channels through characteristic intensity variations in AF and PI, as well as abnormal nuclear density patterns. Our approach leverages a Vision-Language Model (VLM) that jointly integrates LSM image data with structured textual metadata, including species, organ of origin, and imaging protocol. This allows the model to adapt to protocol variations. We further adopt a semi-supervised learning strategy based on sparse annotations, which will be evaluated against fully supervised baselines. By removing dependency on cancer-specific labels at inference, this framework is inherently transferable to human samples. This work demonstrates the potential of VLMs to bridge preclinical imaging and clinical applicability in oncological image analysis. ID: 186
/ Poster No. # 9: 066
Modalities: Graphs, Image, Multimodal Data, Simulation Data, Text Methods: Generative Models, Physics-informed Machine Learning Application Domain: Energy, Matter Physics-Informed Reconstruction of Electron-Bunch Shapes from CTR Spectra using an Untrained Neural Network 1: Helmholtz-Zentrum Dresden-Rossendorf, Germany; 2: Technische Universität Dresden, Germany; 3: Center for Advanced Systems Understanding, Germany; 4: Technische Universität Chemnitz, Germany Coherent transition radiation (CTR) spectroscopy serves as a valuable diagnostic tool for characterizing the structural properties of relativistic electron bunches. However, the phase information is inherently lost in intensity-based observations, rendering the reconstruction of the bunch profile an ill-posed inverse problem. While traditional iterative algorithms, like Gerchberg-Saxton (GS), have been commonly applied to address this challenge, their computational rigidity restricts their adaptability to sophisticated experimental setups. To overcome these limitations, a gradient descent (GD)-based optimization framework can be implemented in conjunction with a differentiable physical forward model of the CTR generation process. However, natively GD exhibit slower convergence than GS-algorithms. To profit from the flexibility of the GD approach, we evaluate how integrating an untrained neural network into the reconstruction loop can improve convergence. This approach will lay the groundwork towards self-supervised pretraining of models without the need for labelled training data. ID: 308
/ Poster No. # 9: 067
Modalities: Text Methods: Generative Models Application Domain: Core Machine Learning Modeling Long Contexts with Hierarchical Compression Transformers Forschungszentrum Jülich, Germany
Transformers have become the prevailing architecture for large language models and beyond. However, their default attention mechanism scales quadratically with sequence length. Consequently, processing long contexts remains challenging for transformers despite its importance in many applications. While several approaches mitigate this issue through sparse or windowed attention, these methods often restrict the model’s ability to reason beyond local context.
To address this, we introduce Hierarchical Compression Transformers (HCT), a transformer architecture for efficient long-context modeling. HCT constructs a multi-level representation of its input using additional registers. Higher levels compress increasingly larger parts of the sequence. Lower levels use windowed attention, while higher levels provide access to global context. This design enables log-linear scaling in sequence length. We evaluate HCT on benchmark tasks to test capabilities such as knowledge memorization or long-range reasoning. Our findings suggest that hierarchical compression provides a promising alternative for long contexts. ID: 340
/ Poster No. # 9: 068
Modalities: Simulation Data Methods: Generative Models Application Domain: Earth & Environment Towards high resolution air quality forecasts for Europe using WGAN-based statistical downscaling with EURAD-IM data 1: Forschungszentrum Jülich, Germany; 2: Quarda Energy GmbH, Germany Up to date, air pollution is one of the largest environmental threats in Europe with about 400,000 premature deaths annually. While severe air pollution is generally linked to local sources, current monitoring and prediction models as the regional model ensemble of the European Copernicus Atmosphere Monitoring Service (CAMS) run on a rather coarse resolution to make operational air pollution forecasts across Europe computationally feasible. Thus, high resolution of air pollutant concentrations is inevitable for accurate public information and mitigation of severe health effects due to air pollution. Although compute resources have increased during the last years, the local extent of air pollution episodes does not justify computing high resolution concentration fields for the entire continent. Except for increasing the horizontal resolution of the air quality model EURAD-IM as part of the regional CAMS ensemble, a statistical downscaling method using WGAN is applied to increase the horizontal resolution of air pollutant concentration. The input data consists of operational forecasting data on two different model resolutions for the years 2012-2018 and 2019-2025: European-wide (Central European) air pollution forecasts are available on a 15 km (5 km) and 9 km (3 km) resolution, respectively. The model is pretrained using the data from 2012 to 2018 to account for the shift in model resolution. We focus our discussion on the importance of emission data as input depending on the area of interest, atmospheric constituent of concern and model resolution. We put special emphasis on three different air pollutants, namely O3, NO2, and PM2.5, which differ in their chemical origin and lifetime, leading to strong differences in their spatial heterogeneity. From the results, we discuss the application of a spatial masking of the input data to evaluate the ability of the model to downscale air pollutant concentrations in untrained areas. We show the importance of ensuring spatial representativity when subsample the spatial domain of the input data. The statistical downscaling model aims to increase the information content provided to the public and policy makers via the regional CAMS ensemble, while keeping the required energy demand at a minimum. ID: 267
/ Poster No. # 9: 069
Modalities: Tabular Data Methods: Other Application Domain: Health Data-Driven Identification of Possible Transmission Sites for Carbapenem-Resistant Gram-Negative Bacteria 1: Helmholtz AI, Helmholtz Munich, Germany; 2: Hessian State Office for Health and Care (HLfGP), Germany; 3: Helmholtz Centre for Infection Research, Germany Carbapenems are last-resort antibiotics, so it is crucial to prevent the spread of carbapenem-resistant Gram-negative bacteria (CRGNB) happening via direct and indirect transmission between patients – often in hospitals. In 2011, the state of Hesse, Germany, has established a notification system where detected cases of CRGNB are reported (10089 cases until Dec 2025). To optimally take advantage of the incoming data and alleviate the effort of manual analysis, we developed a computational tool that predicts outbreak events among the patients reported in a given time window, ranks them according to statistical significance and provides tailored visualizations for individual institutions. It uses a machine learning framework combining supervised and unsupervised techniques with patient-centric and hospital-centric approaches to systematically check for patient clusters with consistent resistance types and possible transmission events in the patient history of hospital stays. Leveraging patient history by custom algorithms is the key novelty of this work, becoming necessary because a colonization with resistant bacteria may remain undiscovered for a long period. Therefore, place or time of notifications do not hold information on the potential transmission source, and common epidemiological surveillance mechanisms for COVID-19 and flu infections are not applicable. In the human-curated annotation of data until Apr 2024, there exist clusters with different characteristics, such as clusters with single or multiple bacterial species, single or double resistances and specific or unspecific resistance types. The tool recovered almost all annotated clusters, even the rare types (omitting only the 6% without any data support, identified by external evidence). In addition to the recovered clusters, a similar quantity of new clusters was predicted. The top significant predictions are investigated via sequencing results. Furthermore, a blind test of the tool will be performed on the most recent data. As a highlight, this real-world data study shows the importance of reporting multiple prior hospital stays of patients. Taking solely the last hospital stay before notification, merely 10% of multi-center notification outbreaks were found. By an adoption and routinely application of automated data analysis and transparent presentation of evidence, we hope to increase awareness among stakeholders, leading to more complete data, precise predictions and timely countermeasures. ID: 336
/ Poster No. # 9: 070
Modalities: Multimodal Data Methods: Other Application Domain: Information Topological Data Analysis of Single-Cell Multimodal Data 1: Biomedical Center (BMC), Physiological Chemistry, Faculty of Medicine, LMU Munich, Munich, Planegg-Martinsried, Germany; 2: TUM, Germany; 3: Institute of Computational Biology, Computational Health Center, Helmholtz Zentrum München, German Research Center for Environmental Health, Neuherberg, Germany; 4: Hospital del Mar Research Institute (HMRIB), Barcelona, Spain Recent advances in single-cell sequencing enable simultaneous profiling of multiple molecular layers within a single cell. While this provides a more comprehensive view of cellular states and regulatory mechanisms, it also generates complex and high-dimensional data. This requires computational methods that can capture the information within the data from multiple perspectives. Persistent Homology, a tool of topological data analysis based on algebraic topology, potentially provides a powerful approach for characterising the structure of such data across multiple scales and modalities. Here, we leverage persistent homology for single-cell multimodal data, in particular scRNA-seq and scATAC-seq, to extract meaningful topological features from the transcriptome and the epigenome. In this framework, individual cells are represented as points in a high-dimensional feature space defined by measurements such as gene expression counts or chromatin accessibility. We hypothesise that the topology of this point cloud reflects the constraints imposed by gene regulation that define the cellular states. By analysing topological features such as connected components, loops, and higher-dimensional voids across different modalities, we aim to characterise cellular states. Lastly, training and explaining an ML classifier on the extracted features may enable inference of the underlying gene regulatory network from the topology of single-cell data. ID: 383
/ Poster No. # 9: 071
Are Soft Labels Worth the Effort? Helmholtz, Germany Learning under label uncertainty is common in real-world classification, yet most datasets rely on hard labels, approximating uncertainty by aggregating many annotations. An alternative is to elicit soft labels directly, which are more expensive per annotator but can reduce the total number of annotations needed. We study this trade-off on a novel dataset of galaxy images with directly elicited soft labels, comparing them to aggregated hard labels in terms of cost-efficiency and downstream model performance. Preliminary results suggest that models trained with soft labels can reach low-loss regimes more quickly, highlighting when explicit soft-label collection is economically and practically advantageous. ID: 237
/ Poster No. # 9: 072
Modalities: Tabular Data Methods: Agentic AI Application Domain: Information Building a Local AI Agent Assisting with Sensor Observation to Archives Framework (O2A) 1: Alfred-Wegener-Institute for Polar and Marine Research, Data Science, Bremerhaven, Bremen, Germany; 2: Alfred-Wegener-Institute for Polar and Marine Research, Software Engineering, Bremerhaven, Bremen, Germany Agents have become one of the most prominent and promising areas of AI research and development. AI agents are increasingly integrated into software applications to automate complex tasks. We have developed an AI agent to assist users with Observation to Archives Framework (O2A). O2A is a powerful framework that supports the entire scientific workflow, from data acquisition all the way to publication. It provides a suite of tools and infrastructure to manage instruments and metadata, handle structured data flows, track physical samples, enable real time monitoring and collaborative analysis, ensure long-term preservation as well as FAIR data publication and presentation. Our goal is to enhance accessibility by providing an intelligent assistant that guides users through the ecosystem and executes tasks on their behalf. Being able to chat with an integrated agent in real time should questions arise, as well as having the agent register or format data for the user, can significantly reduce time and shorten the learning curve for users who are new to O2A. Using Retrieval Augmented Generation (RAG), the agent can answer questions about O2A by referencing the official documentation directly. The Model Context Protocol (MCP) lets us connect the agent to O2A components, giving it the ability to carry out basic operations for users. To ensure sustainability and data integrity, we host the Large Language Model (LLM) that powers our agent locally. This eliminates dependence on third party platforms, providing full operational control, reducing costs, ensuring the agent is future proof and securing data privacy. Going forward our focus will be on making the agent available to the public and working with O2A users to evaluate and optimize its performance. ID: 142
/ Poster No. # 9: 073
Modalities: Text Methods: Agentic AI Application Domain: Earth & Environment SEA2LAND Navigator-GPT 1: Hereon, Germany; 2: HELCOM; 3: GERICS; 4: Tallinn University Coastal and marine planning is difficult due to complex regulations, large and fragmented datasets, and the involvement of multiple authorities across local, regional, and national levels. The SEA2LAND Navigator tool was developed to support this process by organizing governance responsibilities and providing access to a wide range of relevant resources. Given the breadth and depth of information it brings together, users may benefit from additional support when exploring the tool and applying it to specific planning questions. To address this, we introduce SEA2LAND Navigator-GPT, an AI chatbot integrated into the SEA2LAND Navigator website. The chatbot allows users to ask natural-language questions and receive clear, step-by-step guidance supported by user data and official documents. By simplifying access to information and guiding users through complex planning tasks, Navigator-GPT improves usability, supports transparent decision-making, and helps bridge the gap between data, policy, and practical coastal planning. ID: 359
/ Poster No. # 9: 074
Modalities: Multimodal Data Methods: Foundation Models, Generative Models Application Domain: Core Machine Learning, Health Messanger RNA Untranslated Region Design through Transformers 1: Helmholtz Munich, Institute of Computational Biology, Germany; 2: Faculty of Mathematics, Informatics and Mechanics, University of Warsaw Messenger RNA (mRNA) stability and translation efficiency are highly influenced by its untranslated regions, namely the 5′ (5′UTR) and 3′ (3′UTR) untranslated regions located upstream and downstream of the protein-coding region, respectively. As a result, the effectiveness of mRNA-based therapeutics and vaccines is strongly influenced by the 5′UTR sequence. However, there is currently no broadly validated tool for optimizing 5′UTR sequences to enhance therapeutic or vaccine efficacy. The aim of this project is to develop a generative AI–based model to generate novel 5′UTR and 3′UTR sequences with higher stability and translation efficiency than existing ones. The generated mRNA molecules will be synthesized and experimentally tested by collaborators to validate their performance. Our model could then become a resource for the development of new mRNA-based therapeutics and vaccines. ID: 210
/ Poster No. # 9: 075
Modalities: Image, Multimodal Data, Text Methods: Agentic AI, Foundation Models Application Domain: Core Machine Learning, Health Semantic Equilibrium Decoding for Open-Ended Medical Visual Question Answering 1: FAU Erlangen-Nürnberg, Germany; 2: Imperial College London, UK; 3: UCL, UK Small vision-language models (2–8B) are favoured for clinical deployment due to privacy and latency constraints, but their limited capacity amplifies hallucination risks. We extend existing game-theoretic decoding [1, 2] approaches, restricted to text-only, closed-ended tasks, to open-ended Medical VQA with multimodal VLMs. Building on the Bayesian Decoding Game (BDG) [2], we frame inference as a two-player signalling game: a generator produces a distribution over candidate answers conditioned on whether each answer is correct or incorrect, while a verifier scores each candidate's plausibility. Both operate over joint vision-language inputs. Crucially, we replace the lexical order-match stopping criterion with a criterion better suited for an open-ended setting: a Wasserstein-1 distance under a SapBERT-derived biomedical ground metric, weighted by the agents' current σ-separation. This allows convergence at semantic consensus, tolerating swaps between near-synonymous candidates (e.g., liver vs. hepatic region) while requiring resolution of semantically distant disagreements. On VQA-RAD and PathVQA, our Wasserstein BDG variant (BDG-W) yields consistent, statistically significant improvements across different model scales. Using BDG, Qwen3-VL-2B gains +3.5 pp accuracy (LLM-as-judge) over greedy decoding (p < 0.01), notably surpassing the 4B model without BDG. At the same time Qwen3-VL-4B employing BDG-W matches the performance of Qwen3-VL-8B. Additionally, the Wasserstein criterion reduces convergence iterations by ~20% versus classic BDG (p < 0.001). On PathVQA, a standard Gemma-3-4B with our BDG-W extension even outperforms its medically fine-tuned counterpart, MedGemma-4B (24.8 vs. 23.9). Our results indicate that strategic inference-time reasoning can compensate for limited capacity and missing fine-tuning, challenging the belief that reliable medical VQA requires larger or specialised models. [1] Jacob, Athul Paul, et al. "The consensus game: Language model generation via equilibrium search." arXiv preprint arXiv:2310.09139 (2023). [2] Zhang, Weitong, Chengqi Zang, and Bernhard Kainz. "From Self-Check to Consensus: Bayesian Strategic Decoding in Large Language Models." The Thirty-ninth Annual Conference on Neural Information Processing Systems.
ID: 310
/ Poster No. # 9: 076
Modalities: Tabular Data Methods: Other Application Domain: Health Single Cell PBMC Analysis Shows Cell Type Specific Changes On Inflammation In Both Depressed And Non-Depressed Individuals 1: Helmholtz AI, Helmholtz Zentrum München, Germany; 2: Max Planck Institute of Psychiatry, München, Germany The underlying biological mechanisms of depression are poorly understood. A possible mechanism is inflammation. In many studies, C-reactive protein (CRP) is measured as an inflammation marker. However, its effect on another important part of the immune system - peripheral blood mononuclear cells (PBMCs) gene expression - is not well characterized. Also, bulk RNA-seq PBMC studies of depressed patients yielded inconclusive results. Therefore, we analyzed the correlation of CRP with immune cell composition changes and cell type-specific gene expression in both depressed and non-depressed individuals. Integration of the data required conditional variational autoencoders and transfer learning. CRP was measured in serum and 10X single cell PBMC sequencing was performed in 13 patients with depression and other psychiatric disorders from the Max Planck Institute of Psychiatry clinic. The same data modalities were available already preprocessed from 164 non-depressed individuals from the ABF300 study, resulting in 38235 and 976252 cells, respectively. The scvi-tools package was used to correct natch effects in the data from patients with depression and to transfer the cell types from the non-depressed individuals. Cell type composition was analyzed with sccomp, pseudobulk differential gene expression with DESeq2 at 10% FDR. The most abundant cells were CD4+ T cells in both individuals with and without depression. The cell type composition did not change significantly with CRP concentration in either data set. The most differentially expressed genes (DEG) with regard to CRP were in myeloid cells and CD4+ T cells. In individuals with depression, the top DEG in CD4+ T cells was DUSP4 and in myeloid cells FCN1. Gene set enrichment showed enrichment of immune-related gene sets such as “inflammatory response”, “TNFA signaling via NKFB”, “interferon alpha response” or “interferon gamma response”. Of all 19 DEGs, only 5 overlapped with the 3900 DEGs from individuals without depression with different effect directions and the correlation of log fold changes for all genes was around 0. Our results confirm the impact of inflammation status on cell type-specific gene expression in both depressed and non-depressed individuals, including genes related to inflammation in PBMCs. Despite the successful data set integration with variational autoencoders, the higher power of the non-depressed individuals study underlines the need for larger single cell PMBC immuno-psychiatry studies. ID: 334
/ Poster No. # 9: 077
Modalities: Other Methods: Other Application Domain: Aeronautics, Space & Transport, Earth & Environment Advancing Sea Ice Surface Classification by Self-Supervised Contrastive Learning for Radar Altimetry 1: Section for Sea Ice Physics, Alfred Wegener Institute, Germany; 2: Earth Observation Center, German Aerospace Center, Germany; 3: Institute of Software Technology, German Aerospace Center, Germany; 4: Earth and Environmental Engineering, Columbia University, USA; 5: Center for Industrial Mathematics, Bremen University, Germany Accurate estimates of Arctic sea ice thickness from radar altimeter waveforms are essential for climate monitoring but do not yet meet the target accuracy requirements. A sea ice density parametrization resulting from a classification of surface types such as new ice (NI), first-year ice (FYI), multi-year ice (MYI), and open water from the same data could bring an essential improvement. Traditional approaches rely on a small set of hand-crafted waveform parameters, potentially discarding valuable information. At the same time, it is desired to retrieve a low-dimensional representation of the waveforms to facilitate the separation of surface types without relying on class labels. ID: 316
/ Poster No. # 9: 078
Modalities: Simulation Data Methods: Physics-informed Machine Learning Application Domain: Matter Analytic model-informed representations for sample-efficient learning of scattering patterns Forschungszentrum Jülich GmbH, Germany Small-Angle Neutron Scattering (SANS) enables structural characterization of nanoscale matter. Quantitative interpretation commonly relies on iterative fitting of analytic form-factor models, requiring expert model selection and substantial computational effort. Machine-learning approaches have recently been proposed to accelerate analysis by classifying scattering patterns using predictors trained on simulated data. Although inference is fast once trained, deployment in experimental workflows remains challenging because training datasets and candidate model spaces must be adapted to instrument configurations and prior knowledge. These constraints are particularly critical in small-dataset regimes, motivating learning strategies with strong inductive bias and high sample efficiency. Previous work showed that physics-informed transform encodings, such as Fourier representations aligned with isotropic scattering symmetry, substantially improve classification performance for lightweight models. Here we introduce a complementary representation design principle that incorporates prior knowledge directly from analytic scattering models. We construct parameter-eliminated model-consistency features derived from characteristic approximations of form-factor laws in distinct scattering-vector regimes. These features quantify whether measured intensity curves are locally compatible with structural signatures such as Guinier behaviour, power-law scaling, Lorentzian-type correlations, or oscillatory fringe patterns, without explicit parameter estimation or iterative fitting. The resulting analytic model-informed representation enables efficient learning with linear classifiers such as logistic regression and linear support-vector machines, as well as nonlinear kernel extensions. Benchmarks on simulated datasets indicate improved model-classification accuracy in small-data settings compared with previously studied transform-based physics-informed representations. Complementary to approaches that incorporate physical constraints during network training, the proposed strategy embeds analytic model knowledge directly in the feature space, yielding interpretable decision structure and stable estimation in limited-data regimes. More broadly, this work illustrates a pathway for leveraging closed-form scientific models in representation design for sample-efficient scientific machine learning. ID: 138
/ Poster No. # 9: 079
Modalities: Tabular Data Methods: Physics-informed Machine Learning Application Domain: Information Quantum Kernel Alignment for Clinical Data Classification 1: German Cancer Research Center, Germany; 2: Heidelberg University; 3: Fraunhofer IKS, LMU Munich The potential of quantum systems to perform complex computational calculations was recognized by Feynman in the 1980s and with discovery of the two famous algorithms by Shor and Grover one decade after Feynman's proposal the beginning of a paradigm shift towards the quantum era was marked. Out of this computational paradigm and against the backdrop of ML as one of the central pillars of modern research, quantum machine learning (QML) has emerged as an intersectional discipline between quantum computing and machine learning. Quantum kernel methods are among the most promising applications of quantum machine learning for pattern recognition tasks. The utilization of quantum inspired kernels with classical support vector machines (QSVMs) is a non-sparse matrix exponentiation technique for efficiently performing a matrix inversion of the training data inner-product (kernel) matrix. Quantum inspired feature maps define the model's ability to capture meaningful correlations and can lead to increased classification performance while relying less on large amounts of labelled data. We show that the choice of the feature map has a significant impact on the classification model and present a systematic approach to align quantum inspired feature maps by kernel alignment optimization to highly imbalanced and sparse clinical datasets leading to increased AUC scores of ≈0.22 on average compared to a classical radial basis function (RBF)-kernel SVM. We approached 6 clinical datasets of therapeutic and prognostic relevance, intractable for a classical RBF-kernel SVMs, revealing a novel pathway to analyze complex datasets. ID: 232
/ Poster No. # 9: 080
Modalities: Tabular Data, Text, Other Methods: Agentic AI, Physics-informed Machine Learning, Other Application Domain: Core Machine Learning, Earth & Environment Syxplain: An Agentic Framework for Explainable Scientific Hypothesis Generation from Data 1: HUN-REN, Hungary; 2: Obuda University, Hungary Symbolic regression (SR) and large language models (LLMs) each address one half of the scientific reasoning problem: SR extracts mathematical structure from data but cannot explain it, while LLMs reason fluently about scientific concepts but cannot ground that reasoning in empirical observation. Existing approaches that combine the two focus narrowly on using LLMs to steer the SR search. Syxplain extends well beyond this: it is a modular, end-to-end framework in which symbolic regression, machine-learning-based feature analysis, vision-language modelling, and a multilayer agentic interpretation pipeline are jointly orchestrated to help scientists understand their data, contextualize discovered expressions within existing knowledge, and arrive at scientifically grounded hypotheses. Before regression, a vision-language model inspects feature-pair scatter plots while interaction analysis and feature engineering jointly determine which operators, constraints, and hyperparameters the SR engine should use, replacing manual configuration with domain-grounded automation. Syxplain is SR-engine-agnostic, integrating as a modular layer around any framework that exposes explicit search-space settings. The central contribution lies in what happens after regression. A Pareto front of equations is rarely self-explanatory—individual formulas may be concise yet scientifically opaque, and no single expression captures the full picture a dataset encodes. Syxplain subjects the entire frontier to a multilayer interpretation pipeline: it resolves the semantic identity of every symbol; performs deep formula research against the scientific literature; retrieves and reasons over the containing papers; assesses each candidate's compatibility with established theory; and generates plain-language explanations—because a formula claimed to be interpretable still requires interpretation. Every inference is grounded in traceable citations. The result is a comprehensive report that reveals how the dataset behaves as a system, where multiple equations collectively disclose structure no individual expression can convey, and where hypotheses emerge from the interplay of data and domain knowledge. We demonstrate Syxplain in two scientific domains—grassland ecosystem modelling and soil fertilization—and release a full end-to-end software system, from authentication to interactive visualization, that makes the entire pipeline accessible to domain scientists without requiring programming expertise. ID: 357
/ Poster No. # 9: 081
Modalities: Text Methods: Agentic AI Application Domain: Matter Jülich Neutron Agent (JüNA): Vitess AI Agent Jülich Centre for Neutron Science (JCNS), FZ Jülich, Germany Over decades of neutron research, the Jülich Centre for Neutron Science (JCNS) has accumulated extensive knowledge in the form of scientific papers, manuals, wiki-style articles, in-house software, and electronic lab notebooks. Yet, efficient access to and use of these resources remains difficult. In parallel, advanced neutron scattering simulation tools such as VITESS [1] offer powerful capabilities, but their effective use typically requires substantial domain expertise. To reduce this barrier, we are developing the Jülich Neutron Agent (JüNA), an agentic AI framework that supports neutron scientists in everyday research tasks. JüNA combines large language models with reasoning-and-action strategies inspired by ReAct [2] and tool use via a Model Context Protocol (MCP) server. In this way, it enables AI-assisted knowledge retrieval, experimental planning, and code generation for neutron science. A central use case is chatbot-guided support for configuring and running complex simulations. Our first implementation focuses on simulation assistance through the VITESS AI Agent [3], which integrates with the open-source VITESS framework for neutron scattering simulations. The agent is designed to help users set up simulation parameters, construct workflows, and interact with the software more intuitively. Initial results suggest that JüNA can lower the entry barrier for neutron scientists while improving accessibility to JCNS knowledge and software resources. More broadly, it provides a basis for future AI-supported experimentation and autonomous laboratory workflows, establishing JüNA as a step toward next-generation neutron science research. References [1] VITESS. https://vitess.fz-juelich.de [2] Yao, Shunyu, et al. “ReAct: Synergizing Reasoning and Acting in Language Models.” ICLR, 2023. [3] Vitess AI Agent https://github.com/neutron-simlab/Vitess-AI-Agent ID: 134
/ Poster No. # 9: 082
Modalities: Simulation Data Methods: Physics-informed Machine Learning Application Domain: Matter NEOPIC: HPC Strategies for FNO-DSE in Kinetic Plasma Simulation 1: Forschungszentrum Jülich, Germany; 2: Helmholtz Zentrum Dresden-Rossendorf, Germany The NEOPIC project couples Particle-In-Cell (PIC) plasma simulations with Fourier Neural Operators (FNOs) to construct surrogate models for predicting particle fields, replacing conventional mesh- or tree-based field solvers. To mitigate memory pressure at large particle counts, we investigate distributed training and inference strategies that combine data and operator parallelism. This poster presents performance insights and scaling behavior of these parallel approaches, together with the impact of ID: 292
/ Poster No. # 9: 083
Modalities: Graphs, Image, Multimodal Data, Simulation Data, Text, Time Series Methods: Foundation Models Application Domain: Core Machine Learning, Energy Diagnosing and Fixing Training Bottlenecks: A Case Study of a Foundation Model 1: Jülich Supercomputing Centre, Forschungszentrum Jülich, 52428 Jülich, Germany; 2: ECMWF, 53175 Bonn, Germany WeatherGenerator is designed as a single machine-learning tool that can be adapted to many downstream weather-related tasks. Its goal is not to serve only one prediction problem, but to provide a general framework that can learn from a broad range of weather and Earth-system datasets and then be specialized for different applications. This broad scope makes WeatherGenerator scientifically valuable, but it also creates substantial computational challenges. A model that is trained on diverse datasets, deployed across several high-performance computing systems, and intended for repeated retraining and adaptation must be efficient not only in terms of model architecture, but also in terms of the full software and hardware pipeline. In practice, large-scale machine-learning workflows sometimes fail to reach their expected performance because the bottleneck is not the model itself, but the surrounding infrastructure. GPUs may remain idle while waiting for data. Preprocessing may be implemented in ways that are convenient for development but inefficient for execution at scale. Memory transfers between host and device may not exploit the capabilities of modern systems if data is not prepared in the right form. These issues become especially important in WeatherGenerator because the project operates on multiple HPC systems, including JUWELS Booster, Alps–Santis, Leonardo, and HPC2020, each with different architectural characteristics. As a result, performance portability and end-to-end efficiency are central requirements. This work demonstrates that optimizing the end-to-end training pipeline of WeatherGenerator improves performance across modern HPC systems. Through bottleneck analysis with NVIDIA Nsight Systems, vectorization of preprocessing operations, and manual memory pinning for composite batches, the workflow achieved better GPU utilization and improved transfer efficiency, particularly on GH200-based hardware. The resulting recurring savings of roughly 1.55 million core-hours per month show that software and pipeline optimization can be just as important as model innovation in large-scale scientific machine learning. ID: 141
/ Poster No. # 9: 084
Modalities: Tabular Data Methods: Foundation Models, Uncertainty Quantification Application Domain: Health Advancing Explainability as a Differentiator in Machine Learning prediction models for Alzheimer’s Disease 1: Modular HPC and AI, German Center for Neurodegenerative Diseases (DZNE), Bonn, Germany; 2: Systems Medicine, German Center for Neurodegenerative Diseases (DZNE) e.V., 53127 Bonn, Germany; 3: PRECISE Platform for Genomics and Epigenomics at DZNE and University of Bonn, 53127 Bonn, Germany Alzheimer’s disease (AD) is difficult to detect early because its progression is gradual, heterogeneous, and often clinically subtle. Identifying individuals at risk before substantial decline is therefore not only a predictive challenge, but a clinical one. In this study, we analyse clinical metadata from the DELCODE cohort to perform multiclass classification across four cognitive stages: cognitively normal, subjective cognitive decline (SCD), mild cognitive impairment (MCI), and AD. Our work is guided by two objectives: first, to handle missing clinical data in a statistically principled way by combining multiple imputation with appropriate pooling of model predictions, and second, to evaluate models not only by performance metrics, but by how well their attribution aligns with clinical plausibility across AD stages. We address missingness by generating m imputed datasets, train four classifiers (CatBoost, XGBoost, Random Forest, and TabPFN) on each, and pool predictions using Rubin’s Rules to ensure uncertainty is properly propagated into the final estimates. We further assess the stability of imputations through partial dependence analyses, observing that cognitive measures remain comparatively robust across imputations, whereas biomarker-related variables show greater variability. Beyond predictive performance, we centre our evaluation on explainability. Although models achieve comparable results on imbalance-aware metrics such as balanced accuracy and weighted F1, attribution analyses reveal meaningful differences in how they rely on clinical features. By examining these divergences, using global and regional explainability methods, we move beyond metric-based comparison to assess whether models align with clinically plausible disease mechanisms. Our findings not only directly provide valuable insights for AD classification to the DELCODE study but also move towards developing models that are disease stage-appropriate and clinically trustworthy. ID: 212
/ Poster No. # 9: 085
Modalities: Video Methods: Physics-informed Machine Learning, Other Application Domain: Aeronautics, Space & Transport Improving Pedestrian Detection through Temporal Movement Estimations German Aerospace Center (DLR), Germany Perception is an integral component of automated driving systems, primarily relying on deep neural networks (DNNs). We focus on the critical use case of pedestrian detection in an urban environment. However, the limited reliability of DNNs poses a challenge, often leading to false-positive detections that can cause disruptive over-braking of the automated vehicle. In this work, we aim to reduce this error without additional sensors by leveraging temporal information. To do so, we propose integrating state-of-the-art gradient-flow-based monitors derived from information theory into the AI system. These monitors assess whether the temporal movement detected in the input stream at the location of a DNN detection matches the motion pattern of a typical pedestrian. They rely on the fact that false-positive detections usually lack realistic and continuous motion. We demonstrate that this integrated approach significantly reduces the number of false positives and is fully real-time-capable, thereby improving the trustworthiness of the component and passenger comfort. ID: 338
/ Poster No. # 9: 086
Modalities: Image, Text Methods: Other Application Domain: Core Machine Learning Neural Networks as Varieties Forschungszentrum Jülich, Germany It has recently been empirically demonstrated that training large deep neural networks on large-scale data with purely polynomial activations is feasible. A key factor in achieving this was the variance-preserving initialization of these activations. This enabled consideration of purely polynomial representations of neural networks, i.e., polynomial mappings. Recent results also show that such networks are mathematically identifiable, meaning their parameters can be uniquely recovered under mild conditions. Algebraic geometry extensively studied such mappings, defining solutions to polynomial systems as intersections of algebraic sets represented by curves in dimension one, surfaces in dimension two, and algebraic varieties in higher dimensions. Furthermore, mathematical guarantees exist on the maximum number of connected components a variety can have. For plane algebraic curves, this is given by the Harnack curve theorem. For higher-dimensional varieties, similar upper bounds exist in terms of sums of topological invariants, such as Betti numbers (Thom-Milnor bounds). Generally, this upper bound grows proportionally to the polynomial degree raised to the dimension of the space. In this work, we analyze neural networks with polynomial, trigonometric, and tropical activations optimized for supervised classification. We empirically demonstrate that the connected regions formed by these varieties correspond to clusters of data points in each classification class. By considering polynomial mappings that are products of irreducibles, we demonstrate how to construct a polynomial mapping that matches a union of connected components. In this sense, the considered polynomial mappings can be viewed as a multivariate counterpart to the fundamental theorem of algebra (D’Alembert-Gauss decomposition), where the polynomial is expressed as a product of linear and irreducible quadratic functions. Although a general multivariate factorization theorem is lacking, this approach allows us to mimic the 1D decomposition and analyze connected components similarly. This perspective further allows us to establish a correspondence between a neural network viewed as a composition of polynomial mappings and an alternative representation as a product of irreducible polynomials. Ultimately, this framework could inspire future work in interpreting a connectome of neurons as an algebraic variety projected in three-dimensional space, where synapses correspond to intersection points. ID: 176
/ Poster No. # 9: 087
Modalities: Multimodal Data, Simulation Data, Time Series Methods: Physics-informed Machine Learning Application Domain: Energy, Earth & Environment Surrogate Models for Severe Nuclear Accidents 1: Helmholtz AI, Germany; 2: Karlsruhe Institute of Technology, Germany Severe nuclear accidents, though rare, present catastrophic risks that demand sophisticated simulation capabilities for emergency preparedness, operator training, and safety assessment. Current state-of-the-art severe accident codes such as ASTEC and MELCOR provide high-fidelity physics-based simulations but suffer from prohibitive computational costs that limit their application in real-time scenarios, uncertainty quantification, and probabilistic safety assessments. We present a comprehensive framework for developing machine learning-based surrogate models, specifically leveraging physics-informed transformers to accelerate severe accident simulations while maintaining physical fidelity. The effectiveness of this approach is demonstrated on a multi terabyte-spanning data-driven surrogate for a nuclear reactor's vessel core, addressing the unique challenges of severe accident modeling, including high dimensionality, strong nonlinearities, boundary effects, and multi-physics numerical code coupling. We also demonstrate its computational effectiveness, thus proofing to be a path towards real-time severe accident simulation. ID: 270
/ Poster No. # 9: 088
Modalities: Image, Multimodal Data, Text, Video Methods: Foundation Models, Generative Models Application Domain: Core Machine Learning Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs 1: Helmholtz Zentrum München, Germany; 2: Technical University of Munich, Germany; 3: Munich Center for Machine Learning, Germany; 4: Télécom Paris, France; 5: Google, Switzerland Multimodal Large Language Models (MLLMs) often struggle with fine-grained perception, such as identifying small objects in high-resolution images or detecting key moments in long videos. Existing methods typically rely on complex, task-specific fine-tuning, which reduces generalizability and increases system complexity. In this work, we propose an effective, training-free framework that uses an MLLM's intrinsic uncertainty as proactive guidance. Our core insight is that a model's uncertainty decreases when provided with relevant visual information. We introduce a unified mechanism that scores candidate visual inputs by response uncertainty, enabling the model to autonomously focus on the most informative data. We apply this simple principle to three challenging visual tasks: Visual Search, Long Video Understanding, and Temporal Grounding, allowing off-the-shelf MLLMs to achieve performance competitive with specialized, fine-tuned systems. Our results demonstrate that leveraging intrinsic uncertainty is a powerful strategy for improving fine-grained multimodal performance. ID: 192
/ Poster No. # 9: 089
Modalities: Other Methods: Probabilistic Methods, Other Application Domain: Earth & Environment Helixer – a tool for ab initio gene calling combining Deep Learning and a Hidden Markov Model 1: Forschungszentrum Jülich, Germany; 2: Heinrich Heine Univsersität Düsseldorf, Germany; 3: Cluster of Excellence on Plant Sciences (CEPLAS), Germany; 4: Recursion, Valence Labs, Canada Plant breeding of crops plays an increasingly important role in addressing challenges such as climate change and food security. To modify crops in a targeted approach, the genes responsible for important traits such as drought resistance need to be identified first. Historically, eukaryotic genome annotation relied primarily on (generalized) Hidden Markov Models (HMM)s, however in recent years Deep Learning (DL) approaches ranging from smaller supervised models to large self-supervised foundation models became more prevalent. Currently, smaller supervised models with focusing on gene identification are more practical to use compared to the often resource intensive DNA foundation models fine-tuned on a broad variety of tasks. Our tool, Helixer1,2, was one of the first to apply DL to gene annotation. It combines four convolutional layers, each followed by batch normalization, with three bidirectional Long Short-Term Memory layers. A linear layer produces the final class probabilities that are post-processed by an HMM to generate concise gene structures. The inputs for the model are numerical encodings for the genomic sequence (A, C, T, G). The labels are eight classes, the four genic classes “coding sequence”, intron”, “untranslated region”, and “intergenic”, as well as the four coding phases (“no phase” and “phase 0-2”) that are used to determine what part of the gene is translated into a protein. The model was trained on 21 kilo base pairs context windows over the entire genomic sequence for multiple species, predicting all eight classes simultaneously. Four models for fungal, plant, vertebrate and invertebrate genome annotation are available. Helixer outperforms two commonly used HMM gene predictors, GeneMark3,4 and AUGUSTUS5,6, in gene-level evaluation for plants, vertebrates and invertebrates, but both HMM tools gain an edge in fungi. Across 45 test species, Helixer yields a mean precision of 0.3270 and recall of 0.4017 compared to GeneMark’s 0.1534 and 0.2198. In a subset of 16 species for which AUGUSTUS models were available, Helixer maintains comparable precision to it with 0.3825 vs. 0.3847 while achieving a superior recall of 0.4613 vs. 0.3549. For prediction Helixer takes the genome in standard FASTA format and outputs gene annotations in the widely used GFF3 format. Helixer’s design enables immediate use on genomes without retraining, providing an efficient, accessible solution for genome annotation in both research and applied settings. ID: 287
/ Poster No. # 9: 090
Modalities: Simulation Data, Tabular Data Methods: Probabilistic Methods, Uncertainty Quantification Application Domain: Core Machine Learning metabeta - A fast neural model for Bayesian mixed-effects regression Human-Centered AI, Helmholtz Munich, Germany Hierarchical data with multiple observations per group is ubiquitous in empirical sciences and is often analyzed using mixed-effects regression. In such models, Bayesian inference gives an estimate of uncertainty but is analytically intractable and requires costly approximation using Markov Chain Monte Carlo (MCMC) methods. Neural posterior estimation shifts the bulk of computation from inference time to pre-training time, amortizing over simulated datasets with known ground truth targets. We propose metabeta, a transformer-based neural network model for Bayesian mixed-effects regression. Using simulated and real data, we show that it reaches stable and comparable performance to MCMC-based parameter estimation at a fraction of the usually required time, enabling new use cases for Bayesian mixed-effects modeling. ID: 236
/ Poster No. # 9: 091
Modalities: Other Methods: Probabilistic Methods Application Domain: Health Meta-Learning for Phenotypic Drug Discovery with Transformer Neural Processes 1: Institute of AI for Health, Helmholtz Munich; 2: School of Computation, Information and Technology, Technical University of Munich; 3: Helmholtz AI; 4: Munich Center for Machine Learning; 5: Faculty of Mathematics, Informatics and Mechanics, University of Warsaw; 6: Department of Computer Science and Artificial Intelligence, University of Technology Nuremberg Phenotypic drug discovery captures much desired system-level biological effects, but phenotypic bioactivity data remains difficult to exploit with modern deep learning methods due to assay heterogeneity, noisy measurements, and censored activity labels. Meta-learning approaches have been proposed to address this by modeling each assay as a separate task, enabling information sharing across assays. However, these approaches rely on tasks being sufficiently related, an assumption that might be violated in highly diverse phenotypic assays, causing naive meta-learning strategies to perform poorly in practice. Here, building on Transformer Neural Processes, we investigate a meta-learning approach enabling richer conditioning on task-specific context and the incorporation of auxiliary signals from related assays, allowing the model to selectively extract useful information from heterogeneous sources. This represents a step toward machine learning models better suited for complex phenotypic data. ID: 266
/ Poster No. # 9: 092
Modalities: Graphs, Multimodal Data, Text Methods: Agentic AI, Generative Models Application Domain: Energy Agentic Data Extraction and Dynamic Knowledge Graph Construction for Hydrogen Technology Research 1: Theory and Computation of Energy Materials (IET-3), Institute of Energy Technologies, Forschungszentrum Jülich GmbH, 52425 Jülich, Germany; 2: Centre for Advanced Simulation and Analytics (CASA), Simulation and Data Science Lab for Energy Materials (SDL-EM), Forschungszentrum Jülich GmbH, 52425 Jülich, Germany; 3: Chair for Theory and Computation of Energy Materials, Faculty of Georesources and Materials Engineering, RWTH Aachen University, Germany Research data management in the hydrogen technology field is obstructed by the the heterogeneity of data originating from multiple scales and methods. In electrochemical systems such as PEM fuel cells or electrolysers, data on fabrication, characterization, and performance are often siloed, reported with inconsistent terminology, and lacking standardized metadata. These shortcomings limit cross-study comparability and AI-readiness. To address them, we present an AI-enhanced, ontology-driven framework for constructing searchable, provenance-aware knowledge graphs that follow FAIR principles. The methodology comprises a three-phase pipeline applied to the use case of the effect of catalyst layer composition on cell performance. Phase I initiates the workflow by formulating a specific research question and developing an initial domain ontology supported by large language models. This process uses curated templates to constrain concepts and relations, defining the classes, predicates, metadata, and value requirements needed for the knowledge model. In Phase II, we assemble a literature PDF corpus and develop an LLM-based tool to extract entities, values, units, and contextual metadata guided by the domain ontology, followed by terminology alignment and unit harmonization while preserving traceable links to the source documents. In Phase III, we map the curated dataset into a Neo4j graph database to enable graph queries. The resulting knowledge graph integrates data from 148 papers, encompassing 12,425 nodes and 56,623 edges, based on an extended ontology with 219 measurements, 583 parameters, and 545 properties. Exploratory data analysis of this graph demonstrates the ability to identify critical trends. By transforming unstructured literature into standardized, machine-readable data, this framework provides a strong foundation for downstream analytics and advanced machine learning in hydrogen technology.
ID: 385
/ Poster No. # 9: 093
Diagnosing Shortcuts in Traffic Video Question Answering Datasets 1: Technical University of Munich, Germany; 2: Helmholtz Zentrum München, Germany; 3: Munich Center for Machine Learning, Germany Vision–language models (VLMs) are increasingly evaluated on multiple-choice video question-answering (VideoQA) benchmarks that require reasoning about real-world events. Traffic accident scenarios are challenging; they require models to identify agents, interpret complex interactions, and reason about possible prevention. Initially, we aimed to study whether specialized training strategies could improve the identification of accidents. During early experiments, we observe a pattern: model accuracy remains high when the video input is removed. In some instances, performance is decreased when visual information was provided. This suggests that models exploit textual cues within the multiple-choice options rather than performing true visual grounding. The focus of the study shifts to dataset-level analysis. The goal is to determine whether traffic video question answering datasets require visual grounding or whether they allow language shortcuts that enable correct answers without taking visual signals. To study this, we analyze multiple datasets, including VRU-Accident, LOTVS-MM-AU, AccidentBench, and SUTD-TrafficQA. The analysis introduces two dataset-level measures: Blind Gap, which compares text-only accuracy with random chance, and Visual Gain, which measures the change in accuracy when video input is added. Our findings show that while some datasets are robust, others demonstrate high text-only accuracy and little or negative visual gain. These indicate that certain benchmarks may allow language shortcuts that weaken their ability to evaluate visual reasoning. This work aims for a systematic analysis of shortcut structures in traffic video datasets and to identify subsets of questions that require visual evidence for correct answers. Furthermore, we examine how training objectives influence reliance on visual evidence. We train Qwen2.5-VL on TrafficQA using supervised fine-tuning (SFT), direct preference optimization (DPO), and group relative policy optimization (GRPO), and evaluate their effects on blind accuracy and visual gain across multiple datasets. This analysis aims to understand both dataset properties and training dynamics that affect visual grounding in traffic video question answering. ID: 172
/ Poster No. # 9: 094
Modalities: Image, Other Methods: Other Application Domain: Core Machine Learning, Health DREAMS: Preserving both Local and Global Structure in Dimensionality Reduction 1: Hertie Institute for AI in Brain Health, Germany; 2: University of Tübingen, Germany Dimensionality reduction techniques are widely used for visualizing high-dimensional data in two dimensions. Existing methods are typically designed to preserve either local (e.g., t-SNE, UMAP) or global (e.g., MDS, PCA) structure of the data, but none of the established methods can represent both aspects well. We present DREAMS (Dimensionality Reduction Enhanced Across Multiple Scales), a method that combines the local structure preservation of t-SNE with the global structure preservation of PCA via a simple regularization term. Our approach generates a spectrum of embeddings between the locally well-structured t-SNE embedding and the globally well-structured PCA embedding, efficiently balancing both local and global structure preservation. We benchmark DREAMS across eleven real-world datasets, showcasing qualitatively and quantitatively its superior ability to preserve structure across multiple scales compared to previous approaches.
ID: 227
/ Poster No. # 9: 095
Modalities: Multimodal Data, Simulation Data, Tabular Data, Other Methods: Probabilistic Methods Application Domain: Core Machine Learning, Health Factor dependencies and their associations with covariates in multimodal factor analysis 1: University of Warsaw, Poland; 2: Helmholtz Center Munich, Germany Factor analysis methods for multimodal data integration are widely used in medical applications. These methods aim to identify common latent factors underlying an observation for which multiple measurements are available, in an unsupervised manner - without prior knowledge of what these factors might correspond to, such as clinical features such as disease subtypes. Inference typically seeks to disentangle factors by enforcing orthogonality or sparsity constraints, or proceeds without imposing any specific structural assumptions. In (semi-)supervised settings, factor analysis models additionally focus on capturing associations between factors and covariates. In this work, we focus on multimodal factor analysis and explore different ways of modelling factor dependencies - both internal relationships between factors and their relationships with additional covariates. We present a unified framework that describes these approaches from a new perspective, providing both mathematical insights and illustrative examples. ID: 352
/ Poster No. # 9: 097
Modalities: Image, Multimodal Data Methods: Generative Models, Other Application Domain: Health Discovering Interpretable Visual Concepts in Monkey Visual Cortex Using Sparse Autoencoders 1: Institute of Computational Biology, Computational Health Center, Helmholtz Munich, Germany; 2: Department of Experimental Psychology, University of Oxford, UK Understanding how the visual system processes natural stimuli and which visual features are encoded where in the brain, has been a long-standing goal in neuroscience. Advances in neural recording techniques have enabled the measurement of population responses to natural visual stimuli across multiple brain regions and species, generating datasets suitable for addressing this question. Yet, existing analyses of these datasets often rely on a priori human-defined semantic labels of stimuli, limiting discovery to predefined concepts. Sparse autoencoders (SAEs), recently developed for interpreting large-scale foundation models, offer a data-driven alternative. Here, we build on these advances and propose NeuroSAE for learning sparse latent features of neural population activity. We apply NeuroSAE to electrophysiological recordings from the macaque monkey visual ventral stream during presentation of natural images (Papale et al., 2025) and decompose neural activity into linear combinations of distinct concepts. We automatically quantify the alignment between learned concepts and image properties spanning a broad range of complexity, from low-level visual features such as color, edge orientation, and spatial frequency to high-level semantic content. Alignment is computed using a CLIP embedding-based approach that enables testing of arbitrary hypotheses about what each concept represents. Because each concept can be linearly attributed to specific neural populations, this framework enables an unbiased analysis of what and where visual information is encoded in the brain. At the image level, NeuroSAE enables identification of the concepts encoded for each image, which often deviate from the semantic image label. At the level of brain regions, NeuroSAE allows for the comparison of encoded concepts across areas, revealing systematic differences: Qualitatively, V1 concepts represent low-level visual features such as color and texture, V4 concepts showed more complex, colorful patterns, and IT captured semantically-aligned concepts such as animals and faces. ID: 196
/ Poster No. # 9: 098
Modalities: Image, Time Series Methods: Foundation Models, Other Application Domain: Earth & Environment Incremental learning for continuous forest disturbance mapping using AlphaEarth embeddings GFZ Helmholtz Centre for Geosciences, Germany Forest disturbance mapping involves the continuous monitoring of the Earth’s surface to detect spatio-temporal anomalies. These disturbances can be caused by logging, wildfires, storms, or pest outbreaks. Detecting these changes in advanced is essential for understanding ecosystem dynamics, supporting forest management, and informing climate mitigation policies. As monitoring efforts had scale from (ground-based) regional to continental analyses, data-driven modeling frameworks has been employed by leveraging large volumes of remote-based observations. These observations commonly include multi-spectral optical and radar satellite imagery, as well as weather variables. ID: 194
/ Poster No. # 9: 099
Modalities: Multimodal Data, Tabular Data, Time Series Methods: Other Application Domain: Health Multimodal Machine Learning for Postoperative Atrial Fibrillation Prediction 1: Karlsruhe Institute of Technology, Germany; 2: Medical Faculty Mannheim, Germany; 3: University Hospital Heidelberg, Germany Postoperative atrial fibrillation (POAF) affects up to 15% of non-cardiac surgery patients and substantially increases morbidity. Existing feature-based models achieve only moderate discrimination, relying primarily on preoperative clinical variables. The HeiPoDD dataset offers uniquely deep multimodal phenotyping, including continuous postoperative ECG recordings. However, utilizing these postoperative signals requires defining a strict temporal boundary to ensure the model provides a clinically actionable early warning, rather than merely detecting the arrhythmia as it begins. The HeiPoDD dataset (1600 non-cardiac surgery patients) will be used to develop a multimodal feature-based machine learning model for POAF prediction. Clinical and procedural variables will be integrated with electrophysiological features extracted from continuous preoperative and early postoperative ECG recordings using the open-source ECGdeli software. A gradient boosting ensemble classifier will be trained on this feature set. Simultaneously, a systematic temporal ablation study will explore the postoperative observation window for the model’s prediction horizon. This step is critical to prevent information leakage and ensuring the model predicts AF rather than just detecting its onset, while maximizing the early warning capability required for clinical actionability. SHAP analysis will quantify the relative contribution of postoperative ECG features versus standard clinical variables. The temporal ablation study will define the evidence-based postoperative observation window for use in subsequent synthetic-data deep learning development. This work will simultaneously establish two essential foundations for deep learning-based POAF prediction: a feature-based performance baseline and an empirically defined, leak-free postoperative observation window to guide future model training and evaluation. ID: 256
/ Poster No. # 9: 100
Modalities: Graphs, Multimodal Data, Tabular Data, Text, Time Series Methods: Agentic AI, Foundation Models, Generative Models Application Domain: Earth & Environment FrevaGPT: An LLM-Powered Assistant for Reproducible Climate Analysis 1: GFZ Helmholtz Centre for Geosciences, Germany; 2: German Climate Computing Center DKRZ, Germany; 3: Universität Hamburg, Germany; 4: Climate Service Center Germany (GERICS), Germany Large language models (LLMs) have the potential to transform how climate scientists interact with data by lowering technical barriers and enabling more intuitive analysis workflows. Here, we present FrevaGPT, an LLM-powered scientific assistant integrated into the Freva climate data search and analysis platform. FrevaGPT interprets natural language queries and automatically generates traceable, editable, and reusable analysis scripts that can be executed within established scientific environments. It retrieves relevant datasets and literature, performs analyses, and visualises results, allowing researchers to focus on scientific interpretation rather than coding intricacies. By leveraging a large repository of climate observations and model output, FrevaGPT ensures transparent and reproducible workflows that adhere to best practices in climate research. It also integrates seamlessly into Jupyter-AI and, by making use of the Freva library, combines the code-generating capabilities of LLMs with contextual understanding of how to access relevant datasets on the HPC cluster. As a “co-pilot” for geoscientists, FrevaGPT not only responds to user queries but also proactively suggests relevant climate modes, events, and next analytical steps, helping to uncover insights that might otherwise be overlooked. Practical use cases demonstrate how the system supports interactive exploratory analysis and hypothesis refinement across complex climate datasets. By embedding LLM-assisted natural language interaction into real-world climate research workflows, this work highlights methodological considerations and opportunities for enhancing scientific productivity, promoting broader adoption of NLP and AI tools among Earth system scientists. We provide scientific evaluation of FrevaGPT’s capability through a benchmark suite. A live demonstration will be presented and can be used by the audience to perform climate analysis on an HPC with access to petabytes of Earth system data using simple prompts. ID: 272
/ Poster No. # 9: 101
Modalities: Image Methods: Foundation Models Application Domain: Core Machine Learning, Health Mind The Gap: Continuous Magnification Sampling For Pathology Foundation Models 1: Technical University Berlin, Germany; 2: Berlin Institute for the Foundations of Learning and Data; 3: Aignostics Gmbh; 4: Charité, Universitätsmedizin Berlin Introduction Pathology foundation models are integrated in diagnostic tools that operate across the continuous magnification spectrum. Yet how magnification sampling during pretraining affects representation quality remains poorly understood. Current models are trained on patches from a discrete set of resolutions (0.25, 0.5, 1.0, 2.0 microns per pixel). We show that this strategy leads to performance degradation at magnifications outside the training set and propose “Continuous magnification sampling” as a principled alternative. We develop a theoretical framework to derive optimal sampling strategies and validate our findings through unsupervised embedding analysis and new multi-scale benchmarks (TCGA-MS, BRACS-MS). Methods We model patch sampling from the magnification spectrum as a multi-source domain adaptation problem, where each magnification constitutes a domain and a similarity kernel captures how training signal transfers across scales. Based on this framework, we derive and compare multiple sampling strategies: discrete uniform (the current standard), continuous uniform, and two optimized continuous distributions that maximize average or worst-case representation quality. Continuous sampling is implemented via crop-and-resize operations, requiring no architectural changes or computational overhead. Results In controlled experiments, discrete uniform sampling produces dips in representation quality at intermediate magnifications. Continuous sampling eliminates these and improves balanced accuracy by up to 4 pp at intermediate scales. Optimized distributions yield additional gains of 1.8 pp. We observe similar patterns in state-of-the-art models (Virchow2, UNI, Prov-GigaPath) where magnification is a driver of performance variation: single-scale models degrade at distant magnifications, with drops of up to 12 pp relative to multi-scale models, and multi-scale models exhibit performance drops at intermediate scales. Conclusion Our results demonstrate that controlling the resolutions at which patches are sampled is a consequential yet overlooked pretraining design decision. We show that discrete sampling strategies leave systematic blind spots, and that state-of-the-art model rankings change depending on the evaluation scale. Our work addresses these problems, offering “Continuous magnification sampling” as a drop-in replacement for discrete protocols, and paves the way towards pathology foundation models that perform reliably across magnifications. ID: 153
/ Poster No. # 9: 102
Modalities: Text Methods: Agentic AI Application Domain: Information Grounding Large Language Models in Formal Solvers via Unified Planning: Towards High-Accuracy Natural Language Planning Hybrid Methods in Artificial Intelligence and Machine Learning, University of Rostock, Germany Recent studies have showed that Large Language Models (LLMs) and Large Reasoning Models (LRMs) do not possess genuine reasoning and planning capabilities. For example, GPT-o3, Deepseek R1, and Claude LRMs can only output accurate optimal solutions for Blocks World (BW) and Tower of Hanoi problems involving up to 10 blocks/disks. This phenomenon extends to other planning benchmarks, where plan accuracy sharply declines as problem complexity increases. Nevertheless, these models can serve as accessible natural-language interfaces in hybrid systems that use formal reasoning solvers in the background. A common approach involves instructing LLMs to translate user descriptions into PPDL problems. Although this strategy is more robust, recently proposed pipelines still strugle with syntatic and semantic errors increasing with problem complexity. We propose a pipeline where the LLM models planning problems using Unified Planning (UP) objects rather than writing direct PDDL. UP is a uniform framework to model planning problems. By exposing the UP interface through Model Context Protocol (MCP), the LLM can call specific functions to incrementally build the problem from a natural language description. This approach effectively mitigates the syntactic errors common in PDDL generation. The LLM has access to other tools to inspect the domain, undo previous operations, and call formal solvers. We evaluated this pipeline in the BW and Logistics benchmark domains. Because the LLM calls formal solvers, we measured accuracy by comparing the final problem states with ground-truth PDDL files. Using 'qwen3' (30B) and 'gpt-oss' (120B), the pipeline outputs up to 99.7% exact matches with ground-truth PDDL for BW descriptions with up to 8 blocks and 98.6% exact matches with descriptions with up to 15 blocks. Specifically, we investigated the trade-off between providing LLMs with fine-grained versus coarse-grained UP interfaces. A fine-grained setup allows the model to declare each aspect of the problem in an arbitrary order and at varying depth levels, though it requires more tool calls. We found that smaller problems are best described all at once, whereas larger ones benefit significantly from this adjustable granularity. Once proven that this pipeline is domain-independent and scales well with problem complexity, future work looks toward evaluating it in real-life planning scenarios, such as trip planning, automated data processing, and scientific experiment planning. ID: 171
/ Poster No. # 9: 103
Modalities: Image, Multimodal Data, Video Methods: Other Application Domain: Core Machine Learning MLArray: Unifying Machine Learning-Optimized Array Storage and Imaging Metadata German Cancer Research Center (DKFZ), Germany Domain science image formats such as DICOM, NIfTI, and NRRD provide rich spatial metadata and a mature visualization ecosystem, but they are poorly suited for modern machine learning training. In particular, they typically lack efficient partial reading and writing as well as chunked compression, making patch-based sampling slow and memory-intensive. In contrast, array formats such as NumPy, Zarr, or Blosc2 are highly efficient for training workloads due to chunked storage, compression, and memory mapping, yet they provide no standardized metadata layer. This severely limits interoperability with analysis and visualization tools. MLArray fills this gap by combining machine learning (ML) optimized array storage with a standardized, extensible metadata schema in form of a Python library.. Built on top of Blosc2 N-dimensional arrays, MLArray leverages its efficient storage and access model to enable efficient random patch reads without loading full volumes into memory. A central feature is patch size-driven layout optimization: users specify the expected training patch size, and MLArray derives chunk and block sizes to align the on-disk layout with training time access patterns, improving cache efficiency and I/O predictability. In addition, MLArray defines a structured yet flexible schema that standardizes common imaging concepts such as spacing, origin, direction, axis semantics, statistics, bounding boxes, and segmentation flags, while preserving original source metadata (e.g., from DICOM or NIfTI). MLArray can therefore act as an ML-optimized alternative storage layer without breaking downstream analysis workflows. The library integrates naturally into Python workflows: arrays can be created from NumPy, converted to MLArray, and later accessed via memory mapping for patch-based training or interactive visualization. A lightweight command line interface supports header inspection and conversion from existing imaging formats, facilitating adoption in established pipelines. MLArray targets large N-dimensional scientific images such as medical imaging, remote sensing, or segmentation masks, where both metadata fidelity and training time I/O performance are critical. By unifying efficient array storage with standardized imaging metadata, MLArray provides a foundation for interoperable, ML-ready scientific image data and tooling. The code and how to use MLArray can be found here: https://github.com/MIC-DKFZ/mlarray ID: 379
/ Poster No. # 9: 104
Post-training makes large language models less human-like Helmholtz Munich, Germany Large language models (LLMs) such as ChatGPT, Claude, and Gemini have rapidly transformed the landscape of artificial intelligence, serving as powerful tools for writing, coding, and reasoning. One of their most far-reaching promises lies in their ability to simulate human-like behavior. However, the extent to which LLMs actually resemble human behavior remains disputed. We address this gap by introducing Psych-201, a large-scale dataset of natural language transcripts from behavioral experiments. Psych-201 was collected through an open research collaboration and crowdsourced effort, resulting in a dataset with a much more diverse participant population and a broader range of experimental paradigms. We evaluate a large set of LLMs on their ability to predict human responses in Psych-201, and find that: 1. Post-training consistently reduces human-likeness. This effect holds across model families and applies to all post-training objectives, including instruction-tuning, reasoning, code, and vision. 2. Base models continue to improve across generations (i.e., newer models are generally more aligned). 3. The post-training misalignment gap, however, widens in newer models. 4. The largest post-training misalignment gaps occur in moral reasoning. Taken together, these findings have important implications for using LLMs as behavioral surrogates and suggest new approaches to post-training that maintain the usability of aligned models while preserving the behavioral alignment found in base models. ID: 180
/ Poster No. # 9: 106
Modalities: Image, Multimodal Data, Text Methods: Foundation Models, Generative Models Application Domain: Aeronautics, Space & Transport, Earth & Environment A Domain-Specialized Multimodal Foundation Model for Interpreting Sentinel-2 True Color Imagery in Flood Disaster Response DLR, Germany Rapid mapping and disaster response operations rely heavily on the interpretation of complex satellite imagery. While standard Vision-Language Models (VLMs) have shown remarkable capabilities in natural image understanding, their performance remains limited in Earth Observation (EO) due to differences in spatial scales, top-down perspectives, and textural patterns. This work addresses the interpretability gap by developing a specialized VLM that generates professional, context-aware textual descriptions of flood events from Sentinel-2 (RGB) imagery. ID: 363
/ Poster No. # 9: 107
Modalities: Image Methods: Foundation Models Application Domain: Health A Hematology Foundation Model Enables Blood Cancer Screening from Peripheral Blood Smears 1: Helmholtz Munich, Germany; 2: Munich Leukemia Laboratory Background: Peripheral blood smear review is a fast and minimally invasive component of hematological diagnostics, but manual assessment is labor-intensive and subject to inter-observer variability. Methods: We developed cAItomorph, a transformer-based model for predicting hematological malignancies from peripheral blood smears. The study included digitized white blood cell images from 2043 individuals (1,548 patients and 495 healthy stem cell donors) collected at the Munich Leukemia Laboratory between 2021 and 2022. Diagnostic labels were grouped into 8 coarse disease classes. cAItomorph uses the DinoBloom hematology foundation model to encode single-cell images, a transformer aggregator to integrate information from 500 cells per patient, and a classifier to predict the 8 coarse classes together with hemoglobin values. Results: cAItomorph achieved an overall 8-class accuracy of 0.72 on the held-out internal test set. Class-wise performance was highest for acute leukemia (sensitivity/specificity 0.74/0.79) and myeloproliferative neoplasms (0.85/0.76), with moderate performance for myelodysplastic syndromes (0.71/0.57). For binary malignant versus non-malignant prediction, the model reached AUROC 0.97. At a malignancy threshold of 0.5, it reduced potentially unnecessary bone marrow aspirations from 13.5% to 8.7% in the test set while correctly identifying all acute leukemia cases. Hemoglobin prediction correlated with measured values (Pearson r = 0.67, p = 6 × 10^-55; MAE 1.63). External validation showed robust generalization despite staining differences. Attention-based visualization highlighted diagnostically relevant cells such as myeloblasts, promyelocytes, and giant platelets. We built a web-based tool to visualize model outputs and help clinical decision making. Conclusion: cAItomorph demonstrates that AI-based analysis of peripheral blood smears can support initial diagnostic triage across a broad range of hematological malignancies in a real-world laboratory cohort. The model shows strong performance for acute leukemias and myeloproliferative neoplasms, generalizes to external datasets, and may help reduce unnecessary invasive procedures while preserving sensitivity for malignant disease. Prospective multicenter validation is needed to establish real-world clinical utility. Visualization tool, model weights, and the test set (409 individuals; 201,560 blood cell images) are available at github.com/marrlab/cAItomorph. ID: 313
/ Poster No. # 9: 108
Modalities: Simulation Data Methods: Generative Models, Graph Neural Networks, Physics-informed Machine Learning, Other Application Domain: Energy, Matter Surrogate Modeling of Scale-Bridging Dislocation-Driven Plasticity through Latent Dynamics and Microstructure Descriptors 1: Institute for Advanced Simulations – Materials Data Science and Informatics (IAS-9) Forschungszentrum Jülich GmbH 52425 Jülich, Germany; 2: Chair of Materials Data Science and Materials Informatics, Faculty 5 Georesources and Materials Engineering, RWTH Aachen University 52056 Aachen, Germany Plastic deformation in crystalline solids is fundamentally a multiscale phenomenon: the macroscopic mechanical response arises from the collective motion and interaction of dislocations at smaller scales. While continuum plasticity models are efficient for large-scale engineering applications, they often rely on phenomenological assumptions to represent microstructural effects. In contrast, high-fidelity methods such as discrete dislocation dynamics (DDD) explicitly resolve the underlying mechanisms, but at a higher computational cost. A bridging approach called continuum dislocation dynamics (CDD) promises cost efficient predictions while it still captures microstructural dislocation network information. A major challenge in developing this kind of predictive and scalable dislocation-based “continuum” models is the closure problem, namely how to recover microstructure-dependent information needed by continuum formulations from a set of low-order field quantities. In this work, we present a machine-learning framework that learns compact representations of dislocation microstructures and their temporal evolution directly from high-fidelity simulation data, enabling data-driven closure modeling and efficient surrogate prediction. DDD simulations are first coarse-grained into continuum field descriptions, from which we extract two complementary classes of descriptors: low-dimensional latent variables learned by autoencoders operating on the coarse-grained fields, and topology-sensitive representations learned by graph neural networks (GNNs) applied to graph-based descriptions of discrete dislocation networks. Analysis of the resulting latent trajectories reveals a reduced microstructural state space whose dynamics can be modeled explicitly. Based on this learned latent evolution, we build surrogate models that reproduce the main trends observed in the reference DDD simulations, capturing the evolution of key microstructure-dependent quantities as well as the associated macroscopic response with good qualitative fidelity, while significantly lowering computational cost. Overall, the proposed methodology provides a systematic route for transferring discrete-scale dislocation physics into continuum-scale plasticity models through learned descriptors and latent dynamics, and thereby contributes to the broader development of physics-informed machine learning strategies for multiscale computational mechanics. ID: 148
/ Poster No. # 9: 110
Modalities: Multimodal Data, Simulation Data, Time Series Methods: Foundation Models, Generative Models, Uncertainty Quantification Application Domain: Earth & Environment NVIDIA Earth-2: Open Models, Accessible Hardware, and the Future of AI for Weather and Climate NVIDIA Corporation Running a global weather forecast used to require a supercomputer and hours of wall-clock time. Today, AI models can do it in seconds. Over the past several years, NVIDIA's Earth-2 platform has been building the tools behind that shift: open-source AI models for global forecasting, regional downscaling, and long-range climate modeling, all accessible through Earth2Studio, a Python framework with 30+ models ready to run. We are now focused on making these capabilities widely available. Generative climate models like Climate in a Bottle can compress massive simulation datasets into compact learned representations that run on affordable desktop hardware like DGX Spark, putting interactive ensemble generation and km-scale impact analysis within reach of individual researchers. At the same time, emerging approaches to AI-based data assimilation and direct prediction from raw observations are shortening the path from sensor to forecast, and opening a path to real-time interactive updates. The goal is a stack of AI models spanning street-level to global, hourly to centennial, that any scientist can access with a GPU. And the broader AI revolution in coding, agents, and research automation is making it possible for small teams to build and deploy these systems faster than ever.
ID: 156
/ Poster No. # 9: 111
Modalities: Multimodal Data, Simulation Data, Tabular Data Methods: Physics-informed Machine Learning, Probabilistic Methods Application Domain: Health Amortized Bayesian Multilevel Model of Bacterial Cell-Cycle with BayesFlow 1: Helmholtz Institute for RNA-based Infection Research (HIRI), Helmholtz Centre for Infection Research (HZI), Würzburg, Germany; 2: Department of Biology, University of Toronto, Mississauga, ON, Canada; 3: Department of Cell and Systems Biology, University of Toronto, Toronto, ON, Canada Transcriptional heterogeneity enables bacterial populations to survive environmental changes and stress by supporting cooperative behavior to optimize nutrient uptake or survival under antibiotic treatment. Unlike bulk RNA sequencing, which averages over populations, single-cell RNA sequencing (scRNA-seq) captures transcriptomes of individual cells. In genetically identical, exponentially growing cells, a major driver of transcriptional variability is the cell cycle, which induces periodic changes due to ongoing DNA replication1, masking other signals. Separating cell-cycle effects from other sources of variability is essential for integrating single-cell data across conditions and for studying processes like antibiotic tolerance. While cell-cycle signatures have been described in bacterial scRNA-seq data1, a principled and scalable model applicable across transcriptomic technologies is lacking. We address this gap with a Bayesian parametric multilevel model that explicitly encodes periodic changes in DNA content during the bacterial cell cycle. Modern single-cell technologies capture roughly 200 mRNAs per E. coli cell and scale to over 100,000 cells1,2, rendering posterior inference via Markov Chain Monte Carlo (MCMC) infeasible. We overcome this limitation using amortized Bayesian inference (ABI), which trains neural networks on simulated data. While training is computationally intensive, it pays off (“amortizes”) at inference, allowing rapid sampling of cell-cycle states for thousands of bacteria. However, the multilevel structure of the full model, combining global parameters such as growth rate with thousands of local cell-cycle states, makes direct ABI impractical. To extend ABI to multilevel models, we factorize the posterior into global and local components similar to Habermann et al.3. We implement a two-stage ABI workflow using BayesFlow. First, global parameters are inferred from simulated and real scRNA-seq data comprising 50,000 E. coli cells1. These estimates are then used for amortized inference of local cell-cycle states for individual bacteria. We validate our approach against a Bayesian multilevel model implemented in Stan, using MCMC on a subset of 2,000 cells. Our method enables efficient regression of cell-cycle effects from bacterial transcriptomics data, enabling downstream analysis of bacterial heterogeneity. 1 Pountain et al. Nature, 2024 2 Sarfatis et al. Science, 2025 3 Habermann et al. Bayesian Anal. Advance Publication, 2025 ID: 320
/ Poster No. # 9: 112
Modalities: Image Methods: Other Application Domain: Information, Matter Exploring the Effectiveness of Pretraining Task Alignment for Self-Supervised Learning in a Low-Data Regime 1: Institute for Advanced Simulations — Materials Data Science and Informatics (IAS-9), Forschungszentrum Jülich GmbH, 52425 Jülich, Germany; 2: Chair of Materials Data Science and Materials Informatics, Faculty 5 — Georesources and Materials Engineering, RWTH Aachen University, 52056 Aachen, Germany Self-supervised learning (SSL) has proven to be effective at addressing challenges posed by limited annotated data. This is achieved by designing pretraining tasks that leverage raw data instances. Common pretraining strategies include masked image modeling, contrastive learning, and various predictive pretext tasks. These approaches enable models to learn useful representations that can later be adapted to downstream tasks through fine-tuning on relatively small annotated datasets. Furthermore, pretraining on domain-relevant datasets has been shown to be more beneficial than using general, real-world datasets such as ImageNet, especially when significant distribution shifts are present, as is often the case with scientific imagery. In many scientific domains, large-scale annotation efforts remain prohibitively expensive despite the availability of raw data, primarily due to time constraints and the need for expert knowledge. Recent studies suggest that in scenarios with limited fine-tuning data, designing targeted, task-aligned pretraining strategies may be more effective than relying on task-agnostic pretraining methods. Based on this observation, we investigate the role of pretraining task alignment in a low-data regime, focusing on the segmentation of dislocations in optical microscopy images of KOH-etched 4H-SiC wafer. We evaluate multiple pretraining strategies, including a novel segmentation-aligned approach based on Mask Image Modeling. Additionally, we systematically vary the size of the labelled dataset used for fine-tuning to analyze how different pretraining methods perform as the annotation budget increases. Our main objective is to determine whether the choice of pretraining strategy is of critical importance when data required for fine-tuning is scarce. ID: 190
/ Poster No. # 9: 113
Modalities: Image Methods: Probabilistic Methods, Uncertainty Quantification Application Domain: Health, Matter Taming Multimodal Posteriors over Rotations: A Geometric Approach to Bayesian Alignment in Cryo-EM and Beyond 1: Helmholtz Institute Jena, Germany; 2: GSI Helmholtzzentrum für Schwerionenforschung, Germany; 3: Friedrich Schiller University Jena, Germany; 4: Fraunhofer Institute for Applied Optics and Precision Engineering, Germany; 5: University Hospital Jena, Germany; 6: Max Plank Institute for Multidisciplinary Sciences, Germany Determining the 3D structure of macromolecules from cryo-electron microscopy (cryo-EM) images requires estimating the unknown orientation of each particle projection, a problem known as pose estimation (or 3D-2D rigid alignment). This is one of the central computational challenges in cryo-EM, where orientation estimates for thousands of particles must be resolved reliably across iterative reconstruction cycles. The difficulty is compounded by the fact that the rotation posterior is often highly multimodal, arising from structural symmetries and repetitive molecular features. Exhaustive angular search methods address this but scale poorly with dataset size. Recent developments in deep learning methods based on amortized inference offer high throughput, but face convergence issues rooted in the same multimodality problem. Thus, the need for a reliable orientation estimation procedure that directly addresses this remains important. We address this problem with a Bayesian inference approach specifically targeting the rotation group. Molecular structures and their projections are represented as Gaussian mixture models providing a compact, meshless alternative to voxel grids that admits analytic projection computations. Rotations are parametrized as unit quaternions, mapping the problem onto the 3-sphere and enabling the use of our recent work geodesic slice sampling on the sphere (Habeck et al., JMLR 2025), a gradient-free, tuning-free Markov chain Monte Carlo method that exploits the geometry of the spherical manifold for efficient sampling. Here, we extend its utility in estimating the unknown poses in cryo-EM projections. We demonstrate this on 80S ribosome class averages from a public cryo-EM benchmark dataset (EMPIAR-10028). Results show that our inference method identifies dominant posterior modes every time (100% success rate) over several runs and test cases, whereas the well-tuned state-of-the-art spherical-variant of Hamiltonian Monte Carlo, a gradient-based sampler, achieves only 25-27% success, as it often gets trapped in local optima. We further explore extensions in this work towards projection matching tasks in conventional tomography settings, including ptychotomography and laminography. Orientation estimation over rotations is ubiquitous across many scientific domains such as medical imaging, robotics, cosmology, and structural bioinformatics, and geodesic slice sampling offers a principled path toward reliable inference in all these settings.
ID: 125
/ Poster No. # 9: 114
Modalities: Other Methods: Generative Models Application Domain: Core Machine Learning, Health Sample Efficient Generative Molecular Optimization with Joint Self-Improvement 1: Helmholtz Munich, Germany; 2: Technical University of Munich, TUM School of Computation, Information and Technology; 3: Technical University of Munich, Campus Straubing for Biotechnology and Sustainability; 4: University of Applied Sciences Weihenstephan Triesdorf; 5: Faculty of Mathematics, Informatics and Mechanics, University of Warsaw Generative molecular optimization aims to design molecules with properties surpassing those of existing compounds. However, such candidates are rare and expensive to evaluate, yielding sample efficiency essential. Additionally, surrogate models introduced to predict molecule evaluations, suffer from distribution shift as optimization drives candidates increasingly out-of-distribution. To address these challenges, we introduce Joint Self-Improvement, which benefits from (i) a joint generative-predictive model and (ii) a self-improving sampling scheme. The former aligns the generator with the surrogate, alleviating distribution shift, while the latter biases the generative part of the joint model using the predictive one to efficiently generate optimized molecules at inference-time. Experiments across offline and online molecular optimization benchmarks demonstrate that Joint Self-Improvement outperforms state-of-the-art methods under limited evaluation budgets. ID: 324
/ Poster No. # 9: 115
Modalities: Multimodal Data, Text Methods: Agentic AI, Generative Models Application Domain: Health Mind the Gap – From expectation to reality in agentic AI 1: Helmholtz Munich, Germany; 2: German Center for Diabetes Research, Germany At HAICON 2025, I predicted that "Agentic systems are upon us." Not even one year later, they seem to pervade many aspects of our scientific and day-to-day lives. We are slowly getting used to the fact that their deployment fundamentally involves an inversion of operational control: humans must step back and allow the system to take command. However, the enthusiasm they are often met with is heavily contrasted by the apparent lack of concrete benefits in controlled studies. Software engineers expected to be 24% faster with coding tools—they were 19% slower. Ambient clinical scribes promised savings of minutes per encounter, but didn't show any significant improvement in real-world deployments. We term this the "expectation–realisation gap." Complicating the matter, the gap is not uniform. The same tool that slows down experienced developers by 19% can accelerate novices by 56% on constrained tasks. A steep decline in hiring for entry-level positions suggests that employers bet on AI productivity gains. But despite $30–40b of spending, 95% of surveyed companies report no measurable returns. In agentic AI, who benefits and who doesn't depends on context that is often not planned for. To close this gap, we propose the Agentic Automation Canvas (AAC), a structured communication framework for users and developers of agentic systems (https://aac.slolab.ai). Before control is handed over to AI, the canvas requires both sides to make their assumptions explicit. Which benefits are expected and for whom? What are current baseline capabilities? How much human oversight is needed, and what does it cost? By guiding the conversation in early phases, the AAC surfaces misalignments before they become costly disappointments. The canvas is free and open source, implemented as an AI-ready specification and web application that is private by design. Because the AAC produces machine-readable artifacts, completed canvases can be aggregated across projects to surface common components, reduce redundancies, and recognise opportunities for shared infrastructure. As AI is increasingly woven into research and business infrastructure, structured planning becomes integral to our productivity and success. The AAC aims to make this planning FAIR, transparent, and effective. ID: 1401
/ Poster No. # 9: 116
Modalities: Multimodal Data, Text, Other Methods: Foundation Models Application Domain: Health Genomic language model based prediction of transcription start sites in twelve yeast species 1: Computation Molecular Medicine, School of Computation, Information and Technology, TU Munich, Germany; 2: Computational Health Center, Helmholtz Center Munich, Germany We lack genome annotations for many species, especially for features beyond coding sequences. For example, while the transcription start site (TSS) determines the 5' untranslated region, which is crucial for post-transcriptional regulation, its position in non-model organisms is often unclear. Accurate mapping of the TSS across the tree of life is difficult due to the evolution of the regulatory code and lack of experimental data. Recently developed genomic language models (gLM) from AI trained on increasing numbers of sequenced genomes across the Tree of Life seem promising for tackling this problem, but methods still need to be established. Here, we predict CAGE-seq coverage across fungi, spanning 500 million years of evolution. While only using experimental data for 12 species, we leverage a gLM pre-trained on hundreds of fungi to achieve good accuracy at predicting the TSS signal for held-out genes, outperforming a BPNet in direct comparison. On held-out species,using a gLM does improve over a BPNet, but the problem remains hard. We show that multimodal pre-training of the gLM on DNA and RNA jointly improves generalization to new species with available RNA-seq data. | |||||||||||||||||||||||||||||||
