Learning from Many Eyes: Uncertainty-Aware and Calibrated AI for Trustworthy Pancreatic Cancer Analysis in Clinical Contexts

By: Meritxell Riera i Marín
Supervisors: Dr. Miguel Ángel González Ballester, Dr. Javier García López & Dr. Adrian Galdrán Cabello

Date: September 14, 2026 - 11:00 h
Room: 55.309

Abstract

This dissertation addresses the fundamental ”Ground Truth Problem” in medical imaging Artificial Intelligence (AI), where the conventional practice of collapsing divergent expert opinions into a single consensus mask creates overconfident and untrustworthy systems. In the context of high-stakes clinical decision-making, such as the analysis of Pancreatic Ductal Adenocarcinoma (PDAC), this deterministic approach is particularly hazardous. PDAC is characterized by ill-defined infiltrative margins and a complex relationship with adjacent vascular structures, leading to profound inter-rater variability that reflects inherent biological ambiguity rather than mere observer error. This research advocates for a paradigm shift, transitioning from rigid deterministic models toward probabilistic, multi-rater aware frameworks that treat expert disagreement as a valuable diagnostic signal.

The work is structured around three primary pillarss: evaluation, benchmarking, and training. First, we address the instability of current reliability assessments by introducing the Multi-Rater Expected Calibration Error (MR-ECE). This metric virtually expands test sets by considering each expert’s perspective independently, providing a stable measure of model calibration even in data-limited clinical scenarios. We further provide a thorough study on the complementarity of multi-rater metrics, demonstrating that a multi-faceted evaluation suite is essential to capture the nuances of model behavior.

Second, to enable the systematic study of diagnostic uncertainty, this thesis led the creation and public release of the CURVAS and CURVAS-PDACVI datasets through a MICCAI challenge series. These FAIR-compliant benchmarks provide the community with independent annotations from multiple experts, serving as a gold standard for evaluating algorithms under real-world human variability.

Third, we propose ordinal consensus learning utilizing a Ranked Probability Score (RPS) loss. By interpreting annotator consensus as an ordered hierarchy of confidence levels, our models produce inherently calibrated segmentations that preserve the ”gray zones” of clinical interpretation. Clinical validation in PDAC staging reveals that macroscopic volumetric overlap (Dice Score) is a poor proxy for the high-resolution precision required at the tumor-vessel interface.

Ultimately, this dissertation demonstrates that uncertainty-aware models provide a vital clinical ”safety net.” By flagging ambiguous cases for multidisciplinary review, these frameworks can help to bridge the trust gap between AI developers and clinical specialists, facilitating the integration of transparent and reliable AI tools into the clinical decision-making process of pancreatic cancer.


Computational Analysis of Piano Practice: Towards Modeling Practice Mistakes and Repetitions for Understanding Learning Behaviours

By: Alia Ahmed Morsi Moustafa
Supervisor: Dr. Xavier Serra

Date: September 29, 2026 - 14:00 h
Room: 55.309

Abstract

The main goal of this research is to understand the process of musical instrument learning in the Western Classical tradition, through developing computational methods for the analysis of piano practice recordings. We focus on mistakes and the organizational structure of practice sessions, as both can reveal information about one’s musical expertise. Specifically, our research (i) models mistakes in a manner that reflects educational understanding beyond deviations from a reference, (ii) trains classifiers that can automatically detect common piano learning mistakes, and (iii) develops methods to understand the organization of musical content in practice sessions based on similarity matrix analysis.

We propose treating piano mistakes as sequences that include both initial error and subsequent recovery phases, with each composed of low-level operations (note insertions, deletions, and time shifts) and release a toolkit that generates synthetic labelled data according to the proposed framework. Considering mistakes as behavioral phenomena with a contextual and qualitative aspect, we investigate whether the locations of salient errors in a performance can be detected through neighbouring context without music score comparisons, through supervised learning experiments combining synthetic and real data. Our results show feasibility while revealing generalization challenges across mistake types and performance contexts.

As for the computational analysis of practice organization, we develop a similaritybased segmentation method that identifies repeated attempts at passages, with emphasis on grouping together attempts of the same passage despite the existence of mistakes. We discuss how self-similarity analysis enables close examination of mistake patterns across repetitions to provide insights into planning and learning strategies, and to enable future comparative studies that track the evolution of learning.

This dissertation operates in a domain where computational measurement and educational meaning are still being aligned, and takes that gap seriously as a methodological starting point. The contributions above are components of a bigger picture towards formalizing realistic and useful computational analysis targets for piano learning. The choices of approach in mistake modelling, simulation, detection, and practice structure analysis collectively illustrate how methodological transparency can deepen the analysis itself.


Translational Computational Cardiac Safety: Comprehensive Modeling for Predicting Drug-induced Proarrhythmic Risks Across Diverse Populations

By: Paula Domínguez Gómez
Supervisors: Dr. Oscar Cámara, Dr. Jazmín Aguado & Dr. Borje Darpo

Date: October 26, 2026 - 15:00 h
Room: 55.309

Abstract

Cardiac safety assessment is one of the most consequential challenges in drug development. Drug-induced arrhythmias have historically driven late-stage clinical failures and even postmarket withdrawals, and the field has responded with increasingly stringent preclinical and clinical screening requirements. The underlying mechanism (block of cardiac ion channels, leading to QT interval prolongation and potentially fatal ventricular arrhythmias) is well characterized, yet the translation of preclinical safety signals into reliable predictions of clinical risk remains an unsolved problem. Traditional methodologies, such as in vitro ion channel assays and animal models, provide important early warning signals but are limited in their ability to capture the complexity of human cardiac electrophysiology, the heterogeneity of patient populations, and the range of clinical exposure conditions encountered in practice.

Computational modeling has emerged as a powerful complement to existing frameworks, offering mechanistic understanding of underlying processes, enhanced predictive accuracy, and the ability to explore scenarios that are inaccessible to experimental or clinical investigation. Initiatives such as the Comprehensive in vitro Proarrhythmia Assay have formalized the role of in silico modeling in cardiac safety assessment, promoting a shift toward model-informed drug development in regulatory practice. Within this context, the present thesis develops and validates a series of translational computational frameworks for drug-induced proarrhythmic risk assessment, bridging the gap between preclinical testing and clinical outcomes across diverse and clinically relevant patient populations. 

The work develops virtual cardiac populations for in silico clinical trials, building a framework for predicting exposure-response relationships and validating it against clinical trial data. This foundation is progressively extended to incorporate placebo effects, pathological conditions, and mathematical models of drug interactions, enabling safety stratification across underrepresented subpopulations and risk estimation for combination therapies. A parallel line of work addresses scalability limitations through artificial intelligence cardiac emulators, which replicate the behavior of the underlying electrophysiological models at a fraction of the computational cost, enabling real-time sex-specific proarrhythmic risk assessment and uncertainty quantification, demonstrated through a loperamide overdose case study. Underpinning both lines of work is a structured verification, validation, and uncertainty quantification pipeline for scientific machine learning. This pipeline extends traditional credibility assessment principles to systems whose parameters are inferred from data, rather than derived from physics-based models.

Taken together, these contributions support the view that the incorporation of computational tools into drug cardiac safety assessment represents an evolution rather than a revolution. The principles underlying these approaches are not new; what this thesis provides is a more representative, more scalable, and more integrated cardiac safety assessment built on those same principles, better equipped to meet the demands of modern drug development and the expectations of an increasingly model-informed regulatory environment.