Research
I develop Bayesian inference procedures for problems where the number of variables runs into the millions — and then make them fast enough to actually run on biobank-scale data. Most of this work builds on approximate message passing, a family of iterative algorithms that comes with exact asymptotic guarantees through state evolution, which in turn gives calibrated uncertainty and principled significance testing.
- Genome-wide association studies
- Polygenic risk scores
- Epigenomics
- Proteomics
- Time-to-event models
Projects
-
Transfer learning for cross-ancestry polygenic risk scores
TLgVAMP extends the gVAMP framework with transfer learning, so that a polygenic risk score fitted in a large, well-powered cohort can inform inference in a smaller cohort from an under-represented population without washing out population-specific signal.
-
Joint variable selection for omic biomarkers in time-to-event data
We build vampW, a scalable Bayesian framework based on approximate message passing that models disease onset times under a Weibull model and selects omic biomarkers jointly, conditional on all others. Applied to the UK Biobank Pharma Proteomics Project, it improves onset prediction by 26–33% relative to penalised Cox regression and baseline deep-learning approaches.
-
Joint modelling of whole genome sequence data for human height via approximate message passing
We develop a new algorithmic paradigm based on approximate message passing, gVAMP, to directly fine-map whole-genome sequence (WGS) variants and gene burden scores, conditional on all other measured DNA variation genome-wide. We find that the genetic architecture of height inferred from WGS data differs from that inferred from imputed single nucleotide polymorphism (SNP) variants:common variant associations from imputed SNP data are allocated to WGS variants of lower frequency, and there is a stronger relationship of effect size and variant frequency.
-
Incorporating summary statistics into VAMP framework
We first adapt gVAMP to the summary statistics setup in order to propose a novel method called summary gVAMP (sgVAMP). We demonstrate that compared to other popular summary statistics methods, sgVAMP achieves state-of-the-art out-of-sample prediction accuracy across several traits in the UK biobank using the 2.17 million SNP set. Secondly, we extend sgVAMP to a multi-cohort setting, which allows for the joint estimation of shared and population-specific signals across multiple ancestries.
-
Epigenome-wide association studies using approximate message passing
We develop gVAMPomi, approximate message passing-based paradigm, and apply it to the largest human methylation dataset generated to date, the Generation Scotland study. We find 92 CpG probes whose effects are significantly associated with traits, conditional on all other CpG probes, representing a significant increase over 37 CpG probes discovered by baseline MCMC approach.
-
Detection of age-specific genetic effects for age-at-onset complex traits
We develop an extension of the time-to-event MCMC software called BayesW to two and three epochs by imposing statistical model and deriving Gibbs updates for interaction parameters and epoch-specific parameters.
Beyond genomics
I work with the TU Wien Textile Recycling Group on the experimental design, statistical analysis and data visualisation behind their solvent-based recycling processes — two joint papers so far in Waste Management, listed on the publications page.
During a machine learning and privacy internship at the University of Vienna I built a modular library for benchmarking membership inference attacks on large language models, together with a document-level differential privacy auditing framework.