Logo des Repositoriums
Zur Startseite
  • English
  • Deutsch
Anmelden
  1. Startseite
  2. SuUB
  3. Dissertationen
  4. Applied Machine Learning in Epidemiology: Feature Selection, Benchmarks, and Software
 
Zitierlink DOI
10.26092/elib/6360

Applied Machine Learning in Epidemiology: Feature Selection, Benchmarks, and Software

Veröffentlichungsdatum
2026-06-17
Autoren
Burk, Lukas  
Universität Bremen  
Betreuer
Wright, Marvin N.  
Gutachter
Wright, Marvin N.  
Boulesteix Anne-laure
Zusammenfassung
Predictive modeling is central to epidemiology, biostatistics, and adjacent fields.
The number of available algorithms keeps growing, alongside related techniques to make sense of them, leaving practitioners with three common questions:
Which algorithm should they choose, which features are actually relevant, and if so, how relevant are they, in which way?
This cumulative dissertation addresses these questions with three main contributions focusing on empirical evaluation and open-source software.

The first contribution presents the most comprehensive large-scale neutral comparison study of survival prediction algorithms on low-dimensional data to date, evaluating 19 algorithms on 34 datasets with two distinct tuning measures and a total of six evaluation measures.
After extensive hyperparameter tuning via Bayesian optimization, the study does not find any method to statistically significantly outperform the Cox proportional hazards model in aggregate.
Sensitivity analyses using the Plackett-Luce model indicate that violation of the proportional hazards assumption, or generally misspecification of the Cox model, changes this picture.
Model-based boosting and oblique random survival forests are among the methods that gain an advantage when these assumptions are violated.

The second contribution introduces a novel method for feature selection in competing event settings, named cooperative penalized regression (CooPeR).
The method builds on the feature-weighted elastic net, which is used to fit cause-specific penalized Cox models in a manner that reduces penalization for features with large effect on the respective other event, while also amplifying the penalization of noise features.
The method's effectiveness is demonstrated in a simulation study mimicking gene-sequencing data and in a real-data application.

The third contribution presents \xplainfi, an \rstats package for feature importance analysis with a unified and modular interface.
The package supports perturbation-based, refitting-based, and Shapley-based importance measures while supporting both model- and learner-importance.
The package also supports multiple statistical inference approaches, including the conditional predictive impact with either model-X knockoffs or adversarial random forests.
Empirical benchmarks demonstrate correctness of the results alongside reduced runtime compared to selected reference implementations.

Five secondary contributions complement the three primary research areas, covering survival analysis with an overview of reduction techniques, work on performance evaluation, and algorithmic fairness.
Other contributions address conditional feature importance, and community-driven machine learning algorithm integration in the mlr ecosystem
Schlagwörter
machine learning

; 

survival analysis

; 

neutral comparison study

; 

feature importance

; 

feature selection

; 

open-source software
Institution
Universität Bremen  
Fachbereich
Fachbereich 03: Mathematik/Informatik (FB 03)  
Institute
Leibniz-Institut für Präventionsforschung und Epidemiologie BIPS GmbH  
Dokumenttyp
Dissertation
Lizenz
https://creativecommons.org/licenses/by/4.0/
Sprache
Englisch
Dateien
Lade...
Vorschaubild
Name

Burk-Applied-Machine-Learning-in-Epidemiology.pdf

Size

6.89 MB

Format

Adobe PDF

Checksum

(MD5):73bd2d0a7fbd4fb6f24a0b9d4b768635

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Datenschutzbestimmungen
  • Endnutzervereinbarung
  • Feedback schicken