65;7006;1c Institut für Computerlinguistik
Ruprecht-Karls-Universität Heidelberg
Bilder vom Neuenheimer Feld, Heidelberg und der Universität Heidelberg

Reproducible Machine Learning: Theory and Practice

Module Description

Course Module Abbreviation Credit Points
Bachelor CL AS-CL 8 LP
Master CL SS-CL, SS-TAC 8 LP
Seminar Informatik BA + MA 4 LP
Anwendungsgebiet Informatik MA 8 LP
Anwendungsgebiet SciComp MA 8 LP
Lecturer Stefan Riezler
Module Type Seminar
Language English
First Session 14.04.2026
Time and Place Tuesday, 11:15-12:45 and 14:15-15:45
Mathematikon SR10
Commitment Period tbd.

Participants

All advanced CL Bachelor students and all CL master students. Students from MSc Data and Computer Science or MSc Scientific Computing with Field of Application Computational Linguistics are welcome after getting permission from the lecturer. If the seminar should be oversubscribed, CL students will have priority.

Prerequisite for Participation

Good knowledge of statistical machine learning and experience in experimental work.

Assessment

  • Regular and active participation (discussion of presented papers during seminar sessions)
  • First half of semester: Oral presentation (30min presentation + 15min discussion, commitment for presentation by April 20, 2026, by email stating 3 ranked preferences)
  • Second half of semester: Oral presentation of own reproducibility study (one month after paper presentation)
  • The seminar grade is based on the two oral presentatios (50/50)
  • No written exam, report, or term paper!

Content

Reproducibility of experimental results is one of the fundamental pillars of scientific research. If neither a reliable nor significant evaluation result can be obtained when replicating an experiment, the whole methodological foundation of the research result becomes questionable, even casting doubt on its validity.

In this seminar we will learn about several sources of nondeterminism that hamper reproducibility, and about statistical reliability and significance tests to allow us to analyze the inferential reproducibility of machine learning research. This means that instead of removing all sources of measurement noise, we will incorporate certain types of variance as irreducible conditions of measurement, and analyze their interaction with data properties, with the aim to draw inferences beyond particular instances of trained models.
We will learn how to incorporate meta-parameter variations and data properties into statistical significance testing with Generalized Likelihood Ratio Tests (GLRTs), how to use variance component analysis based on Linear Mixed Effects Models (LMEMs) to analyze the contribution of noise sources to overall variance, and how to compute a reliability coefficient as indicator for reproducibility.

The seminar is structured in two parts: In the first half of the semester, students will present papers on various sources of nondeterminism in machine learning, methods to measure reliability and significance of experimental results, and reproducibility studies exemplifying these tools. Students are expected to conduct a reproducibility study on their own and to present their results in the second half of the semester.

Schedule

Date Material Presenter
14.4. morning session Orga Riezler
21.4. morning session Introduction
Hagmann and Riezler, 2023. Towards Inferential Reproducibility of Machine Learning Research.
Chapter 5 of Riezler and Hagmann, 2024. Validity, Reliability, and Significance: Empirical Methods for NLP and Data Science.
Riezler
slides
28.4. morning session Sources of Nondeterminism: Implementation-Level
[1] Pham et al., 2021. Problems and opportunities in training deep learning software systems: An analysis of variance.
[2] Zhuang et al., 2022. Randomness in neural network training: Characterizing the impact of tooling.
[1] Kento Verlaan
[2] Raziye Sari
28.4. afternoon session Sources of Nondeterminism: Optimizer-Level
[3] Schmidt et al., 2021. Descending through a crowded valley - benchmarking deep learning optimizers.
[4] Ahn et al., 2022. Reproducibility in optimization: Theoretical framework and limits.
[3] Yan Gang
[4] Siqing Cai
5.5. morning session Source of Nondeterminism: Metaparameter Variation
[5] Melis et al., 2018. On the state of the art of evaluation in neural language models.
[6] Reimers and Gurevych, 2017. Reporting score distributions makes a difference: Performance study of LSTM-networks for sequence tagging.
[5] -
[6] Liyang Deng
5.5. afternoon session Sources of Nondeterminism: Evaluation Metrics
[7] Chen et al., 2022. Reproducibility issues for BERT-based evaluation metrics.
[8] Post, 2018. A call for clarity in reporting BLEU scores.
[7] Thea Bartmann
[8] Leander Karp
12.5. morning session Sources of Nondeterminism: Data Splits
[9] Sogaard et al., 2021. We need to talk about random splits.
[10] Gorman and Bedrick, 2019. We need to talk about standard splits.
[9] Jiayi Ji
[10] Mukeshkumar Prajapati
12.5. afternoon session Sources of Nondeterminism: Prompt Variation
[11] Maia Polo et al., 2024. Efficient multi-prompt evaluation of LLMs.
[12] Lior et al., 2025. ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments.
[11] Priya Yadav
[12] Til Gramlich
19.5. morning session Reliability Measures: Bootstrap Confidence Intervals
[13] Agarwal et al., 2021. Deep reinforcement learning at the edge of the statistical precipice.
[14] Paraschakis et al., 2024. Confidence Interval Estimation of Predictive Performance in the Context of AutoML.
[13] Min-Han Yeh
[14] Cilian Kerskens
19.5. afternoon session. Reliability Measures: Variance Component Analysis and Intra-Class Correlation Coefficient
[15] Chapter 3 of Riezler and Hagmann, 2024. Validity, Reliability, and Significance: Empirical Methods for NLP and Data Science.
[16] Geburek et al., 2024. LMEMs for post-hoc analysis of HPO Benchmarking.
[15] Ertugrul Taparci
[16] Tan Ke
26.5. morning session Significance Testing: Score Distribution Comparison
[17] Dror et al., 2019. Deep dominance - how to properly compare deep neural models.
[18] Ulmer et al., 2022. deep-significance - Easy and Meaningful Statistical Significance Testing in the Age of Neural Networks.
[17] Dana Simedrea
[18] Lara Eulenpesch
26.5. afternoon session Significance Testing: Bootstrap and Randomization
[19] Clark et al, 2011. Better hypothesis testing for statistical machine translation: Controlling for optimizer instability.
[20] Sellam at al., 2022. The multiBERTs: BERT reproductions for robustness analysis.
[19] Yunhe Dong
[20] Luca Scavone
2.6. morning session Significance Testing: The Generalized Likelihood Ratio Test
[21] Chapter 4 of Riezler and Hagmann, 2024. Validity, Reliability, and Significance: Empirical Methods for NLP and Data Science.
Examples for Reproducibility Studies
[22] Zoellin et al., 2024. Evaluating the reproducibility of a deep learning algorithm for the prediction of retinal age.
[21] Xinyao Peng
[22] Julian Ring
2.6. afternoon session Examples for Reproducibility Studies
[23] Akacik et al., 2025. ModernTCN Revisited: A Critical Look at the Experimental Setup in General Time Series Analysis.
[24] Pollo et al., 2025. Benchmarking LLM Capabilities in Negotiation through Scorable Games.
[23] Bohdana Ivakhnenko
[24] Ningyun Chen
9.6. Students' Reproducibility Studies Kento Verlaan, Raziye Sari, Yan Gang, Siqing Cai
16.6. Students' Reproducibility Studies Liyang Deng, Thea Bartmann, Leander Karp
23.6. Students' Reproducibility Studies Mukeshkumar Prajapati, Jiayi Ji, Priya Yadav, Til Gramlich
30.6. Students' Reproducibility Studies Min-Han Yeh, Cilian Kerskens, Ertugrul Taparci, Tan Ke
7.7. Students' Reproducibility Studies Dana Simedrea, Lara Eulenpesch, Yunhe Dong, Luca Scavone
14.7. Students' Reproducibility Studies Xinyao Peng, Julian Ring, Bohdana Ivakhnenko, Ningyun Chen
zum Seitenanfang