Data Privacy in the face of Machine Learning
Module Description
| Course | Module Abbreviation |
Credit Points | Module Prerequisites |
|---|---|---|---|
| BA-2010 | AS-CL | 8 LP | 1xCS-CL v 1xBS-CL |
| BA-2010[100%|75%] | CS-CL | 6 LP | FLA, FF-FM, ICL |
| BA-2010[50%] | BS-CL | 6 LP | FLA, FF-FM |
| Master | SS-SC-TAC | 8 LP | - |
| BA-2026 | AS-CL | 8 LP | 1xCS-CL v 1xBS-CL |
| BA-2026[100%|75%] | CS-CL | 6 LP | FLA, FF-M, ICL |
| BA-2026[50%] | BS-CL | 6 LP | FLA, FF-M, ICL |
| Lecturer | Constantin Seibold |
| Module Type |
|
| Language | Englisch |
| First Session | 15.10.2026 |
| Time and Place | Thursday, 09:15 - 10:45 INF 289 / SR 2 |
Participants
All advanced CL Bachelor students and all CL master students. Students from MSc Data and Computer Science or MSc Scientific Computing with Field of Application Computational Linguistics are welcome after getting permission from the lecturer. MSc Scientific Computing students can only take the course as HS for 8 LP. If the seminar should be oversubscribed, CL students will have priority.
Prerequisites for Participation
Module Prerequisites
Basic knowledge of computer science is assumed. Basic knowledge of probability and statistics is an advantage; for topics such as federated learning or privacy-preserving machine learning, foundational knowledge of machine learning is helpful. A willingness to read and work through English-language scientific literature is expected. No formal prior knowledge of data protection is required.
Assessment
Continuous-assessment course. The final grade is composed of several partial assessments:
- Presentation of the prepared topic (talk)
- Active participation in discussion and Q&A
- Written short report / seminar paper
Content
Data protection and the safeguarding of privacy are among the central challenges in handling digital data. This seminar introduces current concepts, methods, and open research questions in the field of Data Privacy. Starting from fundamental threat models, it covers technical and methodological approaches to protecting personal and sensitive data, discussed on the basis of current scientific literature.
Possible thematic directions for the seminar papers include, among others:
- Re-identification threats and attacks (e.g. linkage attacks, membership inference)
- Data safety / data security
- Anonymization (k-anonymity, l-diversity, t-closeness)
- Differential privacy
- Federated learning
- Further topics: synthetic data, privacy-preserving machine learning, homomorphic encryption, secure multi-party computation
At the start, participants are assigned a broad thematic direction. In consultation with the course instructor, they research independently relevant scientific articles, work through them, and prepare them for the group in the form of presentations or short reports. Joint discussion (Q&A) is an integral part of each session. Medical data may be used as an application example but are not the focus.
Literature
The following well-established foundational publications serve as good starting points for topic assignment, on which participants can build:
- Dwork & Roth (2014): The Algorithmic Foundations of Differential Privacy. Foundations and Trends in Theoretical Computer Science 9(3–4), 211–407.
- Sweeney (2002): k-anonymity: a model for protecting privacy. International Journal on Uncertainty, Fuzziness and Knowledge-based Systems 10(5), 557–570.
- Narayanan & Shmatikov (2008): Robust De-anonymization of Large Sparse Datasets. IEEE Symposium on Security and Privacy, 111–125. (Netflix dataset)
- McMahan et al. (2017): Communication-Efficient Learning of Deep Networks from Decentralized Data. AISTATS 2017, 1273–1282. (Federated learning)
- Shokri et al. (2017): Membership Inference Attacks Against Machine Learning Models. IEEE Symposium on Security and Privacy, 3–18.
Further literature for each topic will be selected in consultation with the course instructor.
Further reading will be announced during the course of the seminar


