Ruprecht-Karls-Universität Heidelberg
Bilder vom Neuenheimer Feld, Heidelberg und der Universität Heidelberg
Siegel der Uni Heidelberg

Scaling the Right Things: Data, Grounding, and Knowing When Not to Answer

Abstract

We frame trustworthy LLMs as a problem of provenance and grounding: before training, we must ask where a model’s knowledge comes from, and at inference time, whether its answers are actually supported by evidence. This talk follows two research threads across the model lifecycle. First, in Aleph Alpha GermanWeb, we explore how to train strong German-language models from scratch by treating data as design material and not merely as fuel—something to filter, shape, and increasingly synthesize—because data is the source code of model behavior. Second, in Merlin–Arthur RAG, we focus on post-training when models train to answer using retrieved context: a helpful prover provides evidence while an adversarial one tries to mislead, pushing the model to either ground its answer or admit that the evidence is not enough. Looking forward, these threads point to a shift in the field: from static datasets to evolving data synthesis pipelines, and from passive evaluation to training with adversaries that stress-test reasoning itself. Scaling alone is not enough—the future of language models lies in systems that actively curate what they learn, withstand manipulation, and admit when they do not know.

zum Seitenanfang