Ruprecht-Karls-Universität Heidelberg
Bilder vom Neuenheimer Feld, Heidelberg und der Universität Heidelberg
Siegel der Uni Heidelberg

Finding Cross-Lingual Processing Mechanisms in Multilingual Language Models

Abstract

Today, most language models are created multilingual. These models, though trained on different scales, on different corpora, and at different points in time, are all built on variations of the transformer architecture. In this work, we build on a line of research that wants to identify how transformer models are multilingual. While previous work concentrated on identifying parameter subspaces or neurons shared between languages, we set out to identify shared processing mechanisms within these parameter spaces by learning cross-lingual causal interventions with distributed alignment search (DAS). We concentrate on high-level linguistic features that are known to be represented in the parameter space of transformer models: subject-verb agreement, gender agreement, and filler-gap constructions. For each feature, we learn an intervention model on monolingual sentence pairs of a source language, and evaluate whether the intervention model transfers to similar sentences in a target language. If the transfer succeeds, the intervention model has identified both a shared parameter subspace and a shared processing mechanism. We find evidence for shared processing mechanisms for all 3 features across 12 languages. The effect size is modulated by phylogenetic distance, tokenizer overlap, and language model quality on source and target languages.

zum Seitenanfang