Ruprecht-Karls-Universität Heidelberg
Bilder vom Neuenheimer Feld, Heidelberg und der Universität Heidelberg
Siegel der Uni Heidelberg

Leaving the Cave by Breaking the Symmetry

Abstract

Multilingual language models often transfer knowledge across languages, including between languages with little lexical overlap. Separately, a growing line of research on the Platonic Representation Hypothesis argues that neural networks trained on different data, objectives, or modalities may converge toward shared representations of the same underlying world. Could cross-lingual transfer similarly arise because multilingual models infer shared structure across languages? I study this question in a fully inspectable toy setting: a one-layer Transformer trained on modular addition in two artificial languages with disjoint vocabularies. Because the underlying mechanism is known, we can directly compare the representations and computations learned for both languages. Modular addition alone permits multiple equally valid internal codes, and the model can therefore solve both languages without aligning corresponding numbers. I then add a second task that breaks this symmetry. The resulting experiments suggest that, once the shared world becomes sufficiently fingerprintable, the model can align the multilingual representation space even without shared vocabulary.

zum Seitenanfang