Paper recorded by Signals 4 on 2026-09-10 in cs.CL. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-10 on arXiv · recorded by Signals 4 on 2026-09-11
Category: cs.CL · 自然语言处理 · first seen 2026-09-11
A language model normally begins training with random word embeddings: whatever 'banana' means must be learned from training corpora. I implement St. Augustine's picture of word learning, meaning by ostension, for a small masked language model (DeBERTa) trained on 10M words: before training, visually grounded tokens receive embeddings derived from the image regions they label; other tokens start r