Paper recorded by Signals 4 on 2026-09-10 in cs.LG. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-10 on arXiv · recorded by Signals 4 on 2026-09-11
Category: cs.LG · 机器学习 · first seen 2026-09-11
As the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activated Transformers, but the effects of data repetition remains largely unexplored for recently dominant sparse architectures such as Mixture-of-Experts (MoE), despite their increased compute efficiency. We vary data repetit