Signals 4 · free daily AI digest

Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data

Paper recorded by Signals 4 on 2026-09-10 in cs.LG. Abstract reproduced from arXiv; link to the original below.

Published 2026-09-10 on arXiv · recorded by Signals 4 on 2026-09-11

Category: cs.LG · 机器学习 · first seen 2026-09-11

Abstract

As the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activated Transformers, but the effects of data repetition remains largely unexplored for recently dominant sparse architectures such as Mixture-of-Experts (MoE), despite their increased compute efficiency. We vary data repetit

Read on arXiv →

#67 most recent of 215 cs.LG papers we have recorded · ↑ newer: Dissecting GPU Utilization for LLM Inference on Nvidia Hopper · ↓ older: From Protocols to Evidence: Bounded Claims for AI in Service of the Co
Cite this page: Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data: the #67 most recent of 215 cs.LG papers we have recorded (as of 2026-09-10). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/data-scarcity-and-model-sparsity-mixtures-of-experts-overfit-more-to-repeated-da.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable: papers.json
Related: More cs.LG papers · arXiv signals · All papers · Today in AI
Get 4 AI signals a day by email — free.
Subscribe free → See all plans →
Get 4 AI signals a day by email — free
All models · All repos · By company · Daily editions