Paper recorded by Signals 4 on 2026-09-28 in cs.LG. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-28 on arXiv · recorded by Signals 4 on 2026-09-29
Category: cs.LG · 机器学习 · first seen 2026-09-29
Recent progress in machine learning is driven by large-scale foundation models, where scaling laws and finding optimal scaling prescriptions for architecture, data, and hyperparameters are key in advancing the state-of-the-art. Therefore, it is surprising that no systematic study evaluates the methodology to obtain scaling laws and prescriptions across different model types. To shed light on this