Signals 4 · free daily AI digest

Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?

Paper recorded by Signals 4 on 2026-09-30 in cs.LG. Abstract reproduced from arXiv; link to the original below.

Published 2026-09-30 on arXiv · recorded by Signals 4 on 2026-10-01

Category: cs.LG · 机器学习 · first seen 2026-10-01

Abstract

Differentially Private Stochastic Gradient Descent (DP-SGD) is a leading approach for privacy-preserving fine-tuning of large language models (LLMs). Many decoder-only LLMs employ weight tying between input and output embeddings, a design choice originally introduced for parameter efficiency and improved language modeling performance in the non-private setting. However, the impact of weight tying

Read on arXiv →

#4 most recent of 349 cs.LG papers we have recorded · ↑ newer: Image Classifiers are Efficient Self-Supervised Video Representation L · ↓ older: Compression Footprints as Security Signals for Model-Poisoning Defense
Cite this page: Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?: the #4 most recent of 349 cs.LG papers we have recorded (as of 2026-09-30). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/is-weight-tying-still-beneficial-for-decoder-only-llms-in-private-settings-under.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable: papers.json
Related: More cs.LG papers · arXiv signals · All papers · Today in AI
Get 4 AI signals a day by email — free.
Subscribe free → See all plans →
Get 4 AI signals a day by email — free
All models · All repos · By company · Daily editions