EN Submit a tool
Paper

NVIDIA Research: AdamW Has a Scale Ceiling, SOAP/Muon Are Better

Published: Source: X: Elvis Saravia (@omarsar0, DAIR.AI)

ShareXFacebookTelegramWhatsApp

NVIDIA Research: AdamW Has a Scale Ceiling, SOAP/Muon Are Better

New research from NVIDIA shows that under batch sizes up to 100 million tokens, the AdamW optimizer suffers from training instability, while SOAP and Muon remain stable. The team eliminated loss spikes in SOAP under large batch sizes through per-step QR orthogonalization, and verified that both optimizers consistently outperform AdamW on multi-billion parameter models trained on trillions of tokens. They also proposed a layer-wise distributed optimizer compatible with Megatron-LM, balancing memory and communication without sacrificing convergence gains.

Read the original (opens in a new tab)

News stream data aggregated by AI HOT

Related newsLatest in this category