Muon and its follow-ups: matrix-based optimisation

teaching

Motivation

Muon updates weight matrices by approximately equalising the singular values of a momentum matrix. Experiments on language models have reported improved training efficiency relative to AdamW. Muon is normally applied to hidden-layer weight matrices; other parameters use an optimiser such as AdamW.

Equalising singular values increases the relative contribution of weak directions, which may contain useful signal or stochastic noise. Whether this improves training depends on the problem and on choices such as update scaling, momentum, and approximation accuracy. Small transformers and multilayer perceptrons (MLPs) provide settings in which these effects can be studied with repeated experiments.

Project goal

Investigate whether Muon and related methods offer a gain over established optimisers, under which conditions, and how they can be tuned or modified. The comparison may concern training speed, predictive performance, stability, or sensitivity to hyperparameters. The models, tasks, and variants are open choices; the aim is to understand differences in behaviour and test explanations for them.

Compare optimisers on the same models and data, with comparable tuning effort and clearly stated training budgets. Distinguish improvements per training step from improvements per unit of computation or elapsed time. Specify which parameters receive Muon updates and how the remaining parameters are optimised.

Possible directions

  • Small language models and transformers. Compare Muon and selected variants with AdamW or other baselines. Do relative gains change with model size, architecture, batch size, or training duration?
  • Small MLP tasks. Investigate classification, regression, or a synthetic learning problem. Do gains observed in transformer training persist in simpler networks? Such settings may help isolate the role of conditioning, noise, or matrix shape.
  • Implementation from scratch. Derive and reimplement Muon and selected modern variants, such as NorMuon or Dion. Comparing a minimal implementation with a reference can clarify which details affect the update and its performance.
  • Tuning and modification. Explore learning rates, momentum, weight decay, update scaling, or parameter grouping. How sensitive are the methods to these choices, and can an observed limitation motivate a modification?
  • Comparing variants. Which benefits come from normalisation, compression, or error feedback? Does the additional complexity improve performance under the same computational and tuning budgets?
  • Mechanisms and numerical accuracy. Study update spectra, gradient noise, or the accuracy of the orthogonalisation. Does a more accurate matrix transformation improve training, and when is its additional cost justified?

These directions are optional; a different question arising from the literature or experiments is equally appropriate.

Reading starter kit