Training small language models under computational constraints
Project goal
Train a small language model from randomly initialised weights and investigate how to obtain the best performance within a fixed compute budget per training run. State the following at the beginning of your report:
- Your compute budget. Choose and declare the budget before running experiments. Specify the hardware and the amount of computation allowed for each training run.
- Your initialisation. Every model must start from scratch. Do not use pretrained model weights, embeddings, or other pretrained neural components. Existing architectures, implementations, and tokenizers are allowed.
- Your evaluation objective. State your evaluation data and metrics, and explain why they are appropriate.
You are free to choose the dataset, tokenizer, architecture, model size, optimiser, and training procedure. Investigate the choices that interest you and explain how they affect performance under your declared budget.
Declaring the budget
For example, you could allow 5 hours of training on a single GPU, specifying its model. A smaller or larger budget, CPU training time, or a measured computational allowance such as FLOPs is also possible.
The budget applies to each complete training run used in your final comparisons. Separate preliminary and tuning experiments, as well as tokenizer training, are excluded. Resuming a checkpoint does not reset the budget.
Designing the model and training
A larger model may process fewer training examples within the same time, while a smaller model may complete more updates or passes through the data but plateau earlier in terms of performance. The dataset, tokenizer, architecture, and implementation all affect this allocation. Possible questions include:
- Model size and training duration. Does a smaller model trained for longer outperform a larger one? How should parameters be allocated across depth and width?
- Data. Would a smaller, carefully selected corpus be more useful than a larger collection? How do repeated examples or mixtures of sources affect the result?
- Tokenization and context. How do the vocabulary and context length affect prediction, sequence length, memory use, and training speed?
- Optimisation and efficiency. Can changes to the optimiser, batch size, learning-rate schedule, numerical precision, or implementation improve performance within the same allowance?
Choosing evaluation criteria
Choose evaluation data that represent the behaviour you want to study, whether broad language modelling or a particular domain. Held-out predictive loss, perplexity, or performance on a specified task may be appropriate. Explain what your chosen metrics measure and what they leave out.
Use validation data for tuning and a separate test set for final evaluation. Compare configurations on the same test data. Per-token losses and perplexities are not directly comparable across tokenizers; consider a common normalisation such as bits per byte on the same text.
Reading starter kit
- Training Compute-Optimal Large Language Models. Allocation between model size and training data.
- Scaling Data-Constrained Language Models. Limited datasets and repeated training examples.
- nanoGPT. Karpathy’s minimal GPT training implementation.
- The Smol Training Playbook A guide for efficiently training small language models.