Thesis & Awards
This project constitutes the focus of my Master’s Thesis, recognized as Distinguished Project.
The full thesis is available to download as a PDF (8 MB).
Abstract
Modern times are distinguished as the Renaissance of Deep Learning, with remarkable advances across a wide range of domains and abundant research constantly pushing its boundaries.
As models and datasets increase in size, people move to distributed clusters, which pose new challenges to training. Distributed training systems usually use the Parallel Stochastic Gradient Descent algorithm to scale out training. This algorithm creates large network traffic and an intensive need for synchronising the entire system. The fundamental limitation is that they inevitably explode the batch size and force the user to use large batches for training, in order to reduce communication traffic and the intensive demands for synchronising the entire system.
Existing systems tend to adopt high-end specialised hardware and network such as InfiniBand, which eventually fail because of large network traffic and large clusters reaching hundreds of nodes. The hardware solutions are very expensive, do not scale and fundamentally, the system will suffer from large batch training, so the user may not be able to converge training when scaling out. In this project, I aim to design a system for Deep Learning that enables flexible synchronisation to address communication and large batch training solutions.
The system brings three new designs: it has an abstraction that allows the user to declare flexible types of synchronisation, much more complex than Parallel SGD, it has a high-performance communication system where workers exchange gradients and models to collaborate for training and it enables monitoring for network and training statistics, providing support for dynamic adaptation of synchronisation strategies. I develop two advanced synchronisation algorithms based on this new system.