Parallel SGD scales training out by exploding the batch size and synchronising every worker at every step and specialised networks only postpone the problem. KungFu, built with the Large Scale Data & Systems Group at Imperial College London and published at OSDI ‘20, lets the user declare how workers synchronise, monitor gradient and network statistics cheaply and change synchronisation strategy or parallelism at runtime. It started as my Master’s thesis, recognised as a Distinguished Project and was presented at the SOSP ‘19 AI Systems Workshop.