Serverless platforms bill per second of inference, but loading a large model before the first token can take tens of seconds. This project follows that cold start end to end. It starts on a laptop, taking a 2 GiB checkpoint apart and measuring every stage until a loader-friendly checkpoint format logically follows from profile measurements in a bottom-up approach. Then it walks through ServerlessLLM (OSDI ‘24), which I co-authored: using the idle storage inside GPU servers as a checkpoint cache, live-migrating inference by moving tokens instead of gigabytes and scheduling for startup time. It ends with where this goes next, from agents where every step is a cold start to RL fine-tuning, where generation dominates the cost.