The setup
Over the last year I have built three prototypes of reinforcement learning fine-tuning on top of ServerlessLLM, a Ray-based system for serving and fine-tuning LLMs with fast checkpoint loading:
- PPO with joint serving and training. A new fine-tuning backend that runs Proximal Policy Optimization while the same model keeps answering inference requests. Policy, frozen reference and reward model all live inside one backend.
- Real-time RL from streaming data. A design for feeding rewards from Kafka, Redis Streams or Deephaven into the trainer, so the model learns from production traffic and pushes updated weights back to serving without downtime.
- GRPO for code generation. Group Relative Policy Optimization where the policy writes code, Ray sandbox actors execute it against tests and the reward and reference models are served by ServerlessLLM behind its OpenAI-compatible API.
These are three different algorithms with three different data sources, but I reached the same conclusion for all of them: the cost of RL is dominated by generating text, not by training on it.
An RL step is mostly inference
Supervised fine-tuning reads a dataset and runs forward and backward passes. RL is different. Before it can compute a single gradient, it has to produce the data by sampling from the model it is training. Every PPO or GRPO iteration looks like this:
- Sample prompts.
- Generate one or more completions per prompt with the policy model.
- Score each completion with a reward model, or by running the code in a sandbox.
- Compute log-probabilities of the same completions under a frozen reference model for the KL penalty.
- Finally, one optimizer step on the policy.
Steps two to four are pure inference. Step five is the only part that resembles classic training and it is short because the batch is small and LoRA keeps the trainable parameters tiny.
GRPO increases this ratio by design. The group baseline needs several completions per prompt so that it can compare them against each other. My GRPO setup uses a group size of eight, so every task costs eight full generations plus eight reward evaluations plus eight reference passes before the trainer learns anything.
In online RL, slow inference affects correctness
The streaming design pushes this further. The goal is a loop where a user request is served, feedback arrives on a stream, the trainer takes a PPO step and the updated weights are pushed back into the serving instance. The design targets a full pipeline under one second, with the model update itself budgeted at a few hundred milliseconds.
If inference is slow, this loop fails in two ways. First, rollouts arrive slowly, so the policy learns from fewer samples per unit time and the reward signal is noisier. Second, serving keeps answering users with stale weights while the trainer catches up, so the model people talk to is not the model being optimized. In joint serving and training, where one GPU handles both, a slow generation path directly hurts user latency as well. For online RL, fast inference is therefore a requirement rather than an optimisation.
Why ServerlessLLM was the right base
I chose to build on a serving system rather than a training framework and all the reasons relate to inference:
- Fast checkpoint loading. RL juggles three models: policy, reference and reward. My GRPO trainer loads the policy through SLLM Store and reaches the other two over the inference API. Being able to load and swap multi-gigabyte checkpoints in seconds instead of minutes is what makes it practical to keep updating the policy and re-serving it.
- A real inference engine. The serving backends can use vLLM, with continuous batching and paged attention. Routing rollout generation through them instead of calling Hugging Face
generateonce per prompt is the single biggest speed-up available to the PPO backend and the GRPO actors already talk to that path. - Horizontal scale for the rest of the loop. Reward models, reference models and sandboxes are Ray actors. When the reward step is the bottleneck, adding actors is a configuration change rather than a rewrite.
Takeaways
- RL fine-tuning of LLMs is a rollout problem. Most wall-clock time goes to sampling, scoring and reference passes, not to gradients.
- Group-based methods like GRPO multiply the inference cost per prompt by the group size.
- For online RL, slow inference also makes the served model diverge from the trained one.
- Start from a serving system with fast loading and a batched inference engine. In my experience, the training loop is the easier part.
These are still prototypes. The next step is to move rollout generation onto the vLLM backend and measure the difference properly.