ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), Heraklion, Greece, 2027
Prior CPU-offloaded training treats the host optimizer as slow compute and hides it. ReFlow shows the delay comes from host-memory traffic and removes it, reaching GPU-resident speed while the optimizer stays fully offloaded.
Up to 4× training throughput over ZeRO-Infinity
Public artifact · Upstream pull request to DeepSpeed in preparation
@inproceedings{bang2027reflow,
title = {{ReFlow}: Exposing Parallelism in {CPU}-Offloaded {LLM} Training via Register-Resident {SIMD} Fusion and Decoupled Update Scheduling},
author = {Bang, Seongtae and Park, Gyeongseo and Kang, Ki-Dong and Shin, Hyunkyun and Kim, Sungju and Kim, Daehoon},
booktitle = {Proceedings of the 32nd ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1 (ASPLOS '27)},
address = {Heraklion, Greece},
year = {2027},
note = {To appear}
}- Prior work
- CPU offloading frees GPU memory, but the GPU then waits for the CPU optimizer every step. Prior systems assume the CPU is simply slow, so they hide the delay: they move work back to the GPU, which costs GPU memory, or speculate and roll back, which can change the training result.
- Key insight
- The CPU is not slow at the math. Adam arithmetic is only about 16% of the host update; the rest is memory traffic from writing intermediates to DRAM, inflating gradients to FP32, and making the parameters the GPU needs wait behind the FP32 state write-back.
- Approach
- Instead of hiding the host path, ReFlow fixes it. Intermediates stay in CPU registers, gradients cross PCIe in compact BF16, and the parameters the GPU needs next are produced first while the FP32 state is written later, off the critical path.
- Result
- The optimizer stays fully offloaded and training stays identical, yet throughput matches GPU-resident ZeRO-3 (681 vs. 682 TFLOPS per GPU on OPT-30B with 4 B200 GPUs): up to 4× ZeRO-Infinity and 3× SuperOffload, with models up to 95B on 8 B200 GPUs.