Efficient GPU servers, from architecture to full systems

Seongtae Bang

Ph.D. Student, DGIST · Member of CASLAB@Yonsei

Open to research internships and full-time roles in GPU systems and AI infrastructure, starting 2027.

I build efficient GPU servers by reducing memory pressure and data movement and exposing more CPU–GPU parallelism in large-scale AI workloads.

GPU memory and GPU cycles are the scarcest resources in an AI server. I find the data and work that do not need to occupy them, and restructure where and when that work runs so it proceeds alongside the GPU instead of stalling it.

Seongtae Bang

Publications

Full list
ASPLOS2027, to appearFirst author

ReFlow

Prior systems hide the slow CPU update; ReFlow removes its real cause, host-memory traffic.

Training throughput, up to
vs ZeRO-Infinity4.0×
vs SuperOffload3.0×

Public artifact · Upstream pull request to DeepSpeed in preparation

  • LLM Training
  • CPU Offload
  • SIMD
  • Optimizer
MICRO2026, to appearFirst author

ReclaimX

A page fault stalls a GPU core anyway, so ReclaimX lets that stalled core free GPU memory.

Geomean speedup over baseline UVM
ReclaimX2.33×
with prefetching3.61×
  • GPU Architecture
  • UVM
  • Memory Management
  • Accel-Sim
Project
IEEE CAL2026First author

ReplayOpt

The GPU needs only the new parameters to continue, so ReplayOpt sends those first and writes the FP32 state later.

Up to 21.7% shorter training step

  • LLM Training
  • CPU Offload
  • Optimizer
  • SIMD
DOIProject
HPCA2026

ARIADNE

Fast UVM under heavy memory oversubscription from the driver alone, with no hardware, compiler, or application changes.

Average speedup over prior state of the art
130% oversub.1.9×
175% oversub.5.0×
300% oversub.4.8×
  • GPU Memory
  • NVIDIA GPU Driver
  • UVM
  • Oversubscription

ASPLOS, MICRO, and HPCA are top-tier conferences in computer architecture.

Research Focus

GPU memory systems

Use the signals the GPU already produces to manage its memory.

When data outgrows GPU memory, existing systems react late and at coarse granularity. I turn information and resources the stack already has but ignores, such as how widely data is shared across GPU cores and the cores stalled by page faults, into better prefetching, placement, and reclamation.

Heterogeneous CPU–GPU execution

The CPU is a partner in GPU training, not a slow fallback.

Offloading systems usually treat the CPU path as slow and try to hide it. I measure where host-side time actually goes, which is mostly memory traffic rather than arithmetic, and reorder work by when its results are needed, so CPU work overlaps GPU compute instead of extending every step. I am now applying this to Mixture-of-Experts and 3D Gaussian Splatting training.

Workload-aware AI scaling

Only the active part of a model needs to live in GPU memory.

Mixture-of-Experts training uses only the routed experts, and 3D Gaussian Splatting touches only the Gaussians visible from the current view. I use this sparsity and locality to keep just the active working set in HBM and prepare the rest on the CPU in parallel, so models can grow past GPU memory and parallelism can be chosen for speed rather than for fit.

Architecture to full systems

Solve each bottleneck at the layer where it lives.

Some fixes can ship today in software (ARIADNE runs entirely inside NVIDIA’s open-source driver), others need new hardware behavior (ReclaimX, evaluated in a GPU simulator), and studying whole servers needs faithful simulators (pNet-gem5). I move between these layers and build the tools when they do not exist.

Across the system stack

  1. AI training systemsPyTorch, DeepSpeed, Megatron-LM, gsplat
    ReplayOptIEEE CAL 2026ReFlowASPLOS 2027MoE trainingongoing3DGS trainingongoing
  2. CPU SIMD kernelsC++, OpenMP, AVX2, AVX-512
    ReplayOptIEEE CAL 2026ReFlowASPLOS 2027MoE trainingongoing3DGS trainingongoing
  3. GPU driver & UVM runtimeNVIDIA open-source GPU driver, CUDA
  4. GPU microarchitectureAccel-Sim, GPGPU-Sim
  5. Full-system simulationgem5, Linux network driver
Contact

st.bang@dgist.ac.kr