Efficient GPU servers, from architecture to full systems
Seongtae Bang
Ph.D. Student, DGIST · Member of CASLAB@Yonsei
Open to research internships and full-time roles in GPU systems and AI infrastructure, starting 2027.
I build efficient GPU servers by reducing memory pressure and data movement and exposing more CPU–GPU parallelism in large-scale AI workloads.
GPU memory and GPU cycles are the scarcest resources in an AI server. I find the data and work that do not need to occupy them, and restructure where and when that work runs so it proceeds alongside the GPU instead of stalling it.
Research Focus
GPU memory systems
Use the signals the GPU already produces to manage its memory.
When data outgrows GPU memory, existing systems react late and at coarse granularity. I turn information and resources the stack already has but ignores, such as how widely data is shared across GPU cores and the cores stalled by page faults, into better prefetching, placement, and reclamation.
Heterogeneous CPU–GPU execution
The CPU is a partner in GPU training, not a slow fallback.
Offloading systems usually treat the CPU path as slow and try to hide it. I measure where host-side time actually goes, which is mostly memory traffic rather than arithmetic, and reorder work by when its results are needed, so CPU work overlaps GPU compute instead of extending every step. I am now applying this to Mixture-of-Experts and 3D Gaussian Splatting training.
Workload-aware AI scaling
Only the active part of a model needs to live in GPU memory.
Mixture-of-Experts training uses only the routed experts, and 3D Gaussian Splatting touches only the Gaussians visible from the current view. I use this sparsity and locality to keep just the active working set in HBM and prepare the rest on the CPU in parallel, so models can grow past GPU memory and parallelism can be chosen for speed rather than for fit.
Architecture to full systems
Solve each bottleneck at the layer where it lives.
Some fixes can ship today in software (ARIADNE runs entirely inside NVIDIA’s open-source driver), others need new hardware behavior (ReclaimX, evaluated in a GPU simulator), and studying whole servers needs faithful simulators (pNet-gem5). I move between these layers and build the tools when they do not exist.
Across the system stack
- AI training systemsPyTorch, DeepSpeed, Megatron-LM, gsplat
- CPU SIMD kernelsC++, OpenMP, AVX2, AVX-512
- GPU driver & UVM runtimeNVIDIA open-source GPU driver, CUDA
- GPU microarchitectureAccel-Sim, GPGPU-Sim
- Full-system simulationgem5, Linux network driver