Publications

Papers on GPU memory systems, CPU–GPU execution, and simulation infrastructure. Open an overview to see each paper’s key insight, how it differs from prior work, and how it works.

2027

ASPLOS2027, to appearFirst author

Seongtae Bang, Gyeongseo Park, Ki-Dong Kang, Hyunkyun Shin, Sungju Kim, Daehoon Kim

ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), Heraklion, Greece, 2027

Prior CPU-offloaded training treats the host optimizer as slow compute and hides it. ReFlow shows the delay comes from host-memory traffic and removes it, reaching GPU-resident speed while the optimizer stays fully offloaded.

Up to 4× training throughput over ZeRO-Infinity

Public artifact · Upstream pull request to DeepSpeed in preparation

  • LLM Training
  • CPU Offload
  • SIMD
  • Optimizer
ArtifactProject

2026

MICRO2026, to appearFirst author

Seongtae Bang, Hyunkyun Shin, Hyungwon Park, Minho Kim, Daehoon Kim

IEEE/ACM International Symposium on Microarchitecture (MICRO), 2026

Prior UVM work frees GPU memory from the host after a fault has stalled the GPU. ReclaimX lets the stalled GPU cores do the reclamation themselves: earlier, at finer granularity, and without taking compute from running work.

2.33× geomean speedup over baseline UVM

  • GPU Architecture
  • UVM
  • Memory Management
  • Accel-Sim
Project
IEEE CAL2026First author

Seongtae Bang, Gyeongseo Park, Kyeonghyeon Ryu, Daehoon Kim

IEEE Computer Architecture Letters, vol. 25, no. 1, pp. 142–145, 2026

Offloaded optimizers finish the whole update, including the FP32 write-back, before the GPU can continue. ReplayOpt reorders the step by deadline so only the parameters the GPU needs stay on the critical path.

Up to 21.7% shorter training step

  • LLM Training
  • CPU Offload
  • Optimizer
  • SIMD
DOIProject
HPCA2026

Hyunkyun Shin, Seongtae Bang, Hyungwon Park, Daehoon Kim

IEEE International Symposium on High-Performance Computer Architecture (HPCA), Sydney, Australia, 2026

Prior fixes for UVM oversubscription help little or require hardware and compiler changes. ARIADNE decides at runtime, inside the driver, which regions to migrate and which to read in place, so any existing GPU binary runs fast.

5.0× average speedup at 175% oversubscription

  • GPU Memory
  • NVIDIA GPU Driver
  • UVM
  • Oversubscription
DOIArtifactProject

2025

IEEE CAL2025

Jongmin Shin, Seongtae Bang, Gyeongseo Park, Daehoon Kim

IEEE Computer Architecture Letters, vol. 24, no. 2, pp. 193–196, 2025

Prior gem5 networking models a single-queue NIC limited to a few Gbps, and the DPDK-based alternative covers only userspace networking. pNet-gem5 models multi-queue NICs with per-queue interrupts, so kernel-networked servers can be simulated at tens of Gbps.

Simulated networking up to 46 Gbps

  • gem5
  • Full-System Simulation
  • High-Performance Networking
  • Linux Driver
DOICodeProject
IEEE CAL2025Co-first author

Hyunkyun Shin*, Seongtae Bang*, Hyungwon Park, Daehoon Kim

IEEE Computer Architecture Letters, vol. 24, no. 1, pp. 117–120, 2025

Prior UVM prefetching applies one setting everywhere and struggles with irregular workloads. SAFE uses a signal the GPU already has, how many SMs share each block, to prefetch aggressively only where it pays off.

Up to 6.5× over the default UVM prefetcher

  • GPU Memory
  • NVIDIA GPU Driver
  • UVM
  • Prefetching
DOIProject

* Equal contribution.