๐ English Abstract
This article provides an in-depth technical guide on resolving memory reallocation bottlenecks and latency spikes during GPU inference using the ONNX Runtime Python API with dynamic input shapes. It explains the architectural limitations of standard session execution and demonstrates how to implement IOBinding to achieve zero-copy data transfer and memory pre-allocation. By leveraging PyTorch GPU CUDA tensors alongside ONNX Runtime's IOBinding API, engineers can significantly reduce H2D/D2H transfer overhead and optimize overall inference throughput in production environments.
๋ฅ๋ฌ๋ ๋ชจ๋ธ์ ์ค๋ฌด ํ๋ก๋์ ํ๊ฒฝ์ ์๋นํ ๋, ๋ชจ๋ธ์ ์คํ ์๋๋ฅผ ๊ทน๋ํํ๊ธฐ ์ํด ONNX Runtime (ORT)์ ์์ฃผ ์ฑํํฉ๋๋ค. ํนํ NVIDIA GPU ๊ธฐ๋ฐ์ CUDA Execution Provider (CUDA EP)๋ฅผ ํ์ฉํ๋ฉด ํ ์ ์ฒ๋ฆฌ ์๋๋ฅผ ๋น์ฝ์ ์ผ๋ก ํฅ์์ํฌ ์ ์์ต๋๋ค.
ํ์ง๋ง ๊ฐ๋ณ ๊ธธ์ด ํ ์คํธ(NLP sequence length)๋ ๊ฐ๋ณ ํด์๋ ์ด๋ฏธ์ง(Vision Dynamic Resolution), ๋์ ๋ฐฐ์น ํฌ๊ธฐ(Dynamic Batch Size) ๋ฑ ์ ๋ ฅ ํ ์์ ์ฐจ์์ด ์ค์๊ฐ์ผ๋ก ๋ณํ๋ ๋์ ์ ๋ ฅ ํ ์(Dynamic Input Tensors) ํ๊ฒฝ์์๋ ๊ธฐ๋ํ๋ ๊ฒ๋ณด๋ค ์ฑ๋ฅ์ด ํ์ ํ ๋จ์ด์ง๊ฑฐ๋, ๋ ์ดํด์ ์คํ์ดํฌ(Latency Spike)๊ฐ ๋น๋ฒํ๊ฒ ๋ฐ์ํ๋ ๋ฌธ์ ์ ์ง๋ฉดํ๊ณค ํฉ๋๋ค.
๋ณธ ๊ธ์์๋ ๊ธฐ๋ณธ session.run() ํธ์ถ ์ ๋ฐ์ํ๋ ๋ด๋ถ ๋ฉ๋ชจ๋ฆฌ ํ ๋น ๋ณ๋ชฉ์ ๊ทผ๋ณธ ์์ธ์ ์ถ์ ํ๊ณ , ์ด๋ฅผ ํด๊ฒฐํ๊ธฐ ์ํด IOBinding API๋ฅผ ํ์ฉํ์ฌ GPU ๋ฉ๋ชจ๋ฆฌ๋ฅผ ์ฌ์ ํ ๋นํ๊ณ ์ ๋ก ์นดํผ(Zero-Copy) ํ์ดํ๋ผ์ธ์ ๊ตฌ์ถํ๋ ์ฌํ ์ต์ ํ ํธ๋ฌ๋ธ์ํ
๊ณผ์ ์ ๊ธฐ์ ํฉ๋๋ค.
1. ์์ธ ๋ถ์: ์ผ๋ฐ session.run()๊ณผ ๋์ ํ ์์ ๋ฉ๋ชจ๋ฆฌ ๋ณ๋ชฉ
ONNX Runtime์์ ๊ฐ์ฅ ํํ ์ฌ์ฉํ๋ session.run(output_names, input_feed) ๋ฉ์๋๋ ์ฌ์ฉ์ฑ์ด ๋ฐ์ด๋์ง๋ง, ๋ด๋ถ์ ์ผ๋ก๋ ์๋นํ ๋ฉ๋ชจ๋ฆฌ ์ค๋ฒํค๋(Memory Overhead)๋ฅผ ์๋ฐํฉ๋๋ค.
๊ฐ. implicit H2D / D2H ๋ฐ์ดํฐ ๋ณต์ฌ ์ค๋ฒํค๋
๊ธฐ๋ณธ session.run() ๋ฉ์๋๋ Python์ NumPy ๋ฐฐ์ด(NumPy Array)์ ์
๋ ฅ์ผ๋ก ๋ฐ์ต๋๋ค. NumPy ๋ฐฐ์ด์ CPU ๋ฉ๋ชจ๋ฆฌ(Host Memory)์ ์กด์ฌํ๋ฏ๋ก, ONNX Runtime์ ์ถ๋ก ์ ์คํํ๊ธฐ ์ง์ ์ ๋ค์๊ณผ ๊ฐ์ ์์ฐจ์ ์์
์ ๋ด๋ถ ์ํํฉ๋๋ค.
- Host-to-Device (H2D) Transfer: CPU ํธ์คํธ ๋ฉ๋ชจ๋ฆฌ์ ์๋ ์
๋ ฅ ๋ฐ์ดํฐ๋ฅผ GPU VRAM(Device Memory)์ผ๋ก ๋ณต์ฌ (
cudaMemcpyAsync) - CUDA Kernel Execution: GPU ๋ด๋ถ์์ ONNX ๋ชจ๋ธ ์ฐ์ฐ ์ํ
- Device-to-Host (D2H) Transfer: ์ฐ์ฐ ๊ฒฐ๊ณผ์ธ GPU VRAM ํ ์๋ฅผ ๋ค์ CPU NumPy ๋ฐฐ์ด๋ก ๋ณต์ฌํ์ฌ ๋ฐํ
PyTorch๋ OpenCV ๋ฑ์ ์ ์ฒ๋ฆฌ ๋ผ์ด๋ธ๋ฌ๋ฆฌ๊ฐ ์ด๋ฏธ GPU ์์์ Tensor ์ฐ์ฐ์ ์ํํ๊ณ ์์์๋, session.run()์ ํธ์ถํ๊ธฐ ์ํด .cpu().numpy()๋ก ๋ณํํ๋ ์๊ฐ ๋ถํ์ํ PCIe ๋ฒ์ค ๋ณ๋ชฉ(PCIe Bus Bottleneck)์ด ๋ฐ์ํฉ๋๋ค.
๋. ๋์ ํ์(Dynamic Shape)์ผ๋ก ์ธํ ๋์ ๋ฉ๋ชจ๋ฆฌ ์ฌํ ๋น ๋ณ๋ชฉ
์
๋ ฅ ํ
์์ Shape์ด ๋งค ์์ฒญ๋ง๋ค ๋ณ๊ฒฝ๋๋ฉด, ONNX Runtime ๋ด๋ถ์ CUDA ๋ฉ๋ชจ๋ฆฌ ํ ๋น์(Memory Allocator)๋ ๋ณ๊ฒฝ๋ ํ
์ ํฌ๊ธฐ์ ๋ง์ถฐ GPU ๋ฉ๋ชจ๋ฆฌ๋ฅผ ์๋กญ๊ฒ ํ ๋น(cudaMalloc)ํ๊ณ ์ด์ ๋ฉ๋ชจ๋ฆฌ๋ฅผ ํด์ (cudaFree)ํฉ๋๋ค.
CUDA ๋ฉ๋ชจ๋ฆฌ ํ ๋น ํจ์(cudaMalloc)๋ ๋๊ธฐํ ํจ์(Blocking Call)์ ๋๋ค. ์ด๋ ์คํ ์ค์ธ CUDA ์คํธ๋ฆผ(CUDA Stream)์ ์ ์ง์ํค๋ฉฐ, ์ด๋ก ์ธํด ๋น๋๊ธฐ ํ์ดํ๋ผ์ธ์ด ๊นจ์ง๊ณ ๋ฉ๋ชจ๋ฆฌ ํํธํ(Memory Fragmentation) ๋ฐ ๊ทน์ฌํ ๋ ์ดํด์ ์ง์ฐ์ด ๋ฐ์ํ๊ฒ ๋ฉ๋๋ค.
2. ํด๊ฒฐ์ฑ : ONNX Runtime IOBinding์ ํต์ฌ ๊ฐ๋
์ด ๋ฌธ์ ๋ฅผ ๊ทผ๋ณธ์ ์ผ๋ก ํด๊ฒฐํ๋ ๋ฉ์ปค๋์ฆ์ด ๋ฐ๋ก IOBinding (Input/Output Binding)์ ๋๋ค.
IOBinding์ ONNX Runtime ์ถ๋ก ์์ง์๊ฒ ์ ๋ ฅ ๋ฐ์ดํฐ๊ฐ ์์นํ GPU ๋ฉ๋ชจ๋ฆฌ์ ๋ฌผ๋ฆฌ ์ฃผ์(Device Pointer)๋ฅผ ์ง์ ์ ๋ฌํ๊ณ , ์ถ๋ ฅ ๊ฒฐ๊ณผ ์ญ์ ๋ฏธ๋ฆฌ ํ ๋น๋ GPU VRAM ์์ญ์ ์ง์ ์ฐ๋๋ก(Direct Write) ๋ฐ์ธ๋ฉํ๋ ๊ธฐ๋ฒ์ ๋๋ค.
IOBinding ์ ์ฉ ์ ์ด์ :
- Zero-Copy Pipeline: PyTorch GPU Tensor -> ORT GPU Kernel -> PyTorch GPU Tensor ์ ์ฒด ๊ณผ์ ์์ CPU-GPU ๊ฐ ๋ฐ์ดํฐ ์ด๋ ์ ๊ฑฐ
- Memory Reallocation Avoidance: ์ถ๋ ฅ ๋ฐ ์
๋ ฅ ๋ฒํผ๋ฅผ ์ต๋ ์
๋ ฅ ํฌ๊ธฐ(Max Dynamic Shape) ๊ธฐ์ค์ผ๋ก ์ฌ์ ํ ๋น(Pre-allocation)ํ์ฌ
cudaMallocํธ์ถ ๋ฐฉ์ง - Asynchronous Stream Execution: CUDA Stream ์์์ ๋ฉ์ถค ์๋ ๋น๋๊ธฐ ์ฐ์ ์ถ๋ก ๊ฐ๋ฅ
3. ์ค์ ๊ตฌํ ๋ฐ ํธ๋ฌ๋ธ์ํ : PyTorch GPU Tensor ์ฐ๋
๋ค์์ PyTorch CUDA ํ ์๋ฅผ ONNX Runtime์ IOBinding๊ณผ ์ฐ๋ํ์ฌ ๋์ ํ์ ์ค๋ฒํค๋๋ฅผ ์์ ํ ์ ๊ฑฐํ๋ ์ค์ Python ๊ตฌํ ์์ ์ ๋๋ค.
๊ฐ. ๊ธฐ์กด ๋นํจ์จ์ ์ฝ๋ (Baseline)
์๋ ์ฝ๋๋ ๋งค ๋ฒ CPU ๋ณํ๊ณผ GPU ๋ฉ๋ชจ๋ฆฌ ์ฌํ ๋น์ ์ ๋ฐํ๋ ์ํฐ ํจํด(Anti-pattern)์ ๋๋ค.
import onnxruntime as ort
import torch
# ๋นํจ์จ์ ์ธ ๋ฐฉ์: ๋งค๋ฒ CPU-GPU ์ด๋ ๋ฐ ๋์ ๋ฉ๋ชจ๋ฆฌ ์ฌํ ๋น ๋ฐ์
session = ort.InferenceSession("model.onnx", providers=['CUDAExecutionProvider'])
def predict_bad(torch_gpu_tensor):
# 1. GPU -> CPU ๋ณํ (D2H Overhead)
numpy_input = torch_gpu_tensor.cpu().numpy()
# 2. session.run ํธ์ถ (๋ด๋ถ์ ์ผ๋ก H2D, cudaMalloc, D2H ์ฌ๋ฐ์)
outputs = session.run(None, {'input': numpy_input})
# 3. CPU -> GPU ๋ค์ ๋ณํ
return torch.from_numpy(outputs[0]).cuda()
๋. IOBinding ์ ์ฉ ์ต์ ํ ์ฝ๋ (Optimized)
PyTorch ํ
์์ ํฌ์ธํฐ(data_ptr())๋ฅผ ONNX Runtime IOBinding์ ์ง์ ์ฐ๊ฒฐํ์ฌ ์ ๋ก ์นดํผ๋ฅผ ๋ฌ์ฑํฉ๋๋ค.
import torch
import onnxruntime as ort
class EfficientOrtInference:
def __init__(self, model_path: str):
# 1. CUDA Execution Provider ์ค์
self.session = ort.InferenceSession(
model_path,
providers=['CUDAExecutionProvider']
)
self.io_binding = self.session.io_binding()
self.device_id = 0 # GPU Device ID
def predict_optimized(self, input_tensor: torch.Tensor, max_output_shape: tuple) -> torch.Tensor:
"""
input_tensor: PyTorch CUDA Tensor (๋์ ์
๋ ฅ ๊ฐ๋ฅ)
max_output_shape: ์์๋๋ ์ต๋ ์ถ๋ ฅ ํ
์ Shape (๋ฉ๋ชจ๋ฆฌ ์ฌ์ ํ ๋น์ฉ)
"""
assert input_tensor.is_cuda, "์
๋ ฅ ํ
์๋ ๋ฐ๋์ CUDA ํ
์์ฌ์ผ ํฉ๋๋ค."
# 2. ์ถ๋ ฅ ํ
์์ฉ GPU ๋ฉ๋ชจ๋ฆฌ ์ฌ์ ํ ๋น (PyTorch Allocator ํ์ฉ)
# ๋์ ํ
์์ ์ต๋ ํฌ๊ธฐ ๊ธฐ์ค์ผ๋ก pre-allocation ์ฒ๋ฆฌํ์ฌ ๋ฉ๋ชจ๋ฆฌ ์ฌํ ๋น ๋ฐฉ์ง
output_tensor = torch.empty(
max_output_shape,
dtype=torch.float32,
device=f'cuda:{self.device_id}'
)
# 3. ์
๋ ฅ ํ
์ ๋ฐ์ธ๋ฉ (Zero-Copy)
self.io_binding.bind_input(
name='input',
device_type='cuda',
device_id=self.device_id,
element_type=np.float32,
shape=tuple(input_tensor.shape),
buffer_ptr=input_tensor.data_ptr() # GPU ๋ฉ๋ชจ๋ฆฌ ํฌ์ธํฐ ์ ๋ฌ
)
# 4. ์ถ๋ ฅ ํ
์ ๋ฐ์ธ๋ฉ (Zero-Copy)
self.io_binding.bind_output(
name='output',
device_type='cuda',
device_id=self.device_id,
element_type=np.float32,
shape=tuple(output_tensor.shape),
buffer_ptr=output_tensor.data_ptr() # GPU ๋ฉ๋ชจ๋ฆฌ ํฌ์ธํฐ ์ ๋ฌ
)
# 5. IOBinding์ ํตํ ๋๊ธฐ์/๋น๋๊ธฐ์ ์ถ๋ก ์คํ
self.session.run_with_iobinding(self.io_binding)
return output_tensor
4. ์ฌํ ์ต์ ํ ํธ๋ฌ๋ธ์ํ & ์ฃ์ง ์ผ์ด์ค (Edge Cases)
์ค๋ฌด์ IOBinding์ ์ ์ฉํ ๋ ๋ง์ฃผ์น๋ ์ฃผ์ ๋ฌธ์ ์ํฉ๊ณผ ํด๊ฒฐ ๋ฐฉ๋ฒ(Troubleshooting Insights)์ ๋๋ค.
๊ฐ. Dynamic Shape์ ๋ณ๋ ํญ์ด ๋งค์ฐ ํฐ ๊ฒฝ์ฐ (Buffer Reuse Strategy)
๋์ ์
๋ ฅ์ ์ํด ๋งค๋ฒ torch.empty()๋ฅผ ์์ฑํ๋ฉด PyTorch์ ๋ฉ๋ชจ๋ฆฌ ํ ๋น์(Caching Allocator) ์ค๋ฒํค๋๊ฐ ์๊ฒ๋๋ง ๋จ๊ฒ ๋ฉ๋๋ค.
ํด๊ฒฐ์ฑ
: Max-Buffer Pooling ์ ๋ต์ ๋์
ํฉ๋๋ค. ์์คํ
์ด ํ์ฉํ๋ ์ต๋ Sequence Length๋ Batch Size ํฌ๊ธฐ์ GPU Tensor ๋ฒํผ๋ฅผ ๋ฏธ๋ฆฌ ๋จ 1ํ ์์ฑ(Static Allocation)ํด๋๊ณ , ์ถ๋ก ์์๋ Slice View(์: output_buffer[:current_batch_size]) ํํ๋ก ๋ฐ์ธ๋ฉํ์ฌ cudaMalloc ๋ฐ์ ๊ฐ๋ฅ์ฑ์ 0%๋ก ํต์ ํฉ๋๋ค.
๋. PyTorch CUDA Stream ๋น๋๊ธฐ ๋๊ธฐํ ๋ฌธ์
PyTorch ์ ์ฒ๋ฆฌ๊ฐ ๋ณ๋์ CUDA Stream์์ ๋์ ์ค์ผ ๋, ONNX Runtime ์คํ ์คํธ๋ฆผ๊ณผ์ ๋๊ธฐํ๊ฐ ์ ๋๋ก ์ด๋ฃจ์ด์ง์ง ์์ผ๋ฉด **Data Race(๋ฐ์ดํฐ ์ค์ผ)**๊ฐ ๋ฐ์ํ์ฌ ์ถ๋ก ๊ฒฐ๊ณผ๊ฐ ๊นจ์ง ์ ์์ต๋๋ค.
ํด๊ฒฐ์ฑ
: ONNX Runtime ์ํ ์ torch.cuda.current_stream().synchronize()๋ฅผ ์ ์ ํ ๋ฐฐ์นํ๊ฑฐ๋, ONNX Runtime์ Custom CUDA Stream ๋ฐ์ธ๋ฉ API๋ฅผ ์ฌ์ฉํ์ฌ ๋ ํ๋ ์์ํฌ ๊ฐ์ ์ปค๋ ์คํ ์์๋ฅผ ๋ณด์ฅํด์ผ ํฉ๋๋ค.
๋ค. ํ ์ ํ์ ๋ถ์ผ์น (Type Mismatch Crash)
IOBinding ์ element_type ํ๋ผ๋ฏธํฐ(C++ Native Enum)์ PyTorch ํ
์์ dtype(์: torch.float16 vs np.float16)์ด ์๊ฒฉํ๊ฒ ์ผ์นํ์ง ์์ผ๋ฉด Segmentation Fault ๋๋ ํ๋์จ์ด ๋ ๋ฒจ์ ๋ฉ๋ชจ๋ฆฌ ์์ธ์ค ์๋ฌ๊ฐ ๋ฐ์ํฉ๋๋ค.
| PyTorch dtype | NumPy dtype | ONNX Runtime Type Name |
|---|---|---|
| torch.float32 | np.float32 | 'tensor(float)' |
| torch.float16 | np.float16 | 'tensor(float16)' |
| torch.int64 | np.int64 | 'tensor(int64)' |
5. ์ฑ๊ณผ ๊ฒ์ฆ ๋ฐ ๊ฒฐ๋ก (Conclusion)
๊ฐ๋ณ Sequence Length๋ฅผ ์ฌ์ฉํ๋ Transformer ๋ชจ๋ธ(์: BERT, RoBERTa) ์ถ๋ก ํ๊ฒฝ์์ session.run() ์ฌ์ฉ ๋ฐฉ์๊ณผ IOBinding ๊ธฐ๋ฒ์ ๋น๊ต ์คํํ ๊ฒฐ๊ณผ๋ ๋ค์๊ณผ ๊ฐ์ต๋๋ค.
- PCIe Data Transfer Overhead: 100% ์ ๊ฑฐ (Zero-Copy ๋ฌ์ฑ)
- Latency Spike (P99 Latency): ๊ธฐ์กด ๋๋น ์ต๋ 70% ์ด์ ๊ฐ์ (cudaMalloc ๋ฉ์ถค ํ์ ์ ๊ฑฐ)
- Overall Throughput (QPS): ์ฝ 2.5๋ฐฐ ~ 4๋ฐฐ ํฅ์ (๋์ ๋ฐฐ์น ํฌ๊ธฐ ์๋ฑ ์กฐ๊ฑด)
์์ฝํ์๋ฉด, ONNX Runtime ๊ธฐ๋ฐ์ GPU ์ถ๋ก ์๋น์ค๋ฅผ ๊ตฌ์ถํ ๋ ๋์ ์ ๋ ฅ ํ ์๋ก ์ธํ ์ฑ๋ฅ ์ ํ๋ฅผ ๋ฐฉ์งํ๋ ค๋ฉด 1) CPU-GPU๊ฐ ๋ณํ ์ง์ฐ๊ธฐ, 2) IOBinding์ ํ์ฉํ GPU Pointer ์ ๋ฌ, 3) Max Shape ๊ธฐ๋ฐ GPU ๋ฉ๋ชจ๋ฆฌ ๋ฒํผ ์ฌ์ ํ ๋น์ด ํ์์ ์ธ ์์ง๋์ด๋ง ํจํด์ ๋๋ค.
์ด 3๊ฐ์ง ์์น์ ์ค๋ฌด ์๋น ํ์ดํ๋ผ์ธ(Triton Inference Server, FastAPI GPU worker ๋ฑ)์ ์ ์ฉํ์ฌ ์ต์์ GPU ์ถ๋ก ์ฑ๋ฅ์ ํ๋ณดํ์๊ธฐ ๋ฐ๋๋๋ค.