인텔 연구진이 파이토치 기본 기능만으로 BERT류 모델을 서버 CPU에서 최대 5.8배 빠르게 만들었다
arXiv:2608.181822026-08-20
Efficient INT8 Inference of Small NLP Models on Server CPUs with PyTorch Native Stack
인텔 연구진이 파이토치 기본 기능만으로 BERT류 모델을 서버 CPU에서 최대 5.8배 빠르게 만들었다
대형 언어모델이 주목받는 시대에도 BERT 계열의 작은 자연어처리 모델은 분류, 순위매기기, 검색 같은 실무 작업에 여전히 널리 쓰인다. 이 연구는 모델 숫자를 8비트 정수(INT8)로 압축해 속도를 높이는 SmoothQuant 기법을 별도 외부 도구 없이 파이토치 자체 컴파일러 스택에 직접 통합했다. 그 결과 인텔 제온 서버 CPU에서 정확도 손실은 거의 없이 처리량을 최대 5.8배까지 끌어올렸다.
METAL MEDIA 해설 도표
인텔 연구진이 파이토치 기본 기능만으로 BERT류 모델을 서버 CPU에서 최대 5.8배 빠르게 만들었다
01활성화값의 계산 난이도를 가중치 쪽으로 옮겨 INT8 압축을 쉽게 만드는 SmoothQuant 기법을, 파이토치 공식 양자화 라이브러리인 TorchAO에 통합해 외부 도구 없이 바로 쓸 수 있게 했다
02파이토치 컴파일러인 TorchInductor에 그래프 융합 기능을 추가해, INT8 행렬곱과 스케일 조정·편향 더하기 같은 후속 연산을 하나의 효율적인 연산으로 합치고 런타임에서 데이터 배치를 바꾸는 낭비를 없앴다
03AVX512_VNNI와 AMX(정수 연산을 빠르게 처리하는 인텔 CPU 전용 명령어 세트)를 활용한 저수준 커널을 직접 구현해, 구형 Ice Lake 칩과 신형 Granite Rapids 칩 모두에서 최적 속도를 내도록 했다
04특정 명령어 지원이 없는 Ice Lake CPU에서도 INT8 연산이 효율적으로 돌아가도록 s8s8에서 u8s8로 바꾸는 변환 기법을 고안했다
05BERT-large, DistilBERT, XLM-RoBERTa로 실험한 결과, 기존 32비트 부동소수점(FP32) 대비 Ice Lake에서는 처리량이 1.9~2.6배, Granite Rapids에서는 4.2~5.8배 향상됐고, SQuAD와 MultiNLI 과제에서 정확도 손실은 1% 미만이었다
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.
무엇을 했나
활성화값의 계산 난이도를 가중치 쪽으로 옮겨 INT8 압축을 쉽게 만드는 SmoothQuant 기법을, 파이토치 공식 양자화 라이브러리인 TorchAO에 통합해 외부 도구 없이 바로 쓸 수 있게 했다
파이토치 컴파일러인 TorchInductor에 그래프 융합 기능을 추가해, INT8 행렬곱과 스케일 조정·편향 더하기 같은 후속 연산을 하나의 효율적인 연산으로 합치고 런타임에서 데이터 배치를 바꾸는 낭비를 없앴다
AVX512_VNNI와 AMX(정수 연산을 빠르게 처리하는 인텔 CPU 전용 명령어 세트)를 활용한 저수준 커널을 직접 구현해, 구형 Ice Lake 칩과 신형 Granite Rapids 칩 모두에서 최적 속도를 내도록 했다
특정 명령어 지원이 없는 Ice Lake CPU에서도 INT8 연산이 효율적으로 돌아가도록 s8s8에서 u8s8로 바꾸는 변환 기법을 고안했다
BERT-large, DistilBERT, XLM-RoBERTa로 실험한 결과, 기존 32비트 부동소수점(FP32) 대비 Ice Lake에서는 처리량이 1.9~2.6배, Granite Rapids에서는 4.2~5.8배 향상됐고, SQuAD와 MultiNLI 과제에서 정확도 손실은 1% 미만이었다
Figure 1: SmoothQuant workflow built on TorchAO and TorchInductor. TorchAO prepares and quantizes the model, TorchInductor fuses INT8 subgraphs, and the CPP backend selects between oneDNN and template-based GEMM kernels. C++ code is generated and compiled for deployment.
Table 1: Hardware platforms used for baseline collection.
Xeon Generation
Codename
Model
Physical Cores
Instruction Support
3rd Gen
Ice Lake
8358
64
AVX512F, AVX512_VNNI
6th Gen
Granite Rapids
6980P
256
AVX512F, AVX512_VNNI, AMX
Table 2: Baseline collection setup.
Item
Value
Model
BERT-large
Input shape
Batch size 1, sequence length 256
Precision modes
FP32 on both platforms; BF16 on Granite Rapids
Launch strategy
Multi-instance with weight-sharing on each socket
Cores per instance
4
Number of instances
Number of physical cores / 4
Deployment path
AOTI enabled
Input data
Synthetic
Table 3: Baseline hotspot breakdown for the BERT-large analysis case under the same multi-instance setup used for throughput benchmarking. Operator shares are aggregated from representative profiler logs. The linear-block consists of GEMM and fused post-operations such as bias addition, activation, and residual addition.
Baseline
Linear-Block
Attention
LayerNorm
Other
FP32 on Ice Lake
88.5%
10.4%
0.6%
0.5%
FP32 on Granite Rapids
86.9%
11.8%
0.8%
0.5%
BF16 on Granite Rapids
72.70%
19.75%
5.99%
1.56%
Table 4: Estimated peak compute and measured memory bandwidth per socket, balance point for different data types on the evaluated platforms.
Platform
FP32 Throughput (TFLOPS)
BF16 Throughput (TFLOPS)
INT8 Throughput (TOPS)
Memory Bandwidth (GB/s)
FP32 Balance Point (FLOPs/byte)
BF16 Balance Point (FLOPs/byte)
INT8 Balance Point (OPs/byte)
Ice Lake
6.76
N/A
27.03
175.00
38.63
N/A
154.46
Granite Rapids
19.66
262.14
524.29
700.00
28.09
374.49
748.99
Table 5: BERT-large GEMM shapes and their arithmetic intensity.
Arithmetic Intensity (ops/byte)
GEMM Shape M × K × N
Occurrences per forward pass
FP32
BF16
INT8
256×1024×1024
96
85.33
146.29
227.56
256×1024×4096
24
97.52
163.84
248.24
256×4096×1024
24
97.52
186.18
341.33
Table 6: Effective memory bandwidth and balance point for INT8 GEMM on Granite Rapids.
GEMM Shape
Effective Memory Bandwidth (GB/s)
Arithmetic Intensity (OPs/byte)
Balance Point (OPs/byte)
256×1024×1024
807.69
227.56
649.12
256×1024×4096
819.15
248.24
640.04
256×4096×1024
875.00
341.33
599.19
Table 7: Estimated GEMM throughput and speedup on Granite Rapids.
GEMM Shape
FP32 Throughput (TFLOPS)
BF16 Throughput (TFLOPS)
INT8 Throughput (TOPS)
INT8 vs FP32 Speedup
INT8 vs BF16 Speedup
256×1024×1024
19.66
102.40
183.79
9.35×
1.79×
256×1024×4096
19.66
114.69
203.35
10.34×
1.77×
256×4096×1024
19.66
130.33
298.67
15.19×
2.29×
Weighted Average
19.66
114.69
218.87
11.13×
1.91×
Table 8: Models used for the performance benchmark.
Model
Model ID
BERT-large
bert-large-uncased
DistilBERT
distilbert-base-uncased
XLM-RoBERTa
xlm-roberta-base
Table 9: Task-finetuned checkpoints used for accuracy evaluation. The model marked by * is finetuned on SQuAD 2.0 but still workable for the test on SQuAD 1.1.
Table 10: Quantization methods used in the experiments.
Method
Short Name
No quantization, no AMP
FP32
No quantization, with AMP
BF16
SmoothQuant, static quantization
Smooth-Static
SmoothQuant, dynamic quantization
Smooth-Dynamic
Table 11: INT8 speedup over FP32 on Ice Lake.
Model
Method
Throughput Speedup
Linear-Block Speedup
BERT-large
Smooth-Dynamic
2.22×
3.59×
BERT-large
Smooth-Static
2.58×
3.87×
DistilBERT
Smooth-Dynamic
1.90×
3.31×
DistilBERT
Smooth-Static
2.30×
3.35×
XLM-RoBERTa
Smooth-Dynamic
1.86×
3.05×
XLM-RoBERTa
Smooth-Static
2.17×
3.71×
Table 12: INT8 speedup over FP32 on Granite Rapids.
Model
Method
Throughput Speedup
Linear-Block Speedup
BERT-large
Smooth-Dynamic
5.02×
9.92×
BERT-large
Smooth-Static
5.82×
8.67×
DistilBERT
Smooth-Dynamic
4.24×
7.60×
DistilBERT
Smooth-Static
5.37×
8.84×
XLM-RoBERTa
Smooth-Dynamic
4.17×
7.58×
XLM-RoBERTa
Smooth-Static
5.41×
8.71×
Table 13: INT8 speedup over BF16 on Granite Rapids.
Model
Method
Throughput Speedup
Linear-Block Speedup
BERT-large
Smooth-Dynamic
1.37×
1.84×
BERT-large
Smooth-Static
1.59×
1.61×
DistilBERT
Smooth-Dynamic
0.98×
1.45×
DistilBERT
Smooth-Static
1.24×
1.69×
XLM-RoBERTa
Smooth-Dynamic
0.99×
1.50×
XLM-RoBERTa
Smooth-Static
1.29×
1.67×
Table 14: Realized speedup of BERT-large over estimated speedup ceiling on both evaluated platforms.
Throughput Speedup
Linear-Block Speedup
Platform
Method
Baseline
Realized
Estimated
Gap
Realized
Estimated
Gap
Ice Lake
Smooth-Dynamic
FP32
2.22×
3.15×
29.59%
3.59×
4.00×
10.22%
Ice Lake
Smooth-Static
FP32
2.58×
3.15×
18.17%
3.87×
4.00×
3.21%
Granite Rapids
Smooth-Dynamic
FP32
5.02×
7.06×
28.89%
9.92×
11.13×
10.89%
Granite Rapids
Smooth-Static
FP32
5.82×
7.06×
17.56%
8.67×
11.13×
22.12%
Granite Rapids
Smooth-Dynamic
BF16
1.37×
1.53×
10.41%
1.84×
1.91×
3.58%
Granite Rapids
Smooth-Static
BF16
1.59×
1.53×
-3.98%
1.61×
1.91×
15.64%
Table 15: Accuracy results on Granite Rapids.
Model
Metric
FP32
BF16
Smooth-Static
Alpha
Loss
Smooth-Dynamic
Alpha
Loss
BERT-large
SQuAD F1
92.64
92.67
92.66
0.6
0.0%
92.63
0.6
0.0%
BERT-large
SQuAD EM
86.41
86.45
86.49
0.6
-0.1%
86.44
0.6
0.0%
BERT-large
MultiNLI Accuracy
64.55
64.58
64.73
0.7
-0.3%
64.52
0.6
0.0%
DistilBERT
SQuAD F1
86.39
86.41
86.00
0.7
0.5%
86.40
0.6
0.0%
DistilBERT
SQuAD EM
79.02
79.04
78.49
0.7
0.7%
79.04
0.6
0.0%
DistilBERT
MultiNLI Accuracy
82.21
82.16
82.26
0.8
-0.1%
82.17
0.6
0.0%
XLM-RoBERTa
SQuAD F1
76.43
76.34
76.47
0.55
-0.1%
76.59
0.6
-0.2%
XLM-RoBERTa
SQuAD EM
69.95
69.89
70.04
0.55
-0.1%
70.30
0.6
-0.5%
XLM-RoBERTa
MultiNLI Accuracy
82.26
82.23
82.31
0.7
-0.1%
82.26
0.6
0.0%
왜 중요한가
많은 기업이 검색 순위 매기기, 텍스트 분류 같은 일상적인 작업에는 여전히 저렴한 CPU에서 돌아가는 작은 NLP 모델을 쓴다. 이 기술이 별도 플러그인 없이 파이토치 표준 기능만으로 제공되면, 어떤 팀이든 코드를 크게 바꾸지 않고도 몇 배 빠르고 저렴한 추론을 얻을 수 있다.
이 논문의 용어
INT8 양자화 · 숫자를 32비트 대신 8비트 정수로 저장·계산해 속도는 빨라지지만 정확도가 떨어질 수 있는 압축 기법
SmoothQuant · 활성화값과 가중치 사이의 수치 균형을 조정해 INT8 압축 시 정확도 손실을 줄이는 기법
TorchAO · 모델에 양자화 등 저정밀 기법을 적용하는 파이토치 공식 라이브러리
TorchInductor · 파이토치 2에서 모델 연산 그래프를 최적화하고 빠른 코드를 생성하는 컴파일러 백엔드
GEMM · 신경망 계층 내부에서 반복되는 핵심 연산인 행렬곱셈
AVX512_VNNI / AMX · 인텔 제온 CPU에 내장된, INT8 정수 연산 속도를 높이는 전용 명령어 세트
루프라인 모델 · 연산이 프로세서 속도에 의해 제한되는지 메모리 대역폭에 의해 제한되는지를 분석하는 성능 평가 방법