수백만 개 흩어진 3D점을 미리 손대지 않고도 AI가 학습할 수 있는 격자 데이터로 바꾸는 방법
arXiv:2608.179882026-08-17
GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation
수백만 개 흩어진 3D점을 미리 손대지 않고도 AI가 학습할 수 있는 격자 데이터로 바꾸는 방법
드론이나 항공 사진으로 만든 3D 공중 장면은 3D Gaussian Splatting(3DGS)이라는 방식으로 표현되는데, 한 장면에 점이 300만 개 넘게 무질서하게 흩어져 있어 그대로는 생성 AI가 학습하기 어렵다. GS-Voxel은 이 점들을 장면별로 다시 최적화하지 않고 결정적인 규칙만으로 정돈된 격자(복셀)에 담아, 이미지 한 장을 보고 넓은 지역의 3D 공중 장면을 생성하는 AI 파이프라인의 기반 표현으로 쓴다. 저자는 Ming Qian 1인이며, 실험에서 200m×200m 타일부터 최대 1,400m×800m 넓은 지역까지 생성해 보였다.
METAL MEDIA 해설 도표
수백만 개 흩어진 3D점을 미리 손대지 않고도 AI가 학습할 수 있는 격자 데이터로 바꾸는 방법
01문제: 미리 만들어진 3DGS 장면은 점 개수가 제각각이고 순서도 없어서, 격자 형태 데이터를 기대하는 생성 AI 모델에 바로 넣을 수 없다.
02해결책: GS-Voxel은 장면을 정해진 크기의 복셀(3D 격자 칸)로 나누고, 각 칸 안에서 불투명도가 높은 점 상위 16개만 남긴 뒤 위치와 색상 등 속성을 8비트 숫자로 압축해 저장한다. 이 과정은 점을 새로 최적화하거나 전체 장면을 하나의 고정 틀에 맞추지 않아 '피팅 프리(fitting-free)'라 부른다.
03구조: 복셀의 있고 없음(형태)을 다루는 Geometry VAE와, 복셀 안 점들의 색·모양 속성을 다루는 Local Attribute VAE로 나누어 각각 압축된 잠재표현을 만든다.
04생성: 위성/항공 이미지를 조건으로 거친 구조 → 세밀한 형태 → 세부 속성 순으로 3단계 흐름(flow) 모델이 생성하며, 타일 경계가 겹치게 추론해 이어붙임으로써 학습 때 쓴 200m×200m보다 훨씬 넓은 지역까지 만들어낸다.
05성능: 검증 타일 10개 기준 변환만으로도 기존 방식 대비 렌더링 품질을 유지했고(PSNR 등), 최종 생성 결과는 자체 평가에서 FID 28.0, KID 0.020을 기록했다(다른 논문 값과는 데이터셋이 달라 직접 비교 불가).
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.
무엇을 했나
문제: 미리 만들어진 3DGS 장면은 점 개수가 제각각이고 순서도 없어서, 격자 형태 데이터를 기대하는 생성 AI 모델에 바로 넣을 수 없다.
해결책: GS-Voxel은 장면을 정해진 크기의 복셀(3D 격자 칸)로 나누고, 각 칸 안에서 불투명도가 높은 점 상위 16개만 남긴 뒤 위치와 색상 등 속성을 8비트 숫자로 압축해 저장한다. 이 과정은 점을 새로 최적화하거나 전체 장면을 하나의 고정 틀에 맞추지 않아 '피팅 프리(fitting-free)'라 부른다.
구조: 복셀의 있고 없음(형태)을 다루는 Geometry VAE와, 복셀 안 점들의 색·모양 속성을 다루는 Local Attribute VAE로 나누어 각각 압축된 잠재표현을 만든다.
생성: 위성/항공 이미지를 조건으로 거친 구조 → 세밀한 형태 → 세부 속성 순으로 3단계 흐름(flow) 모델이 생성하며, 타일 경계가 겹치게 추론해 이어붙임으로써 학습 때 쓴 200m×200m보다 훨씬 넓은 지역까지 만들어낸다.
성능: 검증 타일 10개 기준 변환만으로도 기존 방식 대비 렌더링 품질을 유지했고(PSNR 등), 최종 생성 결과는 자체 평가에서 FID 28.0, KID 0.020을 기록했다(다른 논문 값과는 데이터셋이 달라 직접 비교 불가).
Figure 1: Overview of our structured-latent pipeline. We first convert a compatible pre-optimized 3DGS reconstruction into GS-Voxel, a sparse representation of the retained local Gaussian parameters, and then encode it with a factorized VAE. The Geometry VAE models hierarchical subdivision and occupancy, while the Local Attribute VAE models Gaussian attributes on the decoded support.
Table 1: Taxonomic comparison of representative outdoor 3D scene creation systems.
Method
Training Data
Condition
Beyond Buildings
Large-Scale
Core Paradigm
Sat2Scene 13
Images & Height Map
Layout
✗
✗
Geo. Coloring
Sat2City 7
Mesh
Height Map
✗
✓
Building Object Gen.
Sat2Density++ 24
Images
Single Satellite
✓
✗
Feedforward
Sat3DGen 25
Images
Single Satellite
✓
✗
Feedforward
UrbanWorld 29
Mesh + UV
OSM or Layout
✓
✓
Multi-Stage
SynCity 3
N/A (training-free)
Text Prompt
✓
✓
Multi-Stage
SkyFall-GS 11
Images
Multi-view Satellite
✓
✓
Iterative Optimization
Orbit2Ground 40
Images
Multi-view Satellite
✓
✓
Iterative Optimization
XCube 27
Mesh
LiDAR Scan
✓
✓
3D Generative
EarthCrafter 17
Mesh
Satellite + Depth
✓
✓
3D Generative
Ours
3DGS
Single Satellite-View Image
✓
✓
3D Generative
Figure 2: Deterministic conversion from a compatible pre-optimized 3DGS reconstruction to GS-Voxel.
Table 2: Per-layer configuration of the virtual aerial camera rig. Pitch is measured as the deviation from the nadir direction. For pitch=0∘ a single yaw is used (top-down view); for pitch>0∘ four compass yaws {0∘,90∘,180∘,270∘} are sampled. Views/cell counts cameras per xy grid cell; Views/window assumes a 10×10 grid.
Layer
z
FOV
Pitches (deg)
Views/cell
Views/window
0
0.0
26∘
{0,30,45,60}
13
1,300
1
0.3
18∘
{15,30,45}
12
1,200
2
0.6
14∘
{0,15,30}
9
900
3
0.9
14∘
{15}
4
400
4
1.2
14∘
{0}
1
100
Total
39
3,900
Figure 4: Examples of retained and rejected aerial renders. Filtering removes unreliable or severely degraded supervision; only the retained views supervise the Local Attribute VAE.
Table 3: Direct GS-Voxel conversion before VAE encoding, evaluated on 10 validation tiles. We report average retained-primitives and active-voxel counts rounded to the nearest thousand (K), together with rendered reconstruction quality. Bold settings denote our default configuration.
R
Kin
Avg. Retained Primitives
Avg. Active Voxels
SSIM ↑
PSNR ↑
LPIPS ↓
256
1
377K
377K
0.55
18.52
0.425
256
4
971K
377K
0.85
26.82
0.166
256
8
1,340K
377K
0.95
33.26
0.058
256
16
1,602K
377K
0.98
40.04
0.023
256
32
1,737K
377K
0.98
40.94
0.021
512
1
980K
980K
0.87
27.38
0.156
512
4
1,501K
980K
0.98
39.20
0.025
Figure 5: Tile-level 3DGS generations conditioned on rendered satellite-view images.
Table 4: Choice of Kout for the Local Attribute VAE, evaluated using rendered-image reconstruction quality.
Method / Setting
PSNR ↑
SSIM ↑
LPIPS ↓
Kout=1
22.81
0.61
0.350
Kout=4 (Ours)
23.09
0.62
0.331
Kout=8
22.12
0.58
0.354
Figure 6: Large-area 3DGS generation (1/2) from constructed satellite-view conditions.
Table 5: Separate evaluation of the Geometry VAE and Local Attribute VAE. Shape IoU measures support reconstruction; PSNR, SSIM, and LPIPS measure attribute reconstruction on the target support.
Model
Shape IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Single-stage VAE
0.76
21.01
0.54
0.418
Geometry VAE
0.99
–
–
–
Local Attribute VAE
–
23.09
0.62
0.331
Figure 7: Large-area 3DGS generation (2/2). Overlap-aware tiled inference produces a large-area 3DGS primitive set from the satellite-view condition.
Table 6: Cross-paper reference for image-level FID/KID. Baselines use different ground-truth sets and camera protocols; values are not directly comparable and no ranking is implied.
Method
FID
KID
CityDreamer 37
97.3
0.096
GaussianCity 38
86.9
0.090
EarthCrafter 17
69.5
0.061
Ours
28.0
0.020
왜 중요한가
드론 매핑, 도시 시뮬레이션, 재난 대응처럼 넓은 지역을 사실적인 3D로 빠르게 만들어야 하는 작업에서, 기존처럼 장면마다 별도 최적화나 고정된 점 개수 제한 없이 실제 3DGS 데이터를 그대로 생성 AI 학습에 쓸 수 있게 해준다는 점에서 의미가 있다. 다만 저자가 밝히듯 아직 SH0(단순 색상) 항공 장면에 한정된 초기 연구다.
이 논문의 용어
3D Gaussian Splatting(3DGS) · 3D 공간에 반투명한 타원형 점(가우시안)들을 흩뿌려 사실적인 장면을 표현하는 렌더링 방식
복셀(Voxel) · 3D 공간을 격자 모양으로 나눈 하나의 칸, 2D의 픽셀에 해당하는 3D 단위
VAE(변분 오토인코더) · 데이터를 압축된 저차원 표현(잠재값)으로 바꿨다가 다시 복원하도록 학습하는 신경망
플로우 매칭(flow matching) · 노이즈에서 목표 데이터로 점진적으로 변환하는 방식을 학습하는 생성 모델 기법, 디퓨전과 유사한 계열
피팅 프리(fitting-free) · 장면마다 별도의 최적화나 고정 틀 맞춤 없이 정해진 규칙만으로 변환하는 방식
본문에 싣지 못한 그림
Figure 3: Conditional generation pipeline. Given a satellite-view image, a coarse sparse-structure stage predicts scene support; GS-Voxel-specific geometry and attribute stages then generate fine support and Gaussian parameters, yielding a standard 3DGS primitive set.