컴백부터 K-뷰티까지 — K-컬쳐의 모든 것을 메일로 받아보세요메일로 받아보기

METAL MEDIA

수백만 개 흩어진 3D점을 미리 손대지 않고도 AI가 학습할 수 있는 격자 데이터로 바꾸는 방법

arXiv:2608.179882026-08-17

GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation

수백만 개 흩어진 3D점을 미리 손대지 않고도 AI가 학습할 수 있는 격자 데이터로 바꾸는 방법

드론이나 항공 사진으로 만든 3D 공중 장면은 3D Gaussian Splatting(3DGS)이라는 방식으로 표현되는데, 한 장면에 점이 300만 개 넘게 무질서하게 흩어져 있어 그대로는 생성 AI가 학습하기 어렵다. GS-Voxel은 이 점들을 장면별로 다시 최적화하지 않고 결정적인 규칙만으로 정돈된 격자(복셀)에 담아, 이미지 한 장을 보고 넓은 지역의 3D 공중 장면을 생성하는 AI 파이프라인의 기반 표현으로 쓴다. 저자는 Ming Qian 1인이며, 실험에서 200m×200m 타일부터 최대 1,400m×800m 넓은 지역까지 생성해 보였다.

METAL MEDIA 해설 도표

수백만 개 흩어진 3D점을 미리 손대지 않고도 AI가 학습할 수 있는 격자 데이터로 바꾸는 방법

  1. 01문제: 미리 만들어진 3DGS 장면은 점 개수가 제각각이고 순서도 없어서, 격자 형태 데이터를 기대하는 생성 AI 모델에 바로 넣을 수 없다.
  2. 02해결책: GS-Voxel은 장면을 정해진 크기의 복셀(3D 격자 칸)로 나누고, 각 칸 안에서 불투명도가 높은 점 상위 16개만 남긴 뒤 위치와 색상 등 속성을 8비트 숫자로 압축해 저장한다. 이 과정은 점을 새로 최적화하거나 전체 장면을 하나의 고정 틀에 맞추지 않아 '피팅 프리(fitting-free)'라 부른다.
  3. 03구조: 복셀의 있고 없음(형태)을 다루는 Geometry VAE와, 복셀 안 점들의 색·모양 속성을 다루는 Local Attribute VAE로 나누어 각각 압축된 잠재표현을 만든다.
  4. 04생성: 위성/항공 이미지를 조건으로 거친 구조 → 세밀한 형태 → 세부 속성 순으로 3단계 흐름(flow) 모델이 생성하며, 타일 경계가 겹치게 추론해 이어붙임으로써 학습 때 쓴 200m×200m보다 훨씬 넓은 지역까지 만들어낸다.
  5. 05성능: 검증 타일 10개 기준 변환만으로도 기존 방식 대비 렌더링 품질을 유지했고(PSNR 등), 최종 생성 결과는 자체 평가에서 FID 28.0, KID 0.020을 기록했다(다른 논문 값과는 데이터셋이 달라 직접 비교 불가).
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. 문제: 미리 만들어진 3DGS 장면은 점 개수가 제각각이고 순서도 없어서, 격자 형태 데이터를 기대하는 생성 AI 모델에 바로 넣을 수 없다.
  2. 해결책: GS-Voxel은 장면을 정해진 크기의 복셀(3D 격자 칸)로 나누고, 각 칸 안에서 불투명도가 높은 점 상위 16개만 남긴 뒤 위치와 색상 등 속성을 8비트 숫자로 압축해 저장한다. 이 과정은 점을 새로 최적화하거나 전체 장면을 하나의 고정 틀에 맞추지 않아 '피팅 프리(fitting-free)'라 부른다.
  3. 구조: 복셀의 있고 없음(형태)을 다루는 Geometry VAE와, 복셀 안 점들의 색·모양 속성을 다루는 Local Attribute VAE로 나누어 각각 압축된 잠재표현을 만든다.
  4. 생성: 위성/항공 이미지를 조건으로 거친 구조 → 세밀한 형태 → 세부 속성 순으로 3단계 흐름(flow) 모델이 생성하며, 타일 경계가 겹치게 추론해 이어붙임으로써 학습 때 쓴 200m×200m보다 훨씬 넓은 지역까지 만들어낸다.
  5. 성능: 검증 타일 10개 기준 변환만으로도 기존 방식 대비 렌더링 품질을 유지했고(PSNR 등), 최종 생성 결과는 자체 평가에서 FID 28.0, KID 0.020을 기록했다(다른 논문 값과는 데이터셋이 달라 직접 비교 불가).
Figure 1: Overview of our structured-latent pipeline. We first convert a compatible pre-optimized 3DGS reconstruction into GS-Voxel, a sparse representation of the retained local Gaussian parameters, and then encode it with a factorized VAE. The Geometry VAE models hierarchical subdivision and occupancy, while the Local Attribute VAE models Gaussian attributes on the decoded support.
Figure 1: Overview of our structured-latent pipeline. We first convert a compatible pre-optimized 3DGS reconstruction into GS-Voxel, a sparse representation of the retained local Gaussian parameters, and then encode it with a factorized VAE. The Geometry VAE models hierarchical subdivision and occupancy, while the Local Attribute VAE models Gaussian attributes on the decoded support.
Table 1: Taxonomic comparison of representative outdoor 3D scene creation systems.
MethodTraining DataConditionBeyond BuildingsLarge-ScaleCore Paradigm
Sat2Scene 13Images & Height MapLayoutGeo. Coloring
Sat2City 7MeshHeight MapBuilding Object Gen.
Sat2Density++ 24ImagesSingle SatelliteFeedforward
Sat3DGen 25ImagesSingle SatelliteFeedforward
UrbanWorld 29Mesh + UVOSM or LayoutMulti-Stage
SynCity 3N/A (training-free)Text PromptMulti-Stage
SkyFall-GS 11ImagesMulti-view SatelliteIterative Optimization
Orbit2Ground 40ImagesMulti-view SatelliteIterative Optimization
XCube 27MeshLiDAR Scan3D Generative
EarthCrafter 17MeshSatellite + Depth3D Generative
Ours3DGSSingle Satellite-View Image3D Generative
Figure 2: Deterministic conversion from a compatible pre-optimized 3DGS reconstruction to GS-Voxel.
Figure 2: Deterministic conversion from a compatible pre-optimized 3DGS reconstruction to GS-Voxel.
Table 2: Per-layer configuration of the virtual aerial camera rig. Pitch is measured as the deviation from the nadir direction. For pitch=0∘ a single yaw is used (top-down view); for pitch>0∘ four compass yaws {0∘,90∘,180∘,270∘} are sampled. Views/cell counts cameras per x​y grid cell; Views/window assumes a 10×10 grid.
LayerzFOVPitches (deg)Views/cellViews/window
00.026∘{0,30,45,60}131,300
10.318∘{15,30,45}121,200
20.614∘{0,15,30}9900
30.914∘{15}4400
41.214∘{0}1100
Total393,900
Figure 4: Examples of retained and rejected aerial renders. Filtering removes unreliable or severely degraded supervision; only the retained views supervise the Local Attribute VAE.
Figure 4: Examples of retained and rejected aerial renders. Filtering removes unreliable or severely degraded supervision; only the retained views supervise the Local Attribute VAE.
Table 3: Direct GS-Voxel conversion before VAE encoding, evaluated on 10 validation tiles. We report average retained-primitives and active-voxel counts rounded to the nearest thousand (K), together with rendered reconstruction quality. Bold settings denote our default configuration.
RKinAvg. Retained PrimitivesAvg. Active VoxelsSSIM ↑PSNR ↑LPIPS ↓
2561377K377K0.5518.520.425
2564971K377K0.8526.820.166
25681,340K377K0.9533.260.058
256161,602K377K0.9840.040.023
256321,737K377K0.9840.940.021
5121980K980K0.8727.380.156
51241,501K980K0.9839.200.025
Figure 5: Tile-level 3DGS generations conditioned on rendered satellite-view images.
Figure 5: Tile-level 3DGS generations conditioned on rendered satellite-view images.
Table 4: Choice of Kout for the Local Attribute VAE, evaluated using rendered-image reconstruction quality.
Method / SettingPSNR ↑SSIM ↑LPIPS ↓
Kout=122.810.610.350
Kout=4 (Ours)23.090.620.331
Kout=822.120.580.354
Figure 6: Large-area 3DGS generation (1/2) from constructed satellite-view conditions.
Figure 6: Large-area 3DGS generation (1/2) from constructed satellite-view conditions.
Table 5: Separate evaluation of the Geometry VAE and Local Attribute VAE. Shape IoU measures support reconstruction; PSNR, SSIM, and LPIPS measure attribute reconstruction on the target support.
ModelShape IoU ↑PSNR ↑SSIM ↑LPIPS ↓
Single-stage VAE0.7621.010.540.418
Geometry VAE0.99
Local Attribute VAE23.090.620.331
Figure 7: Large-area 3DGS generation (2/2). Overlap-aware tiled inference produces a large-area 3DGS primitive set from the satellite-view condition.
Figure 7: Large-area 3DGS generation (2/2). Overlap-aware tiled inference produces a large-area 3DGS primitive set from the satellite-view condition.
Table 6: Cross-paper reference for image-level FID/KID. Baselines use different ground-truth sets and camera protocols; values are not directly comparable and no ranking is implied.
MethodFIDKID
CityDreamer 3797.30.096
GaussianCity 3886.90.090
EarthCrafter 1769.50.061
Ours28.00.020

왜 중요한가

드론 매핑, 도시 시뮬레이션, 재난 대응처럼 넓은 지역을 사실적인 3D로 빠르게 만들어야 하는 작업에서, 기존처럼 장면마다 별도 최적화나 고정된 점 개수 제한 없이 실제 3DGS 데이터를 그대로 생성 AI 학습에 쓸 수 있게 해준다는 점에서 의미가 있다. 다만 저자가 밝히듯 아직 SH0(단순 색상) 항공 장면에 한정된 초기 연구다.

이 논문의 용어

  • 3D Gaussian Splatting(3DGS) · 3D 공간에 반투명한 타원형 점(가우시안)들을 흩뿌려 사실적인 장면을 표현하는 렌더링 방식
  • 복셀(Voxel) · 3D 공간을 격자 모양으로 나눈 하나의 칸, 2D의 픽셀에 해당하는 3D 단위
  • VAE(변분 오토인코더) · 데이터를 압축된 저차원 표현(잠재값)으로 바꿨다가 다시 복원하도록 학습하는 신경망
  • 플로우 매칭(flow matching) · 노이즈에서 목표 데이터로 점진적으로 변환하는 방식을 학습하는 생성 모델 기법, 디퓨전과 유사한 계열
  • 피팅 프리(fitting-free) · 장면마다 별도의 최적화나 고정 틀 맞춤 없이 정해진 규칙만으로 변환하는 방식

본문에 싣지 못한 그림

  • Figure 3: Conditional generation pipeline. Given a satellite-view image, a coarse sparse-structure stage predicts scene support; GS-Voxel-specific geometry and attribute stages then generate fine support and Gaussian parameters, yielding a standard 3DGS primitive set.
원문에서 그림 보기 →

저자 · Ming Qian

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL MEDIA 최신 기사

그림 출처: Ming Qian et al., arXiv:2608.17988, CC BY 4.0