Everything K-culture — comebacks to K-beauty, straight to your inboxGet it in your inbox

METAL MEDIA

Ego-OSCAR: Egocentric Open source Stereo CAptuRe System

arXiv:2608.082852026-08-07

A $200 head-mounted rig collected 550 hours of first-person stereo video, released as an open dataset for robot-learning research

Ego-OSCAR is a head-worn stereo-camera-plus-IMU capture device built entirely from off-the-shelf parts and 3D-printed pieces, costing under USD 200 per unit. The authors deployed it across 25 contributors in India in 40+ indoor environments over about six months, gathering roughly 550 hours of synchronized stereo video and inertial data per camera, then shipped it with free-form action captions and per-frame 3D hand reconstructions rather than as raw sensor dumps. Instead of matching the fidelity of research-grade devices like Project Aria, the goal is the cheapest defensible way to crowdsource egocentric data at scale.

METAL MEDIA explanatory visual

Ego-OSCAR capture-to-dataset pipeline

Evidence statusMeasured results reported

  1. Head-worn hardwareGlobal-shutter stereo camera, 6-axis IMU, SBC, and ESP32 microcontroller assembled for under USD 200
  2. Time synchronizationESP32 merges the camera's exposure signal with IMU readings and uses an LED flash to anchor frame numbers, achieving 700µs residual offset
  3. Field deployment25 contributors across 40+ Indian indoor environments recorded 1,462 sessions over six months, totaling ~550 hours per camera
  4. Quality filteringWatchdog, per-batch validation, and hand-visibility screening bring the usable-session rate to 96%
  5. Annotation layers209,315 free-form action captions and WiLoR-based per-frame 3D hand reconstructions are attached across the full corpus
An explanatory diagram made by METAL MEDIA, not a figure supplied by the paper's authors.

What they did

  1. The device combines a hardware-synchronized global-shutter stereo camera, a 6-axis IMU, an embedded Linux SBC, and a real-time microcontroller, with a full bill of materials under about USD 200 (INR 19,100) per unit.
  2. An ESP32 microcontroller taps the camera's start-of-exposure signal and merges it with IMU readings, using an LED flash to anchor frame numbers, bringing residual visual-inertial time lag down to 700 microseconds.
  3. Across 13 deployed devices, mean per-pixel epipolar error after rectification was 0.4 px and per-camera reprojection error was under 0.03 px, with SGBM and RAFT-Stereo producing full-field disparity maps without special tuning.
  4. Over a 6-month field deployment, 96% of sessions yielded usable data end-to-end; a hand detector (WiLoR) found hands in 94% of frames corpus-wide; and stereo-inertial odometry (VINS-Fusion) produced stable trajectories on 12 of 20 held-out sequences, versus 15 of 20 for an Intel RealSense, which has the unfair advantage of active stereo depth.
  5. The released Ego-OSCAR-550h dataset spans 1,462 sessions, 25 contributors, and 40+ environments, with 209,315 free-form action segments and per-frame 3D hand reconstructions, showing a long-tailed distribution where the top 20 action expressions account for only 1.5% of instances.
Figure 1: Open-source egocentric capture system.
Figure 1: Open-source egocentric capture system.
Table 1: Bill of Materials.
ComponentINRFunction
Dexcin USB stereo camera6,300Stereo capture, global shutter
Radxa Rock 5C (2 GB)6,500SBC: capture, encode, store
Heatsink (Radxa)800Thermal management
256 GB SD card2,500OS + ∼16–18 hr recording
USB cable, A–C 90°300Camera ↔ SBC
ICM-20948 6-axis IMU800Inertial sensing
Seeed Xiao ESP32-S3650UX + watchdog MCU
Misc. electronics300LEDs, buzzer, button, wiring
3D-printed shells + screws300Enclosure
Visor / cap300Head mount
USB-C PD cable, 90°150SBC ↔ power
10,000 mAh power bank500Power
Total19,100∼USD 200
(b) Mechanical Outline
(b) Mechanical Outline
Table 2: Comparison with existing egocentric datasets. Figures are as reported by each dataset’s own publication. “Cam-h” denotes camera-hours. Ego-Exo4D hours combine egocentric and exocentric video.
DatasetHoursWearersCameraSync. IMUDense action labelsOpen HW
Ego4D [7]3,670931Mono consumer, rolling shutter (stereo in a subset)PartialTimestamped narrationsNo
Ego-Exo4D [8]1,286740Aria: mono RGB + 2 mono SLAMYesNarrations + expert commentaryNo
EPIC-K.-100 [2]10037Mono head-mountedNo90K segments, closed taxonomy (97 verbs / 300 nouns)No
Nymeria [15]300264Aria + body mocapYes301.5K narration sentencesNo
Ours550/cam (1,100 cam-h)25Calibrated RGB stereo, global shutter, per-session calib.Yes 120 Hz209,315 segments, open vocabulary, ≈100% coverageYes ∼USD 200
(c) Device in action
(c) Device in action
Table 3: Dataset overview: measured dimensions and what each supports.
DimensionEvidenceRelevance
Video scale≈550 h per camera (≈1,100 stereo cam-h)Large calibrated stereo RGB corpus
Capture geometrySynchronized pair + per-session calibrationMetric binocular depth cues, not just RGB
Delivery formatMP4 (H.264, 1280×720, 30 fps)Direct ingestion into video pipelines
Label densityMedian 94 segments/sessionDense temporal supervision, not clip-level tags
Action vocabulary460 observed verbsCoverage of manipulation primitives
Object vocabulary32,630 object phrasesWide object, material and tool coverage
Effective breadth57,104 verb–object combinationsCompositional, resistant to rare-label inflation
Temporal structure192,509 ordered task transitionsSequence structure for world models
Contributors25 user IDs / 13 devicesVariation in behavior, routine and execution style
Figure 2: System Overview: Hardware architecture.
Figure 2: System Overview: Hardware architecture.
Table 4: Most frequent observed task expressions and their share of all labeled segments.
Task expressionOccurrencesShare
idle / no manipulation6590.31%
cut sewing thread with scissors2690.13%
close refrigerator door2380.11%
peel garlic clove2120.10%
open refrigerator door2120.10%
turn on kitchen faucet1390.07%
pick up iron from side table1340.06%
turn off kitchen faucet1240.06%
adjust stove burner knob1220.06%
rinse small metal cup under running water1190.06%
adjust stove control knob1150.05%
roll dough on rolling board with rolling pin1090.05%
Figure 3: Data utility results.
Figure 3: Data utility results.
Table 5: Activity-domain composition of the Ego-OSCAR-550h dataset (keyword-derived, approximate).
Activity domainLabeled hoursPrimary in
Cooking and food preparation187 h664 sessions
Dishwashing and kitchen cleanup90 h258 sessions
Textile and craft (sewing, tailoring, flowers)54 h214 sessions
Laundry and clothing care45 h145 sessions
Organizing and storage39 h87 sessions
Cleaning and housekeeping29 h87 sessions
Generic manipulation and transitions106 h7 sessions
(b) Stereo depth maps
(b) Stereo depth maps
Table 6: Per-contributor coverage (percentiles across the 25 contributors).
Contributor-level diversity measure25thMedian75th
Labeled action segments5,1836,25812,075
Distinct task expressions4,0305,6369,115
Distinct action verbs90152178
Figure 4: Task diversity in the Ego-OSCAR-550h dataset.
Figure 4: Task diversity in the Ego-OSCAR-550h dataset.
Table 7: Sequence and composition signals in the dataset.
SignalEvidenceWhy it matters
Verb–object composition57,104 unique combinationsSystematic generalization across skills and objects
Broadly recombined verbs132 verbs with 25+ objects; 66 with 100+Reuse of manipulation primitives across object types
Ordered task transitions192,509 unique transitionsTemporal structure for sequence learning
Per-session sequence richnessMedian 86 transitions/session95.8% of sessions contain 10+ distinct transitions
Cross-contributor support19.4% of instances seen across 2+ contributorsReduces reliance on a single execution style

Findings

  • Across 13 deployed units, mean epipolar error after rectification was 0.4 px and per-camera reprojection error was below 0.03 px, confirming usable stereo geometry.
  • Residual visual-inertial time offset after correction was measured at 700 microseconds using Kalibr's Cam-IMU offset test.
  • Stereo-inertial odometry (VINS-Fusion) produced stable trajectories on 12 of 20 held-out sequences (versus 15/20 for an Intel RealSense with active stereo), and a hand detector (WiLoR) found hands in 94% of corpus frames.
  • Over a 6-month deployment, 96% of sessions produced usable data end-to-end, yielding 1,462 sessions across 25 contributors and 40+ environments, totaling about 550 hours per camera (roughly 1,100 stereo camera-hours).
  • The released dataset contains 209,315 free-form action segments and per-frame 3D hand reconstructions across the corpus, with a long-tailed distribution where the top 20 action expressions account for just 1.5% of all instances.

Where it can be used

  • Collecting large-scale first-person visual-action data for pretraining vision-language-action or world models for robots
  • Building or adapting the open-hardware design to crowdsource egocentric data in specific domains not well covered by existing corpora (e.g., household crafts)
  • Using the calibrated stereo and hand-reconstruction annotations for depth or hand-pose research without needing to build capture hardware from scratch
  • Serving as a reference for teams without access to closed research platforms like Project Aria who want to build their own capture pipeline

Limits and open work

  • The paper does not show that a robot policy trained on this data outperforms one trained on existing datasets; that comparison is left as future work.
  • There is no motion-capture or surveyed ground truth for camera trajectories, so the 12/20 odometry figure reflects convergence rate only, not positional accuracy.
  • The corpus is geographically and demographically concentrated: 25 contributors in India using 13 shared devices, skewed toward domestic activities.
  • The consumer-grade IMU is the dominant source of pose error; it is swappable on the same bus but this upgrade has not yet been benchmarked.
  • The device only captures data and cannot reject a bad session in real time, and durability issues remain, including strap pressure points, visor drooping over long sessions, and no moisture resistance.

Why it matters

Vision-language-action models for robots need large, diverse, multimodal data, but teleoperation is expensive to scale, simulation struggles with the sim-to-real gap, and existing egocentric datasets rely on either uncalibrated consumer cameras or closed research-grade hardware that cannot be freely reproduced. By open-sourcing a cheap, buildable capture device alongside its software and a validated dataset, this work lowers the barrier for any team wanting to collect its own large-scale egocentric data.

Terms in this paper

  • Egocentric video · First-person footage captured from a head-worn camera, showing the wearer's own view of the world
  • Global shutter · A sensor design that exposes all pixels simultaneously, avoiding distortion during fast motion
  • IMU (inertial measurement unit) · A sensor that measures acceleration and rotation to track device motion
  • Stereo calibration · The process of measuring the geometric relationship and lens distortion between two cameras so accurate depth can be computed
  • Visual-inertial odometry (VIO) · A technique that combines camera images and IMU data to estimate how a device has moved over time

Original abstract (English)

We present Ego-OSCAR, an open-hardware, low-cost, head-mounted stereo-inertial capture device for egocentric data collection in the wild. EgoOSCAR pairs a hardware-synchronized global-shutter stereo camera with a 6- axis IMU, an embedded Linux SBC for on-device video encoding, and a realtime microcontroller for user feedback and watchdog functions. The complete bill of materials is under USD 200 per unit, using only commercially available components and 3D-printed parts. Alongside the device, we release a complete software stack (hardware-accelerated recording pipeline, IMU sampling daemon, time-synchronization tooling, and watchdog firmware) and roughly 550 hours of egocentric stereo video per camera with synchronized IMU, collected by a distributed contributor network across everyday indoor environments. The release is annotated rather than raw: free-form action captions cover essentially the entire recorded timeline with an open vocabulary, and per-frame 3D hand reconstructions ship alongside per-session stereo calibration. Ego-OSCAR does not aim to match the per-unit fidelity of research-grade systems such as Project Aria; it aims to be the cheapest defensible substrate for crowdsourced egocentric capture, and to lower the activation energy for any team that wants to collect egocentric data at scale. All hardware designs, software, and the dataset are open-sourced

Authors · Gunjan Paul

Read on arXiv

Latest papers

All papers →

Latest from METAL MEDIA

Figures: Gunjan Paul et al., arXiv:2608.08285, CC BY 4.0