Ego-OSCAR: Egocentric Open source Stereo CAptuRe System
A $200 head-mounted rig collected 550 hours of first-person stereo video, released as an open dataset for robot-learning research
Ego-OSCAR is a head-worn stereo-camera-plus-IMU capture device built entirely from off-the-shelf parts and 3D-printed pieces, costing under USD 200 per unit. The authors deployed it across 25 contributors in India in 40+ indoor environments over about six months, gathering roughly 550 hours of synchronized stereo video and inertial data per camera, then shipped it with free-form action captions and per-frame 3D hand reconstructions rather than as raw sensor dumps. Instead of matching the fidelity of research-grade devices like Project Aria, the goal is the cheapest defensible way to crowdsource egocentric data at scale.
METAL MEDIA explanatory visual
Ego-OSCAR capture-to-dataset pipeline
Evidence statusMeasured results reported
- Head-worn hardwareGlobal-shutter stereo camera, 6-axis IMU, SBC, and ESP32 microcontroller assembled for under USD 200
- Time synchronizationESP32 merges the camera's exposure signal with IMU readings and uses an LED flash to anchor frame numbers, achieving 700µs residual offset
- Field deployment25 contributors across 40+ Indian indoor environments recorded 1,462 sessions over six months, totaling ~550 hours per camera
- Quality filteringWatchdog, per-batch validation, and hand-visibility screening bring the usable-session rate to 96%
- Annotation layers209,315 free-form action captions and WiLoR-based per-frame 3D hand reconstructions are attached across the full corpus
What they did
- The device combines a hardware-synchronized global-shutter stereo camera, a 6-axis IMU, an embedded Linux SBC, and a real-time microcontroller, with a full bill of materials under about USD 200 (INR 19,100) per unit.
- An ESP32 microcontroller taps the camera's start-of-exposure signal and merges it with IMU readings, using an LED flash to anchor frame numbers, bringing residual visual-inertial time lag down to 700 microseconds.
- Across 13 deployed devices, mean per-pixel epipolar error after rectification was 0.4 px and per-camera reprojection error was under 0.03 px, with SGBM and RAFT-Stereo producing full-field disparity maps without special tuning.
- Over a 6-month field deployment, 96% of sessions yielded usable data end-to-end; a hand detector (WiLoR) found hands in 94% of frames corpus-wide; and stereo-inertial odometry (VINS-Fusion) produced stable trajectories on 12 of 20 held-out sequences, versus 15 of 20 for an Intel RealSense, which has the unfair advantage of active stereo depth.
- The released Ego-OSCAR-550h dataset spans 1,462 sessions, 25 contributors, and 40+ environments, with 209,315 free-form action segments and per-frame 3D hand reconstructions, showing a long-tailed distribution where the top 20 action expressions account for only 1.5% of instances.

| Component | INR | Function |
|---|---|---|
| Dexcin USB stereo camera | 6,300 | Stereo capture, global shutter |
| Radxa Rock 5C (2 GB) | 6,500 | SBC: capture, encode, store |
| Heatsink (Radxa) | 800 | Thermal management |
| 256 GB SD card | 2,500 | OS + ∼16–18 hr recording |
| USB cable, A–C 90° | 300 | Camera ↔ SBC |
| ICM-20948 6-axis IMU | 800 | Inertial sensing |
| Seeed Xiao ESP32-S3 | 650 | UX + watchdog MCU |
| Misc. electronics | 300 | LEDs, buzzer, button, wiring |
| 3D-printed shells + screws | 300 | Enclosure |
| Visor / cap | 300 | Head mount |
| USB-C PD cable, 90° | 150 | SBC ↔ power |
| 10,000 mAh power bank | 500 | Power |
| Total | 19,100 | ∼USD 200 |

| Dataset | Hours | Wearers | Camera | Sync. IMU | Dense action labels | Open HW |
|---|---|---|---|---|---|---|
| Ego4D [7] | 3,670 | 931 | Mono consumer, rolling shutter (stereo in a subset) | Partial | Timestamped narrations | No |
| Ego-Exo4D [8] | 1,286 | 740 | Aria: mono RGB + 2 mono SLAM | Yes | Narrations + expert commentary | No |
| EPIC-K.-100 [2] | 100 | 37 | Mono head-mounted | No | 90K segments, closed taxonomy (97 verbs / 300 nouns) | No |
| Nymeria [15] | 300 | 264 | Aria + body mocap | Yes | 301.5K narration sentences | No |
| Ours | 550/cam (1,100 cam-h) | 25 | Calibrated RGB stereo, global shutter, per-session calib. | Yes 120 Hz | 209,315 segments, open vocabulary, ≈100% coverage | Yes ∼USD 200 |

| Dimension | Evidence | Relevance |
|---|---|---|
| Video scale | ≈550 h per camera (≈1,100 stereo cam-h) | Large calibrated stereo RGB corpus |
| Capture geometry | Synchronized pair + per-session calibration | Metric binocular depth cues, not just RGB |
| Delivery format | MP4 (H.264, 1280×720, 30 fps) | Direct ingestion into video pipelines |
| Label density | Median 94 segments/session | Dense temporal supervision, not clip-level tags |
| Action vocabulary | 460 observed verbs | Coverage of manipulation primitives |
| Object vocabulary | 32,630 object phrases | Wide object, material and tool coverage |
| Effective breadth | 57,104 verb–object combinations | Compositional, resistant to rare-label inflation |
| Temporal structure | 192,509 ordered task transitions | Sequence structure for world models |
| Contributors | 25 user IDs / 13 devices | Variation in behavior, routine and execution style |

| Task expression | Occurrences | Share |
|---|---|---|
| idle / no manipulation | 659 | 0.31% |
| cut sewing thread with scissors | 269 | 0.13% |
| close refrigerator door | 238 | 0.11% |
| peel garlic clove | 212 | 0.10% |
| open refrigerator door | 212 | 0.10% |
| turn on kitchen faucet | 139 | 0.07% |
| pick up iron from side table | 134 | 0.06% |
| turn off kitchen faucet | 124 | 0.06% |
| adjust stove burner knob | 122 | 0.06% |
| rinse small metal cup under running water | 119 | 0.06% |
| adjust stove control knob | 115 | 0.05% |
| roll dough on rolling board with rolling pin | 109 | 0.05% |

| Activity domain | Labeled hours | Primary in |
|---|---|---|
| Cooking and food preparation | 187 h | 664 sessions |
| Dishwashing and kitchen cleanup | 90 h | 258 sessions |
| Textile and craft (sewing, tailoring, flowers) | 54 h | 214 sessions |
| Laundry and clothing care | 45 h | 145 sessions |
| Organizing and storage | 39 h | 87 sessions |
| Cleaning and housekeeping | 29 h | 87 sessions |
| Generic manipulation and transitions | 106 h | 7 sessions |

| Contributor-level diversity measure | 25th | Median | 75th |
|---|---|---|---|
| Labeled action segments | 5,183 | 6,258 | 12,075 |
| Distinct task expressions | 4,030 | 5,636 | 9,115 |
| Distinct action verbs | 90 | 152 | 178 |

| Signal | Evidence | Why it matters |
|---|---|---|
| Verb–object composition | 57,104 unique combinations | Systematic generalization across skills and objects |
| Broadly recombined verbs | 132 verbs with 25+ objects; 66 with 100+ | Reuse of manipulation primitives across object types |
| Ordered task transitions | 192,509 unique transitions | Temporal structure for sequence learning |
| Per-session sequence richness | Median 86 transitions/session | 95.8% of sessions contain 10+ distinct transitions |
| Cross-contributor support | 19.4% of instances seen across 2+ contributors | Reduces reliance on a single execution style |
Findings
- Across 13 deployed units, mean epipolar error after rectification was 0.4 px and per-camera reprojection error was below 0.03 px, confirming usable stereo geometry.
- Residual visual-inertial time offset after correction was measured at 700 microseconds using Kalibr's Cam-IMU offset test.
- Stereo-inertial odometry (VINS-Fusion) produced stable trajectories on 12 of 20 held-out sequences (versus 15/20 for an Intel RealSense with active stereo), and a hand detector (WiLoR) found hands in 94% of corpus frames.
- Over a 6-month deployment, 96% of sessions produced usable data end-to-end, yielding 1,462 sessions across 25 contributors and 40+ environments, totaling about 550 hours per camera (roughly 1,100 stereo camera-hours).
- The released dataset contains 209,315 free-form action segments and per-frame 3D hand reconstructions across the corpus, with a long-tailed distribution where the top 20 action expressions account for just 1.5% of all instances.
Where it can be used
- Collecting large-scale first-person visual-action data for pretraining vision-language-action or world models for robots
- Building or adapting the open-hardware design to crowdsource egocentric data in specific domains not well covered by existing corpora (e.g., household crafts)
- Using the calibrated stereo and hand-reconstruction annotations for depth or hand-pose research without needing to build capture hardware from scratch
- Serving as a reference for teams without access to closed research platforms like Project Aria who want to build their own capture pipeline
Limits and open work
- The paper does not show that a robot policy trained on this data outperforms one trained on existing datasets; that comparison is left as future work.
- There is no motion-capture or surveyed ground truth for camera trajectories, so the 12/20 odometry figure reflects convergence rate only, not positional accuracy.
- The corpus is geographically and demographically concentrated: 25 contributors in India using 13 shared devices, skewed toward domestic activities.
- The consumer-grade IMU is the dominant source of pose error; it is swappable on the same bus but this upgrade has not yet been benchmarked.
- The device only captures data and cannot reject a bad session in real time, and durability issues remain, including strap pressure points, visor drooping over long sessions, and no moisture resistance.
Why it matters
Vision-language-action models for robots need large, diverse, multimodal data, but teleoperation is expensive to scale, simulation struggles with the sim-to-real gap, and existing egocentric datasets rely on either uncalibrated consumer cameras or closed research-grade hardware that cannot be freely reproduced. By open-sourcing a cheap, buildable capture device alongside its software and a validated dataset, this work lowers the barrier for any team wanting to collect its own large-scale egocentric data.
Terms in this paper
- Egocentric video · First-person footage captured from a head-worn camera, showing the wearer's own view of the world
- Global shutter · A sensor design that exposes all pixels simultaneously, avoiding distortion during fast motion
- IMU (inertial measurement unit) · A sensor that measures acceleration and rotation to track device motion
- Stereo calibration · The process of measuring the geometric relationship and lens distortion between two cameras so accurate depth can be computed
- Visual-inertial odometry (VIO) · A technique that combines camera images and IMU data to estimate how a device has moved over time
Original abstract (English)
We present Ego-OSCAR, an open-hardware, low-cost, head-mounted stereo-inertial capture device for egocentric data collection in the wild. EgoOSCAR pairs a hardware-synchronized global-shutter stereo camera with a 6- axis IMU, an embedded Linux SBC for on-device video encoding, and a realtime microcontroller for user feedback and watchdog functions. The complete bill of materials is under USD 200 per unit, using only commercially available components and 3D-printed parts. Alongside the device, we release a complete software stack (hardware-accelerated recording pipeline, IMU sampling daemon, time-synchronization tooling, and watchdog firmware) and roughly 550 hours of egocentric stereo video per camera with synchronized IMU, collected by a distributed contributor network across everyday indoor environments. The release is annotated rather than raw: free-form action captions cover essentially the entire recorded timeline with an open vocabulary, and per-frame 3D hand reconstructions ship alongside per-session stereo calibration. Ego-OSCAR does not aim to match the per-unit fidelity of research-grade systems such as Project Aria; it aims to be the cheapest defensible substrate for crowdsourced egocentric capture, and to lower the activation energy for any team that wants to collect egocentric data at scale. All hardware designs, software, and the dataset are open-sourced
Read on arXivLatest papers
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?AI coding agents were tested on fixing real scientific software, and even the best one failed more than half the time
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM ServingMaking sparse attention fast enough and accurate enough for real LLM serving, not just papers
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM AgentsMaking customer-service AI agents follow the whole procedure, not just avoid one bad action
- EXIMO: VLM Guided Exploration of VLA PoliciesTeaching a robot new chores without human teleoperation, by letting a chatty AI supervise it
- EnvHarness: Awakening Static Worlds for Agent LearningInstead of building new training worlds from scratch, this work adds a plug-in layer that reshapes existing ones around each agent's actual weaknesses
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the ModelCompanies that rent AI instead of owning it can only do half of AI safety oversight
- PersonalBench: Measuring the Authorship Gap in LLM PersonalizationAI can be prompted to write 'like someone,' but its own voice never fully disappears
- Automated Summarization of Financial News Using Large Language Models and Retrieval-Augmented Generation: An Early Empirical Study (Fall 2023)Testing AI summaries of stock news, the simple approach beat the trendy retrieval-based one
Latest from METAL MEDIA
Figures: Gunjan Paul et al., arXiv:2608.08285, CC BY 4.0