ResearchRobotics 🇨🇳 03.08.2026 14:01

200 Tasks, 17 Million Frames: Daxia Robot Open-Sources L5-Level Embodied Dataset ACE-Data-0, Turning Real Homes into Physical World Textbooks for Robots

Daxia Robot, together with NTU S-Lab, has open-sourced the L5-level multimodal embodied physical intelligence dataset ACE-Data-0, containing 17 million frames, 200 task categories, 50 participants, and 2 real home scenes. The dataset captures synchronized multimodal data including first-person and third-person video, full-body and hand motion, object trajectories, audio, and tactile signals, all aligned to a common timeline and spatial coordinate system. ACE-Data-0 aims to provide a high-quality data foundation for embodied models, supporting generalization, reflection, and evolution towards scalable real-world deployment.
Daxia Robot, in collaboration with Nanyang Technological University's S-Lab, has open-sourced the ACE-Data-0 dataset, an L5-level multimodal embodied physical intelligence dataset. Built on the world's first information density law for embodied models and a human-centered environmental data collection scheme 2.0, the dataset leverages the environmental collection engine ACE. ACE-Data-0 includes 17 million video frames, 200 task categories, 50 participants, 2 real home scenes, and 75,000 interaction segments, covering atomic-level human-object interactions, long-horizon household activity chains, and human-scene interactions. Data is synchronously captured from first-person video, multi-view third-person video, full-body and hand motion, object six-degrees-of-freedom trajectories, multi-channel audio, and tactile signals, all aligned to a unified timeline and spatial coordinate system. The dataset uses two complementary collection systems: a tabletop-scale system for fine-grained hand-object interactions and a room-scale system for whole-room activities, with shared calibration and synchronization. Unlike scripted collection, ACE-Data-0 uses goal-level instructions, allowing participants to freely choose items, sequences, and paths, preserving natural variations like hesitations and reversals. Based on the dataset, Daxia has established a three-level evaluation benchmark for embodied perception, testing models on contact sensing, state recovery, and hand-object interaction understanding. Evaluations of over 30 representative methods show significant capability gaps in contact, severe occlusion, egomotion, extreme viewpoints, and long-horizon tasks. The dataset is designed to support imitation learning, world models, VLA models, robot policy learning, and transfer from human demonstrations to different robot embodiments.
Abbreviations
VLA = Vision-Language-Action — Vision-Language-Action
Source: InfoQ 中国 — original
Our earlier posts on this topic ↓
Fresh news