During the 2026 World Robot Congress (WRC), Noitom Robotics officially introduced its new HiPHI (pronounced like "hi-fi") dataset on the afternoon of August 23rd. This high-precision optical motion capture dataset is specifically designed for researchers and engineers working in humanoid robot learning, digital human development, and computer graphics fields.
The HiPHI dataset comprises 617.5 total hours of motion data, including mirrored left-right variations. This includes 371.8 hours of full-body human movement data and 245.7 hours of human-object interaction sequences. Within the interaction portion, the trajectory and 3D mesh of each manipulated object are recorded in precise synchronization with the actor's movements.
Data was gathered from 132 motion capture performers at a 90 Hz sampling rate with sub-millimeter optical tracking accuracy. The information is systematically organized into action units based on FrameNet's semantic framework for human movement classification. HiPHI also includes complete documentation, featuring an academic paper authored by the Noiterm Robotics research team. According to the paper, when compared against other benchmark datasets, HiPHI demonstrates the most extensive quantified action coverage and the lowest per-frame physical quality error rate.
Models trained on the HiPHI dataset have already been successfully deployed on Unitree G1 humanoid robots in real-world testing, enabling capabilities such as running, sitting, crawling, box carrying, and suitcase pulling. The dataset is now publicly accessible on Hugging Face.
Dr. Li Ruo Dai, founder and CEO of Noitom Robotics, explained the project's motivation: "The true bottleneck facing Physical AI isn't about how much data exists in the world, but how much of it can actually be learned by machines. HiPHI is our first answer to this challenge. We completed the coverage design before capturing the first frame of data. Our precision levels ensure successful transfer from training to real robots, and when humans interact with objects, those objects are recorded as part of the motion itself."
Dr. Dai further stated, "We chose to open-source HiPHI for two reasons. First, we want to provide the embodied intelligence industry with a reliable public data foundation. Second, we have strong confidence in our ongoing ability to produce high-quality data at scale."
Comments