RoboticsResearch 🇨🇳 26.07.2026 14:02

30,000 hours of tactile data give embodied AI a sense of touch

NeoteAI and Fudan University have released a series of three reports and models under the N0 series, aiming to integrate tactile feedback as a core component of embodied AI. The project includes a 30,000-hour tactile interaction dataset, a vision-language-action model N0-VTLA, and a world model N0-TWAM, all open-sourced.
NeoteAI (Shanghai New Intelligent Embodied AI Technology Co., Ltd.) in collaboration with Fudan University's Trustworthy Embodied AI Research Institute published three technical reports for the N0 series of models. The Neodata dataset contains over 30,000 hours of visuo-tactile interaction data, 1.4 million action segments, 33 billion time steps, 8 billion RGB frames, and 10 billion tactile images, collected from 6 robot platforms (Franka, Piper, UR5e, etc.) covering 450 long-horizon tasks by 90 operators. N0-VTLA is a tactile-enhanced VLA that predicts future tactile states 50 steps ahead, achieving 85% success in plug insertion (vs. 60% for vision-only) and 99% in key extraction (vs. 35%). N0-TWAM is a world model with 7.2 billion parameters using a mixture-of-experts architecture, predicting video, tactile, and action sequences simultaneously; simulation average success rate is 84.5% (baseline 36%), real-world average 46.3% (LingBot-VA 21.9%). The project is open-sourced at research.neoteai.com.
Сокращения
VLA = Vision-Language-Action
RGB = Red Green Blue
Source: QbitAI 量子位 — original
Our earlier posts on this topic ↓
Fresh news