Egocentric human data · SenseXperience第一视角人体数据 · SenseXperience

First-person human data for embodied AI面向具身智能的
第一视角人体数据

Capture what people see — and what their hands actually do. Head-mounted egocentric video, wrist-view observations, audio and IMU, collected in the real world with IO-AI SenseXperience and turned into training-ready datasets by EmbodiFlow.既记录人看到了什么,也记录手实际做了什么。头戴第一视角视频、腕部视角、音频与 IMU,由艾欧 SenseXperience 在真实世界中采集,再经 EmbodiFlow 转化为可直接训练的数据集。

capture streams采集数据流
  • Ego Unit · stereo RGBEgo 单元 · 双目 RGBglobal shutter, 170° FoV全局快门,170° 视场角2048×1536 · 30 Hz
  • Wrist Unit · RGB腕部单元 · RGBpalm or wrist mount, 170° FoV掌部或腕部佩戴,170° 视场角1920×1080 · 30 Hz
  • IMUaccelerometer ±16 g, gyro ±2000°/s加速度计 ±16 g,陀螺仪 ±2000°/s240 Hz
  • Audio & interaction音频与交互extendable to mocap and tactile可扩展动捕与触觉synced时间同步
  • Export导出格式via EmbodiFlow经由 EmbodiFlowLeRobotHDF5MCAP

Why human data为什么是人体数据

Human data covers the long tail of the real world人体数据覆盖真实世界的长尾

As Vision-Language-Action models, Physical AI and world models advance, embodiment-agnostic human data is becoming a critical foundation for training.随着视觉-语言-动作(VLA)模型、物理 AI 与世界模型的快速发展,与本体无关的人体数据正在成为模型训练的关键基础。

Diverse task types多样的任务类型

Human demonstrations span far more tasks than a fixed robot setup can practically cover.人类示范能覆盖的任务范围,远超固定机器人平台的实际能力。

Complex manipulation strategies复杂的操作策略

Record how people actually grasp, adjust and coordinate both hands during real work.记录人在真实工作中如何抓取、调整并协调双手。

Unstructured environments非结构化环境

Collect in homes, workplaces and other everyday settings — and scale more effectively for long-tail tasks.在家庭、工作场所等日常环境中采集,更高效地扩展到长尾任务。

Ego + wrist第一视角 + 腕部视角

Global context and local interaction, in one recording一次录制,同时获得全局语境与局部交互

An egocentric camera alone often misses the moment of contact. Pairing it with wrist cameras captures both what a person sees and what actually happens during manipulation.仅靠第一视角相机,往往会错过手与物体接触的关键瞬间。将其与腕部相机组合,既能记录人看到了什么,也能记录操作过程中实际发生了什么。

Egocentric camera第一视角相机

Where and why在哪里、为什么

Global semantics and task context — helping the model understand where and why a task is happening.提供全局语义与任务语境,帮助模型理解任务在哪里发生、为什么发生。

Wrist camera腕部相机

How it is done具体怎么做

Localized physical interaction details — helping the model understand how the operation is mechanically executed.提供局部物理交互细节,帮助模型理解操作在机械层面如何执行。

Ego-only limitation仅第一视角的局限

Occlusion遮挡

Head-mounted cameras often lose sight of the exact contact point between hand and object during close-up manipulation.近距离操作时,头戴相机常常看不到手与物体的精确接触点。

Ego-only limitation仅第一视角的局限

Lost fine-grained detail细粒度信息丢失

Subtle differences in an object’s local shape, pose, texture and material are hard to capture from a distance.物体局部形状、姿态、纹理与材质的细微差异,远距离难以捕捉。

Ego-only limitation仅第一视角的局限

Background noise背景噪声

A high share of task-irrelevant pixels can lead models to rely on accidental background cues.大量与任务无关的画面,可能让模型依赖偶然的背景线索。

Bimanual bonus: in two-handed tasks, the left and right wrist cameras continuously capture each other’s manipulation from alternating angles — often keeping a clear view when the head camera is occluded.双手协同的额外收益:在双臂任务中,左右腕部相机会从不同角度持续拍到对方手的操作,即使头戴相机被遮挡,往往也能保留清晰视角。 Read the article →阅读文章 →

Capture hardware采集硬件

Lightweight, wearable, modular轻量、可穿戴、模块化

SenseXperience is IO-AI’s integrated hardware and software system for real-world human data. Modules can be deployed independently or configured together as one system.SenseXperience 是艾欧自研的真实世界人体数据采集软硬件一体系统,各模块可独立部署,也可组合成统一系统。

Ego Unit

Egocentric first-person camera第一视角相机

  • Stereo RGB, 65 mm baseline双目 RGB,基线 65 mm
  • 2048×1536 @ 30 Hz, global shutter2048×1536 @ 30 Hz,全局快门
  • 170° FoV · IMU at 240 Hz170° 视场角 · IMU 240 Hz
  • 157 ± 5 g

Wrist Unit

Wrist-mounted camera腕部相机

  • RGB 1920×1080 @ 30 Hz, 170° FoVRGB 1920×1080 @ 30 Hz,170° 视场角
  • Palm configuration: tracks hand articulation掌部佩戴:完整呈现手掌姿态与动作细节
  • Wrist configuration: avoids occlusion on large objects腕部佩戴:抓握大物体时避免遮挡
  • 62 ± 5 g

Compute Unit

Main control module主控模块

  • 1 TB SSD on-device storage1 TB SSD 本地存储
  • 2.5G Ethernet, 5× USB-C2.5G 网口,5 个 Type-C
  • Main control for the baseline kit基线套件主控
  • 214 ± 5 g

From raw capture to training data从原始采集到训练数据

One pipeline, natively integrated with EmbodiFlow一条流程,与 EmbodiFlow 原生集成

  1. Capture采集

    Wearable SenseXperience units record synchronized video, IMU, audio and interaction signals in real-world settings.SenseXperience 可穿戴模块在真实场景中同步记录视频、IMU、音频与交互信号。

  2. Process后处理

    IO-Agent-Lab custom-trained agents run timestamp alignment, pose estimation and gesture recognition.IO-Agent-Lab 定制训练的智能体完成时间对齐、位姿估计与手势识别。

  3. Review & QA审核与质检

    Automated and manual QA in EmbodiFlow filters noisy, incomplete or low-quality samples before training.EmbodiFlow 自动与人工质检流程,在训练前过滤噪声、不完整或低质量样本。

  4. Export导出

    Training-ready datasets in LeRobot, HDF5 and MCAP, compatible with ROS and popular robot-learning workflows.导出 LeRobot、HDF5、MCAP 等可训练格式,兼容 ROS 与主流机器人学习工作流。

Try it on real data: the SenseXperience WristCam sample dataset is on Hugging Face, and you can open recordings in the browser with ROSView, IO-AI’s open-source robotics data viewer.用真实数据体验:SenseXperience 腕部相机样例数据集已在 Hugging Face 公开,可用艾欧开源的机器人数据可视化工具 ROSView 在浏览器中直接查看。

FAQ

Frequently asked questions常见问题

What modalities can be captured?能采集哪些模态的数据?

The baseline kit captures egocentric video, wrist-view observations, gripper interactions, audio and IMU signals, with room to extend into motion capture, tactile sensing and other modalities.基线方案覆盖第一视角视频、腕部视角、夹爪交互、音频与 IMU 等关键信号,并支持按场景扩展动捕、触觉等更丰富的多模态采集能力。

How does captured data become training-ready?采集的数据如何变成可训练格式?

Native EmbodiFlow integration carries raw capture into timestamp alignment, pose estimation, gesture recognition, review and export — supporting formats such as LeRobot, HDF5 and MCAP.硬件与 EmbodiFlow 原生集成,可将原始数据送入时间对齐、位姿估计、手势识别、审核与导出流程,衔接 LeRobot、HDF5、MCAP 等训练与交付格式。

Can the setup be customized?采集系统可以定制吗?

Yes. The modular architecture supports customization across cameras, grippers, sensors and data-processing pipelines — you define the data you need, IO-AI designs the system.可以。模块化架构支持相机、夹爪、传感器与数据处理流程的定制——你定义需要的数据,我们来设计系统。

Where can I find documentation?在哪里查看产品文档?

The SenseXperience product documentation covers safety, hardware connection, setup, the software workbench, data formats, FAQ and camera calibration.SenseXperience 产品文档涵盖安全须知、硬件连接、使用前准备、软件工作台、数据格式说明、常见问题与相机标定等内容。

From wearable capture to training-ready data.从可穿戴采集到可训练数据。

For scaling, less is more. You define, we design — tell us about your tasks and data requirements.规模化采集,少即是多。你定义需求,我们设计系统——欢迎告诉我们你的任务与数据需求。