With all the pieces in place, including the TR01 design enabling precise teleoperation and integrated touch sensing with full fingertip coverage, it was only a matter of time before we put them together for autonomous dexterous skills.
Today we are showing examples of autonomous tasks taking advantage of a) dexterous hand kinematics; b) skilled teleoperation via exact fingertip tracking; c) combination of touch and vision sensing. All these tasks are achieved via Imitation Learning, trained on a few dozen complete teleoperated demonstrations and a similar number of error correction demonstrations, and using both vision and touch data from the Popcorn sensors as input. All executions shown below are representative for their respective tasks, fully autonomous, and shown at 1X speed.
For some of the tasks we’ve tried (Container Flip, SD Card Pickup), the “out of the box” Imitation Learning policy works essentially every time, with 100% success rate in our tests. In other cases, the policy works most of the time, as shown in the graph below. While there could be applications where such an “out of the box” policy could be immediately suitable for deployment, we always expected them to serve as a foundation for additional performance and speed improvements, achieved via on-robot learning and refinement, self-correcting behaviors (both high- and low-level), etc.
For a subset of tasks (Container Flip, Unthread, Thread and Gear Install), we also ran a number of ablations to see what each sensing modality brings to the table. Vision-only results are shown in red in the graph above, and a qualitative example of a visuotactile vs. vision-only policy can also be found in the accompanying video. As expected, the more difficult the task the higher the importance of touch sensing for performance. We expect this phenomenon to become even more pronounced as we start pushing the boundaries in terms of execution speed or control frequency.
What’s Next
We are building on these results in multiple ways. From a technical perspective, we are working on both robustness (success rate) and speed increases for these policies, looking beyond Imitation Learning. In parallel, we are working with our pilot commercial partners to turn these skills into full applications, with particular focus on skilled assembly and manufacturing. We chose all the tasks shown today because they are representative of skills needed in the verticals we are investigating, and elicit similar abilities. This allows us to build performance that is both general and relevant to real world applications.
The TR01 has proven to be an exceptional platform for skilled manipulation — robust, capable and amenable to more dexterous teleoperation than any other platform we know. We can only imagine what TR02 will be capable of. Wait, we can actually do more than just imagine. Stay tuned ;)
Technical Details
Policy Architecture and Training
All examples shown here are achieved via Diffusion Policies with Transformer backbones running at 15 Hz. Image observations are provided at 15 Hz using a pre-trained ViT encoder fine-tuned during policy training. Touch data from the Popcorn sensor is provided at 240 Hz, through a convolution encoder trained from scratch (we are currently investigating the use of pre-trained touch encoders as well). We use a 2-step context for images, a 4-step context for touch and proprioception, an 8-step prediction horizon and a 6-step execution horizon. For each task, we use 30 to 40 teleoperated demonstrations of the complete task and a similar number of shorter demonstrations of task subcomponents showing how to avoid common failure modes. Overall, demonstration collection for each task takes approximately one hour.
We have also observed autonomous task execution to be qualitatively similar in terms of speed to the demonstrations. Thus, speeding up demonstrations, which we haven’t really pushed for, might also result in a similar autonomous speed up, with the caveat that, even with our exact fingertip tracking, any teleoperation interface still slows down highly dexterous manipulation to some degree. We are building into our stack a range of methods to increase both success rates and execution speed without additional or different demonstrations, and we hope to report on those in the near future.