@Adamjhung
Learning from human videos often requires restrictive, carefully choreographed human motions. We propose ✨3PoinTr✨: a scalable way to pretrain from casual human videos. It bridges the embodiment gap by learning 3D scene evolution, enabling learning from natural human motions. https://t.co/B9w7PtDUAt